Blink: CPU-Free LLM Inference by Delegating the Serving Stack to GPU and SmartNIC

2026-04-10T08:52:20Z3aa85a39a1b8f0f099e40372a889e19d672033e4ca22021b3dea6f679eaa4c74
DVFSGPU kernelHPC cost modellingLLM inferenceRDMASmartNICTCO analysisadministrative decentralizationagent safetycoflow schedulingdatacenter architecturediffusion modelsedge-cloudenergy efficiencymicro-servingmodel decompositionmodel servingmulti-agent systemsoptical circuit switchingperformance optimizationreliabilitysatellite computeshared logsspace-based AIthermal design

What happened

This feed collects new systems and architecture research focused on high-throughput AI inference, resource-efficient compute, and distributed coordination for datacenter, edge, and space environments. Highlights include Blink, which eliminates host-CPU involvement in steady-state LLM inference by offloading request handling to a SmartNIC and moving batching/scheduling/KV-cache management into a persistent GPU kernel (up to 8.47x P99 TTFT and 48.6% energy-per-token improvements vs. baselines); a proposal for integrated solar/compute/radiator panels to dramatically improve specific power for on‑

Why it matters

A reviewed impact interpretation has not been published for this record.

Evidence and limitations

Source ID
arxiv_cs_dc
Record identifier
3aa85a39a1b8f0f099e40372a889e19d672033e4ca22021b3dea6f679eaa4c74
Enrichment time
2026-04-10T08:52:20Z
AI-assisted enrichment
Yes

This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.