Blink: CPU-Free LLM Inference by Delegating the Serving Stack to GPU and SmartNIC
2026-04-10T08:52:20Z•3aa85a39a1b8f0f099e40372a889e19d672033e4ca22021b3dea6f679eaa4c74
DVFSGPU kernelHPC cost modellingLLM inferenceRDMASmartNICTCO analysisadministrative decentralizationagent safetycoflow schedulingdatacenter architecturediffusion modelsedge-cloudenergy efficiencymicro-servingmodel decompositionmodel servingmulti-agent systemsoptical circuit switchingperformance optimizationreliabilitysatellite computeshared logsspace-based AIthermal design
What happened
This feed collects new systems and architecture research focused on high-throughput AI inference, resource-efficient compute, and distributed coordination for datacenter, edge, and space environments. Highlights include Blink, which eliminates host-CPU involvement in steady-state LLM inference by offloading request handling to a SmartNIC and moving batching/scheduling/KV-cache management into a persistent GPU kernel (up to 8.47x P99 TTFT and 48.6% energy-per-token improvements vs. baselines); a proposal for integrated solar/compute/radiator panels to dramatically improve specific power for on‑
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_dc
- Record identifier
- 3aa85a39a1b8f0f099e40372a889e19d672033e4ca22021b3dea6f679eaa4c74
- Enrichment time
- 2026-04-10T08:52:20Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.