Blink: CPU-Free LLM Inference by Delegating the Serving Stack to GPU and SmartNIC

arXiv 2604.07609•3aa85a39a1b8f0f099e40372a889e19d672033e4ca22021b3dea6f679eaa4c74
DVFSGPU kernelHPC cost modellingLLM inferenceRDMASmartNICTCO analysisadministrative decentralizationagent safetycoflow schedulingdatacenter architecturediffusion modelsedge-cloudenergy efficiencymicro-servingmodel decompositionmodel servingmulti-agent systemsoptical circuit switchingperformance optimizationreliabilitysatellite computeshared logsspace-based AIthermal design

Paper metadata

arXiv ID
2604.07609
Version
Not specified by this published record
Category
Computer Science — Distributed, Parallel, and Cluster Computing (cs.DC)

The PDF link points to arxiv.org. Baitaphish does not expose a private stored PDF.

Evidence and limitations

Source ID
arxiv_cs_dc
Record identifier
3aa85a39a1b8f0f099e40372a889e19d672033e4ca22021b3dea6f679eaa4c74
Enrichment time
2026-04-10T08:52:20Z
AI-assisted enrichment
Yes

This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.