Blink: CPU-Free LLM Inference by Delegating the Serving Stack to GPU and SmartNIC
arXiv 2604.07609•3aa85a39a1b8f0f099e40372a889e19d672033e4ca22021b3dea6f679eaa4c74
DVFSGPU kernelHPC cost modellingLLM inferenceRDMASmartNICTCO analysisadministrative decentralizationagent safetycoflow schedulingdatacenter architecturediffusion modelsedge-cloudenergy efficiencymicro-servingmodel decompositionmodel servingmulti-agent systemsoptical circuit switchingperformance optimizationreliabilitysatellite computeshared logsspace-based AIthermal design
Paper metadata
- arXiv ID
- 2604.07609
- Version
- Not specified by this published record
- Category
- Computer Science — Distributed, Parallel, and Cluster Computing (cs.DC)
The PDF link points to arxiv.org. Baitaphish does not expose a private stored PDF.
Evidence and limitations
- Source ID
- arxiv_cs_dc
- Record identifier
- 3aa85a39a1b8f0f099e40372a889e19d672033e4ca22021b3dea6f679eaa4c74
- Enrichment time
- 2026-04-10T08:52:20Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.