DAK: Direct-Access-Enabled GPU Memory Offloading with Optimal Efficiency for LLM Inference
arXiv 2604.26074•a0bb816a742d891d69bb66c3bfc75f212c68cd0fbea8985341e418c8a0bbffa1
DAKDMRlibDUAL-BLADEFloatSOMGPU memory offloadingKV cache offloadingLLM inferenceLoRAMPI malleabilityMixture-of-Experts (MoE)NVLinkNVMe-directPCIeSMEMSOMSplitFTTMAclient inferencedynamic resource managementfederated learningmulti-GPUout-of-memory streamingpage cache bypasspipelined shardingxLM
Paper metadata
- arXiv ID
- 2604.26074
- Version
- Not specified by this published record
- Category
- Computer Science — Distributed, Parallel, and Cluster Computing (cs.DC)
The PDF link points to arxiv.org. Baitaphish does not expose a private stored PDF.
Evidence and limitations
- Source ID
- arxiv_cs_dc
- Record identifier
- a0bb816a742d891d69bb66c3bfc75f212c68cd0fbea8985341e418c8a0bbffa1
- Enrichment time
- 2026-04-30T08:52:30Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.