DAK: Direct-Access-Enabled GPU Memory Offloading with Optimal Efficiency for LLM Inference
2026-04-30T08:52:30Z•a0bb816a742d891d69bb66c3bfc75f212c68cd0fbea8985341e418c8a0bbffa1
DAKDMRlibDUAL-BLADEFloatSOMGPU memory offloadingKV cache offloadingLLM inferenceLoRAMPI malleabilityMixture-of-Experts (MoE)NVLinkNVMe-directPCIeSMEMSOMSplitFTTMAclient inferencedynamic resource managementfederated learningmulti-GPUout-of-memory streamingpage cache bypasspipelined shardingxLM
What happened
Collection of recent systems research arXiv announcements (Apr 30 2026) focused on large-model inference scalability, memory-tiering, and high-performance kernels. Key contributions: DAK — direct GPU access to remote memory using Tensor Memory Accelerator (TMA) and SMEM for efficient LLM weight/KV offload; DUAL-BLADE — dual-path NVMe-direct KV-cache offloading that maps KV tensors to contiguous LBAs to bypass filesystem/page cache; pipelined sharding — CPU/GPU hybrid scheduling for VRAM-constrained client xLM inference; SplitFT — adaptive federated split learning for LLM fine-tuning with cut‑层
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_dc
- Record identifier
- a0bb816a742d891d69bb66c3bfc75f212c68cd0fbea8985341e418c8a0bbffa1
- Enrichment time
- 2026-04-30T08:52:30Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.