ITME: Inference Tiered Memory Expansion with Disaggregated CXL-Hybrid Memories
2026-06-12T08:52:20Z•369d0a4b06ed7af0b807c1cde8c201396096a4a6ce6acd8d3e6ff711ca7905cd
AMD MI250XARM SMECXLDPUFPGA prototypeGPU interconnectsHPC optimizationsKV cacheLLM inferenceNVIDIA A100NVMe-oFOpenACCOpenMPPCIe Gen5SK Hynix CMMSPECFEM3DSemantic mappingXRdevice-clouddisaggregated memoryfederated learning (FL)gem5 extensionlow-power edgemulti-GPU communicationperformance portability
What happened
This feed aggregates recent systems and ML-systems research (arXiv Jun 12 2026) focused on scaling and performance of large-model inference, multi-GPU/disaggregated memory architectures, device-cloud pipelines, and HPC optimizations. Highlights: ITME proposes using CXL-hybrid remote byte-addressable memory (validated with SK Hynix CMM, PCIe Gen5 NVMe SSDs and an FPGA prototype) to expand KV-cache capacity for LLM inference; Eidola extends gem5 to model fine-grained inter-GPU peer-to-peer traffic; a study shows limits of directive-based GPU portability (OpenACC vs OpenMP) across NVIDIA/AMD GPUs
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_dc
- Record identifier
- 369d0a4b06ed7af0b807c1cde8c201396096a4a6ce6acd8d3e6ff711ca7905cd
- Enrichment time
- 2026-06-12T08:52:20Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.