ITME: Inference Tiered Memory Expansion with Disaggregated CXL-Hybrid Memories

2026-06-12T08:52:20Z369d0a4b06ed7af0b807c1cde8c201396096a4a6ce6acd8d3e6ff711ca7905cd
AMD MI250XARM SMECXLDPUFPGA prototypeGPU interconnectsHPC optimizationsKV cacheLLM inferenceNVIDIA A100NVMe-oFOpenACCOpenMPPCIe Gen5SK Hynix CMMSPECFEM3DSemantic mappingXRdevice-clouddisaggregated memoryfederated learning (FL)gem5 extensionlow-power edgemulti-GPU communicationperformance portability

What happened

This feed aggregates recent systems and ML-systems research (arXiv Jun 12 2026) focused on scaling and performance of large-model inference, multi-GPU/disaggregated memory architectures, device-cloud pipelines, and HPC optimizations. Highlights: ITME proposes using CXL-hybrid remote byte-addressable memory (validated with SK Hynix CMM, PCIe Gen5 NVMe SSDs and an FPGA prototype) to expand KV-cache capacity for LLM inference; Eidola extends gem5 to model fine-grained inter-GPU peer-to-peer traffic; a study shows limits of directive-based GPU portability (OpenACC vs OpenMP) across NVIDIA/AMD GPUs

Why it matters

A reviewed impact interpretation has not been published for this record.

Evidence and limitations

Source ID
arxiv_cs_dc
Record identifier
369d0a4b06ed7af0b807c1cde8c201396096a4a6ce6acd8d3e6ff711ca7905cd
Enrichment time
2026-06-12T08:52:20Z
AI-assisted enrichment
Yes

This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.