Analyzing Reverse Address Translation Overheads in Multi-GPU Scale-Up Pods
2026-04-06T08:52:24Z•d0a6710cdfe150a466aebfdea762fb9f04db503028c30d1d01ade7855762f8c0
ASTRA-simCIDERCPU-GPU interconnectDawnKV cache sharingLink MMULink TLBMetalNVLinkOmnet++TLB prefetchingTokenDanceUALinkVulkanWebGPUdispatch overheadfused pre-translationheterogeneous memory managementkernel fusionmemory-disaggregated KV storesmulti-GPUpessimistic synchronizationreverse address translationvLLMwgpu-native
What happened
This collection of systems papers (arXiv, 06 Apr 2026) focuses on performance and scalability optimizations for large-scale ML, simulation, and distributed systems. Key results include: (1) a first systematic study of destination-side Reverse Address Translation in multi-GPU scale-up pods showing cold Link TLB misses can cause up to 1.4x latency degradation and proposing fused pre-translation kernels and software-guided TLB prefetching; (2) heterogeneous memory management techniques that use host memory to overcome GPU memory limits for large-scale nonlinear time-history simulations and to生成s
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_dc
- Record identifier
- d0a6710cdfe150a466aebfdea762fb9f04db503028c30d1d01ade7855762f8c0
- Enrichment time
- 2026-04-06T08:52:24Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.