Analyzing Reverse Address Translation Overheads in Multi-GPU Scale-Up Pods

2026-04-06T08:52:24Zd0a6710cdfe150a466aebfdea762fb9f04db503028c30d1d01ade7855762f8c0
ASTRA-simCIDERCPU-GPU interconnectDawnKV cache sharingLink MMULink TLBMetalNVLinkOmnet++TLB prefetchingTokenDanceUALinkVulkanWebGPUdispatch overheadfused pre-translationheterogeneous memory managementkernel fusionmemory-disaggregated KV storesmulti-GPUpessimistic synchronizationreverse address translationvLLMwgpu-native

What happened

This collection of systems papers (arXiv, 06 Apr 2026) focuses on performance and scalability optimizations for large-scale ML, simulation, and distributed systems. Key results include: (1) a first systematic study of destination-side Reverse Address Translation in multi-GPU scale-up pods showing cold Link TLB misses can cause up to 1.4x latency degradation and proposing fused pre-translation kernels and software-guided TLB prefetching; (2) heterogeneous memory management techniques that use host memory to overcome GPU memory limits for large-scale nonlinear time-history simulations and to生成s

Why it matters

A reviewed impact interpretation has not been published for this record.

Evidence and limitations

Source ID
arxiv_cs_dc
Record identifier
d0a6710cdfe150a466aebfdea762fb9f04db503028c30d1d01ade7855762f8c0
Enrichment time
2026-04-06T08:52:24Z
AI-assisted enrichment
Yes

This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.