Dynamic Hierarchical Birkhoff-von Neumann Decomposition for All-to-All GPU Communication

2026-03-04T19:51:06Z1e9870ff6b0e7a09bacc1cdc1c92575e520062350def19503bd83df880e40d15
Birkhoff-von NeumannGPU-fabricsLLM trainingNICRDMAall-to-all communicationavailabilitycross-NIC failoverdistributed trainingdynamic frame sizingexactly-oncefault-tolerancekernel-bypassperformancerdma-corereceiver-opacityresilienceschedulingsystem-designzero-copy

What happened

Two systems papers on GPU/cluster networking and RDMA fault tolerance. The first proposes a dynamic hierarchical Birkhoff–von Neumann decomposition plus dynamic frame sizing for two-tier GPU fabrics to reduce all-to-all completion time, mitigate intra-server GPU/NIC skew, and provide provable stability for online scheduling. The second (SHIFT) proves an impossibility Trilemma for cross‑NIC RDMA failover (Exactly‑Once Execution, Receiver‑NIC Opacity, and Zero‑Copy cannot all be preserved) and presents a user-space rdma-core implementation that achieves cross‑NIC failover by relying on idempotid

Why it matters

A reviewed impact interpretation has not been published for this record.

Evidence and limitations

Source ID
arxiv_cs_ni
Record identifier
1e9870ff6b0e7a09bacc1cdc1c92575e520062350def19503bd83df880e40d15
Enrichment time
2026-03-04T19:51:06Z
AI-assisted enrichment
Yes

This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.