Dynamic Hierarchical Birkhoff-von Neumann Decomposition for All-to-All GPU Communication
2026-03-04T19:51:06Z•1e9870ff6b0e7a09bacc1cdc1c92575e520062350def19503bd83df880e40d15
Birkhoff-von NeumannGPU-fabricsLLM trainingNICRDMAall-to-all communicationavailabilitycross-NIC failoverdistributed trainingdynamic frame sizingexactly-oncefault-tolerancekernel-bypassperformancerdma-corereceiver-opacityresilienceschedulingsystem-designzero-copy
What happened
Two systems papers on GPU/cluster networking and RDMA fault tolerance. The first proposes a dynamic hierarchical Birkhoff–von Neumann decomposition plus dynamic frame sizing for two-tier GPU fabrics to reduce all-to-all completion time, mitigate intra-server GPU/NIC skew, and provide provable stability for online scheduling. The second (SHIFT) proves an impossibility Trilemma for cross‑NIC RDMA failover (Exactly‑Once Execution, Receiver‑NIC Opacity, and Zero‑Copy cannot all be preserved) and presents a user-space rdma-core implementation that achieves cross‑NIC failover by relying on idempotid
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_ni
- Record identifier
- 1e9870ff6b0e7a09bacc1cdc1c92575e520062350def19503bd83df880e40d15
- Enrichment time
- 2026-03-04T19:51:06Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.