StageFrontier: Synchronization-Aware Stage Accounting for Distributed ML Training
2026-06-08T08:52:27Z•d428e9cd032cd913f8a3015665eb5045f68b4bcdbcda7ed855e0eafc3feefa9c
CPMLCUDADDPDxPTAFP8GPU-peer-to-peerGlooNCCLOzaki-IIPCCLPyTorchStageFrontiercollective-communicationdistributed-mlheterogeneous-acceleratorshigh-performance-computing (HPC)layer-variantsmulti-gpuobservabilityphotonic-acceleratorsprocess-groupsprofilingreal-time-dnnschedulingsynchronization
What happened
Collection of recent systems and ML infrastructure papers (June 2026). Key contributions: StageFrontier — a low-overhead, always-on synchronization-aware accounting signal for distributed ML that pinpoints which stage and rank first expose group-visible delays (works with PyTorch, Gloo, NCCL). Terastal — layer-variant design and scheduling to reduce deadline misses on heterogeneous DNN accelerators. Multi-GPU 3D FDTD study — communication strategies showing direct GPU-to-GPU peer exchange dominates. PCCL — process-group-aware, scalable synthesizer for topology-aware collective algorithms. A治理:
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_dc
- Record identifier
- d428e9cd032cd913f8a3015665eb5045f68b4bcdbcda7ed855e0eafc3feefa9c
- Enrichment time
- 2026-06-08T08:52:27Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.