ShardTensor: Domain Parallelism for Scientific Machine Learning
2026-05-13T08:52:25Z•f08d2600b0319cc42e2e8c7f4d8c9249cd927352dcd777c09d3210951c0312e7
Byzantine consensusChakraChunkFlowDeFiDeFiPyGNN trainingGriNNderLLM pre-trainingMLCommonsNVMePCIe contentionReCoVerShardTensorState Twincollective communicationsdiffusion transformersdomain parallelismexecution tracefault tolerancelayerwise offloadingmessage authenticationoff-chain simulationscientific machine learningstorage offloadingvector search on-SSD
What happened
Feed of recent arXiv CS/DC papers (2026-05-13) covering systems and algorithms for large-scale machine learning, graph/vector processing, decentralized systems, and DeFi tooling. Key contributions include: ShardTensor — a domain-parallelism paradigm to scale scientific ML inputs beyond device constraints; ReCoVer — a resilient LLM pre-training system with fault-tolerant collectives and in-step recovery that preserves training trajectory under massive GPU loss; a characterization of Byzantine consensus in directed graphs with message authentication (necessary/sufficient conditions); Chakra — an
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_dc
- Record identifier
- f08d2600b0319cc42e2e8c7f4d8c9249cd927352dcd777c09d3210951c0312e7
- Enrichment time
- 2026-05-13T08:52:25Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.