ShardTensor: Domain Parallelism for Scientific Machine Learning

2026-05-13T08:52:25Zf08d2600b0319cc42e2e8c7f4d8c9249cd927352dcd777c09d3210951c0312e7
Byzantine consensusChakraChunkFlowDeFiDeFiPyGNN trainingGriNNderLLM pre-trainingMLCommonsNVMePCIe contentionReCoVerShardTensorState Twincollective communicationsdiffusion transformersdomain parallelismexecution tracefault tolerancelayerwise offloadingmessage authenticationoff-chain simulationscientific machine learningstorage offloadingvector search on-SSD

What happened

Feed of recent arXiv CS/DC papers (2026-05-13) covering systems and algorithms for large-scale machine learning, graph/vector processing, decentralized systems, and DeFi tooling. Key contributions include: ShardTensor — a domain-parallelism paradigm to scale scientific ML inputs beyond device constraints; ReCoVer — a resilient LLM pre-training system with fault-tolerant collectives and in-step recovery that preserves training trajectory under massive GPU loss; a characterization of Byzantine consensus in directed graphs with message authentication (necessary/sufficient conditions); Chakra — an

Why it matters

A reviewed impact interpretation has not been published for this record.

Evidence and limitations

Source ID
arxiv_cs_dc
Record identifier
f08d2600b0319cc42e2e8c7f4d8c9249cd927352dcd777c09d3210951c0312e7
Enrichment time
2026-05-13T08:52:25Z
AI-assisted enrichment
Yes

This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.