CRAFT: Cost-aware Expert Replica Allocation with Fine-Grained Layerwise Estimations
arXiv 2603.28768•7b7dfb81c7c6e9efc9ed2270906b618ae98e29cce6c707719fd7c48f4241e8ac
BYZANTINE-ROBUSTDelta LakeGPU memoryGPU optimizationKV cacheLLM evaluationMixture-of-ExpertsNextflow monitoringautomatic differentiationcommunication efficiencydistributed evaluationdistributed trainingexpert replicationfederated inferencegradient codinghardware accelerationheterogeneous LLMsload balancingmodel servingmonitoring pipelineobservabilityprivacyray tracing coresresponse cachingstatistical rigor
Paper metadata
- arXiv ID
- 2603.28768
- Version
- Not specified by this published record
- Category
- Computer Science — Distributed, Parallel, and Cluster Computing (cs.DC)
The PDF link points to arxiv.org. Baitaphish does not expose a private stored PDF.
Evidence and limitations
- Source ID
- arxiv_cs_dc
- Record identifier
- 7b7dfb81c7c6e9efc9ed2270906b618ae98e29cce6c707719fd7c48f4241e8ac
- Enrichment time
- 2026-04-01T08:52:18Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.