CRAFT: Cost-aware Expert Replica Allocation with Fine-Grained Layerwise Estimations

2026-04-01T08:52:18Z7b7dfb81c7c6e9efc9ed2270906b618ae98e29cce6c707719fd7c48f4241e8ac
BYZANTINE-ROBUSTDelta LakeGPU memoryGPU optimizationKV cacheLLM evaluationMixture-of-ExpertsNextflow monitoringautomatic differentiationcommunication efficiencydistributed evaluationdistributed trainingexpert replicationfederated inferencegradient codinghardware accelerationheterogeneous LLMsload balancingmodel servingmonitoring pipelineobservabilityprivacyray tracing coresresponse cachingstatistical rigor

What happened

This batch of arXiv submissions covers system, ML-serving, and distributed-compute research with several security-relevant implications. Key points: CRAFT (expert replication for MoE) shows over-replication can waste GPU memory and cause resource contention and throughput degradation—risking resource exhaustion or DoS in model-serving fleets. Spark-LLM-Eval introduces large-scale, distributed LLM evaluation with response caching (content-addressable cache) that improves cost but could expose cached model outputs or enable data leakage if caching/backing stores are misconfigured. FedRefine (f¶d

Why it matters

A reviewed impact interpretation has not been published for this record.

Evidence and limitations

Source ID
arxiv_cs_dc
Record identifier
7b7dfb81c7c6e9efc9ed2270906b618ae98e29cce6c707719fd7c48f4241e8ac
Enrichment time
2026-04-01T08:52:18Z
AI-assisted enrichment
Yes

This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.