Energy Efficient Scheduling of AI/ML Workloads on Multi Instance GPUs with Dynamic Repartitioning
2026-06-25T08:52:26Z•60f023437afe6222249304221710f33afa9bd8a6b0a5eb5f387a559c847c2ddc
BFTBitcoinByzantine-fault-toleranceGEMMGPUMIGPairHMMai-mlarxivdata-centerdecentralized-governanceedge-cloudedge-systemsenergy-efficiencygenomicsgrid-interactivep-bitsprobabilistic-computingreinforcement-learningresearchschedulingservice-placementspeculative-decodingtensor-coresuncertainty-quantification
What happened
This is an arXiv CS digest (multiple new submissions) covering systems and ML infrastructure research. Key highlights: a dynamic-repartitioning scheduler for NVIDIA Multi-Instance GPUs (MIGs) that uses simulations and reinforcement learning to cut energy and tardiness vs. static/no partitioning; an analytical and empirical study of distributed speculative decoding (edge draft model + cloud target) that shows limited per-request latency benefit over co-located speculative decoding except in low-RTT/multi-tenant scenarios; an architecture and real-cluster experiments for grid-responsive, power‑/
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_dc
- Record identifier
- 60f023437afe6222249304221710f33afa9bd8a6b0a5eb5f387a559c847c2ddc
- Enrichment time
- 2026-06-25T08:52:26Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.