Energy Efficient Scheduling of AI/ML Workloads on Multi Instance GPUs with Dynamic Repartitioning

2026-06-25T08:52:26Z60f023437afe6222249304221710f33afa9bd8a6b0a5eb5f387a559c847c2ddc
BFTBitcoinByzantine-fault-toleranceGEMMGPUMIGPairHMMai-mlarxivdata-centerdecentralized-governanceedge-cloudedge-systemsenergy-efficiencygenomicsgrid-interactivep-bitsprobabilistic-computingreinforcement-learningresearchschedulingservice-placementspeculative-decodingtensor-coresuncertainty-quantification

What happened

This is an arXiv CS digest (multiple new submissions) covering systems and ML infrastructure research. Key highlights: a dynamic-repartitioning scheduler for NVIDIA Multi-Instance GPUs (MIGs) that uses simulations and reinforcement learning to cut energy and tardiness vs. static/no partitioning; an analytical and empirical study of distributed speculative decoding (edge draft model + cloud target) that shows limited per-request latency benefit over co-located speculative decoding except in low-RTT/multi-tenant scenarios; an architecture and real-cluster experiments for grid-responsive, power‑/

Why it matters

A reviewed impact interpretation has not been published for this record.

Evidence and limitations

Source ID
arxiv_cs_dc
Record identifier
60f023437afe6222249304221710f33afa9bd8a6b0a5eb5f387a559c847c2ddc
Enrichment time
2026-06-25T08:52:26Z
AI-assisted enrichment
Yes

This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.