Regulating Branch Parallelism in LLM Serving
2026-05-11T08:52:25Z•1a71f712582c1d3d3e143c657f38604a2d328814583cddd65d12e22bb4f037f4
AI-RANGPU accelerationHPCKV cachingLLM servingSpMVaccelerators (Cerebras, Tenstorrent)admission controlbranch parallelismheterogeneous clusterslong-context trainingrecommendation systemsschedulingsparse matrixworkflow scheduling
What happened
Collection of systems and HPC research on efficient execution of LLMs and scientific kernels. Key contributions include TAPER (per-step admission control to regulate branch parallelism in LLM serving, improving goodput up to 1.77× vs baselines while preserving SLOs), FATE (future-state-aware scheduler for heterogeneous multi-stage LLM workflows that reduces makespan and P95 latency by ~32% vs simple heuristics), RcLLM (distributed generative-recommendation inference with beyond-prefix KV caching reducing TTFT 1.31–9.51×), HEXISEQ (heterogeneous CP/HP for long-context LLM training improving avg
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_dc
- Record identifier
- 1a71f712582c1d3d3e143c657f38604a2d328814583cddd65d12e22bb4f037f4
- Enrichment time
- 2026-05-11T08:52:25Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.