Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism

2026-05-08T08:52:26Ze7131658c709b5430d64cd5896e3f0d63ffcc6bc7f31006c669ca0c1a0aef9f3
AscendFHEGEMM optimizationKV cache migrationLLM servingMoESMCcontent-addressed cachingdata-parallel load balancingdeadline-aware schedulingdifferential privacyearly-exit inferenceedge inferencematrix multiplicationmodel extractionmulti-DNN servingonline routingperformance optimizationposition-independent cachingprivacy-preserving MLtensor parallelismthreshold automataverification tools

What happened

Feed of newly announced distributed-systems / ML-systems papers (arXiv) focused on efficient, low-latency LLM and DNN serving, edge inference, MoE and tensor-parallel optimizations, and privacy-preserving edge ML. Key contributions include Nitsum (runtime-adaptive tensor parallelism with fast KV migration and weight reuse), Irminsul (content-addressed, position-independent caching for agentic LLM workloads), BalanceRoute (online data-parallel routing to reduce DP imbalance), EdgeServing (deadline-aware multi-DNN scheduling with early-exit), relay-buffer-free MoE communication for Ascend, and a

Why it matters

A reviewed impact interpretation has not been published for this record.

Evidence and limitations

Source ID
arxiv_cs_dc
Record identifier
e7131658c709b5430d64cd5896e3f0d63ffcc6bc7f31006c669ca0c1a0aef9f3
Enrichment time
2026-05-08T08:52:26Z
AI-assisted enrichment
Yes

This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.

Record · Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism · Baitaphish