Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism
arXiv 2605.05467•e7131658c709b5430d64cd5896e3f0d63ffcc6bc7f31006c669ca0c1a0aef9f3
AscendFHEGEMM optimizationKV cache migrationLLM servingMoESMCcontent-addressed cachingdata-parallel load balancingdeadline-aware schedulingdifferential privacyearly-exit inferenceedge inferencematrix multiplicationmodel extractionmulti-DNN servingonline routingperformance optimizationposition-independent cachingprivacy-preserving MLtensor parallelismthreshold automataverification tools
Paper metadata
- arXiv ID
- 2605.05467
- Version
- Not specified by this published record
- Category
- Computer Science — Distributed, Parallel, and Cluster Computing (cs.DC)
The PDF link points to arxiv.org. Baitaphish does not expose a private stored PDF.
Evidence and limitations
- Source ID
- arxiv_cs_dc
- Record identifier
- e7131658c709b5430d64cd5896e3f0d63ffcc6bc7f31006c669ca0c1a0aef9f3
- Enrichment time
- 2026-05-08T08:52:26Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.