Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
arXiv 2604.21072•b34efecf4c6887a8edcc76b7663138314482d0ea2a37e328f27ce711b00e1257
BM25BloomBeeCantelli boundGPU accelerationGWAS privacy riskLLM plannerOpenSearchSLOTorchGWASWiFi offloadcompressiondistributed LLM inferencedynamic programming optimizationedge computingedge selectiongenomicshuman-in-controlhybrid retrievalhysteresismicro-batchingrisk-aware selectionsemantic embeddingsspeculative decodingtask decompositiontensor offload
Paper metadata
- arXiv ID
- 2604.21072
- Version
- Not specified by this published record
- Category
- Computer Science — Distributed, Parallel, and Cluster Computing (cs.DC)
The PDF link points to arxiv.org. Baitaphish does not expose a private stored PDF.
Evidence and limitations
- Source ID
- arxiv_cs_dc
- Record identifier
- b34efecf4c6887a8edcc76b7663138314482d0ea2a37e328f27ce711b00e1257
- Enrichment time
- 2026-04-24T08:52:20Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.