Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization

arXiv 2604.21072•b34efecf4c6887a8edcc76b7663138314482d0ea2a37e328f27ce711b00e1257
BM25BloomBeeCantelli boundGPU accelerationGWAS privacy riskLLM plannerOpenSearchSLOTorchGWASWiFi offloadcompressiondistributed LLM inferencedynamic programming optimizationedge computingedge selectiongenomicshuman-in-controlhybrid retrievalhysteresismicro-batchingrisk-aware selectionsemantic embeddingsspeculative decodingtask decompositiontensor offload

Paper metadata

arXiv ID
2604.21072
Version
Not specified by this published record
Category
Computer Science — Distributed, Parallel, and Cluster Computing (cs.DC)

The PDF link points to arxiv.org. Baitaphish does not expose a private stored PDF.

Evidence and limitations

Source ID
arxiv_cs_dc
Record identifier
b34efecf4c6887a8edcc76b7663138314482d0ea2a37e328f27ce711b00e1257
Enrichment time
2026-04-24T08:52:20Z
AI-assisted enrichment
Yes

This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.