Cost-Aware Query Routing in RAG: Empirical Analysis of Retrieval Depth Tradeoffs
2026-06-03T08:52:25Z•a999019560edc806f76095753f0102f0bc9aee8fa1d1f3fb65fb540fe6469b67
BAHSDBM25CA-RAGFAISSHNSWRAGSlipstreamattention-calibrationbiasblack-box-distillationcost-aware-RAGdense-retrievalfairnessfindability-gaphybrid-searchlegal-retrievalneural-retrieverspositional-biasreciprocal-rank-fusionrelevance-priorretrieval-augmented-generationsection-aware-retrievalsequential-recommendationstreaming-ANNSvirtual-MLE','LLM-agent'
What happened
Collection of recent IR/ML papers (arXiv June 3, 2026) focused on retrieval-augmented generation, dense retrieval, streaming ANNS, fairness/bias, and practical systems. Key contributions include CA-RAG: a per-query routing framework that trades retrieval depth, latency, token cost, and quality (FAISS + OpenAI evaluation, reproducible CSVs); attention calibration to reduce positional bias in dense retrievers; evidence that supervised neural retrievers learn query-independent document relevance priors causing findability gaps; Slipstream: a locality-aware method to speed up streaming ANNS (Faiss
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_ir
- Record identifier
- a999019560edc806f76095753f0102f0bc9aee8fa1d1f3fb65fb540fe6469b67
- Enrichment time
- 2026-06-03T08:52:25Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.