Long-Term Optimization for Large-Scale Generative Retrieval with Off-Policy REINFORCE

2026-07-07T08:52:18Z985013a86b26c9a9e6c4fb3f7794e46269bdbb7180b165e184ff99ed20309d21
HETERQAKuaishou deploymentOCR refinementPower-of-Noise replicationRAG evaluationSentAttackTRIAGEadversarial attacksarchival searchbenchmarksblack-box attackscandidate retrievaldense retrievalgenerative retrievalheterogeneous sourcesknowledge-graph instrumentationlong-term optimizationmodel robustnessoff-policy REINFORCEreinforcement learningrelevance embeddingsretrieval-augmented generation

What happened

A collection of recent arXiv papers (July 7, 2026) on large-scale retrieval, retrieval-augmented generation (RAG), and evaluation for information access. Key contributions include: an off-policy REINFORCE autoregressive method for long-term optimization of generative retrieval (improves session-level cumulative reward on Yambda-5B); HETERQA, a 5-source heterogeneous-record retrieval benchmark; HGenPush, a heterogeneous generative recommendation architecture deployed at scale for push notifications; an LLM-based OCR refinement plus RAG pipeline improving retrieval on historical archives; TRIAGE

Why it matters

A reviewed impact interpretation has not been published for this record.

Evidence and limitations

Source ID
arxiv_cs_ir
Record identifier
985013a86b26c9a9e6c4fb3f7794e46269bdbb7180b165e184ff99ed20309d21
Enrichment time
2026-07-07T08:52:18Z
AI-assisted enrichment
Yes

This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.

Record · Long-Term Optimization for Large-Scale Generative Retrieval with Off-Policy REINFORCE · Baitaphish