Credit-assigned Policy Gradient for Early Stage Retrieval in Two-stage Ranking
2026-05-27T08:52:24Z•9f99d64ec35737fb2e238e661969185360f12e02ea5e6ab6d4e5de885e2557f1
arxivdense-retrievalin-context-retrievalinformation-retrievalknowledge-graphlarge-language-modelsmultimodal-representationmusic-recommendationpolicy-gradientposition-biasrecommender-systemsreinforcement-learningreproducibilityretrieval-augmented-generationtraffic-allocationworkshop
What happened
Collection of recent IR, recommender, and retrieval-augmented LLM papers: Credit‑assigned Policy Gradient (CA‑PG) for early-stage rankers reduces variance by marginalizing over candidate‑set compositions; a proposed evaluation framework for structured generative search summaries; Uniboost, a posterior value‑alignment and linear‑boosting traffic allocation framework for blending/re-ranking; empirical study showing positional bias in dense retrievers is largely driven by the positional distribution of evidence in training data and that position‑balanced training mitigates bias; L2Rec, which uses
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_ir
- Record identifier
- 9f99d64ec35737fb2e238e661969185360f12e02ea5e6ab6d4e5de885e2557f1
- Enrichment time
- 2026-05-27T08:52:24Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.