Mult-DPO: Multinomial Direct Preference Optimization for Recommender Systems

2026-06-10T08:52:20Z0deb52d42e5c160cd3c9c45b8489778ff664db3d7d90634b2b5a31f4295d46ca
AIRBM25','retrieval','query-rewriting'CGMDPOMult-DPOPlackett-LuceRAGSIDInspectorSTORMSkillResolve-Benchagentic-recommendersbenchmarkscounterfactual-explanationsevaluationhealth-datainference-accelerationlarge-language-modelslexical-query-expansionpersonalizationprivacyrecommender-systemssafetysemantic-id-tokenizersskill-retrievaltau-Rec

What happened

Feed of recent arXiv papers (Jun 10 2026) on recommender systems, LLM alignment, retrieval, diagnostics and benchmarks. Key contributions: Mult-DPO — a multinomial DPO objective that tractably aligns LLMs to set-wise (multi-positive) preferences and bounds a Plackett–Luce marginalized loss; MetaPlate — a counterfactual-guided RAG+LLM system that uses CGM and wearable signals to produce personalized meal recommendations to reduce postprandial hyperglycemia (expert-in-the-loop evaluation; privacy/health-data implications; code-backed RAG retrieval over USDA); τ-Rec — a verifiable benchmark and R

Why it matters

A reviewed impact interpretation has not been published for this record.

Evidence and limitations

Source ID
arxiv_cs_ir
Record identifier
0deb52d42e5c160cd3c9c45b8489778ff664db3d7d90634b2b5a31f4295d46ca
Enrichment time
2026-06-10T08:52:20Z
AI-assisted enrichment
Yes

This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.