Listen to the Latents: Self-Correcting Speech Recognition in Large Audio Language Models Through Hidden-State Interactions
arXiv 2609.02940v1•8d59a0bf168bdf9fde8a8ea684b0fea998785646a5bf0b0d40a340d66d95bfab
Paper metadata
- arXiv ID
- 2609.02940
- Version
- v1
- Category
- cs.AI, cs.CL
- Authors
- Chan-Jan Hsu, Jaeyeon Kim, Chao-Han Huck Yang, Shinji Watanabe, Hung-yi Lee, Carlos Busso
- Publication date
- 2026-09-04T04:00:00Z
- Source identifier
- 2609.02940v1
- Public record ID
- record:sha256:8d59a0bf168bdf9fde8a8ea684b0fea998785646a5bf0b0d40a340d66d95bfab
The PDF link points to arxiv.org. Baitaphish does not expose a private stored PDF.
This is source-provided metadata, not an enriched summary or an impact assessment. Follow the canonical source link for the published material.
Evidence and limitations
- Source ID
- arxiv_research
- Record identifier
- 8d59a0bf168bdf9fde8a8ea684b0fea998785646a5bf0b0d40a340d66d95bfab
- Record type
- Source metadata
This record may overlap with other records. Source metadata can be incomplete or change. Validate consequential decisions against the linked source and your own environment.