Listen to the Latents: Self-Correcting Speech Recognition in Large Audio Language Models Through Hidden-State Interactions

arXiv 2609.02940v18d59a0bf168bdf9fde8a8ea684b0fea998785646a5bf0b0d40a340d66d95bfab

Paper metadata

arXiv ID
2609.02940
Version
v1
Category
cs.AI, cs.CL
Authors
Chan-Jan Hsu, Jaeyeon Kim, Chao-Han Huck Yang, Shinji Watanabe, Hung-yi Lee, Carlos Busso
Publication date
2026-09-04T04:00:00Z
Source identifier
2609.02940v1
Public record ID
record:sha256:8d59a0bf168bdf9fde8a8ea684b0fea998785646a5bf0b0d40a340d66d95bfab

Source license ↗

The PDF link points to arxiv.org. Baitaphish does not expose a private stored PDF.

This is source-provided metadata, not an enriched summary or an impact assessment. Follow the canonical source link for the published material.

Evidence and limitations

Source ID
arxiv_research
Record identifier
8d59a0bf168bdf9fde8a8ea684b0fea998785646a5bf0b0d40a340d66d95bfab
Record type
Source metadata

This record may overlap with other records. Source metadata can be incomplete or change. Validate consequential decisions against the linked source and your own environment.