Multi-Head Attention as Ensemble Nadaraya-Watson Estimation: Variance Reduction, Decorrelation, and Optimal Head Diversity
2026-05-21T07:23:54Z•85bf889e39197c8abd5cef64755246785ac9cb1f3ecdf9465a9e25e28dbb8d31
GAMEHead-Diversity-IndexNadaraya-WatsonVC-dimensionarXivbanditsbayesian-inferencebias-variance-decompositioncontradiction-graphsgaussian-processesgroup-aware-matrix-estimationimportance-samplingintegrated-laplacelatent-gaussian-modelsmachine-learningmatrix-estimationmulti-head-attentionoptimal-transportpseudo-marginalrecommender-systemsscaling-lawsspectral-banditstheorytransfer-learningtransformers
What happened
Collection of recent arXiv ML papers (announced 21 May 2026) spanning theoretical and methodological advances. Highlights include a rigorous statistical theory of multi-head attention as an ensemble of Nadaraya–Watson estimators with a Head Diversity Index and optimal head-dimension scaling; corrected integrated Laplace approximation for latent Gaussian models using importance sampling and pseudo/quasi-Monte Carlo; a graph-theoretic characterization of VC dimension via contradiction graphs; an optimal-transport analysis of transfer learning sample complexity; spectral-bandit algorithms for 추천/
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_stat_ml
- Record identifier
- 85bf889e39197c8abd5cef64755246785ac9cb1f3ecdf9465a9e25e28dbb8d31
- Enrichment time
- 2026-05-21T07:23:54Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.