Multi-Head Attention as Ensemble Nadaraya-Watson Estimation: Variance Reduction, Decorrelation, and Optimal Head Diversity

2026-05-21T07:23:54Z85bf889e39197c8abd5cef64755246785ac9cb1f3ecdf9465a9e25e28dbb8d31
GAMEHead-Diversity-IndexNadaraya-WatsonVC-dimensionarXivbanditsbayesian-inferencebias-variance-decompositioncontradiction-graphsgaussian-processesgroup-aware-matrix-estimationimportance-samplingintegrated-laplacelatent-gaussian-modelsmachine-learningmatrix-estimationmulti-head-attentionoptimal-transportpseudo-marginalrecommender-systemsscaling-lawsspectral-banditstheorytransfer-learningtransformers

What happened

Collection of recent arXiv ML papers (announced 21 May 2026) spanning theoretical and methodological advances. Highlights include a rigorous statistical theory of multi-head attention as an ensemble of Nadaraya–Watson estimators with a Head Diversity Index and optimal head-dimension scaling; corrected integrated Laplace approximation for latent Gaussian models using importance sampling and pseudo/quasi-Monte Carlo; a graph-theoretic characterization of VC dimension via contradiction graphs; an optimal-transport analysis of transfer learning sample complexity; spectral-bandit algorithms for 추천/

Why it matters

A reviewed impact interpretation has not been published for this record.

Evidence and limitations

Source ID
arxiv_stat_ml
Record identifier
85bf889e39197c8abd5cef64755246785ac9cb1f3ecdf9465a9e25e28dbb8d31
Enrichment time
2026-05-21T07:23:54Z
AI-assisted enrichment
Yes

This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.

Record · Multi-Head Attention as Ensemble Nadaraya-Watson Estimation: Variance Reduction, Decorrelation, and Optimal Head Diversity · Baitaphish