research

Token Merging for Multilingual Speech Recognition: A Systematic Study Across Model Scale and Fine-Tuning

Published
Published
Reviewed
Reviewed
Next review due
Review due
Version
Version 1

The work systematically evaluates inference-time token merging across a multilingual speech-recognition model family and examines its interaction with parameter-efficient adaptation in lower-resource settings.

By

MACHINE_LEARNINGEMPIRICAL
Trust and provenance

Editorial record

AI-assistance disclosure

Research Intelligence analysis generated with AI and checked against cited source evidence.

This record says human review did not occur.

Sources

  • arxiv.org2609.13151v1

    Claims attributed to the linked primary source in this content record.

    Version
    2609.13151v1
    Retrieved
    Reuse
    link-only

MACHINE LEARNING · EMPIRICAL

Original research: Token Merging for Multilingual Speech Recognition: A Systematic Study Across Model Scale and Fine-Tuning · 2609.13151v1

Paper authors: Dylan Luke Holyoak

Source license: CC BY 4.0. This article summarizes and interprets the source using AI. Attribution does not imply endorsement by the source authors.

This adapted analysis is shared under the same CC BY 4.0 license. Semantic status: supported by automated evidence review. Human scientific review and independent replication have not been established.

TL;DR

The work systematically evaluates inference-time token merging across a multilingual speech-recognition model family and examines its interaction with parameter-efficient adaptation in lower-resource settings.

Source: E014

For the evaluated lower-resource cohort, the middle-scale model showed broadly stable error under merging, with the strongest language-level improvement occurring for a named member of that cohort.

Source: E004

For the evaluated languages and checkpoints, the authors recommend a common high reduction setting without language-specific recalibration; lower settings offer additional accuracy margin when constraints are tighter.

Source: E018

Significance

The work systematically evaluates inference-time token merging across a multilingual speech-recognition model family and examines its interaction with parameter-efficient adaptation in lower-resource settings.

Source: E014

For the evaluated languages and checkpoints, the authors recommend a common high reduction setting without language-specific recalibration; lower settings offer additional accuracy margin when constraints are tighter.

Source: E018

Research Question

The study examines whether adjacent token merging transfers across model scales, linguistically diverse settings, and decoder-side adaptation.

Source: E009

Contribution

The work systematically evaluates inference-time token merging across a multilingual speech-recognition model family and examines its interaction with parameter-efficient adaptation in lower-resource settings.

Source: E014

Datasets

The evaluation draws from a multilingual speech benchmark with a language set spanning phonological variation, linguistic families, and resource strata.

Source: E007

Training Setup

The study evaluates model variants spanning scale and applies decoder-only low-rank adaptation to the middle-scale variant while retaining its encoder.

Source: E008, E017

Method

The merging procedure scores adjacent representations by key-vector similarity, greedily forms disjoint pairs, and replaces selected pairs with averaged hidden representations.

Source: E006, E020

Merging follows a fixed encoder-layer schedule that begins near the input and omits the final encoder layer to retain a buffer before decoder access.

Source: E016

The evaluation compares an unmerged baseline with a range of token-reduction settings under a fixed merge-layer schedule.

Source: E013

Metrics

Recognition quality is assessed using normalized-transcript word error relative to the corresponding unmerged model, while efficiency is assessed through observed computation time.

Source: E003

Findings

For the evaluated lower-resource cohort, the middle-scale model showed broadly stable error under merging, with the strongest language-level improvement occurring for a named member of that cohort.

Source: E004

For the high-resource anchor cohort, merging produced only a small average error change under the stated evaluation setting.

Source: E010

After decoder-side adaptation, merging retained a small average error effect on trained languages and a bounded largest degradation on held-out languages.

Source: E011

Baselines

Across the tested model scales, average error changes remained small, although the direction of the mean change differed by scale.

Source: E005, E012

Environment Sample

Latency measurements use device-timed paired baseline and merged runs, warm-up handling, synchronized execution, and separate encoder and end-to-end reporting in a fixed hardware and software environment.

Source: E015

Metrics Conditions

At the highest tested reduction setting, encoder latency declines relative to the unmerged baseline, yielding an encoder-local speedup whose deployment value is expected to depend on encoder cost.

Source: E002, E019

Limitations

The study uses a static merge schedule across languages and scales, and its adaptation-composition evidence is confined to the selected low-rank adaptation approach.

Source: E001

Failure Modes

The held-out evaluation covers only part of the pretraining distribution, so the findings support broad generalization without establishing exhaustive coverage.

Source: E001

Practical Implications

For the evaluated languages and checkpoints, the authors recommend a common high reduction setting without language-specific recalibration; lower settings offer additional accuracy margin when constraints are tighter.

Source: E018

Evidence and source

Show evidence locators

Evidence labels locate support in the original paper; they do not establish independent replication.

  1. E001 · page 8Limitations: Evidence E001
  2. E002 · page 8Introduction: Evidence E002
  3. E003 · page 4Introduction: Evidence E003
  4. E004 · page 5Introduction: Evidence E004
  5. E005 · page 6Introduction: Evidence E005
  6. E006 · page 3Introduction: Evidence E006
  7. E007 · page 4Introduction: Evidence E007
  8. E008 · page 4Introduction: Evidence E008
  9. E009 · page 2Introduction: Evidence E009
  10. E010 · page 5Introduction: Evidence E010
  11. E011 · page 6Introduction: Evidence E011
  12. E012 · page 6Introduction: Evidence E012
  13. E013 · page 4Introduction: Evidence E013
  14. E014 · page 1Abstract: Evidence E014
  15. E015 · page 6Introduction: Evidence E015
  16. E016 · page 4Introduction: Evidence E016
  17. E017 · page 4Introduction: Evidence E017
  18. E018 · page 8Limitations: Evidence E018
  19. E019 · page 7Introduction: Evidence E019
  20. E020 · page 3Introduction: Evidence E020