The Triage mark: a strip of audio tokens with a pair of scissors cutting through it and dropped tokens falling away.

Audio token attention is predictable before the language model runs

A linear map on the encoder output predicts which audio tokens the language model will attend to, so tokens can be cut before the language model runs.

TL;DR

A large audio language model turns a minute of speech into 750–1,500 tokens and prefills every one. Unlike image tokens, audio tokens are still heavily attended and not yet ranked in the language model's first layers, so audio needs a ranking before the language model runs. A linear map on the encoder output supplies one, reaching Spearman ρ ≥ .69 on eleven of thirteen models. Triage cuts by it before the language model and, on multiple choice, again at layer 2. At the aggressive budget it beats every baseline in all twelve transcription cases, and leads DART, the strongest baseline on average, by .043 in mean multiple-choice accuracy.

The finding

The encoder output predicts language-model attention before the language model runs

Measured per-token language-model attention and the prior's prediction over one held-out LibriSpeech clip, tracking each other; rho equals 0.65 on this clip.

Blue is the attention each audio token actually received; green was computed before the language model ran, from the encoder output alone. We call this linear map the prior: a ridge regression fitted in closed form, in seconds on a CPU, with no labels.

Qwen3-Omni-30B, a held-out LibriSpeech clip with the median ρ.

The method

Two cuts, one prior, nothing retrained

Stage 1 cuts the audio tokens to a shortlist with the prior, before the language model runs. Stage 2, used on multiple choice, cuts the shortlist again at layer 2, correcting the prior with the attention observed there; the encoder and language model stay frozen.

The Triage pipeline: the audio encoder emits audio tokens, a fitted prior scores them, Stage 1 cuts before the frozen language model, and Stage 2 cuts at layer 2 inside it.

Results

Four models, six benchmarks, two budgets

Two Qwen Omni models, Voxtral and Phi-4-multimodal, on three transcription and three multiple-choice benchmarks. Every selector keeps the same number of tokens on the same clip.

Transcription, aggressive budget12 of 12cases with lower word error rate than every baseline
Multiple choice, aggressive budget+.043mean accuracy over DART, 95% interval [+.032, +.053]
Conservative budget≤ .04word error rate and accuracy stay within this of full audio

Efficiency

The savings grow with the number of audio tokens, so the paper times Triage on long audio. The figures come from different machines and measure different costs.

6.3×
faster prefill on 20 minutes of audio, and 4.5× on 5 minutes (A100, Qwen 3B)
1.28–1.51×
encoder and prefill together, for 5 to 40 minutes of audio (A100)
4×
as many concurrent 5-minute streams on one GH200 as with full audio (Qwen 3B)
21.8 → ~62 min
audio that fits the Qwen 3B's context window, cutting the tokens 2.86×
Show per-model tables

Transcription, aggressive budget — word error rate, lower is better

Benchmark full audiooperating pt. poolenergy VADrandomDARTFastVNHeadRouterNTriage
Qwen2.5-Omni-3B
LibriSpeech.083.80/.50 (2.00×).348.327.496.160.395.173.116
FLEURS.153.95/.85 (1.18×).325.355.236.221.200.215.174
TEDLIUM.215.75/.60 (1.67×).403.496.409.290.364.288.195
Qwen3-Omni-30B (MoE)
LibriSpeech.011.70/.50 (2.00×).160.199.429.134.354.305.060
FLEURS.037.70/.50 (2.00×).140.138.385.119.386.210.081
TEDLIUM.020.75/.50 (2.00×).123.166.378.123.381.258.083
Voxtral-Mini-3B
LibriSpeech.020.50/.35 (2.86×).398.665.627.139.753.190.088
FLEURS.043.50/.35 (2.86×).385.719.646.119.782.109.066
TEDLIUM.032.60/.42 (2.38×).241.317.521.441.304.170.152
Phi-4-multimodal
LibriSpeech.017.50/.35 (2.86×).383.472.597.612.456.550.312
FLEURS.045.50/.35 (2.86×).368.402.583.529.426.472.285
TEDLIUM.085.50/.35 (2.86×).488.531.646.697.422.441.397

Corpus WER; n = 300 per case, 100 on the 30B, and seven talks on TEDLIUM. Ours is bin coverage, a single Stage-1 cut. The operating point is r1/r2, with the compression beside it. FastVN and HeadRouterN cut mid-prefill, at layer 2. Triage has lower WER than every baseline in all twelve cases, the mid-prefill ones included, and is below DART by a median of .064.

Transcription, conservative budget

Benchmarkfull audiocompressionbest baselineTriage
Qwen2.5-Omni-3B
LibriSpeech.0831.54×.123.099
FLEURS.1531.11×.175.177
TEDLIUM.2151.25×.245.197
Qwen3-Omni-30B (MoE)
LibriSpeech.0111.54×.030.025
FLEURS.0371.54×.048.050
TEDLIUM.0201.67×.059.049
Voxtral-Mini-3B
LibriSpeech.0201.54×.051.025
FLEURS.0431.54×.050.043
TEDLIUM.0322.00×.076.071
Phi-4-multimodal
LibriSpeech.0171.54×.051.045
FLEURS.0451.54×.068.064
TEDLIUM.0851.54×.109.102

Best baseline is the lowest WER among the baselines at the same number of tokens. At this budget (1.11–2.00×), bin coverage costs a median of .017 WER against full audio. It has lower WER than every baseline in ten of the twelve cases.

Multiple choice, aggressive budget — accuracy, higher is better

Benchmark full audiooperating pt. poolenergy VADrandomDARTbin coverageFastVNHeadRouterNTriage
Qwen2.5-Omni-3B
MMSU.603.35/.25 (4.00×).517.552.505.547.595.557.557.585
DREAM.892.50/.20 (5.00×).530.620.569.757.767.688.740.863†
RACE.818.50/.35 (2.86×).705.718.677.805.818.730.807.820
Qwen3-Omni-30B (MoE)
MMSU.715.85/.35 (2.86×).588.652.573.620.642.578.588.682†
DREAM.922.85/.45 (2.22×).848.877.759.902.915.818.805.943†
RACE.870.65/.35 (2.86×).755.740.723.807.802.740.752.828
Voxtral-Mini-3B
MMSU.552.35/.20 (5.00×).443.405.415.532.520.502.527.525
DREAM.882.50/.35 (2.86×).735.603.596.800.835.802.860.873†
RACE.765.50/.35 (2.86×).675.610.608.632.657.608.690.655
Phi-4-multimodal
MMSU.550.50/.35 (2.86×).507.505.486.460.510.458.472.492
DREAM.843.50/.35 (2.86×).740.690.604.667.767.675.665.743†
RACE.688.50/.35 (2.86×).647.630.600.610.627.613.635.650

Accuracy against the gold answer, n = 400 per case; RACE is AudioMarathon-RACE. Ours is precision fusion, the two-stage cut. Bin coverage, Triage's transcription cut, selects without the question. † Ahead of DART in a one-sided exact McNemar test, Holm-corrected over the twelve cases.

Listen

Hear what the model keeps

Each example compares Triage with DART on the same clip at the same budget. Click a word to seek.

Transcribing a talk on 60% of the audio with Stage 1 alone (Qwen2.5-Omni-3B, TEDLIUM; each clip's own word error rate).

Transcripts are machine-generated for display only.