A linear map on the encoder output predicts which audio tokens the language model will attend to, so tokens can be cut before the language model runs.
TL;DR
A large audio language model turns a minute of speech into 750–1,500 tokens and prefills every one. Unlike image tokens, audio tokens are still heavily attended and not yet ranked in the language model's first layers, so audio needs a ranking before the language model runs. A linear map on the encoder output supplies one, reaching Spearman ρ ≥ .69 on eleven of thirteen models. Triage cuts by it before the language model and, on multiple choice, again at layer 2. At the aggressive budget it beats every baseline in all twelve transcription cases, and leads DART, the strongest baseline on average, by .043 in mean multiple-choice accuracy.
The finding

Blue is the attention each audio token actually received; green was computed before the language model ran, from the encoder output alone. We call this linear map the prior: a ridge regression fitted in closed form, in seconds on a CPU, with no labels.
The method
Stage 1 cuts the audio tokens to a shortlist with the prior, before the language model runs. Stage 2, used on multiple choice, cuts the shortlist again at layer 2, correcting the prior with the attention observed there; the encoder and language model stay frozen.

Results
Two Qwen Omni models, Voxtral and Phi-4-multimodal, on three transcription and three multiple-choice benchmarks. Every selector keeps the same number of tokens on the same clip.
| Transcription, aggressive budget | 12 of 12 | cases with lower word error rate than every baseline |
| Multiple choice, aggressive budget | +.043 | mean accuracy over DART, 95% interval [+.032, +.053] |
| Conservative budget | ≤ .04 | word error rate and accuracy stay within this of full audio |
The savings grow with the number of audio tokens, so the paper times Triage on long audio. The figures come from different machines and measure different costs.
| Benchmark | full audio | operating pt. | pool | energy VAD | random | DART | FastVN | HeadRouterN | Triage |
|---|---|---|---|---|---|---|---|---|---|
| Qwen2.5-Omni-3B | |||||||||
| LibriSpeech | .083 | .80/.50 (2.00×) | .348 | .327 | .496 | .160 | .395 | .173 | .116 |
| FLEURS | .153 | .95/.85 (1.18×) | .325 | .355 | .236 | .221 | .200 | .215 | .174 |
| TEDLIUM | .215 | .75/.60 (1.67×) | .403 | .496 | .409 | .290 | .364 | .288 | .195 |
| Qwen3-Omni-30B (MoE) | |||||||||
| LibriSpeech | .011 | .70/.50 (2.00×) | .160 | .199 | .429 | .134 | .354 | .305 | .060 |
| FLEURS | .037 | .70/.50 (2.00×) | .140 | .138 | .385 | .119 | .386 | .210 | .081 |
| TEDLIUM | .020 | .75/.50 (2.00×) | .123 | .166 | .378 | .123 | .381 | .258 | .083 |
| Voxtral-Mini-3B | |||||||||
| LibriSpeech | .020 | .50/.35 (2.86×) | .398 | .665 | .627 | .139 | .753 | .190 | .088 |
| FLEURS | .043 | .50/.35 (2.86×) | .385 | .719 | .646 | .119 | .782 | .109 | .066 |
| TEDLIUM | .032 | .60/.42 (2.38×) | .241 | .317 | .521 | .441 | .304 | .170 | .152 |
| Phi-4-multimodal | |||||||||
| LibriSpeech | .017 | .50/.35 (2.86×) | .383 | .472 | .597 | .612 | .456 | .550 | .312 |
| FLEURS | .045 | .50/.35 (2.86×) | .368 | .402 | .583 | .529 | .426 | .472 | .285 |
| TEDLIUM | .085 | .50/.35 (2.86×) | .488 | .531 | .646 | .697 | .422 | .441 | .397 |
Corpus WER; n = 300 per case, 100 on the 30B, and seven talks on TEDLIUM. Ours is bin coverage, a single Stage-1 cut. The operating point is r1/r2, with the compression beside it. FastVN and HeadRouterN cut mid-prefill, at layer 2. Triage has lower WER than every baseline in all twelve cases, the mid-prefill ones included, and is below DART by a median of .064.
| Benchmark | full audio | compression | best baseline | Triage |
|---|---|---|---|---|
| Qwen2.5-Omni-3B | ||||
| LibriSpeech | .083 | 1.54× | .123 | .099 |
| FLEURS | .153 | 1.11× | .175 | .177 |
| TEDLIUM | .215 | 1.25× | .245 | .197 |
| Qwen3-Omni-30B (MoE) | ||||
| LibriSpeech | .011 | 1.54× | .030 | .025 |
| FLEURS | .037 | 1.54× | .048 | .050 |
| TEDLIUM | .020 | 1.67× | .059 | .049 |
| Voxtral-Mini-3B | ||||
| LibriSpeech | .020 | 1.54× | .051 | .025 |
| FLEURS | .043 | 1.54× | .050 | .043 |
| TEDLIUM | .032 | 2.00× | .076 | .071 |
| Phi-4-multimodal | ||||
| LibriSpeech | .017 | 1.54× | .051 | .045 |
| FLEURS | .045 | 1.54× | .068 | .064 |
| TEDLIUM | .085 | 1.54× | .109 | .102 |
Best baseline is the lowest WER among the baselines at the same number of tokens. At this budget (1.11–2.00×), bin coverage costs a median of .017 WER against full audio. It has lower WER than every baseline in ten of the twelve cases.
| Benchmark | full audio | operating pt. | pool | energy VAD | random | DART | bin coverage | FastVN | HeadRouterN | Triage |
|---|---|---|---|---|---|---|---|---|---|---|
| Qwen2.5-Omni-3B | ||||||||||
| MMSU | .603 | .35/.25 (4.00×) | .517 | .552 | .505 | .547 | .595 | .557 | .557 | .585 |
| DREAM | .892 | .50/.20 (5.00×) | .530 | .620 | .569 | .757 | .767 | .688 | .740 | .863† |
| RACE | .818 | .50/.35 (2.86×) | .705 | .718 | .677 | .805 | .818 | .730 | .807 | .820 |
| Qwen3-Omni-30B (MoE) | ||||||||||
| MMSU | .715 | .85/.35 (2.86×) | .588 | .652 | .573 | .620 | .642 | .578 | .588 | .682† |
| DREAM | .922 | .85/.45 (2.22×) | .848 | .877 | .759 | .902 | .915 | .818 | .805 | .943† |
| RACE | .870 | .65/.35 (2.86×) | .755 | .740 | .723 | .807 | .802 | .740 | .752 | .828 |
| Voxtral-Mini-3B | ||||||||||
| MMSU | .552 | .35/.20 (5.00×) | .443 | .405 | .415 | .532 | .520 | .502 | .527 | .525 |
| DREAM | .882 | .50/.35 (2.86×) | .735 | .603 | .596 | .800 | .835 | .802 | .860 | .873† |
| RACE | .765 | .50/.35 (2.86×) | .675 | .610 | .608 | .632 | .657 | .608 | .690 | .655 |
| Phi-4-multimodal | ||||||||||
| MMSU | .550 | .50/.35 (2.86×) | .507 | .505 | .486 | .460 | .510 | .458 | .472 | .492 |
| DREAM | .843 | .50/.35 (2.86×) | .740 | .690 | .604 | .667 | .767 | .675 | .665 | .743† |
| RACE | .688 | .50/.35 (2.86×) | .647 | .630 | .600 | .610 | .627 | .613 | .635 | .650 |
Accuracy against the gold answer, n = 400 per case; RACE is AudioMarathon-RACE. Ours is precision fusion, the two-stage cut. Bin coverage, Triage's transcription cut, selects without the question. † Ahead of DART in a one-sided exact McNemar test, Holm-corrected over the twelve cases.
Listen
Each example compares Triage with DART on the same clip at the same budget. Click a word to seek.
Transcribing a talk on 60% of the audio with Stage 1 alone (Qwen2.5-Omni-3B, TEDLIUM; each clip's own word error rate).
Answering a question on a third of the audio, with Stage 1, both cuts and DART each keeping 35% (Qwen2.5-Omni-3B, AudioMarathon-RACE; single passages, not a sample).
Transcripts are machine-generated for display only.