Emit
The ASR frontend produces an intermediate speech hypothesis as the user speaks.
Turn spoken-form hypotheses into clean, intent-preserving text — then revise what was already emitted when later speech changes the meaning.
Automatic speech recognition (ASR) has achieved substantial gains in transcription accuracy, yet verbatim transcription does not necessarily produce readily usable text. It retains fillers, repetitions, false starts, and self-corrections that increase reading effort, obscure the speaker’s final intent, and propagate unresolved or abandoned content to downstream tasks. Existing spoken-to-written methods process completed audio or transcripts but cannot revise emitted text when later speech changes how preceding content should be interpreted. We therefore formulate Agentic Speech Recognition (AgenticSR), an audio-to-clean-text task that removes disfluencies, resolves self-corrections, and normalizes written form while preserving the speaker’s final intent. AgenticASR implements this task through an ASR–Refiner architecture that repeatedly transforms a bounded active context and replaces its corresponding output span as audio arrives. This enables continual emission and revision over streams of arbitrary duration. We also introduce AASR-Bench, a bilingual benchmark with fine-grained atomic rubrics. Across multiple ASR front ends, AgenticASR attains the highest AASR-Bench scores among evaluated systems. A human–AI agreement study shows that rubric-based judgments align with independent expert assessments. Ablations characterize Refiner capacity, context length, and the quality–latency trade-off between online and offline inference. Together, these results establish AgenticASR as a practical framework for intent-preserving clean transcription during ongoing speech.
Speech is full of abandoned starts, fillers, repetitions, and corrections. AgenticSR keeps the final intent while making the output ready for reading and downstream use.

The ASR frontend produces an intermediate speech hypothesis as the user speaks.
A compact language-model Refiner converts the active context from oral to written form.
New evidence replaces only the corresponding local output span, so an earlier guess can be corrected in place.
Both demonstrations show the same core behavior: spoken-form input becomes readable text while the transcript remains open to evidence-supported revision.
English streaming example: disfluencies and incomplete phrasing are rewritten into clean text.
中文演示:系统在保留最终意图的同时,持续清理口语表达并更新局部结果。
The Refiner is deliberately separated from acoustic recognition, which lets the same text-to-text correction model work across different ASR backbones.

Seed, Oral, and Clean generation are followed by ASR simulation, semantic quality control, and global deduplication.
Online inference refines a local K-chunk source window and replaces its aligned output span rather than waiting for an utterance to finish.
AASR-Bench separates Content, Format, Filter, and Rephrase so clean transcription quality is not reduced to token error alone.
On AASR-Bench, AgenticASR wins the Overall score within the Qwen3-ASR families and improves every Whisper configuration over its API baseline.

With Qwen3-ASR-1.7B, AgenticASR reaches 79.95 Overall and leads all four rubric dimensions. Its advantage over the API transformation baseline ranges from 1.73 to 10.02 points across matched ASR backbones, with substantially lower latency.
With Whisper, the gain over the API baseline grows from 1.73 points at Base to 7.39 points at Large. The strongest improvements come from filtering and final-intent rephrasing.
Traditional token metrics remain useful diagnostics, but AASR-Bench exposes formatting, filtering, and correction-resolution failures that WER, CER, and MER cannot capture.
| ASR model | LM | Content | Format | Filter | Rephrase | WER/CER/MER ↓ | Latency ↓ | Overall |
|---|---|---|---|---|---|---|---|---|
| Qwen3-ASR-0.6B | Qwen3.5-Flash | 87.50 | 28.97 | 73.13 | 49.13 | 26.82/17.01/21.91 | 60.08 | 66.47 |
| FormalASR-0.6B | – | 86.63 | 14.35 | 36.51 | 13.29 | 38.67/28.34/34.38 | 3.42 | 48.76 |
| Qwen3-ASR-0.6B | AgenticASR | 87.30 | 54.94 | 78.80 | 69.16 | 14.64/7.79/10.23 | 6.60 | 76.15 |
| Qwen3-ASR-1.7B | Qwen3.5-Flash | 90.21 | 35.48 | 75.82 | 52.10 | 24.60/15.72/20.29 | 60.89 | 69.93 |
| FormalASR-1.7B | – | 90.11 | 19.69 | 40.59 | 15.70 | 34.07/24.47/30.48 | 3.46 | 52.50 |
| Qwen3-ASR-1.7B | AgenticASR | 90.24 | 65.19 | 78.89 | 72.83 | 12.70/6.86/9.01 | 9.59 | 79.95 |
| Whisper Base | Gemini-2.5-Flash | 47.04 | 6.09 | 62.63 | 16.67 | 53.14/39.83/46.41 | 12.00 | 37.09 |
| Whisper Base | AgenticASR | 38.69 | 6.95 | 71.96 | 32.47 | 55.62/41.33/46.70 | 5.86 | 38.82 |
| Whisper Small | Gemini-2.5-Flash | 58.79 | 29.04 | 65.08 | 27.94 | 56.98/43.74/45.78 | 11.99 | 48.78 |
| Whisper Small | AgenticASR | 52.58 | 29.57 | 72.79 | 47.40 | 58.09/44.09/44.28 | 6.89 | 51.72 |
| Whisper Large | Gemini-2.5-Flash | 80.23 | 51.58 | 63.10 | 36.13 | 31.50/21.45/25.33 | 8.04 | 62.90 |
| Whisper Large | AgenticASR | 76.16 | 55.87 | 77.75 | 63.01 | 27.51/18.19/19.63 | 4.42 | 70.29 |
Best values within each ASR family are shown in the paper in bold. The LM column identifies the downstream transformation system; FormalASR performs direct speech-to-clean-text recognition.
The ablations make the design trade-offs explicit: larger Refiners improve contextual rewriting, while a three-chunk online window recovers most of the useful right context.

Moving from K=1 to K=3 raises Rephrase from 36.17 to 70.47 and Explanation from 19.43 to 74.00, while latency grows by only 0.87 s. K=3 closes the gap to offline inference to 2.36 Rephrase points and 1.20 Explanation points.
This is the mechanism that lets AgenticASR correct a previously emitted destination when a later chunk contains the self-repair.
| Measure | 0.6B | 1.7B |
|---|---|---|
| Spearman ρ | 0.8222 | 0.8064 |
| Quadratic-weighted κ | 0.8313 | 0.7918 |
Double-blind experts and the Gemma-4-31B-IT judge agree strongly across 100 sampled utterances.
The mean Spearman correlation is 0.8222 for Qwen3-ASR-0.6B and 0.8064 for Qwen3-ASR-1.7B; quadratic-weighted agreement is 0.8313 and 0.7918.
With Qwen3-ASR-1.7B fixed, scaling the Refiner from 0.5B to 4B raises Overall by 4.66 points; the largest gains are in Format and Rephrase.
Overall rises from 78.76 to 83.42, while latency increases from 9.21 to 10.77 s. Larger Refiners suit latency-tolerant offline use.
| Refiner | Overall | Cont. | Fmt. | Filt. | Reph. | Lat. (s) |
|---|---|---|---|---|---|---|
| Qwen2.5-0.5B-Instruct | 78.76 | 88.00 | 63.40 | 78.36 | 69.85 | 9.21 |
| MiniCPM-5-1B | 79.95 | 90.24 | 65.19 | 78.89 | 72.83 | 9.59 |
| Qwen2.5-4B-Instruct | 83.42 | 91.00 | 74.43 | 83.31 | 75.68 | 10.77 |
| Setting | Rephrase ↑ | Latency (s) ↓ | Explanation ↑ |
|---|---|---|---|
| Offline | 72.83 | 9.59 | 75.20 |
| Window = 1 | 36.17 | 11.28 | 19.43 |
| Window = 2 | 65.08 | 11.70 | 55.06 |
| Window = 3 | 70.47 | 12.15 | 74.00 |
Moving from K=1 to K=3 raises Rephrase from 36.17 to 70.47 and Explanation from 19.43 to 74.00, while latency grows by only 0.87 s.
At K=3, the gaps to offline shrink to 2.36 Rephrase points and 1.20 Explanation points, recovering nearly all useful right context for online revision.
AASR-Bench is bilingual and atomic: every sample is scored on the specific transformation requirements it contains, rather than a single undifferentiated text metric.
Each sample has at least one Content question. Format, Filter, and Rephrase rubrics are added when those phenomena are present. The benchmark covers ten usage scenes plus a pass-through control.
Explore AASR-Bench| Category | Questions | Share (%) | Coverage |
|---|---|---|---|
| Content | 3,448 | 51.95 | 917 |
| Format | 1,498 | 22.57 | 741 |
| Filter | 882 | 13.29 | 882 |
| Rephrase | 809 | 12.19 | 623 |
| Total | 6,637 | 100.00 | – |
If this project is useful, please cite the paper.
@misc{jiang2026agenticasrrefiningspeechrecognition,
title={AgenticASR: Refining Speech Recognition in Real-World Scenarios via an Agentic Approach},
author={Zixuan Jiang and Binghao Qiang and Jiaying Chi and Yanqiao Zhu and Kai Yu and Xie Chen},
year={2026},
eprint={2607.28175},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2607.28175},
}Page structure inspired by the Academic Project Page Template and Nerfies; visual language adapted for AgenticASR.