From Black Box to Glass Box: Cross-Model ASR Disagreement to Prioto Review in Ambient AI Scribe Documentation

Introduction Ambient AI scribes generate transcripts at scale, but routine quality assurance is constrained by the absence of human-verified reference transcripts in most deployment settings. We evaluated whether disagreement among heterogeneous automatic speech recognition (ASR) systems can serve as an informative signal for localizing transcription uncertainty, using a public English-language medical-speech corpus rather than clinical encounter recordings. Methods Eight commercial and open-source ASR systems were applied to 50 medical-education audio clips (8 h 14 min). Multi-model outputs were aligned, and a leave-one-out consensus procedure was used to score per-model agreement while reducing circularity. Results Disagreement across models was sparse and localized: 72.1% of positions showed strong agreement (7-8 systems concordant), whereas only 2.5% were high-risk positions with minimal agreement (0-3 systems). Low-agreement regions were systematically enriched for meaning-bearing lexical differences, defined as lexical mismatches after excluding punctuation, contraction, numeric, and filler variation. A single-annotator human-corrected (HC) validation layer showed that transcription errors increased monotonically with decreasing agreement. At an illustrative post hoc threshold, flagging positions where six or fewer systems agreed selected 28.6% of tokens while recovering 93.7% of single-annotator HC-verified errors on this proxy corpus. Discussion These findings suggest that cross-model disagreement may help focus human review on a small number of likely error-prone transcript regions. However, agreement among all systems does not guarantee correctness, because shared errors may remain undetected by this approach. Validation on real clinical encounter data is required before operational deployment.

Paper

References (19)

Scroll for more · 7 remaining

Similar papers

© 2026 NYSGPT2525 LLC