Open Editor
Writing Tips
5 min read

Voice Dictation Accuracy in 2026: What the Error-Rate Numbers Actually Mean

A 98 percent accurate dictation engine sounds nearly perfect until you count: two wrong words in every hundred is still eleven errors in a short page. And the benchmark rarely sounds like you.

Speech-to-text has quietly crossed a line in 2026: the best on-device engines now post word error rates in the low single digits on clean test audio, which makes "98 percent accurate" an easy claim to make and a hard one to interpret. Accuracy figures are averages over a test set, and the test set is rarely your voice, your room or your vocabulary. This post explains what the numbers measure, what they look like in a real draft, and one question about where your audio goes.

Cleaning up a dictated transcript? Paste it into ClearText Editor to tidy spacing, line breaks and case. It runs in your browser, and nothing is uploaded.

What Word Error Rate Actually Counts

Word error rate (WER) is the share of words a system gets wrong compared with a reference transcript, counting substitutions, deletions and insertions. A 5 percent WER means about one error per twenty words. The useful step is multiplying it by the length of what you dictate.

In a July 2026 benchmark, an app maker called Lyonesse ran several engines on 5,559 LibriSpeech utterances on one Apple M2 Pro, all processed on the device. The company states plainly that a benchmark from a firm that sells one of the engines should be treated with suspicion, and its results were not independently reproduced here, so treat the figures as indicative rather than definitive.

EngineWER, clean audioWER, harder audioErrors in 500 words (harder)
Apple SpeechAnalyzer2.12%4.56%~23
Whisper Small3.74%7.95%~40
Whisper Tiny7.88%17.04%~85
Apple SFSpeechRecognizer (older)9.02%16.25%~81

The last column is my arithmetic: WER multiplied by 500 words. Even the best result there means a few dozen corrections in a page of dictation, and the older engines would need an edit pass on roughly every sixth word.

The Benchmark Is Not Your Voice

LibriSpeech is volunteers reading audiobooks aloud. The Lyonesse write-up itself lists the limits: English only, read speech rather than conversation, one machine, and no accented, far-field or multi-speaker audio yet. Those are the conditions that tend to raise errors in daily use.

A single accuracy number describes the test set, not the speaker. A 2020 PNAS study of five commercial systems, from Amazon, Apple, Google, IBM and Microsoft, found an average WER of 0.35 for Black speakers against 0.19 for white speakers, and traced the gap to the underlying acoustic models. Those systems are older than today's, so the numbers will have moved, but the lesson hasn't: test dictation with your own voice before you trust a headline.

Where Does the Audio Go?

The 2026 shift worth noticing isn't only accuracy, it's location. In the benchmark above, every engine ran entirely on the device. Browsers are less uniform. MDN's documentation for the Web Speech API notes that on some browsers, like Chrome, speech recognition involves a server-based engine: "Your audio is sent to a web service for recognition processing, so it won't work offline." The API also has a processLocally option to require on-device recognition, but it has limited availability and isn't supported everywhere.

So "dictate in the browser" can mean your voice leaves your computer. For a shopping list that is probably fine. For a draft contract, a medical note or anything confidential, find out which kind of engine you're using before you start talking.

A Practical Dictation Routine

  • Prefer an on-device engine for anything sensitive, and check the app's settings or documentation rather than assuming.
  • Dictate in short passages in a quiet room, with a decent microphone close to your mouth.
  • Say punctuation and paragraph breaks as you go; fixing them later takes longer than saying them.
  • Review proper nouns, numbers and homophones first. These are the errors a spell checker won't catch.
  • Always do a read-through before the text leaves your hands. A fluent wrong word reads as correct.

The honest summary: dictation in 2026 is good enough that errors feel rare, and that is exactly why they slip through. On clean audio the best engines miss two or three words in a hundred; on harder audio and older engines the figure climbs fast, and no benchmark covers every voice. Treat dictation as a fast first draft, keep one editing pass in the routine, and check whether your audio is processed on the device or sent to a server.

For questions or inquiries contact us at info@cleartexteditor.com