Measurement · 23 August 2026

Apple’s SpeechAnalyzer takes a list of your jargon. It recovered 0 of 106 terms.

One speaker, one Mac, fifteen recordings. The script and the raw results are at the bottom.

15
recordings
3
doses
8
target words
106
chances
0
recovered

macOS 26 ships a new on-device speech API. SpeechAnalyzer takes an AnalysisContext, and that context has a contextualStrings field: a list of words you expect the speaker to say. Hand it kubectl and OAuth before the audio starts, and the recogniser is supposed to weight them.

I build a dictation app on top of this API, so the question mattered to me directly: is that list worth assembling? Apple publishes no accuracy numbers for it. So I measured.

Across fifteen recordings and 106 chances, the list recovered nothing. Not one term. At any dose.

Before the audio · the word list
0 / 106
target words recovered

Fifteen recordings, three doses. Being handed the exact eight words about to be spoken did not measurably differ from being handed nothing at all.

After the audio · string repair
+18 / 39
recovered with no model, no network

Four deterministic passes over the finished transcript, sub-millisecond, identical on every machine. Recognition alone got 10 of 39; this took it to 28.

The limits, before anything else. One speaker — me. One machine, an Apple silicon Mac on macOS 26. One sitting, on 23 August 2026. English only. The recordings are of my own voice and I am not publishing them. This is not a study; it is one person’s measurement, and the sample is small enough that you should treat the direction as the finding and the exact figure as noise.

The setup

One passage, deliberately dense with the kind of word recognition destroys:

Send the OAuth token to Claude Code before the stand-up. The API latency spiked after we deployed to Kubernetes. Run kubectl get pods, and paste the output into Notion. I read the swiftinterface file, open the Figma file. Speaklively should learn my vocabulary. Ask Claude to summarize the NativeFlow roadmap.

Eight target words: OAuth, Claude Code, Kubernetes, kubectl, swiftinterface, Figma, Speaklively, NativeFlow. Read fifteen times, in three arms: no list at all, the eight words about to be spoken, and those eight inside a forty-three term glossary.

The eight-term arm is the interesting one. It is not a realistic glossary; it is the recogniser being handed the answer key immediately before the question. Scoring is exact and case-sensitive, because for several of these words the casing is the thing you want back. figma is a miss.

The result

no list4 recordings0 of 248 terms3 recordings0 of 2443 terms8 recordings0 of 58
Each track is the number of target words actually spoken in that arm; the fill is how many came back correctly. The middle row is the recogniser being handed the answer key immediately before the question. Nothing is filled in any row.

Zero, flat, across every arm. Being handed the exact eight words about to be spoken did not measurably differ from being handed nothing.

Why 106 and not 120. Eight targets times fifteen recordings is 120, and that is the number I had written down internally. It is wrong. One take is just the word “Hello?” — a fumbled recording that never reached a single target. Another was cut short after four. Counting those as missed chances inflates the denominator with words nobody said. The honest count of attempted terms is 106. The numerator is zero either way, which is the only reason the mistake did not change a conclusion.

What the failures look like is more interesting than the count. OAuth came back as OR token, Oath token, 0 earth token, awath token. kubectl became Quebectel, Kubaktel, Kubak 10, Ren Ku Bechtel. Claude Code became clot code, cloud code, Claude Gordon, Gladcote. These are not near misses that a nudge would have tipped over. The recogniser is confidently producing a different phrase.

I want to be careful about what this does and does not show. It does not show that contextualStrings is a no-op in general — I tested one speaker, one locale, one passage, one build. It is entirely possible the field does meaningful work for homophone disambiguation, or for names it already has some representation of, or in other locales. What it shows is that for newly-coined technical jargon spoken by one person, it bought nothing I could detect, and that if you are planning to solve your jargon problem by assembling a good list, you should measure before you invest in it.

BEFORE THE AUDIOAFTER THE AUDIOYour word listup to 150 termscontextualStringsbiases the recogniserRecognition · SpeechTranscriberon-device0 of 106 recoveredRuns anyway. It is free.the audio is gone from here onRaw textkept in historyTranscriptRepair1 · taught rules2 · phrase3 · join4 · single10 → 28 of 39No model. No network. Sub-millisecond.
The dashed line is the moment the audio is gone. Both mechanisms aim at the same mistake; only one of them can still see the sound. That turned out not to be the one that helped.

What did work: string surgery, after the fact

Same speaker, same machine, different measurement two days later. This time a 60-sentence script with a known reference, so a fix could be counted automatically rather than judged.

The layer under test does no machine learning and touches no network. Four passes over the finished transcript: rules you taught it by hand, a pass for two real words that together sound like one term (clot code → Claude Code), a pass for one term the recogniser split in two (native flow → NativeFlow), and a pass for a single junk token against a single known term (kubernetties → Kubernetes). Sub-millisecond. Identical on every machine.

Recognition alone10 / 39
After the four repair passes28 / 39
Thirty-nine hard words across sixty spoken lines. Eighteen of the ones recognition lost were recovered by string matching alone — no model, no network, on the same machine.

Eighteen of the thirty-nine terms recognition lost were recovered by deterministic string matching, on the same machine, offline. That is the inverse of where I expected the win to be.

The mechanism that makes it safe is a gate: a real English word is never rewritten by sound alone. Sound-based repair only fires on tokens that are not words at all — kubernetties, Aquas — where a rewrite cannot destroy a sentence that was already correct. Comparison is a crude consonant skeleton rather than a real phonetic algorithm: vowels drop, ph→f, ck→k, doubles fold, so cloud, Claude and clawed all key to kld. About thirty lines, deliberately not Double Metaphone.

cloudClaudeclawed
kld
JasonJSON
jsn
kubernettiesKubernetes
kbrnts
Vowels and glides drop, ph→f, ck→k, doubled consonants fold. Two words that transcribe the same speech land on the same key. That is the entire comparison.

Twenty-seven of the sixty lines — fifteen ordinary sentences and twelve traps built from everyday words that merely sound like jargon — contain no target at all. They exist only to catch the layer changing something it should not. That is the half of the corpus that decides whether any of this is shippable.

one hard word each · 15split words · 8long passages · 5ordinary sentences · 15traps · 12spoken punctuation · 5
Twenty-seven of the sixty lines contain no target at all. They exist only to catch the repair layer changing something it should have left alone — and they are how the one false positive on this page was found.

What it broke

One sentence, in the trap block. I read:

read aloudHe clawed his way back into the conversation.
recognition heardHe clawed his way back into the conversation.correct
after repairHe Claude his way back into the conversation.broken

Recognition heard the first one correctly. Then the repair layer produced the second. A rule I had taught by hand fired on ordinary English and made the transcript worse than the raw recognition it was handed. Hand-taught rules are the one path deliberately exempt from the dictionary gate — you said when it writes this, I meant that, and the layer believes you — so this is the designed trade-off doing exactly what it says on the tin. It still reads like a bug when it lands in your email.

That is the only false positive the layer produced across sixty lines. Two other words my scoring flagged turned out to be recognition mangling the sentence — code review heard as court review, Share it with the team heard as Shared with the team — where the repair layer changed nothing. Those are two different failures, and my first scoring pass conflated them: it asked “is this ordinary word missing from the final text”, which is true whenever recognition dropped it, and charged the difference to the layer under test. Worth mentioning because it is the kind of error that quietly makes a layer look better or worse than it is, and I only caught it because the trap block existed.

What none of this reaches

The failure I cannot fix:

Send the over token to Claude Code.

over is a real English word, correctly spelled, sitting in a grammatical position that makes sense. Every gate that makes the repair layer safe is exactly what forbids it from touching that word. I do not have an answer for this class that does not involve a model reading the sentence for meaning, and when I benched a meaning-level corrector it answered “Cuban, it is” with Claude, which is worse than doing nothing.

A few other things I measured and threw away, so nobody repeats them: harvesting terms off the screen as repair targets recovered nothing and broke real sentences; reranking the recogniser’s own alternative transcriptions was two wins against two regressions; and adding punctuation and number rules to the on-device cleanup prompt made it measurably worse.

Run it yourself

One speaker on one Mac is not enough to conclude much. If you work on dictation or speech tooling, read the same sentences into whatever you use and tell me what you get — particularly if contextualStrings does something for you that it did not do for me, because I would genuinely like to be wrong about this.

These measurements came out of building Speaklively, a macOS dictation app that shows you the text before it types — which is why they exist, and why they are this specific.