How we measure Portuguese accuracy, and why a vendor's WER does not compare
6.34% WER on spontaneous Brazilian speech, measured against human transcripts. The number matters less than the protocol, and the protocol is what almost nobody publishes.
Every transcription vendor publishes a WER. Almost none publishes how they got there, and that is where the number stops meaning anything.
WER (word error rate) is
(substitutions + deletions + insertions) / reference words. The formula is
standard. What is not standard, and moves the metric by several points, is
everything around it: which material was evaluated and how the text was
normalized before comparison.
What we measured
| Measure | Value | Where |
|---|---|---|
| WER in Brazilian Portuguese | 6.34% | 2,000 spontaneous speech excerpts |
| Diarization DER (no collar) | 4.6% | 19.6 h with human annotation |
| Diarization DER (0.25 s collar) | 2.8% | 216 files |
| Median word timestamp error | 69 to 86 ms | real audio, matched words |
The material is spontaneous Brazilian speech: real interviews and dialogue, stratified by audio type and sampled with a fixed seed. Not studio-read sentences, which is the material where every system looks excellent.
The reference is a human transcript, not another system's output. Comparing against another model measures agreement, not correctness: two systems can miss the same word and the number still looks great.
Normalization, identical for every system compared: lowercase, punctuation stripped, digits written out, hyphens become spaces.
Why normalization decides the number
This is the part that rarely reaches a marketing page. The same system, on the same material, scores 7.7% normalized and 11.8% raw. Four points of spread without a line of code changing, purely from the comparison rule.
Which means comparing the WER published by two vendors is comparing two different rulers. If one strips punctuation and the other does not, the first wins by construction. If one evaluates read sentences and the other evaluates call center audio, the first wins again.
The only comparison that holds is the one you run: same material, same normalizer, your audio through both. That is exactly what the credit on a new account is for, roughly 65 hours.
What still goes wrong
We publish the limits next to the wins, because whoever integrates needs to know where to check:
- Overlapping speech is the system's worst spot: 14.8% DER on that slice against 4.6% overall. Two people talking at once get a single label.
- Short turns, below 5 seconds, drop speaker accuracy to the 59 to 75% range, against 95 to 99% on turns longer than 15 seconds.
- Proper nouns, brands and loanwords are the most common text error. They
are the words least present in any training set. The fix is the request's
glossaryfield, which rewrites those terms for that audio. - Background music sometimes becomes text: the lyrics get transcribed.
- Output is not deterministic. The same file can come back with small
differences between runs. There is no
seed.
Timestamps
Every word comes back with start and end in decimal seconds. Median error
against an external reference sits between 69 and 86 ms, with 1.8% of words more
than 1 second off on a podcast and 3.1% on a vlog with music. That is precise
enough for captions and for cutting audio at the word.
What to do with this
If you are evaluating vendors, the checklist is short:
- Ask for the material. "Spontaneous speech" and "read sentences" are not the same thing.
- Ask for the normalizer. Without it, a WER compares to nothing.
- Ask for the reference. Human, or another system's output?
- Run your audio. It is the only test that answers your question.
The numbers on this page, broken down by slice, live in accuracy and quality and are updated whenever the measurement is redone.