Quality

How we measure Portuguese accuracy, and why a vendor's WER does not compare

6.34% WER on spontaneous Brazilian speech, measured against human transcripts. The number matters less than the protocol, and the protocol is what almost nobody publishes.

Every transcription vendor publishes a WER. Almost none publishes how they got there, and that is where the number stops meaning anything.

WER (word error rate) is (substitutions + deletions + insertions) / reference words. The formula is standard. What is not standard, and moves the metric by several points, is everything around it: which material was evaluated and how the text was normalized before comparison.

What we measured

MeasureValueWhere
WER in Brazilian Portuguese6.34%2,000 spontaneous speech excerpts
Diarization DER (no collar)4.6%19.6 h with human annotation
Diarization DER (0.25 s collar)2.8%216 files
Median word timestamp error69 to 86 msreal audio, matched words

The material is spontaneous Brazilian speech: real interviews and dialogue, stratified by audio type and sampled with a fixed seed. Not studio-read sentences, which is the material where every system looks excellent.

The reference is a human transcript, not another system's output. Comparing against another model measures agreement, not correctness: two systems can miss the same word and the number still looks great.

Normalization, identical for every system compared: lowercase, punctuation stripped, digits written out, hyphens become spaces.

Two bar panels on the same scale. On the left, 6.34% WER for trv-1 against 21.9% for the previous generation. On the right, the same system scoring 7.7% on normalized text and 11.8% on raw text.
Both panels share one scale. On the right, the same system measured with two rulers: the gap between them is wider than the distance between many competing systems.

Why normalization decides the number

This is the part that rarely reaches a marketing page. The same system, on the same material, scores 7.7% normalized and 11.8% raw. Four points of spread without a line of code changing, purely from the comparison rule.

Which means comparing the WER published by two vendors is comparing two different rulers. If one strips punctuation and the other does not, the first wins by construction. If one evaluates read sentences and the other evaluates call center audio, the first wins again.

The only comparison that holds is the one you run: same material, same normalizer, your audio through both. That is exactly what the credit on a new account is for, roughly 65 hours.

What still goes wrong

We publish the limits next to the wins, because whoever integrates needs to know where to check:

  • Overlapping speech is the system's worst spot: 14.8% DER on that slice against 4.6% overall. Two people talking at once get a single label.
  • Short turns, below 5 seconds, drop speaker accuracy to the 59 to 75% range, against 95 to 99% on turns longer than 15 seconds.
  • Proper nouns, brands and loanwords are the most common text error. They are the words least present in any training set. The fix is the request's glossary field, which rewrites those terms for that audio.
  • Background music sometimes becomes text: the lyrics get transcribed.
  • Output is not deterministic. The same file can come back with small differences between runs. There is no seed.

Timestamps

Every word comes back with start and end in decimal seconds. Median error against an external reference sits between 69 and 86 ms, with 1.8% of words more than 1 second off on a podcast and 3.1% on a vlog with music. That is precise enough for captions and for cutting audio at the word.

What to do with this

If you are evaluating vendors, the checklist is short:

  1. Ask for the material. "Spontaneous speech" and "read sentences" are not the same thing.
  2. Ask for the normalizer. Without it, a WER compares to nothing.
  3. Ask for the reference. Human, or another system's output?
  4. Run your audio. It is the only test that answers your question.

The numbers on this page, broken down by slice, live in accuracy and quality and are updated whenever the measurement is redone.

All posts