transcrevo docs

Accuracy & quality

6.34% WER in Portuguese and 4.6% diarization DER, measured by us against human transcripts, with the protocol behind every number stated.

Every number on this page was measured by us, against human transcripts and annotations, with the protocol stated below. None comes from a model card: text normalization moves WER by several points, and a number without a protocol compares to nothing.

Summary

MetricValueWhere
Brazilian Portuguese WER6.34%2,000 spontaneous speech clips
Diarization DER (no collar)4.6%19.6 human-annotated hours
Diarization DER (0.25 s collar)2.8%216 files
Median word timestamp error69 to 86 msreal audio, matched words

Brazilian Portuguese WER

A sample of 2,000 spontaneous Brazilian speech clips with human reference transcripts: real interviews and dialogue, stratified by audio type and drawn with a fixed seed. Not read sentences recorded in a studio.

SystemWER
transcrevo (trv-1)6.34%
transcrevo's previous generation21.9%

A 71% relative drop, winning on every slice of the sample and around 10x faster on the same batch.

Normalization, identical for both: lowercase, punctuation stripped, digits spelled out, hyphens turned into spaces. WER is (substitutions + deletions + insertions) / reference words.

Be careful comparing this with WER published by other systems. Changing the normalizer moves the metric by several points (the same system can score 7.7% normalized and 11.8% raw on the same material), and evaluating on read speech yields far better numbers than real speech. Only compare equal material with an equal normalizer.

Diarization

Measured on 216 files, 19.6 hours of speech with human annotation, under the standard DER protocol. Human ground truth, not a comparison against another system.

MetricValue
DER, no collar4.6%
DER, 0.25 s collar2.8%
JER15.4%
DER on overlapped speech only14.8%
Missed speech / false alarm / confusion1.55% / 1.72% / 1.37%
Exact speaker count146 of 216 files

DER (diarization error rate) is missed speech plus false alarm plus speaker confusion, over total speech time. Lower is better.

Word timestamps

Every word comes back with start and end in decimal seconds. Error measured against an external reference, matching words by text:

MetricValue
Median error (p50)69 to 86 ms
Words off by more than 1 s3.1% (music-heavy vlog), 1.8% (podcast)

Accurate enough for captions and for cutting audio at word boundaries.

Known limits

We publish what still fails, because whoever integrates needs to know where to double-check:

  • Overlapped speech is the worst spot: 14.8% DER on that slice against 4.6% overall. Two people talking at once get a single label.
  • Short turns under 5 s ("yes", "mhm", a quick question): speaker accuracy drops to the 59 to 75% range, against 95 to 99% on turns longer than 15 s.
  • Proper nouns, brands and loanwords are the most common text error. The fix is the request's glossary field, which rewrites those terms for that audio.
  • Background music sometimes becomes text: lyrics get transcribed.
  • Output is not deterministic. The same file can come back with small differences between runs. There is no seed.

Getting better accuracy on your audio

  • Send glossary with the proper nouns and jargon of your domain: { "what it hears": "what to write" }.
  • Set language when you already know it, instead of letting auto-detection decide on short or noisy audio.
  • Only ask for speakers when you will use speaker separation: it changes the rate and does not improve the text.

On this page