Accuracy & quality
6.34% WER in Portuguese and 4.6% diarization DER, measured by us against human transcripts, with the protocol behind every number stated.
Every number on this page was measured by us, against human transcripts and annotations, with the protocol stated below. None comes from a model card: text normalization moves WER by several points, and a number without a protocol compares to nothing.
Summary
| Metric | Value | Where |
|---|---|---|
| Brazilian Portuguese WER | 6.34% | 2,000 spontaneous speech clips |
| Diarization DER (no collar) | 4.6% | 19.6 human-annotated hours |
| Diarization DER (0.25 s collar) | 2.8% | 216 files |
| Median word timestamp error | 69 to 86 ms | real audio, matched words |
Brazilian Portuguese WER
A sample of 2,000 spontaneous Brazilian speech clips with human reference transcripts: real interviews and dialogue, stratified by audio type and drawn with a fixed seed. Not read sentences recorded in a studio.
| System | WER |
|---|---|
transcrevo (trv-1) | 6.34% |
| transcrevo's previous generation | 21.9% |
A 71% relative drop, winning on every slice of the sample and around 10x faster on the same batch.
Normalization, identical for both: lowercase, punctuation stripped, digits
spelled out, hyphens turned into spaces. WER is
(substitutions + deletions + insertions) / reference words.
Be careful comparing this with WER published by other systems. Changing the normalizer moves the metric by several points (the same system can score 7.7% normalized and 11.8% raw on the same material), and evaluating on read speech yields far better numbers than real speech. Only compare equal material with an equal normalizer.
Diarization
Measured on 216 files, 19.6 hours of speech with human annotation, under the standard DER protocol. Human ground truth, not a comparison against another system.
| Metric | Value |
|---|---|
| DER, no collar | 4.6% |
| DER, 0.25 s collar | 2.8% |
| JER | 15.4% |
| DER on overlapped speech only | 14.8% |
| Missed speech / false alarm / confusion | 1.55% / 1.72% / 1.37% |
| Exact speaker count | 146 of 216 files |
DER (diarization error rate) is missed speech plus false alarm plus speaker confusion, over total speech time. Lower is better.
Word timestamps
Every word comes back with start and end in decimal seconds. Error measured
against an external reference, matching words by text:
| Metric | Value |
|---|---|
| Median error (p50) | 69 to 86 ms |
| Words off by more than 1 s | 3.1% (music-heavy vlog), 1.8% (podcast) |
Accurate enough for captions and for cutting audio at word boundaries.
Known limits
We publish what still fails, because whoever integrates needs to know where to double-check:
- Overlapped speech is the worst spot: 14.8% DER on that slice against 4.6% overall. Two people talking at once get a single label.
- Short turns under 5 s ("yes", "mhm", a quick question): speaker accuracy drops to the 59 to 75% range, against 95 to 99% on turns longer than 15 s.
- Proper nouns, brands and loanwords are the most common text error. The fix
is the request's
glossaryfield, which rewrites those terms for that audio. - Background music sometimes becomes text: lyrics get transcribed.
- Output is not deterministic. The same file can come back with small
differences between runs. There is no
seed.
Getting better accuracy on your audio
- Send
glossarywith the proper nouns and jargon of your domain:{ "what it hears": "what to write" }. - Set
languagewhen you already know it, instead of letting auto-detection decide on short or noisy audio. - Only ask for
speakerswhen you will use speaker separation: it changes the rate and does not improve the text.
Pricing & limits
US$ 0.03 per hour of audio, US$ 0.05 with speaker separation, billed by the second. Prepaid credits, no subscription, and US$ 2 on a new account.
Errors
The API error envelope and the full code table, from validation and invalid_api_key to insufficient_balance, rate_limited and concurrency_limit.