Diarization

What speaker diarization is, and when you need it

Diarization answers "who said what". It is a separate task from transcription, it is billed separately, and it fails in specific places worth knowing before you integrate.

Transcription answers what was said. Diarization answers who said it. They are different tasks solved by different models, which is why turning the second one on changes the rate and does nothing for the first.

In a five-person meeting, transcription alone returns a block of running text. With diarization, every word comes back with a speaker label:

"words": [
	{ "start": 0.32, "end": 0.78, "text": "Bom", "speaker": "A" },
	{ "start": 0.78, "end": 1.04, "text": "dia", "speaker": "A" },
	{ "start": 2.44, "end": 2.9, "text": "Vamos", "speaker": "B" },
	{ "start": 2.9, "end": 3.37, "text": "começar.", "speaker": "B" }
]

The labels are anonymous by nature: "A" and "B" mean "the first voice" and "the second voice", not names. Tying a label to a person is your job, and it normally comes from outside the audio (who joined the room, who was on the extension).

Comparison of the same audio without and with diarization: without, a single bar of running text; with, four separate turns by speakers A and B, including a hatched band of overlapping speech.
The same audio, without and with the speakers field. The hatched band is overlapping speech, where error triples.

How to ask for it

A single field decides everything:

ValueMeaning
omittedRunning text, speakers not separated.
trueSeparate speakers; the API works out how many.
3There are exactly 3 speakers.
"2-5"There are between 2 and 5 speakers.
curl https://api.transcrevo.com/v1/transcripts \
  -H "Authorization: Bearer $TRANSCREVO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "url": "https://example.com/meeting.mp3", "speakers": 5 }'

Only give a count when you actually know it. A wrong guess is worse than no guess: it forces the separation to the number you supplied, so the system starts splitting one voice in two or merging two into one. If you do not know, true leaves the decision to whoever is holding the audio.

Where diarization fails

No system is always right, and knowing the holes before you integrate saves debugging time. We measured ours on 216 files, 19.6 hours of speech with human annotation, using the standard DER protocol:

MetricValue
DER without collar4.6%
DER with a 0.25 s collar2.8%
DER on overlapping speech only14.8%
Exact speaker count146 of 216 files

DER (diarization error rate) is missed speech plus false alarm plus speaker confusion over total speech time. Lower is better.

The two places where it gets worse are predictable:

  • Overlapping speech. Two people talking at once get a single label. It is the system's worst slice: 14.8% DER against 4.6% overall.
  • Short turns. Below 5 seconds ("yes", "mhm", a quick question), speaker accuracy drops to the 59 to 75% range. Above 15 seconds it sits between 95 and 99%.

That has a practical consequence: if your product depends on counting short interventions (who interrupted whom, how many times each person spoke), expect errors in that slice. If it depends on reasonable blocks of speech (meeting minutes, interviews, two-person podcasts), diarization delivers.

The full breakdown, with JER, missed speech and false alarm separated, is in accuracy and quality.

When not to use it

Diarization costs more: US$ 0.05 per hour of audio against US$ 0.03 for plain transcription. It does not make the text better. So:

  • Do not use it for a lecture, a dictation, a voice note, a video with a single narrator. One speaker has nothing to separate.
  • Do not use it "just in case" on a large batch. On a thousand-hour queue, the difference between the two rates is US$ 20.
  • Use it when the label reaches the product: meeting minutes, interviews, support calls with a customer and an agent, captions that must identify who is speaking.

The speakerCount field

With speakers, the response carries speakerCount, how many distinct voices were found. The difference between null and 0 matters: null means "you did not ask", 0 means "I looked and nobody was speaking". Treating them alike hides the silent file, which is exactly the case you want to see in the log.

All posts