Meetings

Meeting transcription by API, with who said what

From the recording to text split by speaker, with the code that groups words into turns. And where the separation fails on real meetings, which is always the same place.

A meeting recording is the worst audio a transcription system gets. Microphone across the room, people talking over each other, one participant on a bad connection, chairs scraping. It is still the most requested use case, because nobody wants to write minutes.

The path from recording to minutes has three parts: send the audio asking for speaker separation, group the words into turns, and know which parts of the result you can trust without review.

Sending the recording

If the recording already sits in a bucket with a public URL, one request does it. If it is on your machine or in a private bucket, it goes through upload and you use the uploadId in place of url. Both paths are in sending files.

curl https://api.transcrevo.com/v1/transcripts \
  -H "Authorization: Bearer $TRANSCREVO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/meeting.mp3",
    "speakers": 5,
    "language": "pt",
    "webhookUrl": "https://your-server.com/hooks/transcrevo"
  }'

Three things in that request deserve a note.

speakers: 5 should only be given when you know. In a meeting you usually do: the invite has the attendee list. And giving it helps, because automatic counting hits the exact number on 146 of 216 measured files. But a wrong guess is worse than no guess: it forces separation to the count you invented, and the system starts splitting one voice in two. When you do not have the list, send true and leave the decision to whoever is holding the audio.

language: "pt" keeps auto-detection from deciding on its own. On a meeting with a short opening or heavy noise, stating the language errs less.

webhookUrl matters because meetings are long. A two-hour recording takes minutes to process, and asking every three seconds is about a hundred requests just to hear processing.

From words to turns

The response comes back as a list of words, each carrying a speaker label. Minutes are written in turns. The conversion is one reduce:

type Word = { start: number; end: number; text: string; speaker: string | null };
type Turn = { speaker: string | null; start: number; end: number; text: string };

function turns(words: Word[]): Turn[] {
	return words.reduce<Turn[]>((acc, word) => {
		const last = acc.at(-1);
		if (last !== undefined && last.speaker === word.speaker && word.start - last.end < 2) {
			last.text += ` ${word.text}`;
			last.end = word.end;
			return acc;
		}
		acc.push({ speaker: word.speaker, start: word.start, end: word.end, text: word.text });
		return acc;
	}, []);
}

The 2-second slack is what keeps a breath from becoming two turns by the same person. Raise it for meetings with a lot of hesitation, lower it for fast conversation.

With that in hand, the minutes become formatting:

[00:04:12] A: So the supplier moved the deadline to December.
[00:04:31] C: Does that change the whole integration schedule?
[00:04:36] A: It changes the delivery date, not the scope.

The A and B labels are anonymous by nature: they mean "first voice" and "second voice", never names. Tying a label to a person comes from outside the audio. In a meeting recorded by a tool that logs join order, or one that records a separate track per participant, the mapping is reliable. Without that, the honest move is to leave A and B in the minutes and let a reviewer name them.

Where this fails on real meetings

We measured diarization on 216 files, 19.6 hours of speech with human annotation. DER of 4.6% without collar, 2.8% with a 0.25 s collar. That is the good number. The two bad numbers are exactly what meetings produce all day:

SituationAccuracy
Turn longer than 15 seconds95 to 99%
Turn shorter than 5 seconds59 to 75%
Overlapping speech14.8% DER, against 4.6% overall

The practical reading is direct. Content minutes work. Who argued what, what was decided, who took the action item: those are long blocks of speech, and the system gets 95 to 99% of them right.

Behavioral metrics do not work. Counting how many times each person spoke, who interrupted whom, how much airtime each got: those numbers live off the two-second "mhm", "yes", "agreed", which is exactly where accuracy drops to the 59 to 75% range. A product promising that kind of analysis from a room recording is promising what the hardest slice of the problem does not deliver. The full breakdown, with JER and false alarm separated, is in accuracy and quality.

And when two people talk at once, the stretch gets a single label. In a heated discussion, that is the part your reviewer will rewrite.

What improves the result before the recording ends

  • One track per participant, when the meeting tool allows it. Then there is no diarization problem at all: you transcribe each track separately and know whose it is. It costs the same, because billing is per hour of audio, and five one-hour tracks are five hours.
  • glossary with your company's names. Proper nouns, brands and jargon are the most common error of any transcription system, because they are the words least present in training. An internal meeting is made of them: project names, customer names, acronyms. Send { "what it hears": "what to write" }, up to 100 entries, applied to that audio only.
  • Room audio from a single microphone is the worst case. No API trick fixes microphone distance.

Keeping and deleting

Meeting recordings usually have a retention policy, and HR or legal minutes almost always do. DELETE /v1/transcripts/{id} erases the text, the words and the stored audio; the usage record stays, because that is what the invoice comes from, but the content is gone and does not come back. A transcript still in processing cannot be deleted: the result would arrive afterwards and repopulate what you just erased.

Worth remembering that the output is not reproducible. Sending the same file twice can return slightly different transcripts, and there is no seed. To audit a complaint about a set of minutes, keep the transcript you delivered instead of regenerating it.

The math

US$ 0.05 per hour with speaker separation, billed by the second. A weekly one-hour meeting, all year, costs US$ 2.60. A hundred meetings a month at a mid-sized company, say 150 hours, cost US$ 7.50 a month. That is cheap enough that the decision becomes about human review, not about the price of transcription. Limits and the discount above 10,000 hours are in pricing and limits.

All posts