Quality

How to fix wrong proper nouns in a transcript

Names, brands and jargon are the most common error in any automatic transcription. The glossary field fixes them per request, and it has limits worth knowing before you build the list.

Every transcription system fails on the same words: people's names, company names, loanwords, internal acronyms. It is not a coincidence or a defect of one particular engine. Those are exactly the words least present in any training set, and the model writes down whatever sounds similar and is common.

"Bertuol" becomes "bertual". "Gatto" becomes "gatu". "Kubernetes" becomes "cubernetes".

The good part is that you know which words will show up in your audio before you send it. An internal meeting has your project names. A support call has your product names. A podcast has this week's guest. That list is already in your database.

The glossary field

It is an object shaped { "what it hears": "what to write" }, applied after transcription and valid for that audio only:

curl https://api.transcrevo.com/v1/transcripts \
  -H "Authorization: Bearer $TRANSCREVO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/meeting.mp3",
    "glossary": { "gatu": "Gatto", "bertual": "Bertuol", "cubernetes": "Kubernetes" }
  }'

The rules:

  • Up to 100 entries, each side up to 80 characters.
  • The correction works over word sequences, so "mote gruto": "Mateus Gruto" works: two words become two.
  • The time span is split across the replaced words, so timestamps stay valid for captions.
  • What you send wins over the corrections that ship by default.
  • None of this changes the rate. The glossary is free.

Full reference in transcripts.

Finding out what it hears

The left side of the glossary is not the correct name misspelled at random: it is what the system actually writes. Finding that out is empirical, and it takes five minutes.

  1. Transcribe a representative file without a glossary.
  2. Search text for the terms you expected and did not find.
  3. See what showed up in their place.
  4. Build the entry with exactly that.

One detail worth knowing before investing time in it: the output is not deterministic. The same file can come back with small differences between runs, and there is no seed. So the same name may be heard two or three different ways across an archive. In practice that means a single entry rarely covers an important term: map the two or three spellings you saw, all to the same target.

{
	"glossary": {
		"bertual": "Bertuol",
		"bertuau": "Bertuol",
		"bertuol": "Bertuol"
	}
}

The third entry looks useless and is not: it normalizes capitalization.

What the glossary does not do

Worth being explicit, because the wrong expectation here breeds frustration.

It does not change how the model listens. The correction happens after transcription, over the text that came out. Which means a word the system never wrote at all, because the audio was too quiet or someone spoke over it, does not come back through the glossary. There is nothing to rewrite.

It does not fix a bad microphone. If the problem is distance, noise or a degraded codec, the glossary fixes one symptom at a time while the rest of the text stays poor.

It is literal, and that cuts both ways. An entry that is too short matches where you did not want it. Mapping "art" to a customer name will rewrite every "art" in the audio. Prefer two-word sequences when the term is short, and test before running a whole batch.

Where this moves the result most

Our accuracy measurement for Portuguese is 6.34% WER on 2,000 excerpts of spontaneous Brazilian speech, against human transcripts. That number is the average over the whole material. The error is not distributed evenly: proper nouns concentrate a large share of it, and they are precisely the words a reader notices.

That asymmetry is useful. A transcript with 6% error on ordinary words reads fine; a transcript that gets the customer's name wrong on every mention looks useless, even at identical WER. Fixing the slice the reader notices costs one list.

The per-slice numbers, with the known limits, are in accuracy and quality.

One list per request, generated from your own data

Because the glossary applies to that audio only, the right way to build it is dynamically. On a support call, the names that matter are the ones on that ticket. In a meeting, the attendees on the invite. In an episode, the guest.

async function glossaryFor(ticketId: string) {
	const { customer, products } = await context(ticketId);
	return {
		...misheard(customer.name),
		...Object.fromEntries(products.flatMap((p) => Object.entries(misheard(p.name)))),
		...COMPANY_TERMS, // fixed: company name, products, internal acronyms
	};
}

The 100-entry ceiling fits a case like that comfortably, and it exists for a practical reason: a thousand-term list raises the chance of a wrong match a lot, and the text gets worse instead of better.

The rest of the quality story

Two other things move the result before any glossary. Setting language when you already know it keeps auto-detection from deciding on short or noisy audio. And asking for speakers only when you will use the separation, because it changes the rate (US$ 0.05 against US$ 0.03 per hour) and does nothing for the text.

If your problem is knowing who spoke rather than what was said, that is a different subject: what speaker diarization is.

All posts