Languages

Automatic language detection in transcription, and when to turn it off

Over 90 languages with automatic detection. The language field exists because detection errs on short and noisy audio, and its error costs the whole transcript.

If you leave out the language field, the API works out the language on its own. Over 90 languages are supported, and transcribing any of them costs the same: US$ 0.03 per hour of audio.

That solves the case of receiving audio of unknown origin, which is common: a platform with users everywhere, an inherited archive with no metadata, a recording somebody uploaded without filling in any form.

And it creates a specific risk, which is what this post is about.

What happens when detection is wrong

A language error is not a word error. When the system decides that Portuguese audio is Spanish, the result is not 10% worse: it is unusable, because the whole transcript is written with the wrong vocabulary. One wrong word you fix with a glossary; one wrong language you throw away and resend.

That changes the economics of the decision. Letting detection decide is free; being wrong costs the entire transcript, plus the time until someone notices.

When to state the language

The rule is simple: if you know, say so.

curl https://api.transcrevo.com/v1/transcripts \
  -H "Authorization: Bearer $TRANSCREVO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "url": "https://example.com/meeting.mp3", "language": "pt" }'

And you know more often than it seems. The user's account language, the product's region, the show that always records in Portuguese, the customer whose support runs entirely in Spanish: all of that is information already in your database, and passing it along costs nothing.

The three cases where detection errs most, and where stating it pays most:

  • Short audio. An 8-second voice note offers little evidence. The less speech, the more the decision rides on luck.
  • Noisy audio. A bad call, a room recording, music over the top.
  • Close languages. Portuguese and Spanish share a lot of sound. It is the pair that shows up most in detection errors on Brazilian products.

An invalid code is refused at creation with 400 validation and the key language_invalid, not silently. The field reference is in transcripts and the validation keys in errors.

When to leave it automatic

When you genuinely do not know, which is a legitimate situation. An old archive without metadata, an open platform, a file that arrived through a third-party integration.

In that case, use what the response gives back. The transcript's language field carries the detected language (or the one you supplied), and it is null while processing. Storing that value does two useful things: it fills in the metadata your archive was missing, and it lets you audit. If 4% of your Brazilian archive is coming back as another language, you have a measurable problem instead of a one-off complaint.

const { data } = await res.json();
if (data.status === "done" && data.language !== expected(data.id)) {
	await flag(data.id, data.language); // review, and possibly resend with language
}

One file, one language

The response carries one language. A meeting that starts in Portuguese and switches to English halfway, or a lecture quoting passages in another language, does not come back segmented by language: there is one transcript, written under one decision.

If your material is systematically mixed, what works is cutting the audio at the switch points on your side and transcribing each piece with the right language. Not elegant, but it produces usable text, and the cost does not change, because billing is per hour of audio: two thirty-minute pieces cost the same as one full hour.

The language decides whether your glossary works

A consequence that goes unnoticed: the glossary field corrects the text after transcription, matching what the system wrote against what you sent. If the language came out wrong, the whole text is written in another vocabulary, and none of your entries match. The glossary announces nothing, it simply has nothing to replace.

So stating the language is a precondition for proper-noun correction to pay off. In an integration that depends on getting customer or product names right, sending language alongside glossary is the pair that works; the glossary alone, on short audio of uncertain language, is the half that cannot hold the other. Details of the correction are in how to fix wrong proper nouns.

What we measured and what we did not

This is the honest part of the post. Our accuracy measurement is in Brazilian Portuguese: 6.34% WER on 2,000 excerpts of spontaneous Brazilian speech, against human reference transcripts, with the protocol published. Diarization was measured on 216 files with human annotation, 19.6 hours.

For the other supported languages we publish no number measured by us, and so we claim nothing about them here. The support exists and works; what does not exist is a measurement of ours worth citing. If your main volume is in another language, the test that answers is yours: a new account comes with US$ 2 in credit, no card, which is roughly 65 hours of audio.

That holds generally, by the way. A WER published by one vendor does not compare to another's, because text normalization moves the metric by several points: the same system scores 7.7% normalized and 11.8% raw on the same material. The full protocol is in accuracy and quality and in how we measure accuracy.

The summary

State language whenever you know it, especially on short, noisy or close-language audio. Leave it automatic when you do not, and store the language that comes back, because it is the metadata you were missing and the basis of your audit.

And if your archive is genuinely multilingual, know that the decision is per file, not per passage.

All posts