Automatic language detection in transcription, and when to turn it off
Over 90 languages with automatic detection. The language field exists because detection errs on short and noisy audio, and its error costs the whole transcript.
If you leave out the language field, the API works out the language on its
own. Over 90 languages are supported, and transcribing any of them costs the
same: US$ 0.03 per hour of audio.
That solves the case of receiving audio of unknown origin, which is common: a platform with users everywhere, an inherited archive with no metadata, a recording somebody uploaded without filling in any form.
And it creates a specific risk, which is what this post is about.
What happens when detection is wrong
A language error is not a word error. When the system decides that Portuguese audio is Spanish, the result is not 10% worse: it is unusable, because the whole transcript is written with the wrong vocabulary. One wrong word you fix with a glossary; one wrong language you throw away and resend.
That changes the economics of the decision. Letting detection decide is free; being wrong costs the entire transcript, plus the time until someone notices.
When to state the language
The rule is simple: if you know, say so.
curl https://api.transcrevo.com/v1/transcripts \
-H "Authorization: Bearer $TRANSCREVO_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "url": "https://example.com/meeting.mp3", "language": "pt" }'And you know more often than it seems. The user's account language, the product's region, the show that always records in Portuguese, the customer whose support runs entirely in Spanish: all of that is information already in your database, and passing it along costs nothing.
The three cases where detection errs most, and where stating it pays most:
- Short audio. An 8-second voice note offers little evidence. The less speech, the more the decision rides on luck.
- Noisy audio. A bad call, a room recording, music over the top.
- Close languages. Portuguese and Spanish share a lot of sound. It is the pair that shows up most in detection errors on Brazilian products.
An invalid code is refused at creation with 400 validation and the key
language_invalid, not silently. The field reference is in
transcripts and the validation keys in
errors.
When to leave it automatic
When you genuinely do not know, which is a legitimate situation. An old archive without metadata, an open platform, a file that arrived through a third-party integration.
In that case, use what the response gives back. The transcript's language
field carries the detected language (or the one you supplied), and it is null
while processing. Storing that value does two useful things: it fills in the
metadata your archive was missing, and it lets you audit. If 4% of your
Brazilian archive is coming back as another language, you have a measurable
problem instead of a one-off complaint.
const { data } = await res.json();
if (data.status === "done" && data.language !== expected(data.id)) {
await flag(data.id, data.language); // review, and possibly resend with language
}One file, one language
The response carries one language. A meeting that starts in Portuguese and
switches to English halfway, or a lecture quoting passages in another language,
does not come back segmented by language: there is one transcript, written under
one decision.
If your material is systematically mixed, what works is cutting the audio at the
switch points on your side and transcribing each piece with the right
language. Not elegant, but it produces usable text, and the cost does not
change, because billing is per hour of audio: two thirty-minute pieces cost the
same as one full hour.
The language decides whether your glossary works
A consequence that goes unnoticed: the glossary field corrects the text after
transcription, matching what the system wrote against what you sent. If the
language came out wrong, the whole text is written in another vocabulary, and
none of your entries match. The glossary announces nothing, it simply has
nothing to replace.
So stating the language is a precondition for proper-noun correction to pay off.
In an integration that depends on getting customer or product names right,
sending language alongside glossary is the pair that works; the glossary
alone, on short audio of uncertain language, is the half that cannot hold the
other. Details of the correction are in
how to fix wrong proper nouns.
What we measured and what we did not
This is the honest part of the post. Our accuracy measurement is in Brazilian Portuguese: 6.34% WER on 2,000 excerpts of spontaneous Brazilian speech, against human reference transcripts, with the protocol published. Diarization was measured on 216 files with human annotation, 19.6 hours.
For the other supported languages we publish no number measured by us, and so we claim nothing about them here. The support exists and works; what does not exist is a measurement of ours worth citing. If your main volume is in another language, the test that answers is yours: a new account comes with US$ 2 in credit, no card, which is roughly 65 hours of audio.
That holds generally, by the way. A WER published by one vendor does not compare to another's, because text normalization moves the metric by several points: the same system scores 7.7% normalized and 11.8% raw on the same material. The full protocol is in accuracy and quality and in how we measure accuracy.
The summary
State language whenever you know it, especially on short, noisy or
close-language audio. Leave it automatic when you do not, and store the
language that comes back, because it is the metadata you were missing and the
basis of your audit.
And if your archive is genuinely multilingual, know that the decision is per file, not per passage.