The transcrevo blog
Technical notes on audio transcription, diarization and what we measured along the way.
- Integration
How to transcribe audio with an API, from zero to text
Three HTTP requests turn an audio file into text with timestamps. No SDK, no queue to configure, no GPU server to run.
- Diarization
What speaker diarization is, and when you need it
Diarization answers "who said what". It is a separate task from transcription, it is billed separately, and it fails in specific places worth knowing before you integrate.
- Pricing
How much audio transcription costs with an API
US$ 0.03 per hour of audio, billed by the second. What that means on a thousand-hour queue, and the three pricing traps that show up on the invoice of whoever skipped the fine print.
- Quality
How we measure Portuguese accuracy, and why a vendor's WER does not compare
6.34% WER on spontaneous Brazilian speech, measured against human transcripts. The number matters less than the protocol, and the protocol is what almost nobody publishes.
- Languages
Automatic language detection in transcription, and when to turn it off
Over 90 languages with automatic detection. The language field exists because detection errs on short and noisy audio, and its error costs the whole transcript.
- Timestamps
Word-level timestamps: what you can build with them
Every word comes back with a start and end in decimal seconds, with a median error of 69 to 86 ms. Search inside audio, clip cutting and synced highlighting all come from that.
- Support
Transcribing support calls: what works and what does not
A recorded call is the most predictable audio there is: two voices, one topic, always the same format. That changes what is worth asking the API for, and what is worth checking in the result.
- Quality
How to fix wrong proper nouns in a transcript
Names, brands and jargon are the most common error in any automatic transcription. The glossary field fixes them per request, and it has limits worth knowing before you build the list.
- Infrastructure
Running a speech model yourself, or using a transcription API
The cost of self-hosting is not the GPU price, it is everything around it. Where each path wins, and where a transcription API simply does not serve.
- Scale
Batch transcription: thousands of files without breaking the queue
What changes when a batch goes from ten files to ten thousand: rate ceilings, reserved balance, duplicate billing, and which failures are worth retrying.
- Integration
Webhooks or polling on an async transcription API
Polling works for one file and turns into a hundred requests per job in production. What changes with webhooks, how to verify the signature, and what to do when your server was down.
- Podcast
Podcast transcription by API, from the RSS feed to the whole archive
The feed enclosure is already a public URL, which makes podcasts the easiest audio to transcribe. The work is in the back catalog, the guest names and what you do with the text.
- Meetings
Meeting transcription by API, with who said what
From the recording to text split by speaker, with the code that groups words into turns. And where the separation fails on real meetings, which is always the same place.
- Captions
How to generate SRT captions automatically from audio
The API returns every word with a start and end in seconds. Here is the code that turns that into SRT and VTT, and the places automatic captions tend to go wrong.