Podcast transcription by API, from the RSS feed to the whole archive
The feed enclosure is already a public URL, which makes podcasts the easiest audio to transcribe. The work is in the back catalog, the guest names and what you do with the text.
Podcasts are the easiest audio to transcribe by API, for a dull reason: the RSS
feed already hands you a public URL for every episode. No upload, no bucket, no
local file. The item's enclosure is exactly what the request's url field
expects.
curl https://api.transcrevo.com/v1/transcripts \
-H "Authorization: Bearer $TRANSCREVO_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://cdn.example.com/ep-142.mp3",
"speakers": true,
"language": "pt",
"glossary": { "gatu": "Gatto", "bertual": "Bertuol" },
"webhookUrl": "https://your-server.com/hooks/transcrevo"
}'The rest of this post is about everything around that: the back catalog, the proper nouns, and the text that comes out.
The back catalog, in a loop
A weekly show running for three years is around 150 episodes. At forty minutes each, that is 100 hours of audio. At US$ 0.03 per hour, US$ 3.00 for the whole archive. With speaker separation, US$ 5.00.
The loop is plain, and the only care needed is the ceiling of 120 creations per minute:
const API = "https://api.transcrevo.com";
const headers = {
Authorization: `Bearer ${process.env.TRANSCREVO_API_KEY}`,
"Content-Type": "application/json",
};
for (const episode of episodes) {
const res = await fetch(`${API}/v1/transcripts`, {
method: "POST",
// The idempotency key is the episode id: rerunning the script does not pay twice.
headers: { ...headers, "Idempotency-Key": `ep-${episode.guid}` },
body: JSON.stringify({
url: episode.enclosure,
speakers: true,
language: "pt",
glossary: episode.glossary,
webhookUrl: "https://your-server.com/hooks/transcrevo",
}),
});
if (res.status === 429) {
await new Promise((r) => setTimeout(r, Number(res.headers.get("Retry-After") ?? 5) * 1000));
continue;
}
const { data } = await res.json();
await save(episode.guid, data.id);
}The Idempotency-Key is what separates a script you can rerun from a script
that bills twice. Same key, same body, returns the same transcript, however
many times you repeat it. A reused key with a different body returns
409 idempotency_conflict, which is the warning that you reused it by mistake.
The 429 rate_limited carries Retry-After with the real wait in seconds. A
whole archive unblocking at once hits the ceiling easily; spreading the burst
across that wait is simpler than any queue you would write. Limits are in
pricing and limits, codes in errors.
Guest names are where the text fails
Proper nouns, brands and loanwords are the most common error of any transcription system. It is not a specific defect: they are the words least present in any training set. And podcasts are made of them. Every episode has a guest with a surname the system has never seen, a company named fifteen times, a niche term that exists nowhere else.
You know this before you send the audio. The glossary field is the fix, and it
applies to that episode only:
{
"url": "https://cdn.example.com/ep-142.mp3",
"glossary": {
"mote gruto": "Mateus Gruto",
"cubernetes": "Kubernetes",
"fintéquis": "fintechs"
}
}Up to 100 entries, 80 characters each. The correction works over word sequences, so swapping two words for two works, and the time span is split across the replaced words. In practice a podcast's glossary is two lists: a fixed one with the hosts' and the show's names, and a per-episode one with that week's guest. Full reference in transcripts.
What to do with the transcript
The running text serves the episode page and search. The more valuable part is
words, which carries every word with start and end in decimal seconds.
- Search inside the episode. Indexing words with their timestamps lets a search result drop the listener at the right minute rather than at the whole episode. That is the difference between "this episode covers that" and "covers that at 34:12".
- Cutting clips. The stretch that becomes a thirty-second clip is a pair of timestamps. Median word timestamp error sits between 69 and 86 ms, so cutting at the word is safe.
- Chapters. A long pause between turns is usually a topic change. Not perfect, but a better draft than a blank page.
What not to trust unreviewed
Publishing a transcript unread is a cost decision, and it needs numbers. Ours: 6.34% WER on 2,000 excerpts of spontaneous Brazilian speech, measured against human transcripts. Six wrong words in a hundred. For search and discovery, that does not get in the way. To publish as the episode's official text, someone reads it.
Two podcast-specific cases:
- Background music sometimes becomes text. Sung intros and scored beds get transcribed as speech. If your show opens with music, the first seconds of the transcript will need trimming.
- Overlapping speech gets a single label. Host and guest laughing over each other become one speaker. Diarization runs 14.8% DER on that slice against 4.6% overall. A two-person show with individual microphones is the good case; a four-person table on one microphone is the bad one. The breakdown is in accuracy and quality, and what speaker separation does and does not do is in what speaker diarization is.
A new episode, on publication day
Whoever publishes weekly does not run an archive script: they run a hook. The
feed updates, you take the new item's enclosure and create the transcript with
webhookUrl. When the notice arrives, you fetch the transcript once by id and
publish it alongside the episode.
The webhook body carries id, status, durationSeconds and cost, but not
text or words: on a two-hour episode those fields run past 10 MB, and a
webhook is not the place for that. Every delivery is signed in the
Transcrevo-Signature header, signed with the account's webhook secret, which
the dashboard shows and rotates without touching your API keys.
A forty-minute episode costs US$ 0.03 with speakers, or US$ 0.02 without. That is less than storing the mp3.