How to generate SRT captions automatically from audio
The API returns every word with a start and end in seconds. Here is the code that turns that into SRT and VTT, and the places automatic captions tend to go wrong.
Automatic captions are not an output format of transcription: they are what you
do with its timestamps. The API returns every word with start and end in
decimal seconds, and SRT is a plain text file with simple rules. The whole job
is grouping words into blocks and printing them.
The part worth discussing is the grouping. It decides whether the captions read well or flicker across the screen.
What the API returns
"words": [
{ "start": 0.32, "end": 0.78, "text": "Bom", "speaker": null },
{ "start": 0.78, "end": 1.04, "text": "dia", "speaker": null },
{ "start": 1.04, "end": 1.51, "text": "a", "speaker": null },
{ "start": 1.51, "end": 2.06, "text": "todos.", "speaker": null }
]start and end are seconds, in decimal: 1.51 is a second and a half, not
1510 milliseconds. That detail has broken integrations written by experienced
people, because most caption libraries expect integer milliseconds. Multiply by
1000 before formatting.
The rest of the response is documented in transcripts.
The rules that decide readability
No captioning rule is law, but the conventions almost every player and platform tolerates well are these four:
- At most two lines per block, around 42 characters each.
- Each block on screen for 1 to 7 seconds.
- Break the block on a noticeable pause between words, somewhere around 0.4 seconds.
- Break after a period, question mark or exclamation mark too.
Spontaneous speech pushes against those limits constantly, because speaking rate swings inside a single recording. That is why the pause matters more than the character count: cutting at silence produces a block the viewer reads in one go, cutting mid-phrase produces a block they read twice.
The code
Here is the full generator, no dependencies:
type Word = { start: number; end: number; text: string; speaker: string | null };
const MAX_CHARS = 84;
const MAX_SECONDS = 7;
const PAUSE = 0.4;
function group(words: Word[]): Word[][] {
const blocks: Word[][] = [];
let current: Word[] = [];
for (const word of words) {
const previous = current.at(-1);
const tooLong = current.reduce((n, w) => n + w.text.length + 1, 0) + word.text.length > MAX_CHARS;
const tooSlow = current.length > 0 && word.end - current[0].start > MAX_SECONDS;
const paused = previous !== undefined && word.start - previous.end > PAUSE;
const ended = previous !== undefined && /[.!?]$/.test(previous.text);
if (current.length > 0 && (tooLong || tooSlow || paused || ended)) {
blocks.push(current);
current = [];
}
current.push(word);
}
if (current.length > 0) blocks.push(current);
return blocks;
}
function stamp(seconds: number, separator: string) {
const ms = Math.round(seconds * 1000);
const pad = (n: number, size = 2) => String(n).padStart(size, "0");
return `${pad(Math.floor(ms / 3600000))}:${pad(Math.floor(ms / 60000) % 60)}:${pad(Math.floor(ms / 1000) % 60)}${separator}${pad(ms % 1000, 3)}`;
}
function wrap(text: string) {
if (text.length <= MAX_CHARS / 2) return text;
const middle = text.lastIndexOf(" ", Math.floor(text.length / 2));
return middle === -1 ? text : `${text.slice(0, middle)}\n${text.slice(middle + 1)}`;
}
export function toSrt(words: Word[]) {
return group(words)
.map((block, index) => {
const text = wrap(block.map((w) => w.text).join(" "));
return `${index + 1}\n${stamp(block[0].start, ",")} --> ${stamp(block.at(-1)!.end, ",")}\n${text}\n`;
})
.join("\n");
}Output:
1
00:00:00,320 --> 00:00:02,060
Bom dia a todos.
2
00:00:02,440 --> 00:00:03,370
Vamos começar.VTT changes three things
WebVTT is the format the HTML track element consumes natively. From the same
grouping:
- The file opens with a
WEBVTTline and a blank line. - The millisecond separator is a dot, not a comma (
stamp(x, ".")). - The block's sequence number is optional.
export function toVtt(words: Word[]) {
const cues = group(words).map((block) => {
const text = wrap(block.map((w) => w.text).join(" "));
return `${stamp(block[0].start, ".")} --> ${stamp(block.at(-1)!.end, ".")}\n${text}\n`;
});
return `WEBVTT\n\n${cues.join("\n")}`;
}Captions that name the speaker
Ask for speakers on creation and every word comes back with a label ("A",
"B"). Two changes to the grouping: break the block whenever the label changes,
and prefix the line with the name.
const paused = previous !== undefined && (word.start - previous.end > PAUSE || word.speaker !== previous.speaker);The labels are anonymous: "A" means "the first voice", not a person. Mapping a
label to a name comes from outside the audio, and the most reliable basis is
usually speaking order combined with who you know was in the room. If that is
central to your product, read
what speaker diarization is first,
because separation degrades on short turns.
Where automatic captions go wrong
The timestamps themselves are good: median error against an external reference sits between 69 and 86 ms, about two frames at 24 fps. Nobody sees that.
What people do see is the tail. On a podcast, 1.8% of words land more than a second off; on a vlog with background music, 3.1%. A seven-second block dilutes that, but on a short block the drift shows up as a caption arriving early or late. The per-slice numbers are in accuracy and quality.
Three cases worth knowing before you publish captions unreviewed:
- Background music sometimes becomes text. Lyrics get transcribed as speech. On video with a sung soundtrack, review.
- Overlapping speech gets a single label. Two people talking at once become one speaker in the captions.
- Proper nouns and brands are the most common error. Fix them with
glossaryon the request, passing what the system hears and what it should write. It is the largest quality gain per line of code you will write.
What it costs
US$ 0.03 per hour of audio, billed by the second. A 12-minute video costs US$ 0.006. Captioning a channel of 300 ten-minute videos is 50 hours, US$ 1.50. With speaker labels, US$ 2.50. The full table is in pricing and limits.
There is no endpoint that hands you finished SRT, and that is deliberate: line breaks and block duration are editorial decisions belonging to whoever publishes, and every product that has actually shipped captions ended up wanting its own. The 40 lines above are the format; the rest is your rule.