Word-level timestamps: what you can build with them
Every word comes back with a start and end in decimal seconds, with a median error of 69 to 86 ms. Search inside audio, clip cutting and synced highlighting all come from that.
A transcription that returns only text solves half the problem. The other half
needs to know where each word sits in the audio, and that is what the words
field delivers:
"words": [
{ "start": 0.32, "end": 0.78, "text": "Bom", "speaker": null },
{ "start": 0.78, "end": 1.04, "text": "dia", "speaker": null },
{ "start": 1.04, "end": 1.51, "text": "a", "speaker": null },
{ "start": 1.51, "end": 2.06, "text": "todos.", "speaker": null }
]start and end are seconds, in decimal, counted from the start of the
audio. 1.51 is a second and a half, not 1510. Worth checking that first,
because a good share of media libraries expect integer milliseconds and the bug
never announces itself: the player just syncs badly.
How precise is it
We measured the error against an external reference, matching words by text:
| Measure | Value |
|---|---|
| Median error (p50) | 69 to 86 ms |
| Words more than 1 s off, on a podcast | 1.8% |
| Words more than 1 s off, on a vlog with music | 3.1% |
The median translates easily: at 24 frames per second, a frame is 41.7 ms. The typical error is under two frames. Nobody sees that in a caption or hears it in a cut.
The number that matters for your product is the other one. Between 1.8% and 3.1% of words land more than a second off, and the worse slice is audio with music. A product that cuts clips automatically and publishes them unreviewed will be wrong on that slice. A product that uses the timestamp to take a user to the passage will not. The difference is whether the error costs a second of scrubbing or a wrongly published video.
The per-slice numbers are in accuracy and quality.
Search that lands on the right minute
The most direct use. Indexing words with their times turns a search result from "this episode covers that" into "covers that at 34:12".
The core is finding the word sequence that matches the term and returning the
first one's start:
function locate(words: Word[], query: string) {
const terms = query.toLowerCase().split(/\s+/);
for (let i = 0; i <= words.length - terms.length; i++) {
const hit = terms.every((term, j) => words[i + j].text.toLowerCase().replace(/[.,!?]/g, "") === term);
if (hit) return { start: words[i].start, end: words[i + terms.length - 1].end };
}
return null;
}When you play it, back off about two seconds from start. Jumping exactly onto
the word drops the listener mid-sentence, and they lose the context.
Highlighting in sync with playback
The effect where the word lights up as the audio plays is a binary search on
timeupdate:
function activeIndex(words: Word[], time: number) {
let low = 0;
let high = words.length - 1;
while (low <= high) {
const mid = (low + high) >> 1;
if (time < words[mid].start) high = mid - 1;
else if (time > words[mid].end) low = mid + 1;
else return mid;
}
return high; // between two words: keep the last one that already passed
}That return high at the end is the detail separating a stable highlight from a
flickering one. The end of one word is not the start of the next: there is
silence between them, and in such a gap the search finds nothing. Returning the
last finished word keeps the highlight still during the pause, which is what the
eye expects.
Cutting audio and video at the word
A pair of timestamps is a cut command:
ffmpeg -i episode.mp3 -ss 2412.31 -to 2447.88 -c copy clip.mp3Two recommendations that save rework. Cut at the pause, not at the exact
word boundary: look for the largest gap between one word's end and the next
word's start within half a second of your target, and cut there. And leave 100
to 200 ms of slack on each side, which is larger than the typical error and
inaudible.
What timestamps do not solve
Overlap. When two people speak at once, the stretch becomes a single word sequence. The times survive, the attribution does not: diarization runs 14.8% DER on that slice against 4.6% overall.
Music. Lyrics sometimes get transcribed as speech, and that is where the over-a-second drift concentrates (3.1% of words on a vlog with a soundtrack).
Reproducibility. The output is not deterministic. Sending the same file
twice can return different segmentation across dozens of blocks. If your product
stores a clip cut by timestamp, store the numbers you used; do not regenerate
the transcript to check. There is no seed.
Aligning existing text. The API transcribes audio; it does not take a finished script and return its timings. If you already have the script and need an alignment, the path is to transcribe and match the two word sequences on your side.
With speakers
Ask for speakers and every word gets a label ("A", "B") alongside the
times. That is what lets you build turns rather than just words, which is what
interview captions and meeting minutes need. The rate changes to US$ 0.05 per
hour, and the details are in transcripts.
Without speakers, speaker comes back null on every word. That is not an
error: it is the answer to a question you did not ask.