Batch transcription: thousands of files without breaking the queue
What changes when a batch goes from ten files to ten thousand: rate ceilings, reserved balance, duplicate billing, and which failures are worth retrying.
Transcribing one file is a POST and a GET. Transcribing an archive of ten
thousand is a different problem, and the difference is not in the transcription
code: it is in what happens around it when everything fires at once.
Four things break in that transition, always the same four.
1. The rate ceilings
Per-account limits, on a one-minute sliding window:
| Path | Per minute |
|---|---|
POST /v1/transcripts | 120 |
POST /v1/uploads | 120 |
PUT /v1/uploads/{id} | 600 |
A backlog unblocking at once hits those ceilings easily, because a customer's
peak looks nothing like their average traffic. When it does, the response is
429 rate_limited with a Retry-After header carrying the real wait in
seconds.
The classic mistake is treating 429 as a failure and pushing the item to a
dead-letter queue. It is not a failure: it is scheduling. Nothing was lost, only
postponed.
async function create(body: object, key: string): Promise<{ id: string }> {
for (let attempt = 0; attempt < 8; attempt++) {
const res = await fetch(`${API}/v1/transcripts`, {
method: "POST",
headers: { ...headers, "Idempotency-Key": key },
body: JSON.stringify(body),
});
if (res.status === 429) {
const wait = Number(res.headers.get("Retry-After") ?? 5);
await new Promise((r) => setTimeout(r, wait * 1000));
continue;
}
const json = await res.json();
if (!res.ok) throw new Error(json.error.code);
return json.data;
}
throw new Error("rate_limited");
}Two ceilings per minute do not sound like much until you run the math backwards: 120 creations per minute is 7,200 per hour, and a ten-thousand-file batch is submitted in under two hours while respecting the ceiling the whole way. If you need more, talk to us before the migration starts, not halfway through it.
2. The balance that stalls the batch
With no credit, new transcripts are refused with 402 insufficient_balance.
Nothing is deleted, but the batch stops. There is a second way this shows up: an
audio file longer than the remaining balance covers ends in failed with
error.code: "insufficient_balance", unbilled.
What to monitor is not total:
{
"data": {
"currency": "USD",
"total": 187.4062,
"reserved": 40.5,
"available": 146.9062
}
}reserved is what in-flight transcripts have already held. available is
total minus reserved, and it is what the next transcript can actually spend.
In a large batch, the gap between the two is exactly the size of the queue in
flight, so an alarm on total fires far too late.
Two things prevent the stall: checking GET /v1/balance inside the loop every
few hundred creations, and switching on auto top-up in the dashboard, which
handles it without somebody awake at three in the morning. The full table is in
pricing and limits.
Worth knowing too: an account running on the welcome credit alone processes at
most 2 transcripts at a time, and a third concurrent one gets
429 concurrency_limit. That is abuse protection on the free balance, not a
plan limit: the first top-up removes it. Anyone testing a batch script on a
brand-new account and concluding the API is slow is measuring that cap.
3. Paying twice for the same audio
An orchestrator that redistributes work sends the same file more than once. A worker that dies mid-job puts the item back on the queue. A script you rerun because the first half went wrong resends the first half. In batch, this is not an exception: it is normal behavior for any queue with retries.
Every submission is a billed transcript unless you say otherwise. There are two mechanisms, and they solve different problems.
Idempotency-Key protects the request. Same key and same body return the
same transcript, however many times you repeat it. It is what saves you when the
connection drops after the request arrived. Same key with a different body
returns 409 idempotency_conflict.
reuseIfIdentical protects the content. Upload the file declaring its
sha256 and create the transcript with that field: if this account already has
a transcript for that same content with the same parameters, the response is
200 with it instead of 201 with a new one, and nothing is billed. The lookup
reaches even one still in processing, which is exactly the burst case.
If the audio came from an upload, use the uploadId itself as the
idempotency key. The instinct is to key on your own media id, but every
attempt uploads the file again, so the body's uploadId changes each time:
same key, different body, 409 idempotency_conflict on every attempt,
forever. What solves "I sent the same video twice" is reuseIfIdentical.
Without sha256 on the upload there is no way to recognize the content, and a
request with reuseIfIdentical is refused with 400 validation
(reuse_needs_sha256) instead of becoming a silent charge. Both upload paths
are in sending files.
4. The failures, sorted by who fixes them
In a batch of ten thousand, some files will fail. What decides whether the batch finishes on its own is classifying each failure correctly. Failed transcripts are not billed, so retrying the ones worth retrying is free.
error.code | Retry? |
|---|---|
engine_failed | Yes. It failed every attempt on our side; resending usually goes through. |
insufficient_balance | Yes, after topping up. |
audio_unreachable | Maybe. The URL may have expired; mint a new one and resend. |
audio_invalid | No. The audio could not be decoded. Fix the file. |
audio_too_long | No. Over 10 hours. Split before sending. |
audio_too_large | No. Over 2 GiB. |
audio_mismatch | No. Stored audio does not match the declared sha256; the upload corrupted. |
Retrying audio_invalid ten thousand times is the fastest way to turn a batch
that would have finished in two hours into an infinite loop. The full list is in
errors.
One more care on chunked uploads: without declaring expectedBytes, a chunk
lost in the middle of 2 GiB is accepted silently. The transcript comes out
done, with plausible text, billed, covering only the part that arrived. In a
large batch that is the worst possible error, because it shows up nowhere.
Declare the size and the digest.
Closing the books afterwards
GET /v1/transcripts takes since and until, and the range is applied in the
database: closing out a day's cost does not mean paginating the whole account.
curl "https://api.transcrevo.com/v1/transcripts?since=2026-07-01T00:00:00Z&until=2026-07-02T00:00:00Z&perPage=100" \
-H "Authorization: Bearer $TRANSCREVO_API_KEY"One detail that bites during reconciliation: cost comes back with fractions
of a cent (0.0167). If you store it in an integer
column, or run it through parseInt, the truncation is silent and it erases
almost every short-file batch. Store decimal.
And the bottom line: a thousand hours of audio cost US$ 30, or US$ 50 with speaker separation. Past 10,000 hours transcribed by the account, the rate drops on its own to US$ 0.02, with no contract and no negotiation.