Transcripts
Create and read transcripts over the API: URL or uploadId, language, speaker separation, glossary, webhook, and the response format with timestamps.
Create a transcript
POST /v1/transcripts| Field | Type | Description |
|---|---|---|
url | string | Public (http/https) URL of the audio. |
uploadId | string | Id of a sent file. Send url or uploadId, never both. |
language | string | Optional. Language code (pt, en, …). Omit it to detect the language. |
speakers | boolean, integer, range | Optional. Separates speakers. Changes the rate. |
webhookUrl | string | Optional. Public http(s) URL that gets a POST when the transcript finishes. |
glossary | object | Optional. Corrections that apply to this audio only (proper names, jargon). |
reuseIfIdentical | boolean | Optional. Returns the transcript this account already has for this exact audio. |
The speakers field
One field decides everything about speakers:
| Value | Meaning |
|---|---|
| omitted | One block of text, speakers not separated. |
true | Separate speakers; the API works out how many there are. |
3 | There are exactly 3 speakers. |
"2-5" | There are between 2 and 5 speakers. |
Only give a count when you know it: a wrong guess forces the separation into the
wrong number. With speakers, every item in words carries a speaker label
("A", "B", …) and the response carries speakerCount.
curl https://api.transcrevo.com/v1/transcripts \
-H "Authorization: Bearer $TRANSCREVO_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com/meeting.mp3",
"speakers": 5
}'201 response:
{
"data": {
"id": "6b9f2b81-1c9a-4f4e-9d5f-8f2a7c1e3b90",
"status": "processing",
"text": null,
"words": null,
"error": null,
"language": null,
"durationSeconds": 0,
"speakerCount": 0,
"cost": 0,
"currency": "USD",
"model": "trv-1",
"url": "https://example.com/meeting.mp3",
"createdAt": "2026-08-24T18:00:00.000Z"
}
}Accepted formats: mp3, wav, m4a/aac, ogg, opus, flac, webm, wma, amr and the
audio track of video files (mp4, mov, mkv). Anything that cannot be decoded
comes back as failed with audio_invalid, never charged. Maximum length:
10 hours per audio.
What it looks like when it completes
The same object, now with text, words, durationSeconds and cost filled
in. start and end are SECONDS, as decimal numbers (1.5 is one and a
half seconds, not 1500), counted from the beginning of the audio:
{
"data": {
"id": "6b9f2b81-1c9a-4f4e-9d5f-8f2a7c1e3b90",
"status": "done",
"text": "Good morning everyone. Let us begin.",
"words": [
{ "start": 0.32, "end": 0.78, "text": "Good", "speaker": "A" },
{ "start": 0.78, "end": 1.04, "text": "morning", "speaker": "A" },
{ "start": 1.04, "end": 1.51, "text": "everyone.", "speaker": "A" },
{ "start": 2.44, "end": 2.9, "text": "Let", "speaker": "B" },
{ "start": 2.9, "end": 3.1, "text": "us", "speaker": "B" },
{ "start": 3.1, "end": 3.37, "text": "begin.", "speaker": "B" }
],
"error": null,
"language": "en",
"durationSeconds": 3,
"speakerCount": 2,
"cost": 0.0042,
"currency": "USD",
"model": "trv-1",
"url": "https://example.com/meeting.mp3",
"createdAt": "2026-08-24T18:00:00.000Z"
}
}text is the words in words joined together, and words is what subtitles,
search-by-passage and video cutting are built from.
The output is not reproducible
Sending the same file twice can return slightly different transcripts: a word
comes out differently, the segmentation shifts by a few dozen blocks, informal
speech normalises one way in one run and another way in the next. This is not a
bug and there is no seed for it: transcription is not a deterministic function
of the file.
In practice: do not diff transcripts across runs. To audit a complaint, keep the
transcript you delivered (its id is permanent) instead of generating it again.
Being told when it finishes (webhook)
A two hour VOD takes minutes to transcribe. Polling every three seconds is ~100
requests per job just to hear processing. Send webhookUrl and the API tells
you once, when it is done:
curl https://api.transcrevo.com/v1/transcripts \
-H "Authorization: Bearer $TRANSCREVO_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"uploadId": "0f3c2d11-8a4b-4c6e-b2d9-5e7f1a9c3b42",
"webhookUrl": "https://your-server.com/hooks/transcrevo"
}'The notification is a POST carrying the summary of the transcript, in the
usual envelope, and it fires for failed as well as done:
{
"data": {
"id": "6b9f2b81-…",
"status": "done",
"durationSeconds": 8700,
"language": "en",
"speakerCount": 6,
"cost": 12.08,
"currency": "USD",
"error": null,
"model": "trv-1",
"url": null,
"createdAt": "2026-08-27T18:00:00.000Z"
}
}text and words are not in the body: on a two hour audio they go past
10 MB. Once you get the notification, fetch the transcript once from
GET /v1/transcripts/{id}.
Verifying the signature
Every delivery carries a Transcrevo-Signature header, shaped
t=<unix>,v1=<hex>. v1 is the HMAC-SHA256 of <t>.<raw body>, and the
secret is your account's webhook secret (whsec_...), shown at
/dashboard/webhooks.
It is deliberately independent from your API keys: the key authenticates your requests and is rotated when it leaks; the secret verifies signatures and is rotated when your verifier changes. Tied together, rotating the key broke your endpoint's verification at the same instant.
When you rotate it, the previous secret keeps verifying for 24 hours: ship the new one, confirm, and let the old one expire. Accept either during that window.
import { createHmac, timingSafeEqual } from "node:crypto";
const secret = process.env.TRANSCREVO_WEBHOOK_SECRET;
export function valid(header, rawBody) {
const { t, v1 } = Object.fromEntries(header.split(",").map((part) => part.split("=")));
// An old notification is a replayed one: reject anything outside five minutes.
if (Math.abs(Date.now() / 1000 - Number(t)) > 300) return false;
const expected = createHmac("sha256", secret).update(`${t}.${rawBody}`).digest("hex");
return timingSafeEqual(Buffer.from(expected), Buffer.from(v1));
}Answer 2xx to acknowledge. Anything else (or no answer within 10s) is another
attempt, with growing waits: 30s, 2min, 10min, 30min, 2h and 6h. After the sixth
the API stops trying; the transcript is still there to be fetched.
Per-request glossary
Proper names and jargon are where every ASR engine slips, and you know in advance
which ones show up in your audio. glossary is { "what it hears": "what to write" }, applied after transcription and only to that audio:
curl https://api.transcrevo.com/v1/transcripts \
-H "Authorization: Bearer $TRANSCREVO_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com/podcast.mp3",
"glossary": { "gatu": "Gatto", "bertual": "Bertuol" }
}'Up to 100 entries, up to 80 characters each. Matching is by sequence of words, so
"mote gruto": "Mateus Gruto" works too, and the time span is split across the
replaced words. What you send wins over the built-in corrections.
Not paying twice for the same audio
An orchestrator that redistributes its queue sends the same video more than once,
and every send is a charged transcript. To close that off without keeping your
own bookkeeping, send the file with sha256 and create the
transcript with reuseIfIdentical:
curl https://api.transcrevo.com/v1/transcripts \
-H "Authorization: Bearer $TRANSCREVO_API_KEY" \
-H "Content-Type: application/json" \
-d "{ \"uploadId\": \"$UPLOAD_ID\", \"reuseIfIdentical\": true }"If this account already has a transcript of that content with the same
parameters, the answer is 200 with it (instead of 201 with a new one) and
nothing is charged. The lookup reaches transcripts still in processing, which
is exactly the burst case. Without sha256 on the upload there is no way to
recognise the content, so the request is refused with 400 validation
(reuse_needs_sha256) rather than quietly becoming a charge.
Two exactly simultaneous requests can still create two transcripts: reuse looks at what already exists, and at that instant neither of them does.
Retrying safely
If the connection drops after the request arrived, retrying would create a
second transcript (and a second charge). To avoid that, send an
Idempotency-Key header with any string up to 64 characters:
curl https://api.transcrevo.com/v1/transcripts \
-H "Authorization: Bearer $TRANSCREVO_API_KEY" \
-H "Idempotency-Key: meeting-2026-08-25-01" \
-H "Content-Type: application/json" \
-d '{ "url": "https://example.com/meeting.mp3" }'The same key returns the same transcript, however many times you retry.
Using the same key with a different body returns 409 idempotency_conflict,
which means the key was reused by mistake.
If the audio came from an upload, use the uploadId itself as the key. The
instinct is to key by your own media id, so that retrying the whole job does
not pay twice, but every retry uploads the file again, so the uploadId in the
body changes each time. Same key, different body: 409 idempotency_conflict on
every retry, forever. What solves "I sent the same video twice" is
reuseIfIdentical, above, not the idempotency key.
Get a transcript
GET /v1/transcripts/{id}curl https://api.transcrevo.com/v1/transcripts/6b9f2b81-1c9a-4f4e-9d5f-8f2a7c1e3b90 \
-H "Authorization: Bearer $TRANSCREVO_API_KEY"Poll it every few seconds until status leaves processing. At volume, send
webhookUrl when you create it instead: one read per job rather than a hundred.
List transcripts
GET /v1/transcripts| Parameter | Description |
|---|---|
page | Page, starting at 1. Default: 1. |
perPage | Items per page, up to 100. Default: 20. |
status | Optional. Filter by processing, done, failed. |
since | Optional. Only what was created at or after this ISO 8601 date-time. |
until | Optional. Only what was created before this ISO 8601 date-time. |
curl "https://api.transcrevo.com/v1/transcripts?status=done&perPage=50" \
-H "Authorization: Bearer $TRANSCREVO_API_KEY"The date range is applied in the database, so closing out one day of cost does not mean paging through the whole account:
curl "https://api.transcrevo.com/v1/transcripts?since=2026-08-27T00:00:00Z&until=2026-08-28T00:00:00Z" \
-H "Authorization: Bearer $TRANSCREVO_API_KEY"Newest first. List items carry every field except text and words, which
go past 10 MB per item on a long audio; fetch the single transcript to get them.
{
"data": {
"transcripts": [{ "id": "6b9f2b81-…", "status": "done", "durationSeconds": 1834 }],
"total": 128,
"page": 1,
"perPage": 20
}
}Delete a transcript
DELETE /v1/transcripts/{id}Erases the text, the words and the stored audio. The usage record stays, since it is what the invoice for the period is built from, but the content is gone and cannot be recovered.
A transcript still in processing cannot be deleted (409 transcript_processing): the result would arrive later and refill what you just
erased. Wait for it to finish.
Response fields
| Field | Description |
|---|---|
id | Transcript id. |
status | processing, done or failed. |
text | The transcribed text. null until it completes. |
words | Every word with start, end (in seconds, decimal), text and speaker. null until done. |
error | { code, message } when status is failed; null otherwise. |
language | Detected (or requested) language. null while processing. |
durationSeconds | Real audio duration. 0 while processing. |
speakerCount | How many distinct speakers were found. null when speakers was not requested, which is not the same as 0 (asked, found nobody). |
cost | Cost charged, in the minor unit of currency (cents), fractional: 0.0167 is a sixth of a cent. 0 while processing, always 0 on failure. |
currency | Currency of cost, ISO 4217 (USD). |
model | Model that produced the transcript. |
url | The URL you sent. null when the audio came from a sent file. |
createdAt | Creation date (ISO 8601). |
Status
| Status | Meaning |
|---|---|
processing | Queued or being processed. |
done | Finished: text available and cost charged. |
failed | Failed: see error. Failures are never charged. |
Authentication
Authenticate every request with the Authorization header and a tk_ key. Up to 10 keys per account, rotation with a 24 h overlap, and revocation.
Sending files
Send local audio to the API: one request up to 100 MB, or resumable chunks of up to 16 MiB for files up to 2 GiB, then transcribe by uploadId.