transcrevo docs

Transcripts

Create and read transcripts over the API: URL or uploadId, language, speaker separation, glossary, webhook, and the response format with timestamps.

Create a transcript

POST /v1/transcripts
FieldTypeDescription
urlstringPublic (http/https) URL of the audio.
uploadIdstringId of a sent file. Send url or uploadId, never both.
languagestringOptional. Language code (pt, en, …). Omit it to detect the language.
speakersboolean, integer, rangeOptional. Separates speakers. Changes the rate.
webhookUrlstringOptional. Public http(s) URL that gets a POST when the transcript finishes.
glossaryobjectOptional. Corrections that apply to this audio only (proper names, jargon).
reuseIfIdenticalbooleanOptional. Returns the transcript this account already has for this exact audio.

The speakers field

One field decides everything about speakers:

ValueMeaning
omittedOne block of text, speakers not separated.
trueSeparate speakers; the API works out how many there are.
3There are exactly 3 speakers.
"2-5"There are between 2 and 5 speakers.

Only give a count when you know it: a wrong guess forces the separation into the wrong number. With speakers, every item in words carries a speaker label ("A", "B", …) and the response carries speakerCount.

curl https://api.transcrevo.com/v1/transcripts \
  -H "Authorization: Bearer $TRANSCREVO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/meeting.mp3",
    "speakers": 5
  }'

201 response:

{
	"data": {
		"id": "6b9f2b81-1c9a-4f4e-9d5f-8f2a7c1e3b90",
		"status": "processing",
		"text": null,
		"words": null,
		"error": null,
		"language": null,
		"durationSeconds": 0,
		"speakerCount": 0,
		"cost": 0,
		"currency": "USD",
		"model": "trv-1",
		"url": "https://example.com/meeting.mp3",
		"createdAt": "2026-08-24T18:00:00.000Z"
	}
}

Accepted formats: mp3, wav, m4a/aac, ogg, opus, flac, webm, wma, amr and the audio track of video files (mp4, mov, mkv). Anything that cannot be decoded comes back as failed with audio_invalid, never charged. Maximum length: 10 hours per audio.

What it looks like when it completes

The same object, now with text, words, durationSeconds and cost filled in. start and end are SECONDS, as decimal numbers (1.5 is one and a half seconds, not 1500), counted from the beginning of the audio:

{
	"data": {
		"id": "6b9f2b81-1c9a-4f4e-9d5f-8f2a7c1e3b90",
		"status": "done",
		"text": "Good morning everyone. Let us begin.",
		"words": [
			{ "start": 0.32, "end": 0.78, "text": "Good", "speaker": "A" },
			{ "start": 0.78, "end": 1.04, "text": "morning", "speaker": "A" },
			{ "start": 1.04, "end": 1.51, "text": "everyone.", "speaker": "A" },
			{ "start": 2.44, "end": 2.9, "text": "Let", "speaker": "B" },
			{ "start": 2.9, "end": 3.1, "text": "us", "speaker": "B" },
			{ "start": 3.1, "end": 3.37, "text": "begin.", "speaker": "B" }
		],
		"error": null,
		"language": "en",
		"durationSeconds": 3,
		"speakerCount": 2,
		"cost": 0.0042,
		"currency": "USD",
		"model": "trv-1",
		"url": "https://example.com/meeting.mp3",
		"createdAt": "2026-08-24T18:00:00.000Z"
	}
}

text is the words in words joined together, and words is what subtitles, search-by-passage and video cutting are built from.

The output is not reproducible

Sending the same file twice can return slightly different transcripts: a word comes out differently, the segmentation shifts by a few dozen blocks, informal speech normalises one way in one run and another way in the next. This is not a bug and there is no seed for it: transcription is not a deterministic function of the file.

In practice: do not diff transcripts across runs. To audit a complaint, keep the transcript you delivered (its id is permanent) instead of generating it again.

Being told when it finishes (webhook)

A two hour VOD takes minutes to transcribe. Polling every three seconds is ~100 requests per job just to hear processing. Send webhookUrl and the API tells you once, when it is done:

curl https://api.transcrevo.com/v1/transcripts \
  -H "Authorization: Bearer $TRANSCREVO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "uploadId": "0f3c2d11-8a4b-4c6e-b2d9-5e7f1a9c3b42",
    "webhookUrl": "https://your-server.com/hooks/transcrevo"
  }'

The notification is a POST carrying the summary of the transcript, in the usual envelope, and it fires for failed as well as done:

{
	"data": {
		"id": "6b9f2b81-…",
		"status": "done",
		"durationSeconds": 8700,
		"language": "en",
		"speakerCount": 6,
		"cost": 12.08,
		"currency": "USD",
		"error": null,
		"model": "trv-1",
		"url": null,
		"createdAt": "2026-08-27T18:00:00.000Z"
	}
}

text and words are not in the body: on a two hour audio they go past 10 MB. Once you get the notification, fetch the transcript once from GET /v1/transcripts/{id}.

Verifying the signature

Every delivery carries a Transcrevo-Signature header, shaped t=<unix>,v1=<hex>. v1 is the HMAC-SHA256 of <t>.<raw body>, and the secret is your account's webhook secret (whsec_...), shown at /dashboard/webhooks.

It is deliberately independent from your API keys: the key authenticates your requests and is rotated when it leaks; the secret verifies signatures and is rotated when your verifier changes. Tied together, rotating the key broke your endpoint's verification at the same instant.

When you rotate it, the previous secret keeps verifying for 24 hours: ship the new one, confirm, and let the old one expire. Accept either during that window.

import { createHmac, timingSafeEqual } from "node:crypto";

const secret = process.env.TRANSCREVO_WEBHOOK_SECRET;

export function valid(header, rawBody) {
	const { t, v1 } = Object.fromEntries(header.split(",").map((part) => part.split("=")));
	// An old notification is a replayed one: reject anything outside five minutes.
	if (Math.abs(Date.now() / 1000 - Number(t)) > 300) return false;
	const expected = createHmac("sha256", secret).update(`${t}.${rawBody}`).digest("hex");
	return timingSafeEqual(Buffer.from(expected), Buffer.from(v1));
}

Answer 2xx to acknowledge. Anything else (or no answer within 10s) is another attempt, with growing waits: 30s, 2min, 10min, 30min, 2h and 6h. After the sixth the API stops trying; the transcript is still there to be fetched.

Per-request glossary

Proper names and jargon are where every ASR engine slips, and you know in advance which ones show up in your audio. glossary is { "what it hears": "what to write" }, applied after transcription and only to that audio:

curl https://api.transcrevo.com/v1/transcripts \
  -H "Authorization: Bearer $TRANSCREVO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/podcast.mp3",
    "glossary": { "gatu": "Gatto", "bertual": "Bertuol" }
  }'

Up to 100 entries, up to 80 characters each. Matching is by sequence of words, so "mote gruto": "Mateus Gruto" works too, and the time span is split across the replaced words. What you send wins over the built-in corrections.

Not paying twice for the same audio

An orchestrator that redistributes its queue sends the same video more than once, and every send is a charged transcript. To close that off without keeping your own bookkeeping, send the file with sha256 and create the transcript with reuseIfIdentical:

curl https://api.transcrevo.com/v1/transcripts \
  -H "Authorization: Bearer $TRANSCREVO_API_KEY" \
  -H "Content-Type: application/json" \
  -d "{ \"uploadId\": \"$UPLOAD_ID\", \"reuseIfIdentical\": true }"

If this account already has a transcript of that content with the same parameters, the answer is 200 with it (instead of 201 with a new one) and nothing is charged. The lookup reaches transcripts still in processing, which is exactly the burst case. Without sha256 on the upload there is no way to recognise the content, so the request is refused with 400 validation (reuse_needs_sha256) rather than quietly becoming a charge.

Two exactly simultaneous requests can still create two transcripts: reuse looks at what already exists, and at that instant neither of them does.

Retrying safely

If the connection drops after the request arrived, retrying would create a second transcript (and a second charge). To avoid that, send an Idempotency-Key header with any string up to 64 characters:

curl https://api.transcrevo.com/v1/transcripts \
  -H "Authorization: Bearer $TRANSCREVO_API_KEY" \
  -H "Idempotency-Key: meeting-2026-08-25-01" \
  -H "Content-Type: application/json" \
  -d '{ "url": "https://example.com/meeting.mp3" }'

The same key returns the same transcript, however many times you retry. Using the same key with a different body returns 409 idempotency_conflict, which means the key was reused by mistake.

If the audio came from an upload, use the uploadId itself as the key. The instinct is to key by your own media id, so that retrying the whole job does not pay twice, but every retry uploads the file again, so the uploadId in the body changes each time. Same key, different body: 409 idempotency_conflict on every retry, forever. What solves "I sent the same video twice" is reuseIfIdentical, above, not the idempotency key.

Get a transcript

GET /v1/transcripts/{id}
curl https://api.transcrevo.com/v1/transcripts/6b9f2b81-1c9a-4f4e-9d5f-8f2a7c1e3b90 \
  -H "Authorization: Bearer $TRANSCREVO_API_KEY"

Poll it every few seconds until status leaves processing. At volume, send webhookUrl when you create it instead: one read per job rather than a hundred.

List transcripts

GET /v1/transcripts
ParameterDescription
pagePage, starting at 1. Default: 1.
perPageItems per page, up to 100. Default: 20.
statusOptional. Filter by processing, done, failed.
sinceOptional. Only what was created at or after this ISO 8601 date-time.
untilOptional. Only what was created before this ISO 8601 date-time.
curl "https://api.transcrevo.com/v1/transcripts?status=done&perPage=50" \
  -H "Authorization: Bearer $TRANSCREVO_API_KEY"

The date range is applied in the database, so closing out one day of cost does not mean paging through the whole account:

curl "https://api.transcrevo.com/v1/transcripts?since=2026-08-27T00:00:00Z&until=2026-08-28T00:00:00Z" \
  -H "Authorization: Bearer $TRANSCREVO_API_KEY"

Newest first. List items carry every field except text and words, which go past 10 MB per item on a long audio; fetch the single transcript to get them.

{
	"data": {
		"transcripts": [{ "id": "6b9f2b81-…", "status": "done", "durationSeconds": 1834 }],
		"total": 128,
		"page": 1,
		"perPage": 20
	}
}

Delete a transcript

DELETE /v1/transcripts/{id}

Erases the text, the words and the stored audio. The usage record stays, since it is what the invoice for the period is built from, but the content is gone and cannot be recovered.

A transcript still in processing cannot be deleted (409 transcript_processing): the result would arrive later and refill what you just erased. Wait for it to finish.

Response fields

FieldDescription
idTranscript id.
statusprocessing, done or failed.
textThe transcribed text. null until it completes.
wordsEvery word with start, end (in seconds, decimal), text and speaker. null until done.
error{ code, message } when status is failed; null otherwise.
languageDetected (or requested) language. null while processing.
durationSecondsReal audio duration. 0 while processing.
speakerCountHow many distinct speakers were found. null when speakers was not requested, which is not the same as 0 (asked, found nobody).
costCost charged, in the minor unit of currency (cents), fractional: 0.0167 is a sixth of a cent. 0 while processing, always 0 on failure.
currencyCurrency of cost, ISO 4217 (USD).
modelModel that produced the transcript.
urlThe URL you sent. null when the audio came from a sent file.
createdAtCreation date (ISO 8601).

Status

StatusMeaning
processingQueued or being processed.
doneFinished: text available and cost charged.
failedFailed: see error. Failures are never charged.

On this page