Integration

How to transcribe audio with an API, from zero to text

Three HTTP requests turn an audio file into text with timestamps. No SDK, no queue to configure, no GPU server to run.

Transcribing audio with an API is simpler than most integrations make it look. It takes three requests: one to create the account and get the key, one to send the audio, one to fetch the text. No SDK, no new dependency in your project, no GPU machine to keep alive.

This post walks the whole path in curl, then covers the three details that usually bite: audio that is not on a public URL, waiting for the result, and cost.

Diagram of the three steps: POST to create the transcript, an immediate response with processing status, and the finished text with status done. At the bottom, the two ways to learn it finished: poll the endpoint or receive a webhook.
The three steps. All that separates a simple integration from a good one is the last: polling in a loop, or receiving a webhook.

1. The key

Create the account and copy the key that comes with it. No credit card, and the account starts with US$ 2 in credits, enough for roughly 65 hours of audio. The key goes in the Authorization header on every request:

export TRANSCREVO_API_KEY="your-key"

2. Send the audio

If the file already sits on a public URL (a bucket, a CDN, a podcast enclosure), one request is enough:

curl https://api.transcrevo.com/v1/transcripts \
  -H "Authorization: Bearer $TRANSCREVO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "url": "https://example.com/podcast.mp3" }'

The response comes back immediately, before the audio is processed:

{
	"data": {
		"id": "6b9f2b81-1c9a-4f4e-9d5f-8f2a7c1e3b90",
		"status": "processing",
		"text": null,
		"language": null,
		"cost": 0,
		"model": "trv-1",
		"createdAt": "2026-09-01T18:00:00.000Z"
	}
}

Keep the id. That is how you fetch the result.

Audio that is not on a public URL

A local file, or one in a private bucket, goes through upload. The API takes the whole file in one request, and takes it in chunks when it is large or the network is bad: a chunked upload resumes where it stopped instead of restarting 800 MB. Both paths are covered in sending files. Either way you get an uploadId, which replaces url:

curl https://api.transcrevo.com/v1/transcripts \
  -H "Authorization: Bearer $TRANSCREVO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "uploadId": "upl_9f2b81c9a4f4e9d5" }'

3. Fetch the text

Processing is asynchronous. Poll the id until status turns done:

curl https://api.transcrevo.com/v1/transcripts/6b9f2b81-1c9a-4f4e-9d5f-8f2a7c1e3b90 \
  -H "Authorization: Bearer $TRANSCREVO_API_KEY"
{
	"data": {
		"id": "6b9f2b81-1c9a-4f4e-9d5f-8f2a7c1e3b90",
		"status": "done",
		"text": "Bom dia a todos. Vamos começar.",
		"words": [
			{ "start": 0.32, "end": 0.78, "text": "Bom", "speaker": null },
			{ "start": 0.78, "end": 1.04, "text": "dia", "speaker": null }
		],
		"language": "pt",
		"durationSeconds": 3,
		"cost": 0.0025,
		"currency": "USD"
	}
}

text is the running transcript. words is what you need for captions, for search inside a recording, and for cutting video at the exact word: every word comes back with start and end in decimal seconds.

Do not sit in a polling loop

Polling is fine for one file. At volume it becomes request cost and delay: you learn it finished on the next tick, not when it finished. Pass webhookUrl on creation and the API calls you once, when it completes:

curl https://api.transcrevo.com/v1/transcripts \
  -H "Authorization: Bearer $TRANSCREVO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/podcast.mp3",
    "webhookUrl": "https://your-server.com/hooks/transcrevo"
  }'

The webhook body carries the id and the status, but not text or words: on a two-hour file those fields run past 10 MB, and a webhook is not the place for that. When the call arrives, fetch the transcript by id.

The parameters that change the result

Four optional fields account for most of the distance between an integration that works and one that disappoints:

FieldWhen to use it
languageWhen you already know the language. On short or noisy audio, auto-detection errs more than you would.
speakersWhen you need to know who said what. Changes the rate, does not improve the text.
glossaryWhenever the audio carries proper nouns, brands or domain jargon. The biggest quality gain per line of code.
webhookUrlWhenever there is volume.

glossary deserves attention. Proper nouns and loanwords are the most common error of any transcription system, because they are exactly the words missing from training. You pass what it hears and what it should write:

{
	"url": "https://example.com/meeting.mp3",
	"glossary": { "gatu": "Gatto", "bertual": "Bertuol" }
}

What it costs

US$ 0.03 per hour of audio, billed by the second: two minutes cost US$ 0.001, not one cent. With speakers, US$ 0.05 per hour. The charge lands when the transcript turns done, and while it processes cost is 0. The details, including the automatic discount above 10,000 hours, are in pricing and limits.

What goes wrong

  • 402 insufficient_balance: credit ran out. Nothing is deleted, and auto top-up exists so this does not happen in the middle of the night.
  • audio_too_long: the limit is 10 hours per file. Split before sending.
  • 429 rate_limited: you passed 120 creations per minute. Wait and retry.
  • 429 concurrency_limit: while the account runs on the welcome credit only, at most two transcripts process at once. The first top-up lifts the cap.

The full list of codes is in errors.

Summary

One key, one POST with url or uploadId, and one GET (or a webhook) to pick up the text. Everything else is optional and exists for specific cases: speakers to know who spoke, glossary to get the names right, language to stop depending on auto-detection.

All posts