transcrevo docs

Migrating here

From another transcription API or from your own open-model script: the concept map, what changes in your code, and what to check before you switch.

Two origins cover almost every migration: a transcription API you already use, or a script of yours running an open model. Both change little code; what really changes is where the decisions live.

Coming from another transcription API

The shape is the one almost every asynchronous API uses, so the translation is close to mechanical:

What you have todayHere
Create a job with an audio URLPOST /v1/transcripts with url
Create a job with a filePOST /v1/uploads (raw body), then uploadId
Polling the jobGET /v1/transcripts/{id}, or ?wait=30 to stop asking
Completion callbackwebhookUrl at creation
Speaker separationspeakers: true, 3 or "2-5"
Custom vocabulary / term boostingglossary, { "what it hears": "what to write" }
Your id on the jobexternalId (and metadata for the rest)
Timed wordswords[], in decimal seconds
Subtitle-sized blocksutterances[]

Five differences that usually bite:

  1. Times are decimal seconds, never milliseconds. 1.5 is one and a half seconds. A client that multiplies by 1000 produces subtitles a thousand times out of place, and the mistake only shows up in the player.
  2. Money is in the currency's minor unit, with a fraction. cost: 0.0167 is a sixtieth of a dollar. Storing it in an integer column zeroes it.
  3. The output is not deterministic and there is no seed. The same file can come back slightly different. Do not diff runs: keep the transcript you delivered, whose id is permanent.
  4. One language per file. Detection is automatic, but it picks ONE language for the whole audio: there are no per-span language marks.
  5. No SDK. It is plain HTTP with the key in the Authorization: Bearer header, and the OpenAPI spec generates a client in your language if you want one.

Before you switch, check it with your own audio (not with a sample of ours): the dashboard playground puts a transcript from here next to the one you already have, on the same audio, with the sound playing along. It is the only comparison that answers for YOUR material.

Coming from your own open-model script

If you run the model on a machine of your own today, what changes is not the quality of the text: it is who carries the operations.

What leaves your side:

  • The queue, and the retry when a machine dies mid-file.
  • Turning GPUs on and off with demand (and paying for them idle).
  • Decoding odd containers, truncated audio, files that are not audio.
  • Diarization, punctuation and word alignment, each with its own model.
  • Keeping all of that running while you work on something else.

What becomes yours:

  • One HTTP call and a webhook (or ?wait=).
  • A published cost per audio hour, with no idle GPU.

The shortest path is to keep your script running and send the same batch here for a few days, comparing on your own material before switching anything off. A new account starts with enough credit for that.

Checklist

  • Key on the server, never in a browser or a distributed app.
  • Idempotency-Key on creation, keyed on the first uploadId (why).
  • Branch on error.retryable instead of keeping your own retry table.
  • Keep requestId in the log for failures.
  • webhookUrl (with signature verification) or ?wait= instead of a polling loop.
  • Read GET /v1/limits instead of hardcoding the ceilings.
  • Decide retention: retentionDays per job, if your material needs a period.
  • Check the estimated cost against the first real invoice: billing is by measured audio duration, not per file and not per request.

On this page