Running a speech model yourself, or using a transcription API
The cost of self-hosting is not the GPU price, it is everything around it. Where each path wins, and where a transcription API simply does not serve.
There are open speech recognition models good enough for production, and they run on a GPU you can rent by the hour. So the question is fair: why pay per hour of audio when you can pay per hour of machine?
The answer depends on numbers only you have. What can be done here is list what belongs in the calculation, because the comparison most people run leaves half of it out.
What the self-hosting math usually forgets
GPU price per hour is the easy number, and in most comparisons it is the only one. The others:
Idle time. You rent by machine hour and charge yourself by audio hour. If your queue arrives in bursts (and transcription queues do arrive in bursts: meetings end at 5pm, lectures upload at night, import batches run overnight), the GPU sits there warm and waiting. A card at 30% utilization costs three times what the spreadsheet said.
Scaling up and down. Bringing up a machine with the model loaded is not instant. Pull the image, load the weights, warm up. Either you keep idle capacity for the burst, or the burst waits. Both cost, one in money and one in latency.
The machine that fails quietly. This is the cost nobody predicts and everybody pays once. A GPU without the right kernel compiled does not throw an error: it returns worse transcripts. Health checks answer 200, the queue moves, and the problem only surfaces in the ear of whoever reads the text, days later. Catching that requires measuring quality continuously, not uptime.
A retry that goes back to the same machine. When a job fails, the naive retry hands it to the same worker, which fails again, and the job burns all three attempts in one place. That is an afternoon to discover and another to fix.
Keeping current. Speech models improve. Every swap is a migration: test, compare, prove the quality went up on your material and not just on somebody else's benchmark.
None of this is impossible. It is platform work, and it costs engineer-months, which is the most expensive line on the whole spreadsheet.
What the API math is
US$ 0.03 per hour of audio processed, or US$ 0.05 with speaker separation, billed by the second. No subscription, no minimum, no plan. Past 10,000 hours transcribed by the account, it drops to US$ 0.02 automatically.
You do not pay per request, per stored file, or per language. Failed transcripts are not billed. The full table is in pricing and limits.
A thousand hours of audio a month is US$ 30. Twenty thousand hours is US$ 500. Compare against what you have: GPU hours billed, divided by hours of audio actually processed. Not the card's hourly price: the card's price divided by the audio it digested, idle time included.
When running it yourself is the right call
Three cases, and they are real.
The audio cannot leave. A compliance requirement, a customer contract, an internal policy that does not allow third-party processing. No price fixes that, and the discussion ends there.
The GPU is already there and idle. If you already run a fleet for something else and it has slack, the marginal cost of putting transcription on it is low.
You need real-time streaming. This API is asynchronous by construction: you create the transcript, it processes, and you collect the result by polling or by webhook. There is no endpoint that transcribes live, word by word, while someone speaks. Live captioning and an assistant that answers during the call are not cases it serves. If that is what you need, the path is a different one, and there is no point pretending otherwise.
When the API wins
Uneven volume. Bursts are where paying per audio hour beats paying per machine hour, because the idle time is not yours.
Small team. An engineer minding a GPU fleet is an engineer not minding your product. That is the real cost, and it is larger than any rate difference at low or middling volume.
You want the measured number, not the model card's. This is the subtle one. Running the model yourself, the quality you get is the quality your configuration produces, and it is not the published benchmark: accuracy depends on post-processing, on normalization, on how the audio was segmented. Measuring that is a project of its own.
What to actually compare
If you are deciding, the test that answers is not reading a price table. It is running your audio through both paths and comparing the text with the same normalizer.
Here is why that matters. The same system, on the same material, scores 7.7% WER normalized and 11.8% raw. Four points of spread without a line of code changing, purely from the comparison rule. Which means a WER published by one side does not compare to a WER published by the other, and the only trustworthy ruler is the one you apply identically to both.
Our measurement, so you have a starting point: 6.34% WER on 2,000 excerpts of spontaneous Brazilian speech, against human reference transcripts, not against another system's output. Spontaneous speech, not studio-read sentences, which is the material where every system looks excellent. The whole protocol is in accuracy and quality and in how we measure accuracy.
A new account comes with US$ 2 in credit, no card, which is roughly 65 hours of audio. Enough to run your real material and decide on your own number, which is the only one that matters in this choice.