Support

Transcribing support calls: what works and what does not

A recorded call is the most predictable audio there is: two voices, one topic, always the same format. That changes what is worth asking the API for, and what is worth checking in the result.

A support call is well-behaved audio. Two people, one topic, three to twenty minutes, recorded the same way every time. That is a real advantage: the pipeline that works on the first call works on the next hundred thousand without tuning.

What differs between one recording archive and another is a decision made before the API ever sees the file: whether the recorder keeps the two sides separate or mixed.

Two channels solve the hardest problem

If your recording has the agent on one channel and the customer on the other, you do not need diarization at all. Split the channels and transcribe each one. Attribution is exact by construction, and no algorithm competes with that.

ffmpeg -i call.wav -map_channel 0.0.0 agent.wav -map_channel 0.0.1 customer.wav

Two transcripts of a ten-minute call cost the same as one twenty-minute transcript, because billing is per hour of audio processed: US$ 0.03 per hour, without the speakers surcharge. It comes out cheaper than the same call with automatic separation, and more accurate.

To interleave the two in the right order, word timestamps are the key: both transcripts share the same clock, the start of the file.

const turns = [...agent.words.map((w) => ({ ...w, side: "agent" })), ...customer.words.map((w) => ({ ...w, side: "customer" }))].sort(
	(a, b) => a.start - b.start,
);

Median timestamp error sits between 69 and 86 ms, so the interleaving comes out in the right order even on a quick reply.

A single channel, with speakers: 2

When the recording arrives mixed, then it is a diarization case:

curl https://api.transcrevo.com/v1/transcripts \
  -H "Authorization: Bearer $TRANSCREVO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "uploadId": "0f3c2d11-8a4b-4c6e-b2d9-5e7f1a9c3b42",
    "speakers": 2,
    "language": "pt",
    "webhookUrl": "https://your-server.com/hooks/transcrevo"
  }'

Giving the count is safe here, because on a call you genuinely know: it is two. This is one of the few cases where the exact number is clearly better than letting true decide.

The labels stay anonymous: "A" and "B" are the first and second voice. Working out which one is the agent is usually trivial in practice, because the side that opens with the standard greeting is always the same one.

Where it fails, and why it matters more here

Support calls are made of short turns. "One moment", "mhm", "that's right", "can you confirm the account number". And short turns are exactly where speaker separation is worst:

TurnSpeaker accuracy
Longer than 15 seconds95 to 99%
Shorter than 5 seconds59 to 75%

On top of that, overlapping speech runs 14.8% DER against 4.6% overall. An irritated customer talking over the agent is the most common overlap case in a support archive.

The practical consequence:

  • Reading the call works. The problem description, the agent's answer, the handoff: those are long blocks, and attribution gets them right.
  • Counting behavior does not work. How many times the agent interrupted, how much airtime each side took, who talked more: those metrics live off the two-second turn, and that is where accuracy drops to two thirds. If your quality scoring depends on it, the fix is recording two channels, not picking a different algorithm.

The full numbers, with the methodology, are in accuracy and quality, and what separation does and does not do is in what speaker diarization is.

The glossary builds itself

Support is where a per-request glossary pays off most, because you already know what the call is about: the customer's name is on the ticket, the product name is in your catalog, the internal acronyms are fixed.

{
	"uploadId": "0f3c2d11-…",
	"glossary": { "bertual": "Bertuol", "essetá": "SST", "silver plan": "Silver Plan" }
}

Up to 100 entries per request, applied to that audio only. Proper nouns and acronyms are the most common error of any transcription, and in a support archive they are exactly what search will look for later. Details in how to fix wrong proper nouns.

What the API does not do

It returns text with timestamps and, if you ask, speaker labels. It does not return summaries, sentiment, topics, quality scores or banned-word alerts. That is not a roadmap omission: it is where the line was drawn. The text is the raw material, and the analysis on top of it belongs to your product.

There is also no live transcription. Processing is asynchronous: you create it, it runs, and you collect the result by polling or webhook. A supervisor following a call in real time is not a case this API serves.

Volume, retention and the bill

A mid-sized contact center produces a lot of audio, so three operational things matter more than in the one-off case.

Balance. With no credit, new transcripts are refused with 402 insufficient_balance and the queue stops. Monitor available (which already subtracts what in-flight transcripts hold, not total) or switch on auto top-up in the dashboard.

Ceilings. 120 transcript creations per minute. An overnight batch unblocking at once hits that; spread the burst across the Retry-After wait instead of retrying immediately.

Retention. DELETE /v1/transcripts/{id} erases the text, the words and the stored audio. The usage record stays, because that is what the invoice comes from, but the content is gone with no recovery. On a 90-day retention policy, it is a cron job.

The math: 1,000 hours of calls cost US$ 30 without speaker separation and US$ 50 with. Past 10,000 hours transcribed by the account, the rate drops on its own to US$ 0.02. The whole table is in pricing and limits.

All posts