Insights  /  Call transcription

Seven thousand calls. All of them searchable.

How every answered call becomes speaker-separated text on the customer record, transcribed on hardware you own.

The old way

A recording is a file. Somebody has to remember the call happened, find it by timestamp, and listen to it at one times speed to learn one fact. So nobody does. The recordings accumulate, the storage bill is real, and the archive gets opened twice a year: once for a dispute, once for a complaint.

The question a support manager actually has is never "play me this call." It is "what did we promise that customer in June", or "how often are we saying the thing we told the team to stop saying". Neither of those is answerable by an archive of audio, no matter how complete it is. And the moment you decide to solve it with a transcription API, you have handed every customer conversation you own to somebody else's servers.

What ISPCQ does

Every answered call is transcribed within minutes of hanging up, and it arrives already separated by speaker. That separation is not a guess. Algorithmic diarization is measurably useless on 8 kHz telephony — on a three-person call it invented thirty-five speakers — so the two parties are captured on physically separate legs instead. Which turns belong to which party is a property of the recording. Only the labelling of which one is the agent is inferred, and it is inferred from the greeting.

The transcript lands on the customer timeline beside the notes and tickets it relates to, and the recording plays in the browser next to it. Every turn carries its offset into the call. Click a line and the audio jumps to that second, with the spoken turn highlighting as it plays, so checking a specific sentence costs one click rather than a scrub bar and patience.

The whole corpus is full-text searchable, and searchable by account number, caller number and answering extension, which is what turns it from an archive into a record. Reading a transcript and playing the audio are separate permissions, held independently. The corpus view is deliberately unscoped and gated behind its own permission; the player on a customer’s own timeline is scoped to that customer, and refuses any recording that is not attached to a call of theirs.

Transcription runs on the operator’s own hardware. Of the transcripts on the reference deployment, all but three of seven thousand six hundred and thirty-one were produced locally, and the three that were not say so on the record: the engine is stored per transcript, so whether a given conversation left the building is a fact somebody can check rather than an assurance somebody gave.

The operational outcome

The question stops being expensive.

A customer disputes what they were told about a cancellation fee two months ago. Previously that is an afternoon: work out roughly when they called, pull the recordings around it, listen. Now it is a search for their account number, and the sentence is on screen with the audio cued to the second it was said.

The second-order effect is the one managers notice. When every call is text, coaching stops being anecdotal. You can see how often a particular explanation is being given, and whether the new process anyone agreed in a meeting actually reached the phones. That was always technically possible with a pile of recordings, and practically impossible with anyone’s afternoon.

Read related

More operational deep-dives.

One of twenty-three detailed articles on real ISP workflows. Each walks through the problem, what teams used to do, what ISPCQ does, and the operational outcome. The architecture is the same; the workflows differ.

Audience
Support managers
Runs
On your own hardware