What the output looks like
The strings a transcript actually renders, and the rules behind them. Every value on this page is one you can check against the file you get back.
A speaker label carries where it came from
Every speaker in a transcript has a provenance, and the provenance decides what the label may say. There are six:
attorney_identified— a person at the firm named this speaker.self_identified— the speaker gave their own name on the recording.system_stamped— the name came from the vendor's own metadata.channel_derived— the label comes from which channel the audio was on.diarization_only— the system separated a voice but has nothing to call it.multiple_or_unknown— the span could not be attributed to one speaker.
The last three carry no name at all, and this is a database constraint rather than a convention: a channel_derived speaker stores a role and is structurally incapable of holding a proper name, and a diarization_only or multiple_or_unknown speaker has no name field to fill. The transcript renders those as Channel 1 or Speaker 3 — a designation, never an identity. What kind of designation it is travels with the data rather than in the printed string: the provenance above is returned on every speaker over the connector, alongside a verified flag, and a name field that is simply absent unless a person supplied one.
What that buys you: a name in the output means a person put it there or the recording said it. It cannot mean the software guessed.
A gap is a thing on the page
When a stretch of audio produces no transcript, that stretch is recorded with a cause. There are five:
no_transcription_returned— the recognizer returned nothing for that span.part_failed— a piece of a long recording failed to process.channel_not_transcribed— one side of a two-channel recording was not transcribed.redacted— the span was deliberately removed.detected_silence— there was nothing there to transcribe.
Only the last one is benign, and it is the only one that does not render. Every other cause produces this line, in the same place in the transcript as the audio it stands for:
[NO TRANSCRIPTION RETURNED HH:MM:SS–HH:MM:SS]
[NO TRANSCRIPTION RETURNED ON CHANNEL 2 HH:MM:SS–HH:MM:SS]
Coverage is measured against the source file
After a recording is transcribed, the system compares what it produced against the duration of the file you sent. That measurement is stored with the recording, and it is what the gap list is checked against.
If the measurement comes back below threshold, the recording is marked as needing attention and no PDF is produced for it. A transcript the system cannot account for does not get exported as a finished document.
Honest limit, stated because it matters: the measurement is a total. It catches a part or a channel that went missing. It does not catch audio that was transcribed in the wrong place within a channel.
One shape of loss the total is structurally unable to see is checked separately, by reading the transcript rather than the clock. Transcribed time is counted as the span each transcript entry covers, and an entry's span is drawn from its first word to its last — so an entry that runs for a long stretch while carrying a single word accounts for that whole stretch as transcribed, and the total reads clean over a hole. An entry that runs long while carrying almost no words is therefore a finding in its own right — one that is disclosed rather than blocking. The seconds are named on the review page, in the export and on the certificate, and the transcript is delivered. It does not hold the recording, and that is deliberate: the check reads the shape of the transcript, so it cannot tell a pause that one long entry was drawn over from speech that was dropped. What holds a recording is the coverage measurement, which compares detected speech against transcribed time.
The seconds that check found are reported next to the coverage figure, never subtracted from it. The coverage figure still counts them as transcribed, which is exactly why it can read higher than the transcript supports, and every surface that shows both says so rather than leaving the two numbers to be added.
The limits of that second check, in the same spirit as the first: it catches an entry that is nearly empty, not one that is merely thin, so a stretch that came back with a few words in it passes. It cannot tell speech that was lost from hold music, a recorded announcement, or dead air that the transcriber correctly declined to transcribe. And it is a tripwire on the shape of the transcript — it reports seconds of transcript, not a measurement of what was said in them.
The chain-of-custody record
Each transcript carries a SHA-256 record with three parts: a digest of the produced text, computed over a canonical serialization of the utterances; a digest of the source file exactly as your firm delivered it; and a digest of the audio bytes submitted for transcription, listed part by part.
The limit, stated because this page would be worthless without it: a recording ingested before source hashing existed carries no source digest, and the certificate says so on its face rather than leaving the field blank for you to misread as a match.
A second limit, for anyone who compares the two audio digests: they are not expected to match. The bytes submitted for transcription are re-encoded from the delivered file, and may be resampled, laid out differently by channel and cut into parts, so their digests differ from the source file's by construction. The record states what was done to the audio beside those digests — the sample rate and the channels, as delivered and as submitted — and states it as a record of what happened, not as a procedure that would reproduce those bytes.
Transcripts carry a reference to each line, in one of two styles: a timestamp, or a page and line number. Which one a recording uses is fixed when it is created and does not change underneath a citation you have already written down.
- Recording
- Source file
- Audio, per part
- Transcript digest
- Rendered