What a coverage measurement actually measures

Last updated

A transcription system can tell you what it produced. Whether that is all of it is a different question, and it is not one the transcriber can answer about itself — asking the same process that may have dropped something to report on whether it dropped anything is circular.

Pincite Audio compares the transcript against a measurement taken from the source file instead: an independent witness rather than the provider's own accounting.

What is actually compared

For each channel that was transcribed, two quantities: how much time the transcript accounts for, taken as the merged span of its utterances, and how much voiced energy the original file contained, measured directly from the source.

The difference is time the source says carried voice and the transcript does not account for. When it exceeds a threshold, the recording is flagged.

Two details decide how to read a passing number. The reported ratio is driven by the worst single channel, not an average across channels — one bad side of a call is not diluted by a good one. And if the source-side measurement is unavailable for a channel, that channel is treated as unverifiable and the recording is flagged; it is never quietly treated as zero, because an absent measurement and a measurement of nothing are different facts.

What it cannot catch

This is the part worth stating plainly, because the check is easy to over-trust. There are three documented limits, and they run in different directions.

Displacement inside a single channel. The comparison is a scalar total: the sum of accounted time against the sum of voiced energy. That structure catches a whole missing part and a whole missing channel robustly. It cannot catch speech dropped in one region of a channel and words emitted in another region of the same channel — two errors of opposite sign cancel, and the total looks correct. Catching that would require carrying a map of where the voiced energy actually was, rather than how much of it there was. That has not been built.

An entry that covers time without carrying words. Accounted time is the merged span of the transcript's entries, and an entry's span runs from its first word to its last. So an entry that runs a long stretch on a single word accounts for that whole stretch as transcribed, and the ratio cannot see the hole underneath it. This one is no longer only a limit: a separate check reads the transcript's own shape and flags an entry that runs long while carrying almost no words. That finding is disclosed and does not hold the recording — the seconds are named on the review page, in the export and on the certificate, and the transcript is delivered. It is deliberately not a block: the check reads the shape of the transcript, so it cannot tell a pause that one long entry was drawn over from speech that was dropped, and the measurement that holds a recording for confirmed loss is the coverage comparison itself. What it does not do is change the coverage figure — those seconds still count as transcribed there, which is why the two numbers are reported side by side, are not addable, and why the coverage figure can read higher than the transcript supports.

The threshold is coarse on long, quiet recordings. The shortfall is measured against the whole length of the recording, not against the speech in it. On a long recording that contains little speech, a meaningful amount of missing speech can sit under the threshold and pass.

All three are documented limits, not tuning. A good coverage number does not rule any of them out — and the second check has limits of its own: it catches the nearly empty entry rather than the merely thin one, and it cannot tell speech that was lost from hold music or dead air the transcriber correctly declined to transcribe.

Why it flags recordings that turn out to be fine

The source-side measurement detects voiced energy. It does not know what kind of voice. An automated prompt, hold music, and touch tones all register.

So a call with several minutes of hold music can be flagged because the transcriber correctly declined to transcribe hold music. That contamination pushes the check toward flagging recordings that are fine — which is the more useful direction to be wrong in, though as the previous section says, it is not symmetrical protection.

What a flag does

Most flagged recordings are held for a person to look at, and no PDF renders while a recording is held. There is no retry. One finding does not hold a recording at all: a stretch of transcript that came back nearly empty is disclosed rather than blocking — the seconds are named on the review page, in the export and on the certificate, and the transcript is delivered. A channel that produced nothing at all fails outright and never renders.

That is the whole point of the mechanism. The failure this is built against is not a transcript with a problem — it is a transcript with a problem that looks finished and gets relied on. Withholding the finished-looking artifact is the response that cannot be overlooked, and when a person chooses to deliver it anyway, the certificate carries the finding rather than hiding it.

Passing is not the same as verified

A recording that clears every clause is not marked verified. The state it reaches is named ready, unverified, and the second word is doing real work: clearing the completeness check is not the same as anyone having confirmed the transcript.

Coverage answers one question — is time accounted for. Nothing in it speaks to whether a word is right, whether a speaker label belongs to the person named, or whether the recording is what it purports to be.

Questions

Does coverage tell me the transcript is accurate?

No, and it is important not to read it that way. Coverage compares how much time the transcript accounts for against how much voiced energy the source contained. A recording transcribed into fluent nonsense at the correct timestamps would pass. It is a completeness tripwire, not a quality measure.

Why was my recording flagged when it sounds fine?

The source-side measurement detects voiced energy, and it cannot distinguish speech from an automated prompt, hold music, or touch tones. If a recording contains several minutes of hold music that the transcriber correctly ignored, the measurement sees energy the transcript does not account for.

What happens when a recording is flagged?

Most flags hold the recording for a person to look at, and no PDF renders while it is held; there is no retry. One finding does not hold a recording at all: a stretch of transcript that came back nearly empty is disclosed rather than blocking — the seconds are named on the review page, in the export and on the certificate, and the transcript is delivered. A channel that produced nothing at all fails outright and never renders.

Is a perfect coverage ratio proof that nothing is missing?

On its own, no. A perfect ratio is the formula's identity element — it is also what you get when there is nothing to account for. That is why completeness is reported as the ratio together with the absence of gaps, rather than as the number alone, and why a separate clause exists to catch the empty case. A high ratio can also be counting seconds that a transcript entry covers without carrying words in it; a separate check flags that, and its seconds are shown beside the ratio rather than inside it.