Why a machine guess never gets a name

Last updated

Consumer transcription hands back SPEAKER_00, SPEAKER_01, and leaves you to work out who is who. The obvious improvement is to guess: match a voice against the rest of the file, pick the most likely person, print the name. Pincite Audio does not do that, and the reason is worth stating plainly, because it is the difference between a document you can work from and one you have to re-check line by line.

The label and its provenance travel together

Every speaker label in a Pincite Audio transcript carries a record of where it came from. A label an attorney supplied, a label a speaker gave for themselves, and a label software estimated are three different things, and they are stored as three different things rather than being flattened into one display string.

The exact set of provenance values is pinned on the format page, which is checked against the code by a test rather than maintained by hand. This guide is about the consequence of that design, not a second copy of the list.

A guess renders as a guess

When a label rests on a machine estimate or on a positional fallback, it renders as a designation — Speaker 2, Channel 1 — and never as a name. A name on a Pincite Audio transcript got there because a person at your firm typed it, and the provenance recording who did travels with the label wherever the transcript goes.

That distinction is decided in exactly one place in the codebase, so a label cannot pick up an attested appearance by travelling through a different rendering path.

A channel-derived label is a role, never a person

A two-channel recording — the common shape for a recorded call, where each side lands on its own track — supports a genuinely different kind of statement. The system can say which side of the call a passage came from, because the recording itself separates them. So it labels the role.

It stops there. A channel-derived label states a role and is structurally barred from holding a proper name: the constraint lives in the database schema, not in a formatting convention, so no future code path can quietly promote a role into an identification. Knowing a passage came from one side of a call is not knowing who was speaking, and the transcript does not blur the two.

What this costs, and why it is worth it

The cost is that a Pincite Audio transcript looks less finished than one that prints confident names everywhere. More labels carry a marker. Some say less than you would like them to.

The benefit is that when a label does carry a name, that name came from a person or a system that attested to it, and the transcript records which. A document that never overstates what it knows is one you can hand to someone else without a list of caveats attached.

Questions

Can I put the real name on a speaker myself?

Yes. An attorney-supplied identification is one of the provenances a label can carry, and it is the only kind that puts a person's name on a transcript. The rule is not that names are forbidden — it is that the system will not invent one.

What happens on a single-channel recording?

Nothing changes about the rule. A single-channel recording gives the system less to work with, so more labels end up as machine guesses, and a machine guess is a positional designation rather than a name. The label gets weaker; it does not get a name.

Why not just pick the most likely name?

Because the page would then read identically whether a person confirmed the identity or software estimated it. The whole point of carrying provenance is that those two cases must not look the same.