What the chain-of-custody record covers

Last updated

Every transcript Pincite Audio produces carries a custody record built from three SHA-256 digests — and, where one of them does not exist for a particular recording, a line saying so rather than a blank. It is a short block, it is not the interesting part of the document, and it is the part most likely to be asked about, so it is worth knowing exactly what it holds and, more usefully, what it does not.

A digest is a fixed-length value computed from a run of bytes. Two identical runs of bytes always produce the same digest; two different runs, in any practical sense, do not. That is the entire mechanism. Everything below is a consequence of it.

The three digests

The format page describes the same record against the code that writes it, and is the place to check the details rather than this one.

What each one lets you check

Each digest answers exactly one question, by comparison, and it is a comparison you can run yourself with any SHA-256 tool.

Note what each of these is a statement about: bytes. A digest comparison tells you whether two runs of bytes are the same. It does not carry anything about where a file came from, who produced it, or what happened to it in between — those are questions about handling, and no arithmetic over the file's contents can answer them.

The two audio digests are not supposed to match

This is the one thing about the record that reliably surprises people, so it is stated on the record itself rather than left to be discovered.

The bytes submitted for transcription are re-encoded from the delivered file. They may also be resampled, laid out differently by channel, and cut into parts. So the source digest and the submitted-audio digests describe different runs of bytes by construction, and they will differ. The record carries a note saying so, along with what was actually done to the audio for that recording — the sample rate as delivered and as submitted, and how the channels were laid out.

Those rows record what was done. They are not a recipe that would reproduce the submitted bytes, and nothing here should be read as one — re-running the same steps is not something this product offers or claims.

Where the record has a hole, it says so

Source hashing has not always existed. A recording ingested before it did carries no source digest and never can, and the record prints a line saying that instead of leaving the field blank.

That is a deliberate choice and the reasoning generalises past this one field. A blank space where a value should be reads as "we looked and there was nothing there". The actual situation — "this measurement did not exist when this recording was processed" — is a different statement, and on a document that gets handed to someone else, the difference is the whole point. The same discipline runs through the rest of the record: a recording with no parts recorded says that, and a part whose bytes were not digested says that, rather than any of them rendering as empty.

What is not in it

The record hashes bytes. It stores no person, no transfer, and no handling event. The only time it carries is when the document itself was generated.

So it is narrower than the phrase "chain of custody" tends to suggest in conversation, and it is worth being precise about that rather than letting the phrase do work it cannot do. What the record gives you is a byte-level comparison at three points in one pipeline. Whatever account your firm keeps of who held a file and when is a separate thing, kept elsewhere, and this record neither contains it nor stands in for it.

Two further limits, in the same spirit. The record says nothing about how much of the audio was transcribed — that is a coverage measurement, a different check with its own limits. And it says nothing about who spoke: a digest is indifferent to content, and speaker labels are governed by their own rule about where a name is allowed to come from.

Questions

What are the three digests?

One over the produced text, computed on a canonical serialization of the transcript's utterances. One over the source file exactly as your firm delivered it. And one over the audio bytes submitted for transcription, listed part by part.

Should the source digest and the submitted-audio digests match?

No, and they are not supposed to. The bytes submitted are re-encoded from the delivered file, and may also be resampled, laid out differently by channel, and cut into parts. The record says so on its face, next to the digests, so that a difference is not read as a discrepancy.

What if a recording has no source digest?

Then it was ingested before source hashing existed, and the record states that rather than leaving the field empty. An empty field reads as 'we looked and found nothing', which is a different claim from 'this predates the measurement', and on a document like this the difference shows on its face.

Does the record say who handled the recording?

No. It is a record of bytes. No person, transfer, or handling event is stored in it; the only time it carries is when the document was generated. Whatever else you keep about a file's handling is separate from this and is not replaced by it.