How to transcribe an interview
An interview here means one interviewer and one subject: a recorded witness statement, a client intake, a claims call, an investigation interview. That shape is worth naming at the start, because it decides most of what follows — and because a roomful of people sharing one microphone is a genuinely different problem, one this guide does not cover and this product does not attempt.
One thing this guide is not about: whether or when a conversation may be recorded at all. That is a legal question, and a transcription guide is the wrong place to answer it.
The recording decides more than the tool
By the time any transcription tool sees the file, the ceiling on the transcript is already set. Three capture choices move it more than anything a tool does afterwards:
- Microphone distance. A recorder on the table between two people beats one across the room by more than any later processing recovers. Distance costs consonants, and consonants are what distinguish similar words.
- Separate channels, if you can get them. An interview conducted by phone through a system that records each side on its own channel arrives with the speakers already separated — a property of the file, not a guess. One microphone in a room produces one mixed signal, and no processing honestly un-mixes it.
- One voice at a time. Overlapping speech is the hardest ordinary condition in speech recognition. An interviewer who lets answers finish is doing more for the transcript than any setting on the recorder.
Keep the original, submit the original
The file the recorder wrote is the best version of the recording that will ever exist. Every conversion, trim, "enhancement", or messaging-app share that re-encodes the attachment is a lossy pass over it. Renaming a file is harmless; re-encoding one is not, and the two are easy to confuse because both change what you see in a folder.
So: copy the original somewhere safe, and submit the original. If a tool cannot take the format as-is, prefer one that can to converting the evidence to suit the tool.
What to check in the transcript you get back
Whatever tool you use, three properties separate a transcript you can work from and one you have to re-listen to:
- Gaps are marked on the page. Stretches that produced no transcript should appear as explicit markers with a time range, so "what did it skip" is a question the document answers rather than one you settle by replaying the audio.
- Completeness is checked against the source. Asking the transcriber whether it got everything is circular. A useful check measures the source file independently and compares the result to it.
- Speaker labels say where they came from. In a one-on-one interview the attribution question sounds trivial and is not — a quote is only as good as the label on it. A label should record whether it rests on the recording's own structure, on a person's identification, or on a machine's guess, and those should not look alike on the page.
These are properties of the output. You can check whether a tool has them on a recording that does not matter before trusting it with one that does.
Where Pincite Audio draws its lines
Pincite Audio is built for exactly this shape of recording, and its behaviour on it follows the rules above: untranscribed spans render as explicit gaps, a coverage measurement runs against the source file, and speaker labels carry their provenance — the specifics are on the format page, checked against the code rather than maintained by hand.
Two limits, stated plainly. On a single-microphone interview the speaker tracks come from a machine guess about the voices rather than from the recording, so how many people spoke is not established and neither is who any track is — the reasoning is its own guide. And multi-party room audio is out of scope entirely, not handled badly but declined: a tool that quietly did a poor job of it would be worse than one that says no.
Questions
Can I transcribe an interview recorded on my phone?
A phone's voice-memo recording is ordinary audio and there is nothing exotic about the file it produces. Submit the file the phone wrote rather than a copy that has been converted, trimmed, or shared through an app that re-encodes attachments — each of those is a lossy step applied before anything has had a chance to use what was there.
Will the transcript tell the two speakers apart?
Only if the recording does. A call recorded by a system that keeps each side on its own channel arrives separated, and the transcript can use that. One microphone in a room produces one mixed signal, and every line is attributed to a single speaker you can then name — the transcript will not guess at dividing it.
What about a group interview or a panel?
That is a different problem, and not this tool's. Three or more people sharing one microphone is room audio, which Pincite Audio deliberately does not attempt to split into speakers. This guide is about the one-interviewer, one-subject shape.
Should I clean up the audio before submitting it?
No. Noise reduction, normalization, and format conversion are all lossy passes over the recording, and they run before the transcription step that could have used the detail they discard. Keep the original exactly as the recorder wrote it and submit that.