Transcribe an AMR file to text
Transcribing an .amr file is not technically hard. Getting a transcript you can rely on from one is harder, because the format itself has already thrown away much of what a transcriber would have used.
Submit the original
The most common preparation step is the one to skip. Converting the file to MP3 or WAV before submitting it feels like helping, and it is not: it is a second lossy pass over audio that had very little detail to begin with, applied before anything has had a chance to use what was there.
Where you can, hand over the file you were given and let a single normalization step run from it. One conversion from the original is strictly better than two, and the file's unfamiliar extension is rarely the obstacle people expect — Pincite Audio, for one, applies no extension allowlist to a direct upload and inspects the actual contents instead. (Its Clio browser is the exception — that surface only offers documents whose extension or content type marks them as audio or video.)
Expect it to be difficult, and expect that to show
AMR narrowband is 8 kHz audio compressed hard. Frequencies above roughly 4 kHz are absent, consonants that distinguish similar words are exactly what suffers first, and any background noise present at capture was encoded by a codec that models voices rather than rooms.
A transcriber will find some of it difficult. The question worth asking of any tool you use is not whether it finds difficult audio difficult — they all do — but whether it tells you where.
What a transcript of hard audio should give you
Three things, none of which are about accuracy claims:
- Marked gaps. Stretches that produced no transcript should appear on the page as explicit markers with a timestamp range, so you can go back to those seconds specifically rather than re-listening to everything.
- A completeness check against the source. Asking the transcriber whether it transcribed everything is circular. A useful check measures the source file independently and compares.
- Speaker labels that carry their provenance. Difficult audio produces uncertain attribution, and a label that records how it was arrived at is worth more than a confident-looking name.
Those are properties of the output, and you can check whether a tool has them before trusting it with a case.
If a tool rejects the file
AMR files trip software in ways ordinary formats do not, often because a tool is reading the container's own metadata rather than measuring the audio. So a rejection is not evidence that your file is damaged.
It is also the better failure. A tool that stops on something it cannot handle is behaving well; one that quietly transcribes the part it managed to decode hands you a document that looks complete and is not.
Questions
Do I need to convert the AMR file first?
No, and it is better not to. Converting is a lossy step on audio that has very little detail to spare, and it happens before anything has had a chance to use what was there. Submit the original and let one normalization step run from it.
Will an AMR file transcribe as well as a normal recording?
Expect it to be harder. The audio was sampled at 8 kHz and compressed aggressively, so information a transcriber would use is not in the file. What matters is not pretending otherwise — a transcript of difficult audio should show you where it struggled rather than smoothing over it.
The file is one long recording of two people. Can the sides be separated?
It depends on what the recording system wrote. If each side of the call was captured on its own channel, that separation is present in the file and can be used. If both sides were mixed into one channel, the information is gone and no processing recovers it.
How do I know the transcript covers the whole file?
By the transcript saying so. Spans that produced no transcript should appear on the page as explicit markers, and a completeness check should compare the result against a measurement of the source rather than against the transcriber's own report.