Transcribe an AMR file to text

Last updated

Transcribing an .amr file is not technically hard. Getting a transcript you can rely on from one is harder, because the format itself has already thrown away much of what a transcriber would have used.

Submit the original

The most common preparation step is the one to skip. Converting the file to MP3 or WAV before submitting it feels like helping, and it is not: it is a second lossy pass over audio that had very little detail to begin with, applied before anything has had a chance to use what was there.

Where you can, hand over the file you were given and let a single normalization step run from it. One conversion from the original is strictly better than two, and the file's unfamiliar extension is rarely the obstacle people expect — Pincite Audio, for one, applies no extension allowlist to a direct upload and inspects the actual contents instead. (Its Clio browser is the exception — that surface only offers documents whose extension or content type marks them as audio or video.)

Expect it to be difficult, and expect that to show

AMR narrowband is 8 kHz audio compressed hard. Frequencies above roughly 4 kHz are absent, consonants that distinguish similar words are exactly what suffers first, and any background noise present at capture was encoded by a codec that models voices rather than rooms.

A transcriber will find some of it difficult. The question worth asking of any tool you use is not whether it finds difficult audio difficult — they all do — but whether it tells you where.

What a transcript of hard audio should give you

Three things, none of which are about accuracy claims:

Those are properties of the output, and you can check whether a tool has them before trusting it with a case.

If a tool rejects the file

AMR files trip software in ways ordinary formats do not, often because a tool is reading the container's own metadata rather than measuring the audio. So a rejection is not evidence that your file is damaged.

It is also the better failure. A tool that stops on something it cannot handle is behaving well; one that quietly transcribes the part it managed to decode hands you a document that looks complete and is not.

Questions

Do I need to convert the AMR file first?

No, and it is better not to. Converting is a lossy step on audio that has very little detail to spare, and it happens before anything has had a chance to use what was there. Submit the original and let one normalization step run from it.

Will an AMR file transcribe as well as a normal recording?

Expect it to be harder. The audio was sampled at 8 kHz and compressed aggressively, so information a transcriber would use is not in the file. What matters is not pretending otherwise — a transcript of difficult audio should show you where it struggled rather than smoothing over it.

The file is one long recording of two people. Can the sides be separated?

It depends on what the recording system wrote. If each side of the call was captured on its own channel, that separation is present in the file and can be used. If both sides were mixed into one channel, the information is gone and no processing recovers it.

How do I know the transcript covers the whole file?

By the transcript saying so. Spans that produced no transcript should appear on the page as explicit markers, and a completeness check should compare the result against a measurement of the source rather than against the transcriber's own report.