G.711 vs G.729: what a recorded call actually contains

Last updated

Two codecs account for a great deal of recorded-call audio, and they sit at opposite ends of the same trade-off. Knowing which one you have tells you something useful about what is recoverable from the recording — and, more importantly, about what is not.

G.711: lightly compressed, larger files

G.711 is the older and simpler of the two: 8 kHz audio stored as eight logarithmically companded bits per sample — the μ-law and A-law variants you see named in file metadata — at 64 kbit/s. The companding is the whole codec; no further compression is applied.

The consequence is that a G.711 recording is comparatively faithful to what arrived at the recorder, and comparatively large. Where storage was not the binding constraint, this is often what a system wrote.

G.729: heavily compressed, much smaller files

G.729 targets the same narrowband speech at 8 kbit/s — an eighth of G.711. It gets there by encoding a compact description of the speech rather than a faithful waveform: filter and codebook parameters chosen, at encode time, to minimise the audible difference from the actual input.

The decoder is deterministic — it reconstructs the signal from exactly those parameters and invents nothing. What the encoder did not keep is simply absent, and shows up as coding noise rather than as fabricated content. That distinction matters if the recording is ever argued over: the file is a lossy record of the call, not a synthesis of it.

Both are narrowband, and that is the bigger limit

The compression difference is real, but it is not the main constraint. Both codecs work on narrowband telephone audio, which cuts off well below the range of ordinary speech recording.

So the ceiling on any telephony recording is set before the codec choice matters. G.711 preserves more of a signal that was already narrow. That is worth having, and it is not the same as having a good recording.

Why it matters when the call is evidence

Two practical implications.

The recording is what it is. Heavy compression is not something later processing undoes. Cleanup tools can make a file more pleasant to listen to; they cannot restore information that was never encoded, and a tool that appears to do so is generating plausible detail rather than recovering real detail — which is worth knowing before it ends up in an argument about what someone said.

A transcriber has less to work with under G.729 than under G.711. Not catastrophically less, and not in a way that shows up as an obvious failure. It shows up as more of the recording being genuinely ambiguous — which is exactly why a transcript that marks what it could not make out is more useful on this audio than on any other kind.

Questions

Which one is better?

For preserving a recording, G.711 — it applies far less compression, so more of the original signal survives. For moving a call across a constrained link, G.729, which was the point. Neither was designed for the job of being evidence later.

Can I tell which codec a recording used by listening to it?

Not reliably. Both are narrowband and both sound like a phone call. The encoding is a property of the file, readable from the file itself, rather than something to judge by ear.

Does a heavily compressed codec make transcription worse?

It gives a transcriber less to work with, and the parts that suffer first are the fine distinctions between similar consonants — exactly the ones that separate similar words. It does not make transcription impossible; it makes an honest account of what could not be made out more valuable.

Why would a system use a low-bitrate codec for something important?

Because the recording was a by-product. These systems were built to carry and store calls at volume and low cost, and the encoding choice served that. Nobody chose it with a later evidentiary use in mind.