Skip to content
All common fixes

Common fix

Make an uneven recording listenable

A recording at a consistent, standard loudness with the dead air removed.

By Updated How we write these

Recordings come out uneven for ordinary reasons: people sit at different distances from the microphone, gain was set for the wrong voice, or the room was quiet and the level was left low. The fix is measurement rather than dragging a volume slider until it sounds right.

Do it in the right order — cut, then level — and the loudness measurement is taken on the audio you are actually keeping.

What to do

  1. Step 1

    Cut the top and tail

    Trim the setup chatter off the front and the fumbling off the end. On MP3 and M4A this happens on a frame boundary, so nothing is re-encoded and no quality is lost.

    Trim the recording
  2. Step 2

    Take out the dead air

    Long silences make a recording feel slower than it is and inflate the file. Removing them is the single biggest improvement to most interview recordings.

    Remove silence
  3. Step 3

    Normalize the loudness

    Normalizing measures perceived loudness across the whole file in LUFS and applies one gain change to land it on a standard target — −16 LUFS for a stereo podcast. It does not squash the dynamics.

    Normalize the loudness
  4. Step 4

    Drop it to mono if it is speech

    A single voice has no stereo information worth keeping. Mono halves the file size for the same quality, and it is what most podcast advice means by "smaller file".

    Convert to mono

Worth knowing

  • If one speaker is quiet and the other is not, normalizing the mixed file will not separate them. That needs the individual tracks, or a manual volume change on the quiet section.
  • Work from the most original file you have. Levelling an MP3 export and re-encoding it stacks two rounds of loss.

Questions

What loudness should a voice recording be?

About −16 LUFS for stereo or −19 LUFS for mono, with true peaks no higher than −1 dBTP. What LUFS is explains why.

Is normalizing the same as compressing?

No. Normalizing applies one gain change based on a measurement; compression reduces the distance between loud and quiet parts. Normalizing leaves the dynamics alone.

Will removing silence cut people off mid-sentence?

Not with sensible settings — it looks for stretches of genuine silence, not short pauses. Check the result before you keep it.

The reasoning behind it