Skip to content
streamneo.
Streaming Settings14 min read

How to Normalise Audio Levels Across Videos in an FFmpeg YouTube Playlist

Use FFmpeg loudnorm per video, then fit mixed frame sizes and aspect ratios into a consistent YouTube playlist canvas.

sn.
StreamNeoPublished 4 October 2026
Worth sharing?

When viewers move between separate videos in a YouTube playlist, use FFmpeg's loudnorm filter on each file independently. A measured two-pass workflow gives you more repeatable results than measuring unrelated videos as one combined programme.

Audio is only part of the job. If the playlist also contains mixed frame sizes or aspect ratios, decide on an output canvas, fit each video inside it, and pad the remaining space consistently before you begin the long render.

Frame size, DAR, and SAR explained

A video file has more than one geometry value. Confusing them is a common reason a batch looks stretched even when every output file reports the same width and height.

Frame dimensions are the stored pixel dimensions, such as 1,920 × 1,080 or 1,280 × 720. They describe how many samples exist across and down each frame. They do not, by themselves, tell you how wide the finished image should appear on a display.

Display aspect ratio, or DAR, is the shape the viewer is meant to see. A frame with a 16:9 DAR appears widescreen. A frame with a 4:3 DAR appears closer to a traditional television shape. DAR can be derived from the stored dimensions, but it can also be affected by the pixel shape.

Sample aspect ratio, or SAR, describes the shape of each stored pixel. Square pixels have a SAR of 1:1. Non-square pixels were used by some older formats and workflows, where the stored width and height do not directly match the display shape.

The relationship is broadly:

DAR = frame width / frame height × SAR

That means two files with identical frame dimensions can display differently if their SAR metadata differs. Conversely, two files with different stored dimensions can display with the same DAR.

For a long-running channel, inspect all three before choosing a filter chain. ffprobe can help you inspect width, height, sample aspect ratio and display aspect ratio without opening the files in an editor. You can also inspect the output from FFmpeg's video filters during a short test render.

Do not treat 16:9 as a universal YouTube rule. It is a useful landscape canvas for many channels, but it is a production choice. A devotional channel may choose a 16:9 background for wide artwork, a local news loop may use a different layout, and a portrait-led source may be better presented with deliberate side panels than by being cropped.

Why mixed inputs may look stretched

Stretching usually occurs when a workflow changes the stored frame dimensions without accounting for DAR or SAR. For example, forcing a 4:3 source directly into a 16:9 rectangle makes people appear wider unless the source is cropped or fitted with empty space.

A second problem is inherited metadata. An older file may carry a non-square SAR. If you scale its stored pixels as though they were square, the result can have the wrong display shape. The video may look acceptable in one player and wrong in another because applications do not always interpret metadata in exactly the same way.

Cropping is another source of trouble. A command that scales every input until it fills the output canvas may remove the top and bottom of a speaker, the edge of a news ticker, or text placed near the frame boundary. That can be acceptable for a carefully designed montage, but it is a poor default for mixed material.

There is also a difference between geometry and loudness. Normalising audio does not correct stretched video, and padding video does not make separate recordings sound equally loud. Treat the file as two related but separate problems: a video layout pass and an audio measurement pass.

If you are building a non-stop channel around existing files, first decide whether the source videos are meant to be edited, fitted or simply passed through. The YouTube encoding requirements for pre-recorded content are a useful checklist for the delivery stage, but they do not decide your editorial canvas for you.

Choose an output canvas intentionally

Choose the output canvas from the viewing experience you want, not from whichever input happens to be first in the playlist. Write down the output width, height, frame rate and background treatment before processing the batch.

A 16:9 landscape canvas is often practical for a channel viewed on televisions, laptops and ordinary landscape displays. That does not make it compulsory. You might choose a square or portrait arrangement for a channel designed around mobile artwork, or a custom landscape composition with a branded background for portrait sources.

The important point is consistency. If the first file produces a wide black border and the next uses a blurred background, viewers will notice the change even if the audio level is stable. A single background colour, image, or designed panel system makes the playlist feel like one channel rather than unrelated files.

Consider these questions:

Decision What to check Trade-off
Canvas shape Where will most viewers watch, and what shape suits the programme? A wider canvas can leave larger unused areas around portrait or square sources.
Output size Does the chosen size suit the delivery and the source detail? Upscaling a small source does not create new detail.
Fit or crop Must the whole source remain visible? Fitting preserves content but creates unused space; cropping fills the frame but can remove content.
Background Is one colour, an image, or a designed panel appropriate? A plain background is simple; a branded design takes more preparation.
Frame rate Do the sources have different motion characteristics? Converting frame rates can affect motion and render time.

A 16:9 canvas can therefore be a sensible house style, not a supposed platform command. Document the decision so that future videos are prepared the same way.

Fit each video inside the canvas

For mixed inputs where the full picture must remain visible, use a scale-and-pad approach. The basic logic is:

  1. Reset or account for the input's sample aspect ratio.
  2. Scale the image so it fits within the chosen canvas without changing its shape.
  3. Pad the unused area to reach the exact output dimensions.
  4. Set the output sample aspect ratio deliberately.

A common FFmpeg video-filter shape is:

-vf "scale=w=1920:h=1080:force_original_aspect_ratio=decrease,pad=1920:1080:(ow-iw)/2:(oh-ih)/2,setsar=1"

This is an illustrative filter shape, not a claim that it has been verified on your FFmpeg build or with your particular files. Replace the dimensions with your chosen canvas. force_original_aspect_ratio=decrease keeps the fitted image within the canvas. The pad stage adds the missing width or height and centres the image. setsar=1 requests square output pixels.

You may prefer a background colour other than the default black, for example:

-vf "scale=w=1920:h=1080:force_original_aspect_ratio=decrease,pad=1920:1080:(ow-iw)/2:(oh-ih)/2:color=black,setsar=1"

If the source carries unusual aspect-ratio metadata, test whether you should use setsar=1 before scaling or after the final pad. The correct placement depends on how the input is being interpreted and on the intended display shape. Do not blindly add filters until the result looks right; inspect a representative frame and compare it with the original.

If you deliberately want a crop-to-fill design, use a crop strategy instead. That is suitable for material where the edges are disposable, but not for subtitles, faces or information graphics placed close to the boundaries. Fit-and-pad is the safer starting point for a mixed playlist.

Pad unused space consistently

Padding is not merely empty space. It becomes part of the channel's visual identity and determines whether portrait or 4:3 material looks intentional.

A plain black pad is easy to render and keeps attention on the source. A solid colour can match a logo or devotional artwork. A background image can make the layout more polished, but it may compete with the main video or introduce its own scaling and licensing questions.

For a station that contains repeated artwork, consider designing one background template at the final canvas size. Place the fitted video above it and keep the same margins, logo position and text-safe areas across the playlist. This is often more convincing than allowing every source to choose its own borders.

Avoid adding an aggressive background animation merely to hide unused space. A subtle, consistent layout is easier to monitor overnight and less likely to distract from speech, music or a news ticker. If you use text, check it at the smallest display size you expect viewers to use.

The padding step also gives you a useful technical boundary. Once every output has identical frame dimensions and square-pixel output, later playlist and streaming stages have fewer geometry variations to handle. That does not repair a damaged source, but it makes the delivery set more predictable.

Check sample aspect ratio and output geometry

After rendering a test file, inspect the output rather than trusting the command line. Check the stored width and height, SAR, DAR, codec, frame rate and audio streams. A file that looks correct in one desktop player can still carry metadata that causes a different interpretation elsewhere.

For audio, check which stream was selected. The illustrative mapping below chooses the first video and first audio stream:

-map 0:v? -map 0:a:0

The question mark on the video mapping allows the command to continue when a file has no video, but it does not solve every stream-selection problem. A playlist containing files with multiple language tracks, commentary tracks or unusual layouts needs an explicit mapping policy. Decide whether you want the first programme audio, a named language, or a particular channel layout.

When you filter audio, it must be encoded again. Copying the video stream with -c:v copy can save video encoding time when the geometry is unchanged, but it cannot copy a video unchanged if you have applied scale and pad filters. In a geometry-normalisation pass, encode the video with a codec and settings suitable for your delivery workflow, and encode the filtered audio into a compatible container.

Do not copy every stream blindly. Subtitles, data streams and attachments may not be accepted by the selected output container or may not belong in a public live playlist. Keep a source copy, create a clearly named processed copy, and retain the command and target values used for each file.

Normalise loudness one file at a time

FFmpeg documents loudnorm as EBU R128 loudness normalisation with dynamic and linear modes, and with single-pass and double-pass operation. The filter exposes targets for integrated loudness, loudness range and maximum true peak. These describe different properties and should be selected as a coordinated project decision.

Integrated loudness estimates the overall perceived level. Loudness range describes how much the programme's level varies. Maximum true peak limits peaks, including inter-sample peaks. Raising one value does not substitute for controlling the others.

The filter documentation lists defaults of I=-24.0, LRA=7.0 and TP=-2.0. Those are filter defaults, not a current YouTube creator upload requirement. YouTube's playback controls can also affect what a viewer hears. Its Stable volume help page explains that the viewer-side feature can balance quiet and loud parts, is not available for every video, and is disabled for YouTube Music and official music videos. Check the current YouTube documentation before treating any playback behaviour as fixed.

For offline files, a two-pass workflow is usually the more controlled choice. Run the first pass separately for each video and save the measurements:

ffmpeg -i input.mp4 -af "loudnorm=I=-16:LRA=11:TP=-1.5:print_format=json" -f null -

The -16, 11 and -1.5 values are illustrative project settings only. They are not presented as YouTube requirements. Choose targets with regard to the material and the listening brief. A speech-heavy local news loop and a music-led devotional playlist may need different editorial decisions, even if they ultimately share a channel.

For the second pass, supply the measured integrated loudness, loudness range, true peak, threshold and offset reported for that same input:

ffmpeg -i input.mp4 -map 0:v? -map 0:a:0 \
  -af "loudnorm=I=-16:LRA=11:TP=-1.5:measured_I=...:measured_LRA=...:measured_TP=...:measured_thresh=...:offset=...:linear=true:print_format=summary" \
  -c:v copy output.mp4

This command is an outline, not a verified command for a particular FFmpeg build or stream. If you have also applied video scale and pad filters, -c:v copy is not suitable for that filtered video stream. Build the audio and video stages together, or process geometry and audio in a command that encodes the affected streams.

Repeat the measurement and processing independently for every file. Do not concatenate unrelated playlist entries and measure them as if they were one programme. The first-pass values belong to the file that produced them.

One-pass, two-pass, linear and dynamic choices

One-pass processing is simpler and is supported for files and livestreams. It can suit a live operation or a quick initial batch, but it does not use the explicit measurements from a separate analysis pass.

Two-pass processing takes longer to organise, yet it gives you file-specific measurements and a repeatable record of the selected targets. That makes it useful when you are preparing a playlist overnight or replacing a weak file without changing the treatment of every other entry.

Linear mode aims to preserve the source dynamics through a constant gain when its conditions can be met. It requires measured inputs. FFmpeg documents that it can fall back to dynamic mode if the target loudness range is below the source range or if the adjustment would exceed the true-peak limit.

Dynamic mode can alter level over time to meet the target where a simple gain change is insufficient. It is not the same as turning a file up by one fixed amount. Listen to quiet passages, speech, music with deliberate contrast and transitions between playlist entries. FFmpeg also documents true-peak detection in dynamic mode with upsampling to 192 kHz; treat that as filter behaviour, not as a guarantee about what your complete delivery chain will do.

Your decision should balance workflow time, preservation of dynamics, true-peak headroom and whether fallback to dynamic processing is acceptable. The reported normalisation type matters. If you request linear mode but the filter reports dynamic processing, investigate the reason instead of assuming the file received the treatment you intended.

For the complete filter reference, consult the official FFmpeg loudnorm documentation. Use the documentation for the FFmpeg version you actually run, because command syntax and supported options should be checked against the installed build.

Test the playlist before it goes live

Do not test only the first file. Choose a short section from a quiet source, a loud source, a file with speech, and a file with music or sustained ambience. Include the transition between at least two playlist entries, because that is where inconsistent levels are most obvious.

Measure the rendered outputs with FFmpeg's EBU R128 capability or inspect the loudnorm summary. Confirm that the output integrated loudness and true peak are within the tolerances you chose. Also check the audio stream selection, channel layout and sample rate. A successful encode only proves that a file was written; it does not prove that the right programme audio was processed.

Then listen without changing the volume between clips. Pay attention to spoken introductions, bells, percussion, long quiet sections and any material that was intentionally dynamic. Measurement can bring files closer together, but it cannot decide whether a sudden change is editorially appropriate.

Watch the video at the same time. Check that faces are not stretched, subtitles are not cropped, the pad is symmetrical where intended, and any logo or ticker remains inside the visible area. Test on more than one player or display if possible, since aspect-ratio metadata and viewer-side audio controls can affect the result.

Keep a simple processing log with the source filename, output filename, canvas choice, filter chain, target values, measured values and reported normalisation type. When a viewer reports that one entry is quiet or distorted, the log tells you whether to revisit the source, the measurement, the filter settings or the playback stage.

If the final files will feed an always-on stream, remove avoidable local points of failure before scheduling them. A desktop loop can work, but it depends on the computer, its power settings and the streaming application remaining healthy. For a workflow where the machine being switched off is the main concern, StreamNeo removes the need to keep your own computer running after you upload the prepared file and provide the YouTube stream key; it is still your responsibility to check the rendered media and channel settings before using it.

For a computer-based loop, the OBS encoder settings for a non-stop YouTube loop stream can help you separate encoding decisions from the media-preparation work. If you are preparing a larger devotional playlist, the guide to streaming a church choir concert playlist around the clock covers the operational side of keeping multiple prepared files in rotation.

Before committing, compare the operating options on the pricing page. When the file and channel are ready, start free — 24-hour trial, no card.

FAQ

Should every YouTube playlist use a 16:9 output?

No. 16:9 is a common landscape canvas, but it is a choice based on your programme, artwork and expected viewing context. Use fit-and-pad when the complete source must remain visible, or choose a deliberate crop or custom layout when the edges are not important.

Is one LUFS target required for YouTube uploads?

The official material used here does not establish one universal creator-side LUFS target. Choose integrated loudness, loudness range and true-peak values for the project, then remember that YouTube playback features and viewer settings can affect the final listening experience.

Why should I measure each video separately?

Each file has its own loudness, dynamics and true-peak behaviour. In a two-pass workflow, the measurements supplied to the second pass must belong to that same input, so unrelated videos should not be treated as one programme.

What should I do if linear mode changes to dynamic mode?

Check the filter summary and the measured values. FFmpeg documents fallback when the target loudness range or true-peak constraint cannot be met with linear processing; listen to the result and decide whether the dynamic treatment is acceptable for that particular programme.

YOU’VE REACHED THE END

Keep the ideas coming.

More guides, useful tools and a little help for your next broadcast.

Back to the journal ↗
YOUR NEXT READ

A little more to explore.

More Streaming Settings guides ↗ · All topics ↗