If a video has no audio, FFmpeg can generate a silent audio stream with anullsrc and send it alongside the video. You must map both streams explicitly, because a video-only input does not provide an audio track for YouTube to receive.
For a playlist, do not treat silence as the whole audio solution. Add the track first, then measure and normalise each video as its own programme with loudnorm when it contains audio. That gives you a repeatable workflow without claiming that YouTube has one universal loudness target for every creator.
Why playlist loudness can vary
A playlist can contain files made at different times, on different phones, with different microphones, music libraries, or editing software. One bhajan recording may have a quiet opening and a strong chorus. A local news clip may be mastered more loudly than an older devotional video. A lofi file may be intentionally soft so that rain or room tone remains comfortable over a long period.
The result is a stream that changes level when the next file starts. The change may be obvious even when every file technically contains an audio stream. Conversely, a video with no audio can cause a different problem: the output may have video but no audio stream at all.
These are separate jobs. anullsrc creates an audio track where none exists. loudnorm measures and adjusts an existing audio programme. Adding silence does not make other files equally loud, and normalising a playlist does not create audio for a video-only file.
For a channel that runs overnight, both details matter. A missing audio stream can trigger an ingest warning, while large level changes can make a listener turn the stream down and then miss the next quiet section. If the stream is built from scheduled files, the FFmpeg playlist scheduling guide is useful for the file-ordering part, but it does not replace audio checks.
Add a silent track with anullsrc
The basic pattern is to use the video file as input zero and create a second input with FFmpeg's anullsrc source. Then map the first input's video and the second input's audio:
ffmpeg -i input.mp4 \\
-f lavfi -i anullsrc=r=44100:cl=stereo \\
-map 0:v:0 -map 1:a:0 \\
-c:v copy -c:a aac -shortest output.mp4
Here, r=44100 requests a 44.1 kHz sample rate and cl=stereo requests a stereo channel layout. The first -map selects the first video stream from the file. The second selects the first audio stream from the generated lavfi input. The video is copied, while the generated audio is encoded as AAC.
The explicit maps are important. Without them, FFmpeg's automatic stream selection may not produce the structure you intended, particularly when the input has additional streams. The FFmpeg documentation describes manual stream selection with -map, and its filter documentation describes anullsrc as a null audio source with configurable rate and channel layout.
-shortest is appropriate in this finite-file example because the video ends while the silent source can continue. It tells FFmpeg to finish when the shortest output stream ends. Do not copy this command unchanged into a long-running live command. A live output needs its own protocol, pacing, destination, and stream-key handling, and an unbounded generated source must be considered differently.
If your original file already has audio but you want to replace it with silence, use the same maps: -map 0:v:0 -map 1:a:0. Do not map 0:a:0 as well. YouTube's live error guidance warns about multiple audio streams, so the intended output should contain one audio stream rather than the original track and the generated track together.
Normalise each video as its own programme
Treat every file as an independent programme before it enters the playlist. This is more reliable than measuring a whole folder as though it were one continuous recording, because the listener encounters each file as a separate piece of content.
A programme might be one sermon, one music video, one news bulletin, or one long ambience recording. Measure that file, choose the adjustment for that file, and write a processed copy. Keep the original untouched so you can change your decisions later.
For files that already contain audio, loudnorm can measure integrated loudness, loudness range, and true peak. It can also apply a chosen normalisation target. The target is a production choice. It is not a universal YouTube creator requirement established by the official pages used here, and no target guarantees that every listener will perceive every programme as equally loud.
For a video-only file, there is nothing to normalise. Add the silent stream and check its structure instead. If the file is meant to sit beside spoken or musical programmes, decide whether complete silence is actually the right editorial choice. A black-screen prayer loop, for example, may intentionally have no sound, while a station ident may need a short, designed bed.
This per-file approach also gives you a practical audit trail. Keep the measured values beside the filename, along with the selected target and whether the file was changed. If a particular transition sounds wrong later, you can revisit one file instead of rebuilding the entire playlist.
Measure each file in a first pass
For tighter control, use loudnorm in two passes. The first pass analyses one file and prints measurements. The second pass uses those measurements to apply the chosen correction. This is an offline workflow, so it is best done before the file is placed in a continuous FFmpeg stream.
A first-pass command can look like this:
ffmpeg -i input.mp4 -af loudnorm=I=-16:LRA=11:TP=-1.5:print_format=json -f null -
The values shown are examples of selected settings, not a rule imposed by YouTube. You may choose different targets for the type of material you publish and the amount of variation you are prepared to accept. The useful point is consistency: record the same selected targets for the files you want to compare, then inspect the measurements for each file independently.
The JSON output includes values needed by a measured second pass, such as the input integrated loudness, input loudness range, input true peak, and the offset reported by the filter. Save the complete result for that file. Do not copy measurements from one video into another, even when both files came from the same channel.
A two-pass workflow takes more preparation than adding a filter during the live run. It can be worthwhile when the playlist contains spoken word, music, and long ambience files together, or when you need predictable output after a power cut or overnight restart. For a small set of similar files, a simpler one-pass check may be enough, but listen to the result rather than assuming the command solved every transition.
If FFmpeg is built differently on your machine, confirm the available filter options with the documentation and your installed version. The command is a starting pattern, not a guarantee that every build prints output in exactly the same place or format.
Process the file with its own measurements
The second pass substitutes the first pass's measurements into loudnorm. A typical pattern is:
ffmpeg -i input.mp4 \\
-af "loudnorm=I=-16:LRA=11:TP=-1.5:measured_I=-18.2:measured_LRA=7.4:measured_TP=-2.1:measured_thresh=-28.6:offset=2.2:linear=true:print_format=summary" \\
-c:v copy -c:a aac output-normalised.mp4
The measured values in this example are placeholders showing where the first pass results belong. They are not measurements for your file. Replace them with the values returned for that exact input. If the first pass reports that linear mode cannot be used for the material, follow the filter's output and use the mode supported by that measurement rather than forcing a setting blindly.
Copying the video avoids another video encode when the container and downstream workflow permit it. The audio is encoded because the filter has produced a new audio stream. Check the output after processing: an audio filter cannot be applied to a file that has no audio input. For a silent video, use the generated anullsrc input instead, or combine the silence creation and any required processing in a command designed for that case.
The purpose of the second pass is not to make every file identical in character. A quiet meditation track and a news bulletin may have different dynamics by design. Normalisation can bring their overall levels into a selected working range, but it cannot remove the creative contrast within a programme without changing the material.
If you process a large directory, make the filename and measurement pairing unambiguous. A shell script that accidentally applies one file's values to the next file can produce results that are worse than leaving the original audio alone. Test a few representative files manually before automating the whole library.
Use consistent selected targets across files
Choose targets before processing, then apply the same selected targets to the files you want to compare. This does not mean every file will end with the same perceived loudness. Loudness range, spectral balance, silence, speech density, and arrangement all affect what a listener experiences.
For example, spoken news with little background music may seem more present than a music track at a similar integrated reading. A rain recording with a broad, steady texture may feel louder than its number suggests. If you raise quiet ambience until it matches speech by measurement alone, it may become tiring over several hours.
YouTube's encoder and ingest guidance should shape the output format, but it does not turn one loudness setting into a universal creator mandate. Use the target as part of your channel's editorial policy. Write it down, listen to the result on ordinary headphones and a small speaker, and change it if the channel's material requires a different balance.
For RTMP or RTMPS ingest, YouTube lists AAC or MP3 audio and recommends, for stereo, 44.1 kHz and 128 kbps in its encoder guidance. These are YouTube's technical recommendations, not defaults that FFmpeg silently applies and not proof that a particular loudness target is required. Check the current YouTube live encoder settings before publishing because ingest guidance can change.
A practical policy might be: files with audio are measured individually, files that need correction are rendered to a separate directory, and video-only files receive one generated stereo track. That policy is more useful than chasing a single number without listening to the actual playlist.
Do not measure unrelated uploads as one programme
Do not concatenate a prayer, a commercial, a music performance, and a silent title card into one analysis file merely to obtain one set of loudness values. The combined measurement can hide a problem in one component. Long silence can pull down an overall reading, while a short loud section can influence true peak and make the result look different from the way the file is normally heard.
Measuring unrelated uploads together also makes corrections difficult to explain. If the evening news is quiet after processing, you do not know whether the cause was the news file, the preceding ambience, or the combined analysis. Independent measurements let you identify the source of the change.
There is one useful distinction: a single continuous recording can be measured as one programme even if it contains chapters. The question is whether the listener experiences it as one uninterrupted work. If your FFmpeg playlist stops and starts separate files, measure those files separately unless you have a deliberate reason to create a shared programme treatment.
For channels that switch content without ending the broadcast, test the file boundaries as well as the files themselves. The guide to changing videos without ending a YouTube live stream covers the continuity problem, while this workflow covers the audio structure and level decisions inside each file.
Check transitions and perceived loudness
After processing, listen across the joins. Do not check only the middle of each file. Play the final minute of one item and the first minute of the next, including any fade, spoken introduction, or silence. A meter can show compliant-looking values while the transition still feels abrupt.
Check at least these points:
- Does every output contain exactly one intended audio stream?
- Does a video-only file remain silent rather than accidentally receiving an old audio track?
- Does the first spoken word arrive at a sensible level after a quiet opening?
- Does a long ambience file become uncomfortable when it follows speech?
- Does the end of one file leave a sudden gap before the next file begins?
Use ffprobe or another media inspector to examine the output stream structure before you publish. For a live test, watch YouTube Live Control Room's stream health and confirm that the outgoing stream contains audio. YouTube Help says, “Make sure to test before you start your live stream. Tests should include audio and movement in the video similar to what you'll be doing in the stream.”
If the source has no audio and FFmpeg says that no audio stream was selected, check the input indexes. In the example, the file is input 0 and anullsrc is input 1, so the silent stream is 1:a:0. If the stream runs indefinitely in a finite-file workflow, check whether the generated source is unbounded and whether -shortest or an explicit duration is appropriate.
For a channel that must continue while your computer is off, removing the local encoder from the overnight routine can also remove one source of failure. StreamNeo is useful at the point where a prepared file needs to keep running as a YouTube stream with monitoring and automatic restart rather than depending on a computer staying awake, but you still need to validate the file's audio and YouTube's ingest response first.
Match the ingest protocol to the output
The silent-track command above creates a finite output file. It is not a complete YouTube live command, and an MP4 file command should not be treated as an HLS configuration.
For RTMP or RTMPS, check the video and audio codec settings, output format, real-time behaviour, destination URL, and credential handling. YouTube recommends RTMPS for secure transmission in its encoder guidance. Keep the stream key outside scripts that may be shared, published, or pasted into troubleshooting posts.
HLS has separate requirements. YouTube's developer documentation says that HLS ingestion uses muxed audio and video in M2TS and specifies AAC audio, along with segment and transport rules. Read the current YouTube HLS ingestion documentation when HLS is your chosen protocol rather than adapting an RTMP command by changing one output flag.
YouTube also provides live streaming error guidance. Use it when the ingest reports missing audio, unsupported settings, or multiple audio streams. The correct response is to inspect the actual output and compare it with the protocol's current requirements, not to add another audio map until the warning disappears.
Before putting the file into an overnight playlist, test the exact encoded output. Confirm the audio stream with ffprobe, play the beginning and a transition locally, then run a private or otherwise appropriate live test in YouTube's control room. If the stream is meant to repeat continuously, also test the point where the playlist loops. A channel can have a correct first file and still fail at the boundary between the last and first files.
If the workflow is part of a larger 24/7 setup, the overnight church stream checklist covers operational concerns such as continuity and supervision. The audio rule remains simple: generate silence where there is no audio, map one intended track, and measure each real programme on its own.
Before committing, compare the operating options on the pricing page. When the file and channel are ready, start free — 24-hour trial, no card.
FAQ
How do I add a silent audio track in FFmpeg?
Use -f lavfi -i anullsrc=... to create the audio input, then map the original video and generated audio explicitly. For a finite file, encode the audio as AAC and use -shortest so the output ends with the video rather than continuing with generated silence.
Should I map the original audio as well as the silent track?
No. If the source has no audio, map the video from input zero and the silence from input one. If you are replacing existing audio, leave the original audio unmapped so the output contains only the intended track.
Does YouTube require one universal loudness target?
The reviewed YouTube guidance provides ingest and encoder recommendations, but it does not establish one universal creator loudness target that guarantees consistent perceived loudness. Choose a target suitable for your material, apply it consistently where useful, and listen across real transitions.
Why use two passes with loudnorm?
The first pass measures one file, and the second uses that file's measurements to apply the selected correction. Two passes take more preparation, but they give you tighter control when a playlist mixes speech, music, and ambience.