To add background music under a prerecorded YouTube video without removing its speech or ambience, mix the two audio streams with FFmpeg’s amix filter. Then map the original video separately, encode the new audio, and listen to the exported file; the example below is a starting point, not a guaranteed fit for every set of files.
The key decisions are whether the original sound remains clear, which audio stream determines the ending, and whether the music is licensed for your intended YouTube use. Test the inputs and mix level on a short section before committing to a full export.
What the audio mix should do
The aim is to combine two sounds, not replace one with the other. For a devotional talk, for example, the voice may carry the message while a quiet instrumental bed fills pauses. For a lofi or ambience video, the original recording might be rain, room tone or a distant bell, with music adding another layer. Decide which sound is meant to lead before choosing a level.
FFmpeg’s amix filter combines multiple audio inputs into one output. That is different from simply putting two tracks next to each other, or combining channels with a filter intended for that purpose. In this workflow, the video's audio and the separate music file enter amix; the result becomes the output audio track.
A useful mix should preserve the sound you want viewers to notice. If a spoken phrase becomes hard to understand when music is playing, the bed is too prominent for that particular material, regardless of what a command-line value suggests. If the source is an ambience recording, listen for details that the music may mask. There is no universal setting that suits every voice, recording or musical arrangement.
Keep a clean copy of the original video and audio. That lets you compare the result, change the music level or choose another track without having to reconstruct the source. If you publish loops or a long prerecorded stream, preparing the finished file before scheduling also gives you a chance to review transitions and the ending; the practical considerations in planning YouTube Live content are useful once the edit is ready to become part of a channel schedule.
Check the video and music inputs
Before writing a command, confirm that the video actually contains an audio stream. A video may be silent, may have more than one audio track, or may contain a track you do not want to retain. FFmpeg stream selectors such as [0:a:0] refer to the first audio stream in the first input; [1:a:0] refers to the first audio stream in the second input. The numbering follows the input order in the command, so swapping the filenames changes what those selectors refer to.
Also check what is in the files, not just their names. Play the video near the beginning, in a representative middle passage and near its end. Listen for dialogue, room noise, existing music, sudden changes in level, and any section where the original sound disappears. Play the music file on its own and note whether it has an introduction, an abrupt ending, or a level change that will become distracting when repeated or placed underneath speech.
The example later uses input.mp4 and music.mp3 as placeholders. Replace them with your own filenames, quoting paths that contain spaces. For instance, a path such as ~/Videos/sermon edit.mp4 needs shell quoting so that it is treated as one filename. Keep the output filename distinct from the input filename; overwriting a source is an avoidable risk.
Consider durations before mixing. If the video has several minutes of picture after its audio ends, a mix that follows the audio may not provide sound for that remaining picture. If the music is shorter than the original audio, the music will stop before the original sound does unless you deliberately make the music longer or loop it. Do not assume that the two files end together simply because they appear to have similar durations in a file browser.
If you are not sure which stream is which, inspect the file’s streams with FFmpeg’s input information output and test a short excerpt. A source with no audio cannot preserve an original audio stream: in that case, mix only the music, and map that result as the output audio. Do not use the two-input command unchanged when one of its referenced audio streams does not exist.
Understand the FFmpeg amix workflow
The command has four jobs: open the video, open the music, combine their audio, and write video and audio to an output file. This is a representative pattern:
ffmpeg -i input.mp4 -i music.mp3 \
-filter_complex "[0:a:0][1:a:0]amix=inputs=2:duration=first:dropout_transition=0[aout]" \
-map 0:v:0 -map "[aout]" -c:v copy -c:a aac -b:a 192k output.mp4
This is an illustrative starting point, not a tested recipe for every file. The official FFmpeg filter reference describes amix and its options. Here, inputs=2 says to combine two audio inputs. The labels in square brackets identify those inputs and the [aout] label names the result of the filter chain so it can be mapped later.
duration=first makes the mixed audio end when the first audio input ends. In this command, the first audio input is the video's original sound, because [0:a:0] appears before [1:a:0]. This is a choice, not a default that should be accepted without thought. FFmpeg also documents shortest and longest; they end the mix at the shortest or longest input respectively. Choose according to the intended sound and check how that audio duration relates to the picture duration.
dropout_transition controls how long the filter takes to renormalise volume when an input ends. The documented default is two seconds; setting it to 0 avoids that transition. Leaving the option out is also reasonable when you want the documented default behaviour. Neither choice decides which stream ends first; that is what duration controls.
The filter’s normalisation is enabled by default. If you turn it off, FFmpeg’s documentation warns that clipping can be heavy when input levels are not controlled. For a first pass, keep the default and adjust the music input separately if it competes with the original sound. Do not infer the final loudness or quality from the command alone: playback, source levels and encoding all matter.
The command copies the video stream with -c:v copy and encodes the filtered audio with AAC. Filtering creates a new audio stream, so the original audio cannot simply be copied as-is into the mixed result. Video stream copying may be suitable when the source codec is compatible with the MP4 container and intended playback; if it is not, choose a compatible video encoding method instead. The FFmpeg Cookbook’s background music example illustrates the separate mapping pattern, but an example elsewhere is not evidence that a particular command works unchanged on your files.
Map the original video and mixed audio
The output mapping is what keeps the picture while replacing its audio with the mix. -map 0:v:0 selects the first video stream from the first input. -map "[aout]" selects the labelled result of the audio filter. Because the command maps those two streams explicitly, it does not copy the source audio as a third track or accidentally select the music as the only audio.
This separation is useful when the source video is already in a format that can be copied into the desired container. Stream copying avoids re-encoding the picture, but it does not repair an incompatible codec, change the frame size, or make an unsuitable source format suitable. If FFmpeg reports that the copied video stream cannot be written to the selected container, use a compatible video codec and encoding settings for your target rather than removing the audio filter or mapping the wrong stream.
If the video contains multiple audio tracks, [0:a:0] selects only the first. Confirm that it is the track you intend to preserve. If the intended track has a different index, change the selector after checking the input’s stream listing. Similarly, if a file has multiple video streams, be deliberate about which one is mapped.
Duration needs separate attention because audio and video are mapped as separate streams. Suppose the picture lasts longer than the original audio and duration=first ends the mixed audio with that original track: the remaining video may have no audio. Conversely, if you choose an audio duration that extends beyond the picture, decide whether the output should stop with the video. An output-level option such as -shortest can end the file when its shortest mapped stream ends, but it changes the result, so inspect it rather than adding it automatically.
A short test export is a good way to check stream selection and duration before rendering a long programme. If you are preparing a prerecorded file for a scheduled broadcast, distinguish the editing step from the later live-streaming setup. The article on scheduling a prerecorded YouTube stream with OneStream Live covers that separate publishing workflow; this FFmpeg command only prepares a video file.
Adjust and test the music balance
Apply the music level before mixing so that you can change it independently. For example, replace the filter expression in the command with:
-filter_complex "[1:a:0]volume=0.25[music];[0:a:0][music]amix=inputs=2:duration=first[aout]"
The volume filter here is an example setting, not a recommendation for every file. It reduces the music input before amix combines it with the original audio. If speech is meant to stay prominent, this is easier to reason about than changing both files or altering the final mix without knowing which sound caused the problem. Use a small test export and adjust the value by listening.
Listen on the equipment your viewers are likely to use as well as on your usual editing setup. Check at a comfortable level on headphones or speakers, and try a phone speaker if that is a common way your audience watches. You are checking whether words remain understandable, whether music covers important ambience, and whether the mix changes noticeably when a speaker pauses. A setting that sounds balanced in one passage can overwhelm a quieter passage elsewhere.
Review more than one point in the timeline: an opening, a section with the fullest speech or sound, a quiet passage, and the ending. If the music has a strong beat or a prominent melody, it may compete with speech even when its overall level seems modest. If the source has a loud section, lowering the music may be more useful than raising the voice, particularly when you want to retain the source’s character.
Listen for clipping or harsh distortion, especially where both inputs become loud at once. The default amix normalisation is not a promise that the result will sound right or be free from every problem. If you disable normalisation, take extra care with input levels; FFmpeg notes the clipping risk when it is disabled. Do not treat a numerical volume value as a loudness target, because the perceived balance depends on the particular recordings and their changes over time.
For a longer video, you may need a more involved mix than one fixed music level. You can make a separate edit or use filters and automation suited to the material, but keep the simple mix as a clear baseline. If you need the music to fade in or out, a hard cut at the beginning or end may be audible; add and test a deliberate fade rather than assuming amix will create one for you.
Export and inspect the output file
After the test sounds acceptable, export to a new file and let FFmpeg finish. Read the output for errors and warnings rather than assuming that a file with the expected name is complete. A successful command should produce a playable file, but a process interruption, missing stream, incompatible codec or inadequate disk space can leave you with an incomplete or unusable result.
Open the exported file in a player and check the picture and sound together. Confirm that the video starts where expected, the original audio is still present, the music is audible at the intended level, and there is no unintended silence or extra track. Skip to the middle and near the end; check whether the audio and picture stop where you planned. Listening to the output matters because a command that parses correctly cannot tell you whether a voice has become difficult to follow.
If you find a mismatch, make one change at a time. For example, if the music is too prominent, adjust the music volume value, export a short section, and compare it with the previous test. If the sound ends too early, reconsider the chosen duration and the video’s ending rather than increasing the music level. If picture is missing or playback fails, check the stream mapping and container compatibility.
Keep the original file and a note of the filter expression used for the approved export. This makes it possible to recreate the result if you revise the music or correct a caption or visual edit later. If the file will be part of a continuous YouTube broadcast, a clean source export also makes it easier to diagnose whether an issue lies in the prepared video or in the later streaming setup. For example, a channel loop may need a separate check for OBS encoder overload; changing the audio mix will not by itself resolve an overloaded live encoder.
Publish the finished video on YouTube
Before upload, check the music’s actual permission terms. A file being freely downloadable or labelled “royalty-free” does not, by itself, establish permission for every use or remove attribution requirements. Keep a record of the source and the licence or terms that applied when you selected the track, and include any required credit in the video description.
YouTube’s Audio Library guidance says tracks and sound effects in the library are known to YouTube to be copyright-safe for use on the platform. It also explains that some Creative Commons tracks require attribution. YouTube says it cannot give legal advice and is not responsible for issues with music from other libraries, so check the current official guidance and the terms for the track you use. Using a track in a mix does not grant permission to publish it.
After upload, check the video on YouTube and review any notices or restrictions shown in the account. Do not assume an export test establishes how a platform will treat a music licence or a particular upload. If your channel is a 24/7 station, think about how the track will be heard in context and whether you have permission for the intended repeat or broadcast use. A channel’s other operational decisions are separate from this edit; for example, the guidance on copyright claims on an Indian channel’s livestream is relevant to the live context, not a substitute for checking the music’s rights.
Preparing and checking the file is one step; keeping a channel live is another. If your intended workflow is to upload a finished video, provide your YouTube stream key and let the broadcast continue while your own computer is off, StreamNeo can remove the need to keep a local machine running for that job. It is for YouTube, and it does not grant music rights or decide whether your file is suitable; prepare and review the video first.
Before committing, compare the operating options on the pricing page. When the file and channel are ready, start free — 24-hour trial, no card.
FAQ
How do I add background music and keep the original video audio?
Use amix with the video’s audio stream and the music file as inputs, then map the mixed output as the audio track. Map the video stream separately so the picture remains in the export. Confirm that the source really has the audio stream selected in your command.
Should I use duration=first, shortest or longest?
Use first when the first audio input should determine when the mix ends, shortest when the earliest-ending input should control it, and longest when you want the mix to continue until its latest-ending input. Choose based on the intended result, then check how the mixed audio duration lines up with the video. An output-level option may also be needed if the picture should end first.
What volume should I set for the music?
There is no single setting that works for every voice, ambience recording and music file. The example value is only a place to begin; lower or raise it, render a short test, and listen for intelligibility and clipping across different passages. Judge the exported result rather than relying on the command alone.
Does adding music mean I can publish it on YouTube?
No. You need permission for your intended use and should check the track’s current terms, including any attribution requirements. YouTube’s Audio Library is an official source to consider, but music from another library needs its own terms checked, and a mix does not change those rights.