Keep the prerecorded video's picture on its own path, and send its decoded audio and the decoded music file to separate inputs of audiomixer. Lower the music input there, then convert and encode the mixed audio alongside the separately encoded video before muxing the output for YouTube Live.
That design lets you preserve speech or other original sound while adding a bed underneath it. The exact elements and caps depend on the files, installed GStreamer plugins and encoder choices, so treat a pipeline sketch as a plan to adapt and test rather than a guaranteed command.
Plan the video and music audio paths
Start by deciding what should be audible. If the video has narration or useful sound, retain its audio as one mixer input and put background music on another. If the video is silent, the music can be the only input; there is no need to invent a silent video-audio branch. Keeping the inputs separate means you can adjust or mute the music without also reducing the original soundtrack.
The picture path is distinct. A file source and decoder expose the prerecorded video's streams; route decoded video to the video encoder, and route decoded audio to the mixer. The mixer does not mix or alter the picture. Its job is only to combine raw audio streams. For a broader comparison of prerecorded-video workflows, see this guide to OBS or FFmpeg for a 24/7 ambient stream.
Draw the branches before building the pipeline. A practical outline is:
video file -> source / decode -> raw video -> video encoder ----┐
-> raw audio -> convert / queue --┐│
audiomixer -> audio convert -> audio encoder --┤
music file -> source / decode -> raw audio -> convert / queue -┘ mux -> network sink
The drawing is conceptual: actual element names depend on file types and available plugins. A decodebin can discover and decode supported streams, but its output pads are created dynamically. A real application must identify the desired audio and video pads and link them to their intended branches. A gst-launch-1.0 prototype can help test a known pipeline, while an application gives you better control over dynamic pads, errors and playback behaviour.
Put queues on branches that can run at different rates, especially around decoding and before branches meet. They help separate scheduling between branches, but they do not repair an unsupported codec or mismatched format. Before writing the full pipeline, list the input container and codecs, whether video audio exists, the desired output codec and the encoders and muxer installed on the machine.
Decode the prerecorded video's audio
Use a file source for the video file and a decoder path that can expose its streams. GStreamer's filesrc reads a file, while decodebin can select demuxers and decoders for supported content. Its pads are dynamic: do not assume that the first pad is always the audio stream or that every file contains audio. Inspect the stream types and connect the audio pad to an audio-only branch, while sending the video pad to the picture branch.
After decoding, audio should be raw audio before it reaches audiomixer. A conversion element can reconcile format details such as sample representation or channel layout, and a resampler can change the sample rate where required. Use caps filters when you need an explicit format, but choose caps based on the next elements and the actual source rather than copying a generic example blindly.
If the video has dialogue, listen to that branch alone first. Check whether it is already quiet, heavily compressed or uneven between scenes. That establishes a baseline for the music adjustment. If there is no usable audio stream, omit this branch: an absent stream cannot be fixed by changing the mixer gain.
You may be using gst-launch-1.0 while testing, but a continuous channel needs handling for end-of-file and errors. GStreamer's decodebin documentation describes its decoder behaviour; its dynamic pads still need to be routed correctly in your pipeline. For longer-running playback, decide explicitly whether the video repeats, advances to another file or stops, and check for a gap at each transition.
Decode the background music file
Treat music as a second source, not as an effect attached to the video's decoded sound. Read the music file through a suitable source and decoder, then convert it to raw audio. For WAV, GStreamer provides wavparse; other containers and formats may require a demuxer and decoder selected by a decoder bin. The installed plugin set determines what is available, so a pipeline that works for one WAV file may not work for a compressed file with a different codec.
Inspect whether the music is mono or stereo and what sample rate it uses. The mixer can convert input formats as needed, but its output sample rate must suit downstream processing or the first configured pad. If your desired encoder needs a particular format, set and verify caps after the mixer, and make sure the conversions are actually supported by installed elements.
Decide whether the music is a one-shot file or should continue through a long stream. A finite file will not necessarily repeat on its own. If it must loop, arrange repetition in the source or application design and listen at the join: a click, silence or abrupt level change may make the loop obvious. Music licensing is a separate question from pipeline operation; having a technically valid mix does not establish permission to broadcast a recording. You can review the channel-specific concerns in this article about broadcasting recorded songs on YouTube.
Connect both paths to audiomixer
Connect each decoded, raw audio branch to its own audiomixer sink pad. The mixer combines the incoming streams and synchronises them; it is designed to accommodate live inputs as well as other streams. Keep a distinct pad for video audio and a distinct one for music so their gains and mute states remain independently controllable.
The GStreamer audiomixer reference documents raw audio inputs and per-input volume and mute controls. It also notes that the mixer can convert formats, but that sample-rate requirements still matter: the selected rate needs to work with downstream elements or the first configured pad. In practice, normalise branches before the mixer where useful and check the resulting caps rather than expecting the mixer to resolve every incompatibility.
Queues on the two incoming branches can help when file decoding proceeds at different rates. Keep the pipeline clock and timestamps in mind, especially if you mix a file with a live source or if files have unusual timestamps. A branch that is starved, linked to the wrong pad or producing unexpected caps can yield silence or an error even when the diagram looks right.
A useful debugging order is to test each source on its own, then each decoded branch, then both connected to the mixer. Confirm that the video is still flowing to its encoder while checking audio separately. If you change the input files, re-check decoder pads and caps; the same pipeline description need not fit media with different codecs or channel layouts.
Lower the music input volume
Adjust the music pad's volume property rather than turning down the mixed output. That preserves the video audio level while letting you place the background underneath it. The documented pad range is 0.0 to 10.0; that range describes the control, not a recommended listening setting. A value suitable for one recording may be too loud or too quiet for another.
Set an initial gain, then listen to representative sections. Include a quiet spoken passage, a louder passage, and any part where the original video audio changes substantially. Music with strong transients can obscure consonants even if its average level seems modest. If intelligibility matters, check the mix on ordinary speakers or headphones at a sensible listening level, not only on studio monitors.
Avoid assuming that a single loudness balance suits every video. A devotional talk, a lofi loop and a local news segment have different purposes and source levels. If the music is meant to be subtle, use listening and meter readings to find a level that supports rather than competes with the primary sound. Make changes to the music pad first; change video audio gain only when you have a separate reason to do so.
This is also where an automated broadcast workflow can remove a specific operational burden: StreamNeo can keep an uploaded file streaming without leaving your own computer on, so a workstation does not have to remain awake just to sustain the channel. It does not decide whether your mix sounds right, and you should make and check the audio mix before relying on any playback workflow.
Convert, encode, and mux for YouTube Live
Take the mixer's output through the conversions needed by the chosen audio encoder, then encode the audio. Keep the video branch separate through video conversion and encoding. Only after both tracks are encoded should they meet at a muxer that produces a format accepted by the network sink. For an RTMP-family workflow, GStreamer's rtmpsink expects FLV content, so an appropriate muxer must precede it; consult the rtmpsink documentation and the installed element details.
YouTube's current live encoder settings list RTMP and RTMPS ingest and give codec and audio guidance that varies by configuration. For RTMP/RTMPS, it lists AAC or MP3 audio; for stereo it lists 44.1 kHz and 128 kbps. It also recommends a two-second keyframe interval and says not to exceed four seconds. Treat these as platform guidance, not a promise that any pipeline will be accepted: check the current page and Live Control Room settings for your chosen stream.
Do not copy a single video bitrate into every setup. YouTube's recommendations depend on codec, resolution and frame rate, and your upload capacity also matters. Select encoder settings for the actual output and leave headroom for network variation. YouTube recommends a 20% upload bandwidth margin in its streaming tips. If you use RTMPS, obtain the ingest URL and stream key from Live Control Room; YouTube explains the encrypted connection in its RTMPS guidance.
The muxer and sink are separate from the mixer. A correctly mixed audio signal does not ensure that the encoded tracks have compatible timestamps, that the muxer accepts them or that the network path is reachable. Build and test those stages independently, and verify the sink's URL handling and stream-key configuration without publishing credentials in logs or shared pipeline text. For a continuously running service, the operational side matters too; this guide on monitoring a remote 24/7 YouTube stream covers the checks that continue after the initial setup.
Listen for clipping and verify the output
Mixing adds signals. If their combined peaks exceed the representable range, the output can clip; the mixer clamps samples rather than making the result sound clean automatically. A music bed that is unobtrusive in one section can still push a louder passage over the limit. Listen to the encoded output as well as the input branches, and use a meter if one is available in your pipeline.
Check for more than clipping. Listen for dialogue that becomes hard to understand, an unexpected channel balance, silence when one file ends, a discontinuity at a loop point, or audio and picture that drift apart. If the audio disappears, check that the intended decoder pad was linked and that caps negotiation succeeded. If the pipeline stops, inspect bus errors and element messages rather than repeatedly restarting without finding the cause.
Run a representative test before the planned broadcast. Use audio and picture with a similar level of movement and dynamics, preview the stream in Live Control Room, and check stream health while it is running. YouTube's guidance recommends testing and monitoring; a local preview alone does not verify the complete upload and ingest path. For a 24/7 channel, also confirm what happens after a file reaches its end and who or what will notice a dropped stream.
| Check | What to verify | Why it matters |
|---|---|---|
| Source discovery | Expected audio and video pads appear | File formats and stream layouts differ |
| Mixer inputs | Original audio and music reach separate pads | Music gain must not lower the primary sound |
| Output format | Sample rate, channels and encoder caps agree | A downstream element may reject incompatible caps |
| Peak behaviour | Combined signal remains clear in loud sections | Summed audio can clip |
| Live ingest | Preview and stream health look normal | Local playback does not prove the network path works |
| Long run | End-of-file, looping and reconnect behaviour are known | A test of one short passage is not a 24/7 plan |
Before committing, compare the operating options on the pricing page. When the file and channel are ready, start free — 24-hour trial, no card.
FAQ
Can I keep the video's original audio and add music?
Yes. Decode the video's audio and the music to raw audio, then connect them to separate audiomixer inputs. Lower the music pad independently and listen across both quiet and loud sections of the video.
What if the prerecorded video has no audio?
Use the music as the mixer's only input, or omit the mixer if you do not need to combine sources. Confirm the decoded streams before linking branches; the pipeline should reflect what the file actually contains.
Will the same pipeline work with every video and music codec?
No. Decoder and conversion elements depend on the formats and installed plugins, and output caps must suit the encoder and muxer. Inspect pads, test negotiation and adapt the branches to the files you intend to use.
Does lowering the music guarantee a clean mix?
No. Different recordings have different peak and average levels, and the combined signal can clip. Listen to the result, inspect meters where available and adjust the music input for the specific programme.