Skip to content
streamneo.
Troubleshooting13 min read

How to Loop MP4 Files with Different Audio Formats in an FFmpeg YouTube Stream

Prepare MP4 files with mismatched audio for a reliable FFmpeg YouTube Live loop by inspecting, normalising and testing the playlist.

sn.
StreamNeoPublished 4 October 2026
Worth sharing?

When MP4 files use different audio codecs, sample rates or channel layouts, a stream-copy concat playlist may fail or produce glitches. The reliable approach is to inspect each file, convert the tracks to a compatible profile where needed, then loop the prepared playlist and encode it for YouTube Live.

The important distinction is between making files compatible and merely placing them in a list. FFmpeg’s concat demuxer is useful for stream-copying files only when their streams, codecs and time bases match; if they do not, normalise or re-encode instead.

Why mixed audio formats can break concatenation

An MP4 is a container, not a guarantee that every file inside it has the same audio. One file may hold AAC stereo at one sample rate, another may contain MP3, and a third may have a different channel layout or multiple audio tracks. These differences can be invisible when you play each clip on its own, but they matter when FFmpeg tries to treat the files as one continuous input.

The concat demuxer joins packets without decoding and rebuilding the media. That is why it can be efficient when all files already share compatible streams. It does not convert one codec into another, reconcile different audio layouts, or make distinct time bases identical. FFmpeg’s concat demuxer documentation specifies that the files need the same streams, including codecs and time bases.

A mismatch can lead to an error during processing, a change or loss of audio at a boundary, or timing problems. Some files may appear to join but still produce unwanted results because their streams do not end together. The demuxer adjusts timestamps across the sequence; where stream durations differ, it can leave gaps. Incorrect duration metadata can also affect the transition to the next item.

There are two practical routes. If the files already match, use the concat demuxer and avoid re-encoding. If they do not, decode and re-encode them to a shared profile, or use a concat filter workflow that re-encodes as part of joining. FFmpeg’s FAQ distinguishes these approaches: the concat filter is suited to concatenation that requires re-encoding, while the demuxer can avoid it when the inputs meet its requirements.

This is not a rule that every input must become identical in every respect. It is a rule that the streams you intend to join must be compatible for the chosen method. A quiet devotional recording and a spoken introduction might have different source properties, for example, but both can be prepared to the same intended audio codec, sample rate and channel layout before being joined.

Inspect the streams in each MP4

Start by checking every source file, not just the first one. Running ffmpeg -i clip.mp4 prints information about the streams and duration; it may also report that an input has no video or no audio. Record what you intend to preserve, rather than assuming all clips have the same stream order or that the first audio track is the one you want.

For each file, note the video codec, dimensions and frame rate, and the audio codec, sample rate, channel layout and stream order. Compare durations as well. An MP4 can contain multiple audio tracks, such as different languages, commentary or a version with music. If the wrong one is mapped into the output, the concat problem may be mistaken for an encoding problem.

The command-line option -map lets you explicitly select streams. Without it, FFmpeg chooses streams automatically according to its selection rules, which can produce a different result from the one you intended. The FFmpeg command-line documentation explains stream selection and output codec options. A deliberate mapping is especially useful if a file has more than one audio track, or if your visual playlist includes an item that should not contribute its original sound.

Make a small inventory before conversion. A plain note or spreadsheet with one row per file is enough: filename, intended video, intended audio, duration and any unusual properties. If one clip has mono audio while others are stereo, decide whether you want mono preserved or converted to stereo. If a file’s audio ends well before its video, decide whether that silence is intentional. Do not discover these choices after the playlist has been running overnight.

The source duration is useful, but do not assume that the container’s duration tells the whole story. The video and audio streams can have different lengths. Listen to the end of each file and inspect its timeline if boundaries matter. If you are preparing a regional-language playlist, the considerations around selecting and checking each clip are also relevant to this regional-language MP4 loop in OBS, though FFmpeg’s concat methods have their own compatibility requirements.

Choose a shared output profile

A shared profile is a set of output choices, not a universal preset. Decide what is appropriate for your actual files and YouTube output: video dimensions, frame rate, video codec, audio codec, sample rate, channel layout and audio bitrate. Then apply those choices consistently to the prepared files. Do not force a profile simply because it appears in an example; a portrait source, an older low-resolution clip and a high-resolution landscape video may call for different decisions.

For a simple stereo music or speech stream, you might choose a common stereo output and a single audio codec for every prepared item. YouTube’s current live encoder guidance lists AAC or MP3 audio for RTMP/RTMPS workflows. It gives stereo audio guidance of 44.1 kHz and 128 kbps, and different guidance for 5.1 audio. These are platform recommendations, not a guarantee that a particular source or workflow will sound good. Check the current YouTube encoder settings page before choosing your output values.

Choice When it helps What to check
Keep the source streams and use the concat demuxer All intended streams, codecs and time bases already match Audio layout, stream selection, durations and boundaries
Prepare each file to a shared profile, then use the demuxer Sources differ, but you want a reusable playlist of normalised files That the prepared files now have matching streams and compatible timing
Use the concat filter and re-encode You prefer a single FFmpeg joining workflow and can afford the processing Consistent filter inputs, explicit maps and chosen output codecs

The table describes workflow choices rather than quality rankings. Re-encoding involves processing and can change the media; stream-copying avoids another encode but only works for already compatible inputs. Keep an untouched copy of the originals so that you can revise the profile or mappings without starting again from degraded intermediate files.

Video needs attention as well as audio. If frame rates, dimensions or stream layouts differ, making only the audio match may not be enough for the playlist method you have chosen. Standardise the characteristics needed for the concat demuxer, or use a re-encoding filter workflow that handles the differences. A single profile will not suit every set of inputs, so make the choice from the files you have and the intended output.

Normalise or re-encode before concatenation

If the audio tracks differ, decode and re-encode them to a common target before using a stream-copy playlist. For example, your preparation step might produce AAC stereo at a chosen sample rate for every item. The exact values should follow your source material and YouTube’s current guidance, not a one-size-fits-all recipe. Re-encoding the audio gives the files a consistent codec and layout; if the video streams also differ in ways the demuxer cannot handle, prepare the video consistently too.

The key is to set the intended streams explicitly. Decide whether you want the first video and first audio stream, or another track, and use -map accordingly. Set output codecs and audio options for the output file. FFmpeg options apply in relation to inputs and outputs, so ordering matters: input-specific options belong with the input, while output options should appear before the output filename. The command-line documentation is the best reference when adapting a command to your installed version and media.

You can prepare files individually, which makes it straightforward to inspect and replace one failed conversion. Alternatively, when you want to join several inputs in one re-encoding operation, FFmpeg’s concat filter can be appropriate. It requires filter inputs to be made compatible for the filter’s chosen media parameters, then encodes the combined output. This can be a cleaner workflow for a short sequence, but for a reusable 24/7 playlist, prepared files are easier to validate one by one and repeat.

Avoid treating a sample command as a tested universal solution. Inputs may have extra streams, variable frame rate, different dimensions, or timestamps that require attention. Confirm that your FFmpeg build includes the encoders you select. Test a short prepared segment first, including a transition between two sources, before generating a full library of converted files.

For devotional or bhajan material, listen for clipped starts, silence, changes in loudness and abrupt room-tone shifts at the join. For ambience or lofi, check for an obvious change in background noise or stereo width. These are editorial checks as much as technical ones. Normalising the encoding format does not make two recordings sound alike, so choose whether you need a deliberate crossfade or a clean cut in a separate editing step.

Use the concat demuxer only for compatible inputs

Once the prepared files share the streams and timing characteristics you need, the concat demuxer can assemble the playlist without re-encoding those prepared files. Create a text playlist whose first line is exactly ffconcat version 1.0, then add one file 'path' directive for each item. That exact header allows FFmpeg to recognise the format automatically. Use paths that are valid from the working directory where you run the command.

For instance, the list can be structured like this:

ffconcat version 1.0
file 'prepared/clip-a.mp4'
file 'prepared/clip-b.mp4'
file 'prepared/clip-c.mp4'

This playlist is not a mechanism for repairing mismatches. If one prepared file has AAC stereo and another retains a different audio codec or layout, do not assume the demuxer will convert it on the fly. Return to the preparation step, make the streams compatible, or choose a re-encoding concat-filter workflow. The FFmpeg FAQ on concatenation is useful when deciding between the demuxer and filter.

The demuxer adjusts timestamps so that each file follows the preceding one. That does not promise a perfectly seamless transition. If video and audio durations differ within a file, there may be a gap before the next file. Check the reported durations and verify the actual boundary by playing a joined test. The demuxer supports duration directives when stored durations are wrong, but only use them when you have established the correct duration; an invented value can create a new timing problem.

You can avoid re-encoding at the join when the prepared inputs are genuinely compatible, but you may still need to encode the resulting sequence for live delivery. These are separate decisions: how to join the source files, and how to produce the output stream. Do not use -c copy as a substitute for checking the inputs. It instructs FFmpeg to copy encoded streams, not to reconcile their differences.

Loop the prepared playlist into YouTube Live

When the playlist is ready, test a file-based loop locally before sending it to YouTube. FFmpeg’s -stream_loop -1 repeats an input indefinitely; -re reads it at a real-time rate for a live-style input workflow. A schematic command might look like this:

ffmpeg -re -stream_loop -1 -f concat -safe 0 -i playlist.ffconcat \\
  -map 0:v:0 -map 0:a:0 \\
  -c:v libx264 -r 30 -g 60 -b:v <chosen bitrate> \\
  -c:a aac -ar 44100 -ac 2 -b:a 128k \\
  -f flv '<YouTube RTMP(S) ingest URL and stream key>'

Treat this as a pattern, not a command to paste unchanged. Confirm that the selected maps refer to the intended streams, replace the bitrate and frame rate to suit the resolution and content, and confirm that your FFmpeg build has the requested encoders. The sample GOP value corresponds to a two-second interval only at the sample frame rate shown. YouTube recommends a two-second keyframe interval and says it should not exceed four seconds; check its current guidance for your selected workflow and settings.

The example maps the first video and first audio stream in the concatenated input. That is suitable only if every item has the expected layout. If your content is audio-only, has extra tracks, or uses another order, revise the maps rather than accepting unexpected automatic selection. Keep the real stream key private: do not paste it into a public post, support request or shared command transcript.

YouTube recommends RTMPS for encrypted transport. Use the ingest URL and protocol supplied for your own live setup, and check the current platform page rather than assuming that an old saved address or encoder setting remains suitable. This is a technical hand-off from FFmpeg to YouTube, not an approval guarantee. If you are deciding whether a local machine should remain involved in the broadcast, this guide to keeping FFmpeg streaming after a video ends covers the separate problem of keeping a process alive beyond one file.

For a playlist that must repeat while your own computer is off, StreamNeo removes the burden of keeping a local FFmpeg process and machine running by taking an uploaded file and your YouTube stream key for a cloud-run 24/7 broadcast. It is YouTube-only, so it is not a replacement if you need FFmpeg access for another destination or custom filter graph.

Check the output and stream health

Before going live for an audience, play the joined output locally. Listen through at least one boundary and check the beginning and end of each source segment. Confirm that the expected audio is present, there is no unintended silence, and the video is still in sync. If the sound is wrong, inspect the maps and source stream order before trying different bitrate values; if there is a gap, compare audio and video durations and check the file metadata.

Then perform a YouTube stream test using representative content. YouTube recommends testing with audio and representative video movement, then monitoring stream health. Include material that resembles the actual channel: still devotional artwork is not a good test of a later video with movement, and a short speech sample will not expose a problem that occurs only in a longer music track. Check the live dashboard for warnings and listen on the receiving end as well.

If YouTube reports an encoder or stream-health issue, compare the actual output with its current ingest recommendations. Check codec, frame rate, keyframe interval, audio settings and selected protocol. Avoid changing several settings at once. A controlled change helps identify whether the issue lies in the source, the prepared playlist, the live encoder or the ingest configuration.

Keep the prepared files and playlist versioned or clearly named. If you replace an input with a new encode, verify that it still meets the same profile before adding it to the concat list. A channel that alternates between local music videos, spoken announcements and ambient loops can use different source material, but its playlist workflow should make the intended transitions and selected streams clear. For other always-on channel considerations, this guide to Sanskrit mantra streaming on YouTube Live offers a relevant channel-specific context.

Before committing, compare the operating options on the pricing page. When the file and channel are ready, start free — 24-hour trial, no card.

FAQ

Can I concatenate MP4 files with different audio codecs using -c copy?

Not reliably with the concat demuxer. Stream-copy concatenation requires compatible streams, including matching codecs and time bases; -c copy copies encoded data rather than converting it. Re-encode to a shared profile or use a concat-filter workflow when the inputs differ.

Should I use the concat demuxer or the concat filter?

Use the demuxer when the files already have the same required streams and compatible codecs and time bases, and you want to avoid re-encoding at the join. Use the concat filter when you need to re-encode while joining or accommodate varying inputs. In either case, inspect stream layouts and test the boundaries.

Why is there a pause or audio gap between files?

The audio and video streams may not have the same duration, and the demuxer adjusts timestamps across the sequence. Check stream durations and duration metadata, then listen to the joined result; do not assume the container’s overall duration guarantees a clean boundary. A re-encoding workflow or an intentional edit may be needed.

Can I use the same audio settings for every YouTube Live playlist?

No single profile suits every source or channel. Choose settings from your material, intended channel layout and current YouTube guidance, then apply them consistently to the files you are joining. Test the actual output and monitor stream health before relying on it for a long broadcast.

YOU’VE REACHED THE END

Keep the ideas coming.

More guides, useful tools and a little help for your next broadcast.

Back to the journal ↗
YOUR NEXT READ

A little more to explore.

More Troubleshooting guides ↗ · All topics ↗