Use FFmpeg’s concat demuxer when your files have matching stream layouts and you want to read them in sequence, often without re-encoding. Use the concat filter when clips need to be resized, resampled, filtered or otherwise normalised; that route decodes and re-encodes them.
A concat playlist file is an instruction for FFmpeg, not a YouTube playlist URL or an HLS manifest. Once FFmpeg has produced or processed the sequence, a separate encoding and ingest step sends the live feed to YouTube.
Choose the concat method by compatibility
The choice is mainly about what is inside the files, not how many files you have. The demuxer sequences encoded packets and can copy them directly when streams match. The filter works on decoded media, which lets you transform clips into a consistent output, but requires encoding after filtering.
| Method | Use it when | Main constraint |
|---|---|---|
| Concat demuxer | Files have matching stream layouts and codecs, and you want to sequence them; stream copy may be possible | Streams must match, including codec and time base. It cannot apply frame or audio filters. |
| Concat filter | Clips need filtering or normalisation, or the output needs a new encoding configuration | You need a decode, filter and encode workflow; it takes more processing than copying packets. |
| Concat protocol | You are joining formats that support file-level concatenation | This is format-dependent and not the general-purpose choice for a playlist of varied video files. |
FFmpeg’s FAQ on concatenating media distinguishes these approaches. Its FAQ recommends the filter when re-encoding is needed, while the concat demuxer is suited to compatible inputs. The protocol is a narrower option for certain formats, rather than a shortcut that makes arbitrary MP4 files compatible.
For example, three exports from the same editing preset may be good demuxer candidates if their streams really do match. A phone clip, a screen recording and an older archive file may differ in frame rate, resolution, audio layout or timestamps. If you need to bring them to one output format, use a filter/transcode workflow instead of expecting stream copy to repair those differences.
A demuxer script is also not the broadcast destination. It tells FFmpeg which local files to read. YouTube receives an encoded live signal only after you configure an output for the ingest URL and key from Live Control Room.
Inspect the input streams first
Before writing the playlist, inspect every file and compare the video and audio streams. FFmpeg’s ffprobe can report codec names, dimensions, frame rates, time bases, sample rates, channel layouts and stream counts. A compact inspection command is:
ffprobe -v error -show_streams -show_format input-01.mp4
Run it for each candidate file and compare corresponding streams. Do not compare only file extensions. Two .mp4 files can contain different codecs or stream arrangements, and files with different extensions may still contain compatible streams. Check whether each file has the same video and audio streams in the same order, with compatible properties.
Pay attention to time base as well as visible properties such as resolution. The demuxer documentation says the files must have the same streams, including the same codecs and time base. A sequence that looks alike in a media player may still be unsuitable for direct packet concatenation.
Also check duration and the end of each clip. The demuxer uses duration information to place later segments in time. If a container’s stored duration is inaccurate or unavailable, a boundary can develop a gap, overlap or timestamp artefact. Listen and watch across boundaries after producing a test output; do not assume a successful command means the sequence is clean.
If the source files do not match, decide whether to prepare consistent intermediate files or use a filter graph that normalises them together. This is especially useful for a channel that combines landscape devotional footage with still-image lyric cards, or a study stream mixing video clips and music. A good workflow settles the output frame size, frame rate and audio format before a long broadcast.
Build a text file for the concat demuxer
The demuxer reads a plain-text script. A minimal playlist looks like this:
ffconcat version 1.0
file 'segment-01.mp4'
file 'segment-02.mp4'
file 'segment-03.mp4'
Save it as playlist.txt in a location where the paths resolve, or use explicit paths. For automatic format recognition, ffconcat version 1.0 must be the exact first line, without extra spaces or a byte-order mark. If FFmpeg does not recognise the header, check how the editor saved the text file and whether it inserted hidden characters.
Relative paths are resolved from the working directory used to run FFmpeg, which can surprise you when a scheduled job starts in a different folder from your interactive shell. To reduce ambiguity, either run the command from the playlist’s directory or use absolute paths. Paths with spaces and special characters need correct quoting or escaping in the script; test the exact filenames rather than assuming shell quoting rules are identical to concat-script syntax.
A basic demuxer invocation is:
ffmpeg -f concat -i playlist.txt -c copy joined-output.mkv
This illustrates reading the sequence and copying packets into an output container. It is not a universal YouTube command. The container must support the streams, the inputs must be compatible, and this command does not set a YouTube live video bitrate, keyframe interval or ingest destination.
The demuxer presents the listed packets in sequence and adjusts timestamps so the first file begins at zero and later files follow the preceding duration. The timestamp operation is global; streams of unequal lengths can leave gaps. Where the stored duration is wrong, the script can include a duration directive, but that value needs to be accurate. Adding a guessed duration can create a different boundary problem rather than fixing one.
For a practical test, start with two short files. Confirm the order, inspect or play the output, and check the transition in both picture and sound. Then test the full playlist. If the channel uses a daily rotation, keep the playlist readable and verify every path after moving or renaming source files.
Use stream copy only for matching streams
With -c copy, FFmpeg moves encoded packets to the output without decoding and re-encoding them. That can save processing and preserve the existing compressed streams. It also means FFmpeg cannot change the picture size, frame rate, pixel format, audio sample rate or volume by applying filters in that same stream-copy path.
Do not treat -c copy as a general compatibility mode. It will not reconcile a 25 fps clip with a 30 fps clip, convert one codec to another, add a missing audio stream, or normalise different resolutions. If the files need those changes, use a filter and encode the result. FFmpeg’s documentation on stream copy and filtering explains the distinction: filtering requires decoding, so filtered streams cannot simply be copied as encoded packets.
Even when stream copy is technically possible, consider the intended live output. YouTube ingest settings apply to the feed you send, not just to the files as stored. If the source encoding does not suit the chosen live configuration, copying it unchanged may leave you with a joined file that is valid but unsuitable for the broadcast. In that case, encode the output to the required live format rather than preserving the source packets.
The output container matters too. A source sequence that can be copied into one container is not necessarily suitable for another. Choose a container that supports the streams and validate the result with a probe and playback before directing it to a live endpoint.
For channels that need an unattended loop, boundary testing is as important as the initial command. A brief silence, a frozen last frame, or an audio discontinuity may be more noticeable when a video repeats all night. If a transition is unacceptable, re-export or normalise the clips rather than relying on packet copying to conceal it.
Normalise clips with the concat filter
Choose the concat filter when the inputs need a common format or when you need to apply filters. A typical graph has one input per clip, prepares corresponding video and audio streams, then joins those streams with concat. You map the graph’s output labels and encode them into a suitable output. The details depend on which streams actually exist and what properties need to be changed.
For example, if one clip is portrait and another is landscape, you first need to decide how the output should look: crop, pad, or fit the images into a shared frame. If audio sample rates or channel layouts differ, decide how to resample or map them. The filter does not make those editorial decisions for you; it gives you a place to implement them before joining the sequence.
A filter workflow requires multiple -i inputs and a -filter_complex graph. Avoid copying a generic graph without checking its stream indices, because the assumed input layout may not match your files. A clip with no audio stream cannot be mapped as if it had one. Inspect the inputs, define the common output properties, build the graph for those streams, and then encode and mux the output.
The trade-off is processing time and complexity. Decoding, filtering and encoding uses more compute than stream copy, and a careless output profile can add avoidable quality loss. Use a short test to verify the joins and output properties before processing a long programme. For a one-off set of mismatched clips, creating normalised intermediate files first can be easier to troubleshoot than maintaining a complicated live filter graph.
For a 24/7 sequence, keep a known-good output profile and document the decisions: frame size, frame rate, audio layout and codecs. When a new clip arrives, normalise it to that profile and test it at the intended join. This makes a playlist easier to maintain than discovering an unusual source file after the stream is already live.
Encode and mux for YouTube Live ingest
After concatenation, configure the output for YouTube’s live ingest requirements. YouTube’s encoder settings for live streaming list supported video codecs, frame-rate limits, bitrate guidance, keyframe timing and audio recommendations. These are YouTube recommendations, not FFmpeg defaults, and settings differ by codec, resolution and frame rate.
For RTMP or RTMPS ingest, YouTube recommends H.264, H.265/HEVC or AV1 video, up to 60 fps, constant bitrate encoding, and a keyframe interval of two seconds, with no more than four seconds between keyframes. AAC or MP3 audio is supported; the guidance for 5.1 audio over RTMP/RTMPS requires AAC. Check the current official table for the resolution and frame rate you intend to send.
As one concrete reference, YouTube’s table gives 1080p at 30 fps with H.264 a minimum bitrate of 5 Mbps and a recommended bitrate of 14 Mbps. That figure is specifically for this codec and mode; it is not a universal live bitrate. Select the row for your own codec, resolution and frame rate, and consider whether your upload connection can sustain the configured output. These figures and recommendations should be checked on YouTube’s current page before you publish.
The settings you choose for encoding and muxing are separate from the concat script. The script selects files; the encoder creates the video and audio stream in the required form, and the muxer packages them for transmission. If you used stream copy, you have not changed the encoded settings. If you need to meet different ingest settings, encode rather than copying.
YouTube’s live streaming setup instructions direct you to get the Stream URL and stream key in Live Control Room and enter them in your encoder. Treat the key like a password. Keep it out of a shared playlist, public command history, screenshots and source code. Use the current URL and key for the specific live setup, and verify them before starting the encoder.
Send the result through an encoder
The final stage is the encoder output to YouTube’s ingest destination. Whether FFmpeg reads the playlist directly while encoding, or first creates a joined intermediate file, the ingest feed remains a separate part of the workflow. The local text file does not become a YouTube playlist, and it does not create an HLS manifest.
For a short test, encode a few minutes using the intended video and audio settings and send it to a test live setup or an appropriate private workflow. Confirm YouTube receives a stable picture and sound, inspect the stream preview, and check the transition between clips. If you see a timestamp or continuity issue, return to the inputs and concat method rather than changing ingest settings at random.
For a 24/7 operation, plan what happens if the source process stops or the machine loses power. FFmpeg running locally depends on that computer, its power, its network connection and the process being supervised. The OBS versus FFmpeg guide for looping prerecorded video discusses the operational differences, while the guide to setting YouTube ingest bitrate is useful when choosing the outgoing encode.
If the recurring pain is keeping your own computer on to run a fixed video sequence, StreamNeo removes that specific burden: you upload the file, provide your YouTube stream key, and the broadcast can continue with your computer off. It is YouTube-only, so it is not a replacement for an FFmpeg workflow when you need local filtering, custom scene logic or a different destination. For a broader planning view, see scheduling a YouTube live playlist for exam preparation music and creating a store-promotion loop with FFmpeg.
Before committing, compare the operating options on the pricing page. When the file and channel are ready, start free — 24-hour trial, no card.
FAQ
Is an FFmpeg concat playlist a YouTube playlist URL?
No. It is a local text script that tells FFmpeg which files to read and in what order. YouTube receives a live encoded feed in a separate step, using the ingest URL and stream key from Live Control Room.
Can I use -c copy when clips have different resolutions or codecs?
Not as a way to normalise them. Stream copy does not decode or filter, so mismatched inputs that need conversion should go through a filter and re-encode workflow, or be prepared as compatible files first.
Why can a demuxer join have a gap at a clip boundary?
The demuxer uses stream and duration information to place later packets. Unequal stream lengths or inaccurate duration metadata can cause timestamp gaps or artefacts, so inspect durations and test the actual transitions.
Should I use the concat filter for every playlist?
No. For matching streams where packet copy is suitable, the demuxer is simpler and avoids re-encoding. Use the filter when you need to normalise, filter or re-encode, and build its graph around the streams your files actually contain.