Skip to content
streamneo.
Setup Guides13 min read

How to Make FFmpeg Stream Videos in Order from a Text File to YouTube

Use an FFmpeg text playlist for ordered YouTube streaming, then match media properties before filtering, encoding and testing.

sn.
StreamNeoPublished 4 October 2026
Worth sharing?

FFmpeg can play videos in the order listed in a plain text playlist and send the result to YouTube Live. For simple back-to-back playback, use the concat demuxer; for fades between clips, use separate video and audio filters and make their input properties consistent first.

The important distinction is that xfade handles video pictures, while acrossfade handles audio samples. Neither filter makes arbitrary files compatible by itself, so check the clips, choose a common output profile, test privately, and only then leave the stream running.

Use xfade for video and acrossfade for audio

A text playlist and a filtergraph solve different parts of the problem. The concat demuxer reads entries such as video-01.mp4, video-02.mp4 and video-03.mp4 in sequence. It is suitable when one clip should end and the next should begin. FFmpeg describes this demuxer as reading files one after another as though their packets had been muxed together. See the official concat demuxer documentation for the directives and compatibility requirements.

A transition needs a different arrangement. xfade takes two video streams and blends the outgoing picture into the incoming picture over a chosen duration. acrossfade does the equivalent job for audio, cross-mixing the end of one audio stream with the beginning of another. Applying xfade to video does not fade the soundtrack, and applying acrossfade does not affect the pictures.

For example, a one-second transition might use video labels like this:

[v0][v1]xfade=transition=fade:duration=1:offset=29[v01]

The audio needs its own connection:

[a0][a1]acrossfade=d=1[a01]

The offset in the video expression is measured from the start of the first video, so it depends on that clip's duration and the transition length. With several clips, you feed the output of the first xfade into the next one, and do the same with acrossfade. The offsets therefore need to account for the time already consumed by earlier clips and transitions.

This is not a drop-in replacement for the concat demuxer. A playlist containing an unknown number of files is easy to process sequentially with concat, but a filtergraph normally needs each input and each label declared. If your priority is reliable ordering rather than transitions, begin with concat and hard cuts. If fades are important, create a controlled set of inputs or normalise the files before constructing the filtergraph.

If you want an always-on music channel without maintaining a machine overnight, first understand the difference between local playback, a VPS and a managed workflow in this guide to the best free way to run a 24/7 YouTube music stream in India. The choice affects how you deal with a stopped process as much as how you build the command.

Match the properties that the filters require

The concat demuxer has its own compatibility rules, and xfade is stricter in a different way. Do not assume that two MP4 files are compatible because they have the same extension. Inspect the streams with ffprobe or another media-information tool before designing the command.

For an xfade workflow, the video inputs need matching properties such as frame rate, resolution, pixel format and timebase. A 25 fps 1920×1080 file and a 30 fps 1280×720 file should be converted to a common format before they meet at the video filter. Differences in colour range, interlacing or sample aspect ratio can also produce an output that needs inspection even after the main dimensions match.

Audio must be made consistent for acrossfade. In practice, choose one sample rate, sample format, channel layout and timebase for the audio path. A stereo 48 kHz source and a mono 44.1 kHz source should not be connected directly and treated as interchangeable. Normalising them first makes the filtergraph easier to reason about and makes the final encoder settings predictable.

The concat demuxer also expects the files to have the same streams, including codecs and time bases, when you want to avoid re-encoding. Its documentation warns that inaccurate duration information can affect the timestamps of later files and lead to artefacts. A file with a damaged or misleading duration can therefore cause trouble at the point where the next file starts, even when the playlist order is correct.

There are two practical paths:

Path What it does Main advantage Main limitation
Stream copy Passes compatible streams through without re-encoding Low processing demand and no generation loss from a new encode Requires matching streams, codecs and time bases, and does not add transitions
Transcode and filter Decodes, normalises, filters and encodes one output Produces a consistent profile and allows scaling, frame-rate conversion and fades Uses processing capacity and re-encodes the material

For a devotional playlist made from files collected over several years, the second path is usually easier to control. That is not a promise that it will fix every damaged file. It means you are deliberately converting known differences instead of asking -c copy to carry them through.

Create the playlist and verify its order

Start with a plain text file such as playlist.txt:

ffconcat version 1.0
file 'video-01.mp4'
file 'video-02.mp4'
file 'video-03.mp4'

The ffconcat version 1.0 line is optional in some situations, but when used it must be the first line exactly as shown so FFmpeg can recognise the script automatically. A file line controls the playback order. Paths can be relative to the working directory or absolute, and spaces or special characters need appropriate quoting or escaping.

Keep the playlist and media in a predictable directory. Before streaming, open every path manually or run a short FFmpeg check that reports missing files. A typo in the third entry may not appear until the first two videos have already played, which is an especially unpleasant failure for an overnight broadcast.

For trusted, simple relative paths, a basic command is:

ffmpeg -re -f concat -i playlist.txt -c copy -f flv "<YouTube ingest URL and stream key>"

If the playlist contains paths outside FFmpeg's safe subset, -safe 0 permits them:

ffmpeg -re -f concat -safe 0 -i playlist.txt -c copy -f flv "<YouTube ingest URL and stream key>"

Use that option only with a playlist you control. Do not allow untrusted text to decide which local files the process can open. The -re option reads at the media's normal rate, which is appropriate when the output must behave as a real-time live source. -c copy avoids decoding and re-encoding, but it works only when the input streams and output container are suitable.

The concat demuxer can also use a duration directive where detected duration metadata is inaccurate. Treat that as a repair for a known timing problem, not as a general way to guess durations. Check the resulting joins rather than assuming that a duration line has corrected every timestamp issue.

Build and label the filtergraph outputs

A filtered command needs explicit labels so FFmpeg knows which video and audio streams are going to the encoder. Labels are just names inside the filtergraph. For example, [0:v] means the video stream from input zero, while [v0] can be the output after scaling and format conversion.

For two inputs, a normalising stage could look like this:

[0:v]scale=1280:720:force_original_aspect_ratio=decrease,
pad=1280:720:(ow-iw)/2:(oh-ih)/2,
fps=30,format=yuv420p,settb=1/90000[v0];
[1:v]scale=1280:720:force_original_aspect_ratio=decrease,
pad=1280:720:(ow-iw)/2:(oh-ih)/2,
fps=30,format=yuv420p,settb=1/90000[v1];
[0:a]aresample=44100,aformat=sample_fmts=fltp:sample_rates=44100:channel_layouts=stereo[a0];
[1:a]aresample=44100,aformat=sample_fmts=fltp:sample_rates=44100:channel_layouts=stereo[a1]

The line breaks are for readability. In a shell command, you will normally put the filtergraph in one quoted string or use the shell's continuation syntax. The exact filters you need depend on your files. For example, padding preserves the picture without cropping, while a different scaling policy may be preferable if every source already has the same dimensions.

Once the inputs have common properties, connect the video and audio separately:

[v0][v1]xfade=transition=fade:duration=1:offset=29[vout];
[a0][a1]acrossfade=d=1[aout]

The 29 is only an illustration. It must be calculated from the actual first clip duration and the transition duration. If the first clip is not exactly the length you assumed, the transition can start too early or too late. A filtergraph that parses successfully can still produce an undesirable edit.

For three or more inputs, continue the chain with new labels. Conceptually:

[v01][v2]xfade=transition=fade:duration=1:offset=<calculated-offset>[vout];
[a01][a2]acrossfade=d=1[aout]

Do not copy these labels into a command without declaring the corresponding inputs and normalisation filters. The example shows the relationship between the filters, not a universal command for arbitrary media.

Map the filtered video and audio explicitly

After the filtergraph creates [vout] and [aout], map those labels to the output. Explicit mapping prevents FFmpeg from selecting an unintended original stream, an attached picture or an extra audio track.

A filtered output section has this shape:

-map "[vout]" -map "[aout]"

If a source has no audio, acrossfade cannot use it as though it contained silence. You need to add a deliberate silent audio source or choose a playlist policy that rejects clips without audio. Similarly, if one file has multiple video streams, identify the intended stream rather than relying on automatic selection.

Mapping is also useful when you are diagnosing a failed command. If the output says that a label does not exist, inspect the spelling and the filtergraph connections. If the stream is present but silent, confirm that the audio label is mapped and that the audio encoder accepts the format produced by the filter.

For a simple concat command with no filters, explicit mapping is less important, but it can still make the output predictable. With a filtered command, consider it part of the design rather than an optional clean-up step. The final output should have one intentional video path and one intentional audio path.

Encode the filtered result

Filtering requires decoding, so -c copy is no longer appropriate for the filtered streams. A practical starting point for an SDR 720p30 output is:

ffmpeg -re -f concat -safe 0 -i playlist.txt \\
  -vf "scale=1280:720:force_original_aspect_ratio=decrease,pad=1280:720:(ow-iw)/2:(oh-ih)/2,fps=30,format=yuv420p" \\
  -c:v libx264 -preset veryfast -b:v 8M -maxrate 8M -bufsize 16M \\
  -g 60 -keyint_min 60 -sc_threshold 0 \\
  -c:a aac -b:a 128k -ar 44100 \\
  -f flv "<YouTube ingest URL and stream key>"

This is an illustrative starting point, not a tested command for your files or your FFmpeg build. It does not add xfade or acrossfade; those belong in a -filter_complex graph with separately labelled inputs. The video filter shown here only normalises a single concat output before encoding.

YouTube's live encoder settings guidance recommends CBR, a two-second keyframe interval and a keyframe interval no longer than four seconds. For H.264 at 720p30, the page lists 3 Mbps as a minimum and 8 Mbps as recommended. For 1080p30 H.264, it lists 5 Mbps minimum and 14 Mbps recommended. These figures are tied to the stated resolution, frame rate and codec, not universal settings for every channel.

At 30 frames per second, -g 60 represents a two-second GOP. If you choose another frame rate, calculate the GOP from that frame rate instead. Check YouTube's current table for the output profile you actually intend to use, and make sure your upload capacity can sustain the combined video and audio output without repeated congestion.

AAC is a commonly used audio choice for this workflow, but the encoder settings still need to match your source material and YouTube's current guidance. If the machine cannot encode the chosen profile in real time, lower the workload by changing the output profile or use a different hosting arrangement. A command that produces a good file but runs slower than real time will not maintain a live broadcast.

Apply YouTube ingest guidance

In YouTube Live Control Room, create or select the live stream and copy the stream URL and stream key shown for the encoder. Keep the key private. YouTube's encoder setup instructions explain where those values are entered and which live-stream details must be configured in Studio.

Use the exact ingest address YouTube displays rather than inventing a suffix or copying one from an old setup. YouTube supports RTMP and RTMPS, and its guidance recommends RTMPS when supported. The protocol choice is separate from the playlist and filtergraph: FFmpeg still has to produce a stream format that the selected ingest accepts.

First-time live activation may take up to 24 hours according to YouTube's help guidance. Do this account step before the planned broadcast rather than discovering it when the playlist is ready. If the event is unlisted or private for testing, check the visibility setting before you begin.

For a longer discussion of transport choices, see RTMP vs RTMPS vs SRT for always-on streams. It is also worth separating ingest transport from delay settings. A secure transport choice does not, by itself, determine the latency your viewers experience.

Test representative clips and monitor stream health

Do not test only the first file. Choose a short sample that includes movement, quiet audio, louder audio, a typical transition and the largest or least conventional source you plan to use. YouTube specifically advises testing before starting a live stream and recommends a test that includes audio and movement.

A useful test sequence is:

  1. Validate every playlist path and confirm the listed order.
  2. Inspect frame rate, dimensions, pixel format, timebase, sample rate and channel layout.
  3. Run the command for a short period and check that it stays at real time rather than falling behind.
  4. Use an unlisted or private event where appropriate and watch the preview in Live Control Room.
  5. Check both picture and audio at the joins, including the first transition and the return to normal playback.
  6. Read the stream health messages while the test is running.

YouTube's live stream health guidance should be your reference when the preview shows warnings or the incoming bitrate is unstable. A local FFmpeg log can show a missing file or encoder error, while Live Control Room can show problems that occur after the output reaches YouTube. You need both views to diagnose the full path.

If the video freezes at a join, compare the source properties and duration metadata before changing random encoder flags. If the audio clicks or disappears, inspect the audio layouts and the acrossfade connections. If YouTube reports an unstable connection, compare the encoder's actual output with your upload capacity rather than assuming that changing the transition filter will help.

A local computer also introduces failure points: sleep settings, a changing network connection, a terminal window being closed and a process that stops after an error. For creators who want the uploaded file and channel to continue without leaving their own computer switched on, StreamNeo removes the particular burden of keeping the local FFmpeg process running and watching for a restart, while you still remain responsible for the media, channel settings and YouTube checks.

If you are running the command on a VPS, compare its operational trade-offs with this guide to running a 24/7 ASMR stream from a VPS. If the process does stop, the recovery steps in how to reconnect a podcast stream to YouTube after an internet outage are relevant even when your content is not a podcast.

Before committing, compare the operating options on the pricing page. When the file and channel are ready, start free — 24-hour trial, no card.

FAQ

Can I use -c copy with every text playlist?

No. Stream copy depends on compatible stream layouts, codecs and time bases, and it does not create transitions. If the files differ or the joins show artefacts, transcode them to a common output instead of assuming that -c copy will work.

Does xfade fade the audio as well?

No. xfade operates on video frames. Use acrossfade for the audio path, and make the audio sample rate, format and channel layout consistent before connecting the streams.

Can a playlist file automatically create fades between any number of videos?

Not by itself. The concat demuxer is designed to present files sequentially, while a fade filtergraph needs declared inputs, labels and calculated offsets. For a changing playlist, hard cuts are simpler, or you can normalise and generate a filtergraph deliberately.

What should I check before making the stream public?

Confirm the order and paths, inspect the input properties, test clips with movement and audio, and check the YouTube preview and stream-health messages. Also verify the current ingest URL, key, codec and bitrate guidance in YouTube Studio before going live.

YOU’VE REACHED THE END

Keep the ideas coming.

More guides, useful tools and a little help for your next broadcast.

Back to the journal ↗
YOUR NEXT READ

A little more to explore.

More Setup Guides guides ↗ · All topics ↗