A black frame between videos in an FFmpeg YouTube stream is often a sign that the files are not being joined as one continuous, compatible programme. Keep one output process running and choose the concat demuxer or concat filter according to the source files and whether you need to re-encode them.
Careful preparation and testing can reduce transition problems, but no concat command guarantees a clean visual cut for every file. The right workflow depends on the streams inside your files, their timing and the output you intend to send.
Why a video switch can show black
A playlist looks like a simple sequence: finish one file, open the next, continue streaming. In practice, a live encoder and muxer are sending video and audio with timing information, and the next file has to fit into that continuing output. A pause while one process exits and another starts can leave a gap. Even without a process restart, a cut can expose a short interval with no usable picture or sound.
There is no single cause to assume from the symptom alone. A video may have a different frame rate, dimensions, codec settings or timestamp pattern from its neighbour. One file may contain audio while another does not, or their audio layouts may differ. Keyframe placement and packet timing can also matter at an edit point. These are possibilities to investigate, not a diagnosis of what YouTube did in a particular broadcast.
A black frame may also be in the source itself: for example, a fade to black at the end of a clip or a blank frame at its start. Inspect the files before changing the streaming command. If a local combined output already has a gap, focus on the files and FFmpeg workflow; if it does not, continue testing the live path and playback separately.
The distinction between joining encoded packets and decoding and re-encoding is central. FFmpeg documents separate concatenation methods, and its FAQ on concatenation recommends the concat filter when re-encoding is needed. The methods are not interchangeable: their compatibility requirements and trade-offs differ.
Keep one FFmpeg output process
For a fixed playlist, build the sequence as one continuous programme and let a single FFmpeg process send the live output. If a shell script closes one FFmpeg invocation and starts another for the next file, the output connection and encoder state are being stopped and recreated at every switch. That creates an avoidable handover and makes it harder to preserve continuity.
A useful model is: several source files enter a concat workflow, the workflow produces one video and audio output, and that output is encoded and sent continuously. The inputs may be joined before encoding with the demuxer, or decoded, arranged and re-encoded with the filter. In either case, the goal is one output timeline, not a succession of separate live sessions.
This approach suits a known sequence of prerecorded files. It does not, by itself, solve a requirement to choose arbitrary clips interactively at runtime; that is a different control problem. If your priority is a fixed playlist rather than a custom FFmpeg pipeline, the practical trade-offs in FFmpeg versus OBS for looping a playlist may help you choose the production tool before you build the sequence.
Do not restart FFmpeg at each boundary just to make a playlist easier to manage. First prepare the list and test the combined output locally. If the output is later interrupted for a separate reason, treat that as a stream recovery problem; relaunching a YouTube live stream without changing its URL covers that separate situation.
Check whether the files can be joined as packets
The concat demuxer reads a text manifest and presents listed files in sequence. It can avoid decoding and re-encoding the video, but this path is appropriate only when the input streams and formats meet its compatibility requirements and the chosen output can accept the joined packets. Files that look alike in a media player are not necessarily compatible at the stream level.
Inspect every source with ffprobe. Compare video codec and codec parameters, dimensions, frame rate, pixel format, time base and stream layout. Check whether each file has audio, and compare its codec, sample rate and channel layout. Record duration and look for unusual timestamp behaviour. This check is diagnostic; it does not establish that a join will be clean, so the output still needs a boundary test.
The FFmpeg demuxer documentation describes the manifest directives, including durations and edit points, and warns that inaccurate duration information can produce artifacts. It also notes a boundary caveat: with non-intra-frame codecs, packets or decoded frames around an inpoint or outpoint may be included. Setting an edit point in a manifest is not the same as guaranteeing a frame-perfect visual cut.
That caveat explains why packet-level joining can be efficient yet still need inspection. If the next source starts with frames that depend on earlier frames, or if a packet boundary does not correspond neatly to the visible cut you expect, the transition may not look as intended. Do not treat a plausible manifest as proof of a clean result.
Use the concat demuxer when the sources fit
Use the demuxer branch when the files are compatible for packet-level joining, you want to avoid re-encoding, and your output container and stream settings support the sequence. Create a manifest with one file entry per source, in the intended order. Use paths that resolve from the working directory where FFmpeg runs, and quote paths carefully if they contain spaces or special characters.
A minimal manifest has this shape:
file 'opening.mp4'
file 'main.mp4'
file 'closing.mp4'
A command may then use the concat demuxer as its input and copy the selected streams into an output container. This is only an outline, not a universal YouTube command: it assumes compatible inputs and does not choose a live protocol, output codec, stream key or audio policy for your setup. Before using stream copy, confirm the actual stream selection and output requirements. If your sources differ or the output needs a consistent encoded format, use the filter workflow instead.
Do not invent manifest durations to make the playlist appear continuous. If you use duration directives, base them on verified values and remember the demuxer documentation's warning that inaccurate durations can cause artifacts. Inpoint and outpoint directives are useful tools, but for inter-frame codecs their boundaries can include data from nearby frames or packets. Inspect the actual cut rather than relying on the requested timestamp alone.
Explicit stream mapping matters here as well. With several inputs or extra tracks, FFmpeg's automatic selection may not pick the video and audio you intended. The FFmpeg command-line documentation explains mapping and automatic stream selection. Choose the desired streams deliberately, especially if one source has commentary, subtitles or multiple audio tracks that should not enter the programme.
Use the concat filter when re-encoding is needed
Choose the concat filter when source streams need to be brought to a common format, when packet-level joining is unsuitable, or when you need to control the output format consistently. FFmpeg decodes the inputs, combines them through a filter graph and encodes the resulting output. The added processing costs CPU or other available encoding capacity and introduces a generation loss from re-encoding; in return, you can normalise properties that prevent a direct join.
For the filter to make a well-formed sequence, prepare inputs with compatible video and audio characteristics, including the intended resolution, frame rate and audio layout. The filter does not turn arbitrary media into a matching programme by assumption. A workflow can scale and resample inputs before concatenation, but those choices should reflect what you want the final broadcast to look and sound like, not merely silence an error message.
The command structure is a filter graph that feeds each input's video and, where present, audio into the concat filter, then maps the filter outputs to the encoder. Keep the graph explicit. A labeled filter output must be mapped once and exactly once; FFmpeg's command-line manual documents this mapping rule. If some clips have no audio, decide how the programme should behave at those sections rather than assuming an audio stream will appear automatically.
A conceptual graph might connect video and audio from two normalised inputs to a concat filter configured for two segments, then encode its video and audio outputs. The exact syntax depends on how many inputs and streams you have, how you handle missing audio, and the frame rate and resolution you choose. Test a short local render before adapting the graph to the live output. A filter graph can produce a consistent timeline, but it cannot restore picture absent from the source or guarantee that every boundary will appear black-frame-free.
Normalise and inspect the source properties
Before writing commands, make a small inventory for every file. You do not need to become a codec specialist; you need to find differences that affect joining. Use ffprobe output to compare each video and audio stream, then decide whether those differences are acceptable for packet copy or need normalisation and re-encoding.
| Property to compare | Why it matters at a boundary | Practical decision |
|---|---|---|
| Video codec and parameters | A packet-level join expects compatible streams; differing parameters can make the next segment unsuitable for the same output. | If they do not match, test the demuxer cautiously or normalise through decoding and re-encoding. |
| Dimensions and pixel format | A sequence with changing picture sizes or formats may not suit one consistent encoded output. | Pick the intended output format and prepare all clips to it when using the filter path. |
| Frame rate and time base | Different timing can affect how frames line up across an edit. | Inspect the rendered boundary and use one deliberate output timing policy. |
| Audio presence and layout | A clip may have no audio, or use a different sample rate or channel layout. | Decide whether to supply silence, omit audio consistently, or convert streams to a common layout. |
| Duration and timestamps | Incorrect duration metadata or unexpected timestamps can shift joins and create artifacts. | Verify actual duration and timing; do not guess manifest durations. |
Also inspect the first and last seconds of each source. Check whether a clip ends on black, begins with a fade, has a frozen final frame or contains a brief audio pause. If the source has a deliberate fade, removing it may not be desirable; the right edit depends on the programme. For a channel that plays devotional videos, for example, a short visual dissolve may suit the material better than a hard cut, while a news loop may need a direct cut between segments.
Audio deserves its own check. Listen across each join, not just to the middle of the clips. One file ending before its picture, a change in channel layout, or a sudden level difference can make the transition feel broken even if there is no black picture. For further background on output audio choices, see the YouTube Live audio bitrate and sample rate guide.
FFmpeg's online documentation tracks the current development version and is regenerated as the project changes. Check the documentation and options against the version installed on your machine, rather than assuming a command copied from a current web page behaves identically on an older package. The FFmpeg documentation overview explains this version distinction.
Test transitions before going live
Create a local test output that includes the end of one file and the beginning of the next. Inspect every boundary visually, frame by frame if needed, and listen through the audio change. Include boundaries with missing audio, different content types or unusual durations. A test of one easy transition is not evidence that all the joins in a long playlist behave the same way.
If a black frame appears in the local output, isolate the boundary and compare the source frames with the concatenated output. Check the manifest order, timestamps, stream selection, duration assumptions and codec compatibility. If packet-level joining is producing a troublesome boundary, try normalising and re-encoding through the concat filter, then compare the outputs. This is a reasoned troubleshooting step, not a guarantee that the filter will remove every defect.
If the local output looks correct but the live playback does not, preserve a recording or other evidence around the affected transition. Then investigate the live encoder settings, connection, stream timing and playback path as separate parts of the system. Do not infer from one viewer or one playback state that YouTube has a universal black-frame rule, or that the platform will correct a bad join.
For a 24/7 channel, test the exact playlist and FFmpeg version you plan to run, including after you replace or edit a source file. Keep a known-good copy of the manifest and command so you can compare changes. If you are deciding where the playlist should run, the guide to updating videos on a cloud-hosted YouTube stream remotely discusses the operational side of changing content without being at the streaming computer.
When the file sequence is fixed but keeping a computer running and recovering a stopped process are the main burden, StreamNeo can remove that particular workload by taking an uploaded video and running it as a continuous YouTube stream; it does not replace checking your files or guarantee a clean edit.
Before committing, compare the operating options on the pricing page. When the file and channel are ready, start free — 24-hour trial, no card.
FAQ
Should I use the concat demuxer or the concat filter?
Use the concat demuxer when the files are compatible for packet-level joining and you want to avoid re-encoding. Use the concat filter when the streams need normalising or re-encoding, accepting the extra processing and generation loss in exchange for a consistent output. Inspect and test either result.
Does a concat manifest guarantee no black frame?
No. The manifest can specify files and edit points, but FFmpeg documents boundary caveats for inter-frame codecs and warns that inaccurate duration information can cause artifacts. A clean result depends on the actual inputs and should be checked at each transition.
Why keep one FFmpeg process running?
One continuous output process avoids deliberately stopping and reopening the live output at every file change. It gives you a single sequence to encode and send, although it does not correct bad source frames or guarantee what a viewer will see.
Can I copy a command from the latest FFmpeg documentation?
Treat it as a starting point and check it against your installed FFmpeg version and the properties of your own files. The online documentation follows the developing project, and stream mapping, audio handling and filter syntax depend on your particular workflow.