To add a black pause between videos in an FFmpeg YouTube loop, pad each clip before joining them: use tpad to add black frames and, when the clips have audio, apad to add matching silence. Then concatenate the padded video and audio segments with FFmpeg’s concat filter. This filter-and-encode workflow is not stream copying, and the command below assumes each input has one suitable video stream and one audio stream.
The gap belongs to the clips before concatenation; -stream_loop repeats material but does not create a pause. If you want only a gap between clips, leave the final clip unpadded. Inspect your inputs first, because video-only files and files with multiple audio tracks require a different filter graph and stream mapping.
What a pause between clips requires
A visible pause is not just a timestamp change between two files. It is a segment of video with a duration and a defined picture. For a black gap, FFmpeg must produce additional frames whose colour is black. The tpad video filter can add such frames to the end of a segment, or instead repeat its last frame. Those are different choices: a black card creates a clear break, while a held frame keeps the previous image on screen.
If the programme has sound, picture and sound need to be considered separately. Adding black frames does not add silence. Without an audio padding step, the audio may end before the padded picture does, or the audio from the next segment may follow with no silent interval. Applying apad to a segment extends its audio with silence, so the silent duration can match the visual gap.
In this article, “pause” means a deliberate black picture and silent audio between clips. If you want black picture while music continues, pad the video but do not add silence to the audio. If you want silence with a held image, use a clone mode rather than black frames and still decide whether the audio should continue. These are editorial choices, not effects supplied automatically by concatenation.
A pause is added to each segment that should end with one, then the padded segments are concatenated in order. For two clips and a two-second gap, pad the first clip by two seconds and leave the second alone. Padding both adds a gap at the end of the finished programme as well, which may be useful for a repeating playlist but is usually not what “between” means for a single output file.
When to use tpad and apad
Use tpad for the visual part. Its stop_mode=add setting adds frames of a solid colour at a segment’s end; color=black makes those frames black. The stop_duration value sets their duration. The filter also supports cloning the last frame instead, which can be less abrupt for a scene that should linger. See the FFmpeg filter documentation for the filter options available in your installed version.
Use apad when you want silence to extend an audio stream. It does not change the video, and tpad does not change the audio. The two filters therefore need separate branches in a filter graph, with the resulting video and audio segments paired in the concat filter. The FFmpeg filters reference documents both filters and their options; check it alongside the output of ffmpeg -filters if a command fails on an older or differently packaged build.
Neither filter solves every input-layout problem. The example later in this article addresses each input’s first video stream and first audio stream, and expects both to exist. A file with no audio cannot supply the [0:a] branch as written. A file with multiple audio tracks needs an explicit choice: selecting one track, mixing tracks, or creating silence are different outcomes. Do not paste the example unchanged and assume those cases are preserved.
There is also a choice between a filter graph and stream-copy concatenation. FFmpeg’s FAQ recommends the concat filter when re-encoding is required; this pause workflow uses filters, so the output must be encoded. The concat demuxer documentation describes a different route for reading compatible files sequentially. It is useful where no new frames or silence need to be created and the stream properties match, but it cannot create this padded gap merely by joining files.
Prepare and inspect the inputs
Before building a command, establish what each file contains and whether they can be treated alike. Inspect stream layout and duration with ffprobe, for example:
ffprobe -v error -show_entries stream=index,codec_type,codec_name,width,height,sample_rate,channels -show_entries format=duration -of json one.mp4
Run it for each input. The output tells you whether a file has audio, whether it has more than one audio stream, and whether basic video and audio properties differ. It does not prove that two streams are otherwise suitable for a single concat graph, but it catches the most obvious mismatch before you spend time encoding.
The focused command below assumes two inputs, each with one video stream and one audio stream, and a two-second black-and-silent ending on both. That makes the graph easy to inspect, but it also means the final output has a two-second tail. The command has not been tested against your files; treat it as a template and adjust stream selection, encoding options, and duration after inspection.
If you have more than two clips, the same structure can be extended: create a padded video label and audio label for every clip that should end with a gap, then list the pairs in order in the concat filter. Keep the number of input pairs consistent with n. For a long playlist, generating the graph from a checked list of files is safer than editing a long chain by eye.
For repeated or mismatched media, consider normalising the inputs before applying the pause. Differences in dimensions, frame rates, pixel formats, sample rates, channel layouts, or timestamps can make concatenation fail or produce unexpected transitions. The right normalisation depends on the source files and delivery requirements. Do not assume that this example converts every possible source into a standard format merely because the output filename ends in .mp4.
For background on the separate task of selecting and correcting audio streams, see the guide to troubleshooting FFmpeg audio errors in a YouTube podcast stream. It covers a different workflow, but the same practical point applies here: identify the streams first, then map the ones you intend to publish.
Add black video and matching silence
Here is an illustrative command for two inputs with one video and one audio stream apiece. It adds two seconds of black and silence to each clip, including the second clip’s end:
ffmpeg -i one.mp4 -i two.mp4 \
-filter_complex "[0:v]tpad=stop_mode=add:stop_duration=2:color=black[v0];[0:a]apad=pad_dur=2[a0];[1:v]tpad=stop_mode=add:stop_duration=2:color=black[v1];[1:a]apad=pad_dur=2[a1];[v0][a0][v1][a1]concat=n=2:v=1:a=1[outv][outa]" \
-map "[outv]" -map "[outa]" output.mp4
Read the graph from left to right. [0:v] selects the first input’s video stream and sends it through tpad; [v0] is the resulting padded video label. [0:a] sends the first input’s audio through apad, creating [a0]. The same steps produce [v1] and [a1] for the second input. The concat filter then takes the video/audio pair for the first segment followed by the pair for the second, and emits one video and one audio output. -map selects those outputs for the file.
To keep a pause only between the clips, remove the padding from the final input and pass its original streams into the graph. For example, the second input can be labelled directly as [1:v] and [1:a], then those labels can be used for the second concat pair. The first input still needs both padding branches. The graph becomes less symmetrical, so check that every label is defined once and used in the intended order.
For a video-only source, remove the audio padding branches, and configure concat for video only, such as concat=n=2:v=1:a=0. Map only the resulting video label. If one clip has audio and another does not, you need to decide whether to synthesise a silent audio segment or omit audio for all segments; the supplied two-stream command makes neither decision for you. For multiple audio tracks, select or combine the intended track or tracks explicitly before padding. A careless map can choose the wrong language, commentary, or music mix.
If you want a still frame rather than black, replace the padding mode with the clone mode supported by tpad. If you want the music to carry through the gap, do not pad each audio segment with silence; instead design the audio path so it remains continuous. That is not the same as leaving the audio branches as shown, because the concat operation still joins the segments at their respective boundaries.
Concatenate and encode the segments
The concat filter is appropriate here because the filter graph changes the media. It joins the already-padded segments and produces a new output, so this is not stream copying. In practice, you will normally specify codecs and, where needed, settings to make the output suitable for your editor or YouTube upload. Choose those settings based on the input, intended quality, and playback needs rather than copying a generic preset without checking the result.
The example omits explicit codec options to keep attention on the filter graph. FFmpeg chooses output defaults based on the build and container, and those defaults may not suit every use. If you need a broadly playable MP4, select compatible video and audio codecs deliberately, and confirm the output codecs with ffprobe. Do not add -c copy: filters that add frames or silence require processing and encoding those streams.
This distinction matters if you have been using the concat demuxer. That demuxer can read a list of files one after another, and can be useful for stream-copy-oriented concatenation when stream properties are compatible. FFmpeg notes constraints around matching streams and warns that duration information affects timestamps. It does not insert actual black frames or silent samples. For a gap, create the gap in the media first, then concatenate the result by an appropriate method.
If you are preparing a continuous YouTube channel rather than a single edited file, a pause inside the source playlist is only one part of the publishing plan. Decide how the finished video should repeat and whether its beginning and end make a clean loop. The guide to using a black screen and rain audio for a YouTube sleep stream discusses the presentation choices around black picture and sound; it is not a substitute for building the gap in this FFmpeg graph.
Check timing and troubleshoot output
After encoding, inspect the output with ffprobe and play the transition, not just the first few seconds. Check that the black frames last for the duration you intended, that silence begins and ends at the right points, and that the second clip starts cleanly. Also check the very end of the output: if you padded every segment, the final one has a gap too.
A duration discrepancy can come from different stream lengths or inaccurate duration metadata. The concat demuxer documentation notes that timestamps are adjusted using file durations and that differing stream lengths can create gaps; while the filter workflow is different, mismatched or unusual source timings can still complicate the result. Inspect video and audio durations separately rather than relying only on the container’s reported total duration.
If the graph reports that a stream specifier matches no streams, inspect the input again and revise the graph. A common cause is using [0:a] for a video-only file. If FFmpeg reports an unconnected output or an invalid label, check the spelling and order of each bracketed label. If the output has a black gap but no silence, verify that the audio branch is padded and that the output maps [outa], not an unintended input stream.
If the concat filter rejects the segments or the result changes size, frame rate, or sound unexpectedly, compare the inputs’ stream properties and normalise them deliberately. A filter graph that works for two matching clips is not a universal command for every collection of media. Test a short copy of the actual inputs, retain the originals, and change one part of the graph at a time so you can identify which adjustment fixed the issue.
Do not confuse adding a gap with looping an input indefinitely. FFmpeg’s -stream_loop -1 option repeats an input; it does not build a black-and-silent interval between successive clips. The gap in this example comes from tpad and apad before concatenation. See the FFmpeg command-line documentation for the loop option and confirm the syntax for your installed version.
Once the edited file is ready, you can either keep managing the live encode and restart behaviour on your own machine or use a workflow that removes that particular burden. StreamNeo takes an uploaded video and runs it as a 24/7 YouTube live stream, so you do not need to leave your computer on for that broadcast; prepare and verify this pause in the media before uploading it.
Choose the workflow that fits
The right method depends on what must change and what you want to preserve. The following comparison is about the editing step, not a claim that any route is suitable for every source file.
| Need | Approach | Main trade-off |
|---|---|---|
| Add black picture between clips | tpad each clip that should end with a gap, then concatenate with filters |
Creates new video frames and requires encoding |
| Add black picture and silence | Use tpad and matching apad on each relevant video/audio pair |
Requires suitable audio streams or a separate plan for missing tracks |
| Hold the final picture during the gap | Use tpad clone mode |
The previous frame remains visible rather than becoming black |
| Keep sound running under black picture | Pad video only and arrange audio as a continuous programme | Audio timing must be designed separately from the video segments |
| Join compatible files without adding a gap | Consider concat demuxer or concat filter depending on encoding needs | No new black or silent material is created by joining alone |
For a small devotional channel, for example, you might place a short black-and-silent interval between a recorded bhajan and a spoken introduction so the change feels deliberate. If the music should continue across the visual break, silence is the wrong choice. For a lofi station, a brief fade or a held image might be preferable to abrupt black frames; the desired transition should guide the graph.
If the files are already encoded and their stream properties match, there may be no reason to re-encode when simply joining them without a gap. But as soon as you add frames or silence, the particular filter workflow described here processes the streams. Weigh that encoding step against the editing control it gives you, and keep a copy of the original sources in case the timing needs revision.
If your goal is a dependable all-day channel, test the actual finished file and the repeat point before scheduling or streaming it. A technically valid encode can still have an awkward ending, an unintended silent tail, or a jump in loudness between clips. For broader planning around a continuous broadcast, the guide to running an always-on YouTube video stream from a cloud platform sets out operational considerations separately from this media-editing task.
Before committing, compare the operating options on the pricing page. When the file and channel are ready, start free — 24-hour trial, no card.
FAQ
Does -stream_loop create a pause between videos?
No. It repeats an input stream; it does not insert black frames or silence between files. Build the pause in the filter graph with video padding and, when required, audio padding.
Why is the final clip followed by a black screen?
The example pads both inputs, so the second one receives the same ending as the first. Leave the final segment unpadded if you want a gap only between clips, or keep it if a silent black tail is part of the intended repeat point.
Can I use the example with a file that has no audio?
Not unchanged. The example expects an audio stream for each input and includes an apad branch; inspect with ffprobe, then remove or replace the audio branch and adjust concat’s audio setting and output mapping.
Does this preserve the original encoded streams?
No. Adding frames and silence with filters requires processing and encoding the output streams. The concat demuxer is a separate option for compatible files when you want sequential joining without constructing a gap.