If a video-only clip enters an FFmpeg crossfade, create an audio stream for it and map that stream explicitly. Use xfade for the pictures, acrossfade for real audio, and anullsrc only when silence is the intended audio track.
A reliable workflow has three parts: normalise the clips before filtering, calculate each overlap from the real timeline, and encode the labelled filter outputs for YouTube. A command that writes a file is useful for learning the filtergraph, but it is not automatically a suitable 24/7 live command.
Why a video-only source needs an audio stream
A video file can be perfectly valid while containing only a video stream. That is common with slides, visualisers, screen recordings, artwork loops, and exported clips where the soundtrack was removed. The problem appears when your filtergraph expects both video and audio from every input.
xfade blends two video streams. acrossfade blends two audio streams. They are separate operations, so a video transition can work while the audio side fails because one input has no audio pad to connect. FFmpeg cannot crossfade audio that does not exist.
This distinction matters for a continuous YouTube channel. If the next item has no audio, you have to decide what the audience should hear. You could keep the previous soundtrack for longer, provide a separate music bed, mix another source, or make the silent clip genuinely silent. The method in this article covers the last choice. It does not recreate, recover, or restore missing source audio.
Before changing the command, inspect each file. Check whether it has video, audio, duration, frame rate, dimensions, pixel format, and timebase. ffprobe is useful for this inspection, although the exact command you use depends on your operating system and FFmpeg installation. A playlist that looks uniform in a media player can still contain different stream layouts.
For a broader playlist workflow, compare this with the guidance in how to loop videos in a 24/7 YouTube livestream. A loop and a crossfade have different timing requirements: a simple loop can hand one file to the next, while a crossfade overlaps two decoded timelines.
Generate silence with anullsrc
FFmpeg’s anullsrc filter generates silent audio. It is an input to the filtergraph, not a repair mechanism for a source file. When a clip has no audio, create a silent stream with the sample rate and channel layout you intend to use, then trim or otherwise constrain it so it covers the clip’s useful duration.
A simple silent input looks like this:
anullsrc=r=48000:cl=stereo
Here r sets the sample rate and cl sets the channel layout. Those values are examples for a common output choice, not a universal requirement. You should select values that fit the rest of your audio chain and final encoder settings.
For a two-input filtergraph, the silent source normally needs a duration that is long enough for the video-only clip. One conceptual arrangement is:
[1:v]setpts=PTS-STARTPTS[v1];
anullsrc=r=48000:cl=stereo,atrim=duration=CLIP_DURATION,asetpts=N/SR/TB[a1]
Replace CLIP_DURATION with the measured duration of the second clip. Do not copy a guessed duration into a long playlist. If the generated audio ends before the video, the filtergraph can terminate early or leave the streams with different useful lengths. If it runs longer, a file-oriented command may still finish because of another option, but that does not mean the timing is right for a live output.
If the first clip has audio and the second has none, connect the first clip’s actual audio to one side of acrossfade and the generated silent audio to the other. The result fades the existing sound towards silence. It does not make the second clip contain its original soundtrack.
If both clips have no audio, generate two silent streams and decide whether an audio track should be present in the final programme. A YouTube Live output with a deliberate silent track can be easier to handle than an output whose audio stream appears and disappears, but you should still test how your chosen FFmpeg build and encoder behave.
You can also choose a different design. For example, a continuous devotional or ambience channel might use a separate, licensed audio bed instead of silence. That is a content and rights decision as well as a technical one. Do not add a soundtrack merely to make the filtergraph pass if you do not have permission to use it.
Map the source video and generated audio
Explicit mapping is the part that prevents many “missing audio” surprises. A complex filtergraph produces labelled pads such as [v] and [a]; the output command must map those labels. Do not rely on FFmpeg to guess that the generated audio is the audio you intended.
For two clips, the shape is:
[0:v]...normalisation...[v0];
[1:v]...normalisation...[v1];
[v0][v1]xfade=transition=fade:duration=D:offset=O[v];
[0:a]...audio preparation...[a0];
anullsrc=r=48000:cl=stereo,atrim=duration=CLIP_DURATION,asetpts=N/SR/TB[a1];
[a0][a1]acrossfade=d=D[a]
The spaces and line breaks above are for explanation. In a shell command, quote the complete filtergraph and use the syntax appropriate to your shell. The [v] and [a] labels are arbitrary names, but they must match the -map options.
The output side then has the essential structure:
-map "[v]" -map "[a]"
This says to use the filtered video and filtered audio, rather than letting automatic stream selection pick from the original inputs. It is especially important when one input has no audio, because there may be no original audio stream for FFmpeg to select.
The video inputs to xfade must be made compatible first. FFmpeg’s filter documentation requires the two inputs to have matching constant frame rate, resolution, pixel format, and timebase. A typical preparation chain might include fps, scale, format, and settb, but the correct expressions depend on your source dimensions, aspect-ratio policy, and chosen output frame rate.
For example, conceptually:
[0:v]fps=25,scale=1280:720:force_original_aspect_ratio=decrease,
pad=1280:720:(ow-iw)/2:(oh-ih)/2,format=yuv420p,settb=AVTB,setpts=PTS-STARTPTS[v0]
That is a pattern to adapt, not a finished command for every playlist. Padding may be preferable to cropping for devotional artwork, while cropping may be preferable for footage that must fill the frame. Whatever policy you choose, apply it consistently before xfade.
Audio needs preparation too. If a real audio stream is mono and the generated stream is stereo, normalise the channel layout before acrossfade. You may also need to resample or reset timestamps. The goal is a deliberate, matching audio chain, not simply the presence of an [a] label.
Understand the example’s codecs and output options
A supplied output-file example often combines the filtergraph with options such as -shortest, -c:v libx264, -c:a aac, a pixel format, and an output filename. Each option has a job, but none should be treated as a universal live-stream recipe.
-shortest tells FFmpeg to stop when the shortest output stream ends. That can be useful when producing a finite file from clips with known durations. It can also stop a file earlier than expected if generated audio, filtered video, or another mapped stream ends first. A continuous live workflow normally needs a playlist, a repeat strategy, or a process supervisor rather than a one-shot finite output with -shortest copied unchanged.
The video codec option selects the encoder used after filtering. -c:v libx264 is a common FFmpeg choice when that encoder is present in the local build, but availability, speed, profile, level, bitrate, and hardware capacity still matter. Because xfade has decoded and changed the frames, -c:v copy cannot simply pass the filtered video through without encoding.
The same principle applies to the audio side. -c:a aac asks FFmpeg to encode the filtered audio as AAC if the encoder is available. It does not prove that AAC is configured with the channel layout, sample rate, bitrate, or timing appropriate for your output. A video-only source that receives generated silence still needs an actual audio encoder if you want an encoded audio stream.
Options such as -pix_fmt yuv420p can improve compatibility with common playback paths, but they do not replace the earlier format step in a filtergraph when xfade inputs need matching pixel formats. Likewise, an output frame-rate option does not automatically fix mismatched input timebases before the transition.
A file command may also include -movflags, a filename extension, or a container-specific option. Those choices belong to the output container and delivery method. A local MP4 test file, an MPEG-TS output, and an RTMPS live publication do not have identical requirements.
This is why the example should be read as a map of responsibilities: generate or prepare streams, filter them, map the labels, encode the results, and select an output format. It is not a promise that every listed option belongs in your final live command.
Adapt the mapping for a YouTube Live output
For YouTube, map the final filtered video and audio, encode them, and send the result to the RTMPS ingest URL and stream key shown in YouTube Live Control Room. Keep the key private. FFmpeg documents RTMPS as RTMP over a secure SSL connection in its protocol documentation, while YouTube’s encoder settings guidance describes the platform’s current ingest recommendations.
A live-oriented output has to keep producing data. Do not assume that a command which creates one correct test file will also move from clip one to clip two indefinitely. For a sequence of clips, the filtergraph must include the intended playlist, or another part of your workflow must feed successive transitions into the output.
YouTube’s current guidance lists RTMP and RTMPS ingest, H.264, HEVC, and AV1 video options, AAC or MP3 audio, and frame rates up to 60 fps. It recommends constant bitrate and a two-second keyframe interval, with the interval not exceeding four seconds. Choose the bitrate from YouTube’s current table for the codec, resolution, and frame rate you actually use rather than copying a value from an unrelated setup.
The exact FFmpeg flags depend on the encoder available in your build. With a software H.264 encoder, you might deliberately set the target frame rate, bitrate mode, video bitrate, GOP or keyframe interval, audio codec, and audio bitrate. With another encoder, the option names and supported controls may differ. Confirm the options with the documentation for the encoder installed on your machine.
The transition also changes the timeline. If the first clip lasts L1 and the video transition lasts D, the second clip begins contributing during the final D seconds of the first clip. In a chain, the next xfade offset must be calculated from the already overlapped output, not copied from the first offset. The same timing idea applies to acrossfade.
For a continuous channel, this can be more demanding than ordinary looping because FFmpeg must decode and encode overlapping material. A spare computer may handle a modest playlist, but the real test depends on resolution, frame rate, filters, encoder, and other work on the machine. If the computer must remain on all night, also consider the operational steps in how to restart an FFmpeg YouTube stream automatically on a VPS.
If leaving a computer running and recovering a dropped process is the specific pain, StreamNeo removes that part of the workflow by letting you upload the finished video, add your YouTube stream key, and let the channel run with automatic monitoring and restart, without installing software locally. It does not replace designing and testing the FFmpeg edit, and it is for YouTube output rather than a general streaming destination.
Check audio format and stream health
Use ffprobe or the media information view in your editor to verify the final result. Confirm that the output has one video stream and the intended audio stream, and check the audio codec, sample rate, channel layout, duration, and timestamps. If the source was silent by design, confirm that the output contains silence rather than assuming the presence of an audio stream means sound has been restored.
Listen through every join. A correct-looking waveform can still produce a click, a gap, a sudden level change, or a transition that begins too early. acrossfade has audio duration and curve controls, so a short fade may suit spoken announcements while a longer one may suit ambient music. The right duration depends on the material and the point at which the next clip should become clear.
Watch the video joins as well. Look for a black frame, a frozen frame, a change in aspect ratio, motion judder, or a transition that starts at the wrong point. These symptoms usually point to input normalisation, timestamps, duration calculation, or an offset problem rather than to anullsrc itself.
For a live output, check the encoder log and YouTube’s stream-health messages. YouTube advises testing before starting the live stream and monitoring stream health during the event. Its live encoder help page is the place to recheck current ingest settings because platform guidance can change.
Keep a small test playlist that represents the real channel: one clip with normal sound, one video-only clip, one clip with a different frame size, and one with motion similar to the longest or busiest material. A test using only two identical files can hide the exact mismatch that will stop an overnight stream.
If your channel carries devotional songs, local news loops, or study footage, also check the content rights and any platform restrictions separately. A technically valid stream is not evidence that every source, soundtrack, or repeated programme is suitable for publication.
Test a source that has no audio
Start with a short two-clip test, not the complete overnight playlist. Make the first clip contain audio and the second contain video only, or reverse the order so you test both directions. Generate silence for the video-only input, connect the two audio pads to acrossfade, connect the two video pads to xfade, and map [v] and [a] explicitly.
Use measured durations in your offsets. If the source reports a duration that includes a damaged tail, variable timestamps, or an unusual edit list, inspect the decoded result rather than trusting one number. The transition duration must fit inside the usable material on both sides, and the next offset must account for earlier overlaps.
Then make a local output file and inspect it from start to finish. This is where -shortest can be useful as a deliberate finite-test choice, provided you understand which stream determines the stop point. Remove or revise it for the live design if it would end the programme when a generated or auxiliary stream finishes.
Once the local test behaves correctly, run a private or otherwise controlled live test with the same resolution, frame rate, audio arrangement, encoder, and network path you plan to use. Check YouTube’s preview and health messages, listen at the transition, and leave enough time to observe whether the process keeps producing data.
Common failures have straightforward clues:
| Symptom | Likely area to inspect | Practical check |
|---|---|---|
| Filtergraph reports no audio input | One source has no audio stream | Generate a matching anullsrc stream and map it explicitly |
| Audio fades but video cuts early | Audio and video durations or offsets differ | Recalculate the overlap from measured timelines |
xfade rejects the inputs |
Frame rate, size, pixel format, or timebase differs | Normalise both video streams before xfade |
| Output ends during a file test | A stream is shorter or -shortest is active |
Inspect durations and decide whether finite output is intended |
| YouTube reports unstable health | Ingest, bitrate, encoder, network, or keyframe settings | Recheck current YouTube guidance and watch the encoder log |
| The join has a click or silence gap | Audio timestamps, layout, or fade duration differ | Prepare audio consistently and listen through the exact join |
For a channel that needs a simpler recovery routine after a network failure, how to restart a YouTube Live stream automatically after a disconnect covers the operational side. It does not remove the need to validate the filtergraph, but it helps separate media-editing faults from process-recovery faults.
Before committing, compare the operating options on the pricing page. When the file and channel are ready, start free — 24-hour trial, no card.
FAQ
Can anullsrc recover the missing audio from a video-only file?
No. It creates silence with the properties you specify. It gives the filtergraph an audio stream to process, but it cannot recreate a soundtrack that was never present in the source.
Do I need both xfade and acrossfade?
Use xfade when you want the pictures to blend. Use acrossfade when you want two audio streams to overlap smoothly. They are separate filters, so choosing one does not automatically transition the other.
Is -shortest suitable for a 24/7 YouTube stream?
Not by default. It is often useful for a finite output file, where stopping at the shortest mapped stream is intentional. A continuous live workflow needs its own playlist duration and recovery design, so test whether -shortest would stop the process earlier than you expect.
Why must the filtered outputs be mapped explicitly?
A complex filtergraph creates labelled outputs, and FFmpeg does not know which labelled video and audio pads represent your intended programme unless you map them. Using -map "[v]" -map "[a]" makes that choice explicit and prevents automatic selection from overlooking generated silence or selecting an unfiltered stream.