An audio gap between videos in an FFmpeg YouTube livestream is often easier to diagnose if you keep one FFmpeg output running across clip changes instead of stopping and starting the broadcast for each file. Whether the concat demuxer or concat filter is the better approach depends on how compatible the clips are and whether you can re-encode them.
Neither method fixes every dropout. A gap can come from clip timing, audio and video ending at different points, incompatible streams, timestamp discontinuities, or a fault farther along the encoder, network or YouTube ingest path. Start by locating the gap, then choose the method that fits your files.
Locate the gap before changing the command
First establish whether the silence happens at every clip boundary, only between particular files, or at unrelated points during playback. Note the clip names and approximate boundary time. If the same transition repeatedly causes the same gap, inspect those files and their timing first. If audio drops out mid-clip or the whole broadcast freezes, concatenation may not be the cause.
Listen to the outgoing FFmpeg signal or a local recording if your setup provides one, and compare it with YouTube’s preview or playback. The comparison helps separate a fault already present in the encoded output from one introduced later. If the local output contains silence, focus on the source, filter chain, encoder messages and timing. If it sounds intact locally but not to viewers, look beyond the clip join to the outbound connection and ingest path.
Make a small test playlist containing the transition that fails, with enough material on either side to hear whether audio ends early or resumes late. Use clips representative of the real channel: a devotional recording with a quiet tail, for example, may behave differently from a short music bed with a hard cut. A controlled test is more useful than changing several settings during a live broadcast.
If you are still deciding whether FFmpeg’s command-line workflow suits the channel, this comparison of OBS and FFmpeg for a continuous nature-sounds stream covers the operational trade-offs. The question here is narrower: whether the outgoing FFmpeg process and the files it receives preserve the intended audio across each boundary.
Keep the outgoing FFmpeg process running
A continuous live broadcast should have one outgoing FFmpeg process for the programme, with the clips changing inside that process. If the process exits after every clip and a new one starts, each restart can interrupt the output while the next process opens files, initialises encoding and reconnects to the destination. Keeping the output alive removes that particular source of boundary interruption; it does not guarantee that the files themselves join cleanly.
How you organise the inputs depends on the playlist and command you use. A concat input can present a series of files to one FFmpeg invocation, while a filter graph can combine decoded segments before encoding. In either case, check that the programme does not inadvertently launch a separate live output for each item. If a supervisor script or scheduler restarts FFmpeg at every change, inspect that control flow as well as the media files.
A single long-running process also makes diagnosis clearer. Its logs can show whether errors coincide with a clip change, and the stream can be observed across a boundary without confusing an intentional restart with a media timing problem. Keep a record of the process output and the tested file order; changing the list and command at the same time makes it harder to know what changed.
This is distinct from how the channel is hosted. If you want the stream to continue while your own computer is switched off, a hosted workflow can remove the need to keep a local machine running; the uploaded file and channel still need to be prepared correctly. A church channel planning recurring recorded services may also find this guide to streaming pre-recorded sermons continuously useful for the wider operating plan.
Check durations and where each stream ends
A media file can contain audio and video streams whose actual endings do not line up. One clip’s picture might end before its soundtrack, or audio might stop while the final video frames remain. When files are joined, those differences can affect when the next segment is placed and whether silence appears at the boundary.
Do not rely only on a player’s displayed duration. Inspect the streams in each source, and compare their durations, codecs, sample rates, channel layouts and time bases. Also check whether the files have been trimmed, remuxed or produced by different editing exports. A playlist that looks uniform in a media player can still contain differences that matter to FFmpeg’s packet-level joining.
The concat demuxer uses duration information when placing the next file. The FFmpeg documentation notes that the operation is global and may create gaps when streams do not have exactly the same length. In practical terms, a longer video stream can affect the calculated end even if the audio has already finished. Incorrect or truncated duration metadata can also produce misplaced timestamps or other artefacts.
If the duration reported for a file is wrong, FFmpeg’s demuxer documentation describes a duration directive for the concat list that can override the stored duration. Treat that as a targeted correction, not a general-purpose patch: you need to know the intended timing and verify the resulting transition. Adding an arbitrary duration can move the next clip to the wrong point rather than repair the source.
For a channel with songs or talks at different loudness levels, do not confuse a quiet transition with a gap. A separate guide to fixing loudness differences between videos covers level matching; this article concerns silence or missing audio caused by timing and stream joins.
Choose the concat method that fits the files
The two common FFmpeg approaches solve different problems. The concat demuxer joins compatible compressed streams without decoding and re-encoding them. The concat filter works on decoded media and is the documented route when you need to process or re-encode the segments. Avoid choosing solely because one command looks shorter: compatibility and whether re-encoding is acceptable matter more.
| Approach | Better fit | Important constraint |
|---|---|---|
| Concat demuxer | Files with matching streams when you want to avoid re-encoding | Depends on compatible stream properties and usable duration information |
| Concat filter | Files that need processing or are being re-encoded | Segments need to begin at timestamp zero; shorter audio can be padded for non-final segments |
| Audio crossfade | A deliberately blended musical or ambience transition | Changes the intended sound and does not validate timestamps |
The FFmpeg concat demuxer documentation explains that it reads files as though their packets had been muxed together, adjusting timestamps so later inputs follow the calculated end of earlier ones. This is useful when the files are compatible and you want to retain the existing encoded streams. It is not a repair tool for mismatched codecs, stream layouts or inaccurate timing.
The FFmpeg FAQ’s concat section recommends the concat filter when re-encoding is acceptable. The filter lets you decode and align segments, which can be useful when normalising inputs, but it comes with its own rules. In particular, each segment must start at timestamp zero. For non-final segments, the filter uses the longest stream duration and may pad a shorter audio stream with silence to align it.
That padding matters if a clip’s audio ends before its video. The filter may preserve the video duration by extending the audio side with silence rather than pulling the next clip’s audio forward. This can be the correct result when picture and sound are meant to last for the same segment, but it can sound like a gap if the source’s tail was not intended. Inspect what each stream should do before changing the filter graph.
A crossfade is another choice, but it is a creative transition rather than a timing diagnosis. FFmpeg’s acrossfade filter can blend adjoining audio, with controls for duration, curve and overlap behaviour. Use it when the programme should audibly blend, such as two continuous ambience recordings. For spoken announcements or bhajans that should end before the next begins, a crossfade may alter the content undesirably. It does not establish that timestamps are correct.
Check compatibility and timestamp continuity
Before using the demuxer, compare every file in the sequence. Its inputs need the same streams, including compatible codecs and time bases. Check that the clips have the expected number of audio and video streams, consistent channel layouts and sample rates, and matching video characteristics where relevant. Files exported from different sources may differ even if their extensions are identical.
If one clip has stereo AAC and another has a different audio codec or layout, stream-copy concatenation is not a way to reconcile them. Likewise, -c copy does not re-encode packets, change a codec into another codec, or correct a bad source duration. Either make the files compatible before joining or choose a filter workflow that decodes and produces a consistent output.
Timestamp continuity deserves separate attention. A segment may begin with a non-zero timestamp or contain discontinuities introduced during recording, trimming or remuxing. The concat filter’s requirement that segments start at zero is explicit; normalise or prepare inputs appropriately and verify the output. For the demuxer, inspect how the reported durations and calculated packet endings place the following file. A visually continuous picture does not prove the audio timestamps are continuous.
Keep a note of which source file follows which. If a gap moves when the order changes, that is useful evidence about a particular file or its duration. If it stays at a clock time regardless of playlist order, broaden the investigation to the output process, encoder and connection. This is a diagnostic clue, not a definitive test: repeat with a controlled sample and check logs.
Decide whether re-encoding is acceptable
The demuxer’s main advantage is avoiding another encode when inputs already meet its requirements. That can preserve the existing encoded streams and avoid spending time on a full re-encode. The trade-off is that the files must be suitably compatible and their timing must be trustworthy. A quick stream-copy command is not a shortcut around those prerequisites.
Re-encoding takes more processing and introduces another encoding step, but it gives you room to produce a consistent output from varied sources. That may be preferable if a playlist mixes exports with different audio formats or needs processing. Check the machine’s available capacity and test a representative sequence before relying on it for an always-on channel. Do not assume that changing to a filter automatically fixes a bad source or guarantees a clean handoff.
The choice is operational as well as technical. If preserving the current encoded streams matters and the files already match, investigate the demuxer path and correct timing details where evidence supports doing so. If normalisation is needed and the processing load is manageable, the filter path may be a better fit. If the material is intended to overlap musically, add an audio transition deliberately rather than using one to conceal an unexplained dropout.
For YouTube’s RTMP or RTMPS ingest, its encoder settings guidance lists AAC or MP3 as audio codec options, recommends 44.1 kHz for stereo and 48 kHz for 5.1 audio, and gives audio bitrate guidance. These are platform-facing output settings, not a cure for a duration mismatch in the source files. The same guidance recommends testing before going live; use the current official page when configuring a real broadcast because platform documentation can change.
Test the output and separate local faults from ingest faults
Run a test using the actual file order, output settings and transition that have caused trouble. Listen through each boundary, watch the picture, and note whether the FFmpeg process reports errors at the same moment. Keep the test long enough to include normal programme content, not just a brief black screen or static image. A silent image may not exercise the same audio path as the real stream.
Inspect a local recording or other available output from the encoder. If it contains the gap, investigate the source files, filter graph, timestamps, encoder output and machine load. If the local result is clean but YouTube playback has a gap, check the platform preview and stream health messages, then assess outbound connection strength and the ingest path. YouTube’s troubleshooting guidance for live streaming errors distinguishes encoder-side checks from connection checks; follow current guidance rather than assuming that a concat change addresses a platform problem.
YouTube’s ingest guidance also says to provide one audio stream and, for primary and backup configurations, to match their audio sample rates. Confirm those details if the channel uses a backup encoder or has multiple audio tracks in its output. This check is separate from validating that adjacent files’ timestamps and durations line up.
Change one thing at a time and repeat the same transition. First confirm a single outgoing process. Then compare stream properties and endings. Next, test the method appropriate for the inputs and inspect the local result. Finally, when the local output is healthy but the viewer’s experience is not, investigate the connection and ingest path. This order avoids treating every dropout as a concatenation fault.
If the file and channel are ready, choose how to keep the broadcast running in a way that suits your operation. StreamNeo can remove the recurring task of keeping your own computer on for an uploaded-file broadcast; it does not diagnose bad source timing or fix a network fault.
Before committing, compare the operating options on the pricing page. When the file and channel are ready, start free — 24-hour trial, no card.
FAQ
Should I use the concat demuxer or the concat filter?
Use the demuxer when the inputs meet its stream compatibility requirements and you want to avoid re-encoding. Use the filter when processing or re-encoding is acceptable and needed to handle the inputs. Check durations and timestamps whichever route you choose.
Will -c copy remove an audio gap?
Not by itself. Stream copy avoids re-encoding, but it cannot reconcile incompatible streams or make inaccurate duration information correct. Diagnose the boundary and verify the files before relying on it.
Should I add an audio crossfade?
Only when a blended transition suits the programme. A crossfade can make adjoining music or ambience sound intentional, but it changes the sound and does not repair timestamp continuity. For speech or devotional tracks with a deliberate pause, blending may be the wrong effect.
What if the local recording sounds fine but YouTube does not?
That points away from a gap already present in the local encoded output, but it does not identify the fault on its own. Check YouTube’s stream health and messages, confirm the ingest audio configuration, and investigate outbound connectivity. Changing concat methods will not necessarily address an ingest or network problem.