Skip to content
streamneo.
Tools12 min read

FFmpeg Playlist Concat Settings for Seamless 4K 60fps YouTube Live

Prepare compatible FFmpeg playlists, correct timestamps, and YouTube-aligned 4K60 output; diagnose joins rather than relying on settings alone.

sn.
StreamNeoPublished 4 October 2026
Worth sharing?

FFmpeg’s concat demuxer can play a list of files in sequence, but playlist syntax alone cannot guarantee a seamless 4K 60fps YouTube Live stream. Reliable joins depend first on compatible inputs and accurate durations, then on an output encode that meets YouTube’s current ingest guidance.

Use packet copying only when the files genuinely match; otherwise decode and normalize them into a common output. Inspect transition points and run a rehearsal, because black frames, gaps and timestamp jumps are symptoms to diagnose, not problems a particular command can promise to prevent.

Prepare a concat demuxer file list

The concat demuxer reads a text file containing one file directive per input and presents the files sequentially, adjusting their timestamps to follow one another. It is virtual packet concatenation: it does not decode clips and make unlike media compatible. That distinction matters when choosing between stream copy and a re-encode.

A basic list looks like this:

ffconcat version 1.0
file 'clip-01.mp4'
file 'clip-02.mp4'
file 'clip-03.mp4'

Put ffconcat version 1.0 exactly on the first line, without leading spaces or a byte-order mark, so FFmpeg can recognise the format automatically. Use paths that resolve from the working directory where you run FFmpeg, or provide full paths. Spaces and special characters need the escaping rules documented in the FFmpeg concat demuxer documentation; do not assume shell quoting rules are identical to playlist-file quoting rules.

For example, with a prepared list called playlist.ffconcat, a stream-copy test might begin with ffmpeg -re -f concat -i playlist.ffconcat -c copy output.ts. This is only a diagnostic shape, not a universal YouTube command: the container, codecs and output parameters still have to suit the inputs and the ingest route. The -re option reads at a real-time pace, useful when testing a live-style feed rather than processing files as quickly as possible.

By default, the demuxer’s safe option is enabled and rejects some unsafe paths or directives. Turning it off with -safe 0 can be necessary for certain controlled path layouts, but it broadens what the playlist may refer to. Use it only when you have created and checked every entry yourself; do not run an untrusted playlist with relaxed path checking.

If your aim is a repeating station, the list describes one pass, not by itself a guarantee of continuous looping. Plan how the process will restart the list or repeat it, and test the boundary between the last and first clips as carefully as the interior joins. A devotional playlist, for instance, can have individually clean transitions but still sound abrupt where the final bhajan returns to the opening track. The planning advice in how to loop Kannada Christian songs in a 24/7 YouTube stream is relevant when the order and musical boundary matter as much as the FFmpeg syntax.

Check that inputs are structurally compatible

The concat demuxer expects corresponding streams to line up: the files should have the same streams, codecs and time bases, with other relevant properties consistent enough for packet concatenation. Check stream count and ordering, video dimensions, pixel format, frame rate and time base, plus audio codec, channel layout and sample rate. A set of files can all be MP4 and still be structurally unlike.

Use ffprobe to inspect each source before building the final list. Compare its stream information rather than relying on filenames, export presets or a media player’s summary. If one clip has stereo AAC audio and another has no audio stream, or their video codecs and time bases differ, stream copy is not a sensible assumption. Re-encoding decodes the sources and creates a common output format, at the cost of processing load and another lossy generation if the chosen codec is lossy.

The simplest workflow is to separate a packet-copy candidate from a normalization candidate. If the files match and a short test confirms that the output muxer and target accept them, copying avoids a full encode. If sources differ, decode and re-encode to the same video and audio settings. For a small business’s mixed product clips, for example, one export might be variable frame rate and another fixed 30 fps; choosing a common progressive 60 fps output means both have to be transformed, not merely joined by listing them.

Also check whether the selected FFmpeg build actually includes the encoder you intend to use. Software and hardware encoders have different option names, performance limits and rate-control behaviour. A command copied from another machine may fail because an encoder is absent, or may run but fall behind real time at 3840×2160 and 60 fps. Before you commit to a long broadcast, test with the exact build and representative material. If FFmpeg itself is part of a broader radio workflow, running a YouTube radio stream with Liquidsoap and FFmpeg provides adjacent operational context.

Verify durations used for subsequent timestamps

The demuxer calculates where each next file begins using the previous file’s duration. In the FFmpeg documentation, timestamps are adjusted so that each next file starts where the previous one finishes. If a container reports a duration that is wrong, rounded, or unavailable, the calculated next timestamp can be wrong too. A pause, overlap, or timestamp discontinuity at a join may therefore originate in metadata or timing rather than a mysterious encoder defect.

Inspect each file’s reported duration and compare it with its actual end. Seeking near the tail and checking decoded frames and audio can reveal that a nominal duration does not match the material. Pay particular attention to sources with variable frame rate, incomplete recordings, unusual edit lists, or audio and video streams that end at different times. The concat demuxer documentation notes that unequal stream lengths can lead to gaps.

A playlist can include a duration directive after a file entry to override the duration FFmpeg uses for that input. Use this only when you have verified the intended duration against the media timing. An estimated or rounded value is not a repair: if it is too long, it can create an interval before the next file; if too short, timestamps may collide or the transition may behave unexpectedly. Keep a record of why an override is present, so a later replacement file does not inherit a stale value.

Do not judge duration accuracy by looking only at the displayed runtime in a player. Container metadata, stream duration and wall-clock playback can differ. Probe the streams, inspect the ending region, and test the actual concatenated output. When an error recurs at every boundary, compare durations systematically across all inputs rather than inserting arbitrary offsets until one transition looks acceptable.

Inspect joins for black frames and audio gaps

Treat a black frame, frozen picture or audio pause as evidence to investigate. First locate the exact time of the symptom and identify which two files meet there. Then inspect the last decoded frames and audio of the first file, and the first decoded frames and audio of the next. This narrows the problem to missing source content, mismatched streams, timestamp arithmetic, decode behaviour or output encoding.

A brief black picture may already be present at the end of a source clip. Check the original file in a frame-by-frame viewer and use FFmpeg filters or decoded frame inspection when visual review is inconclusive. If the source has a bad tail, repairing or trimming it before concat is more direct than changing playlist timing. Conversely, if both files look correct alone but the combined output has a gap, inspect stream properties and the timestamps around the boundary.

Audio deserves its own pass. Compare the actual end of the first audio stream with the beginning of the next; nominal file durations do not prove that sound is continuous. A music loop can have a valid digital boundary and still sound abrupt because the musical phrasing or room tone changes. For rain, chanting or ambience, listen at normal playback level as well as watching the picture. The practical listening checks in how to make rain audio sound consistent in a YouTube Live loop help distinguish a technical gap from a source transition that simply sounds unnatural.

Make a short test render around each suspect join. Confirm whether the symptom appears in the source, in a stream-copy output, or only after re-encoding. That comparison is useful: if a copy output shows it, the join, stream compatibility or timestamp model deserves attention; if only the encoded version shows it, check filters, frame-rate conversion and encoder output. Do not infer that a clean sample proves every boundary is clean, especially when the playlist has many files or includes a loop point.

Normalize output to progressive 3840×2160 at 60 fps

A 4K60 target means progressive 3840×2160 output with a 60 fps cadence. It does not mean every input must already have that format, but transforming inputs has consequences. Scaling lower-resolution clips to 3840×2160 changes the output raster, not the detail captured in the source. Converting 30 fps footage to 60 fps also does not create the motion information of footage recorded at 60 fps; the conversion may duplicate or interpolate frames depending on the filter chain.

For mixed sources, build an explicit normalization stage: decode each input, set the intended frame rate, scale to the target dimensions, and ensure compatible pixel format and audio properties before encoding. The precise filters depend on aspect ratio and content. A 16:9 clip can fill 3840×2160 directly; a portrait video needs a deliberate choice between cropping, padding or using a background. Avoid silently stretching pictures to fill the frame, and review text, faces and edge details after scaling.

A command template should make its assumptions visible rather than suggest that one line fits every FFmpeg build:

ffmpeg -re -f concat -safe 0 -i playlist.ffconcat \
  -vf "scale=3840:2160,fps=60" \
  -c:v <encoder> -b:v <codec-specific-target> -minrate <target> -maxrate <target> \
  -bufsize <encoder-appropriate-value> -g 120 -keyint_min 120 \
  -c:a aac -ar 44100 -ac 2 -b:a 128k \
  -f flv "<RTMPS-destination>/<private-stream-key>"

This is a template, not a ready-to-run prescription. Replace placeholders with settings supported by the selected encoder, verify that its rate-control options implement constant bitrate as intended, and confirm keyframe behaviour in the output. -g 120 is a nominal two-second GOP at 60 fps for encoders that express GOP length in frames; option semantics vary. The example’s audio settings reflect YouTube’s published stereo guidance for RTMP/RTMPS, but a source with a different channel arrangement needs a deliberate downmix or layout decision.

For SDR, YouTube’s advanced settings guidance specifies Rec. 709 colour and 8-bit depth. Preserve or convert colour deliberately rather than applying a generic pixel-format change that mislabels the source. HDR needs an appropriate supported codec and configuration, so do not treat an SDR example as an HDR command. After conversion, inspect actual output properties with a probe and view both dark and bright material for unexpected levels or colour shifts.

Match encoder settings to YouTube ingest guidance

YouTube’s live encoder settings guidance lists codec-specific targets for 4K/2160p at 60 fps. As listed in YouTube Help in September 2026, H.264 has a 14 Mbps minimum and 50 Mbps recommended bitrate; AV1 or H.265/HEVC has a 10 Mbps minimum and 35 Mbps recommended bitrate. These are platform recommendations, not a guarantee that your encoder or connection can sustain the stream. State the codec alongside the bitrate: a number without its codec is easy to misapply.

The same YouTube guidance recommends constant bitrate (CBR) and keyframes every two seconds, with a maximum interval of four seconds. At 60 fps, two seconds corresponds arithmetically to 120 frames. In an encoder that defines GOP length in frames, 120 is a starting point; check the encoder’s documentation and inspect the emitted keyframes, since option semantics and scene-cut behaviour can vary. Do not borrow bitrate or GOP values from a 1080p or 30 fps example without checking the 4K60 table.

YouTube supports up to 60 fps for live ingest and recommends RTMPS. For 4K/2160 streams, low-latency mode is not available; YouTube uses normal latency, as described in its latency settings guidance. Plan for that trade-off if your format depends on rapid viewer interaction. On a devotional music channel, a few seconds of extra delivery delay may be acceptable; for a live call-in discussion, it changes how you handle replies.

The bitrate target is not the same as upload headroom. A connection that briefly reaches a target may still fluctuate under household use, and an encoder that cannot encode frames in real time will fall behind even with a strong connection. Choose settings your hardware or hosted workflow can sustain, then test the full path. For more on separating encoder-side and connection-side symptoms, see how to fix dropped frames in a 24/7 FFmpeg YouTube stream.

Rehearse, monitor and choose the operating method

Before a public broadcast, run a private or unlisted rehearsal with representative clips, motion and audio. Include the joins most likely to expose differences, such as a source with a different frame rate, a long clip with a questionable duration, and the last-to-first loop boundary. Check that YouTube detects the intended 2160p60 stream and reports healthy ingest; a command completing without an error is not proof of that result.

During the test, watch the FFmpeg log for timestamp warnings, encoder lag and dropped frames, and compare those times with YouTube’s stream health indicators. If the ingest is unhealthy, change one layer at a time: first source and timing, then normalization and encoder load, then connection conditions. Keep a small record of the tested files, FFmpeg build, encoder and working settings. Repeating a known-good rehearsal is safer than changing several variables immediately before an overnight run.

There is an operational choice as well as a technical one. Running FFmpeg on your own computer gives you direct control and suits people who can monitor the machine, power and network. A computer-based setup can be the better choice when you need filters, live source switching or custom processing. If the main burden is keeping a prepared file going while your own computer is off, StreamNeo removes that specific always-on-computer task by letting you upload a video, provide your YouTube stream key and have the broadcast run and restart automatically if it drops. It is YouTube-only, so it does not fit a workflow that needs multiple platforms or live compositing.

Before committing, compare the operating options on the pricing page. When the file and channel are ready, start free — 24-hour trial, no card.

FAQ

Can I concatenate MP4 files without gaps using FFmpeg?

You can sequence compatible files with the concat demuxer, but the playlist does not eliminate gaps by itself. Check that corresponding streams match and that reported durations reflect actual timing, then inspect the resulting boundary. If inputs differ, normalize them through decoding and re-encoding.

What bitrate do I need for 4K 60fps YouTube Live?

YouTube Help’s 4K60 table gives H.264 a 14 Mbps minimum and 50 Mbps recommended target, and AV1/HEVC a 10 Mbps minimum and 35 Mbps recommended target, as listed in September 2026. Treat those as codec-specific ingest guidance, not a promise of performance; your encoder and upload connection must sustain the chosen rate.

What should my FFmpeg keyframe interval be at 60 fps?

YouTube recommends a two-second keyframe interval and says not to exceed four seconds. Two seconds at 60 fps is 120 frames, so that is a nominal GOP value for encoders whose GOP option is measured in frames. Confirm the option’s meaning and check actual output keyframes.

Does a 4K60 command guarantee that YouTube receives 4K60?

No. The source, conversion filters, encoder throughput, bitrate control, connection and YouTube ingest health all affect what arrives. Rehearse with representative material and confirm the detected resolution, frame rate and stream health in YouTube Studio.

YOU’VE REACHED THE END

Keep the ideas coming.

More guides, useful tools and a little help for your next broadcast.

Back to the journal ↗
YOUR NEXT READ

A little more to explore.

More Tools guides ↗ · All topics ↗