Yes, FFmpeg can process a YouTube playlist whose videos have different resolutions, but the right method depends on whether you mean playing items in sequence, combining them into one file, or sending a continuous live output. For a consistent picture size, decode the items, scale them to a common canvas and encode the result; stream copy is not a universal shortcut.
The distinction matters because retrieving a playlist, joining media files and broadcasting them are separate tasks. A workflow that works for local sequential playback may not produce a clean joined file or a reliable live feed without checking the source streams and the destination.
What “stream a playlist” means
People use “stream a playlist” for several different jobs. You might want a playlist-aware tool to fetch each YouTube item, combine local files into one programme, play files one after another, or feed a continuing sequence to YouTube Live. FFmpeg can be part of each workflow, but the inputs and output steps differ.
For example, a devotional channel could have older bhajan recordings at one resolution and newer videos at another. A retriever such as yt-dlp can select playlist items; FFmpeg can then process the resulting media. yt-dlp documents playlist-item and format selection, as well as writing video to standard output. Its format-selection defaults differ when output goes to stdout, so select the format explicitly if you need a repeatable process (yt-dlp documentation).
If you only want sequential playback, you may not need to join every item into a single file. If you need one output file, you need a concatenation method. If you need a live broadcast, you also need to consider the output mapping and YouTube ingest requirements. Those are different decisions from whether the original clips have the same dimensions.
This is why a guide to looping a playlist in OBS for YouTube Live can be useful when your goal is playback automation rather than preparing one joined file. The operational choice changes the points at which you need to inspect and normalise the video.
Can FFmpeg handle mixed resolutions?
Yes. Different input resolutions do not prevent FFmpeg from decoding and processing a sequence. Once the video is decoded, filters can resize or pad frames, and an encoder can produce a new output. If the output should retain each clip’s original dimensions, a workflow may preserve those differences; if it should look like one continuous programme on a fixed canvas, normalise the frames before encoding.
That flexibility has a cost. Decoding, filtering and re-encoding use more processing than copying existing compressed streams, and a lossy output codec introduces another lossy encoding step. A small channel running a long ambience loop may prefer to prepare one output file and check it before broadcast. A channel rotating items live may instead process each item as it is sent, but should test transitions and resource use on its actual machine.
There is no single choice that suits every playlist. A set of files with matching stream layouts and compatible timing may be a candidate for stream copy, while files with different codecs, time bases, audio layouts or dimensions may need a decoded workflow. FFmpeg’s concat documentation describes the requirements for the demuxer and filter approaches; inspect the actual inputs rather than inferring compatibility from their file extensions.
If your end goal is a long-running channel, the computer and playback design matter too. The comparison of a Raspberry Pi 5 and a VPS for a 24/7 FFmpeg playlist stream is relevant after you know what processing the files require. A device that can copy streams may struggle more when asked to decode, scale and encode continuously.
Why stream-copy concatenation may fail
Stream copy means FFmpeg passes compressed packets through rather than decoding and re-encoding the audio or video. It can avoid the processing and quality loss associated with a new encode. But it does not make incompatible source streams compatible. The concat demuxer expects the files to have matching streams, including codecs and time bases, and stream layouts need to line up.
Two files can both be MP4s and still differ in ways that matter. One may contain H.264 video and AAC audio, while another has a different video codec, no audio, or a different stream arrangement. Even when codecs match, a time-base difference or inaccurate duration information can affect how packets are placed at a join. A successful command on one pair of files is not evidence that every item in a playlist will behave the same way.
The demuxer also adjusts timestamps according to the durations it reads. FFmpeg warns that incorrect duration information can cause artefacts. At a transition, a duration error can show up as a pause, an abrupt jump, overlapping or missing material, or audio that no longer lines up with the picture. These symptoms are reasons to inspect the join and the file metadata, not to assume that the resolution alone caused the problem.
Use stream copy only after checking every file’s streams, codecs, time bases and duration information, and only when the files meet the concat demuxer’s requirements. If you cannot establish that compatibility, use a decode-and-filter workflow or prepare the files separately. For a 24/7 channel, a clean short test is preferable to discovering a bad join after the playlist has been running overnight.
Concat demuxer requirements and duration pitfalls
The concat demuxer reads a list of files as if they were one input. It is designed for concatenation without re-encoding when the inputs have compatible streams. The key check is not simply whether the files have the same extension or appear to have the same resolution: compare their stream counts and types, codecs, time bases, and other relevant parameters. If one item lacks audio, for instance, it does not have the same audio stream arrangement as a file that contains it.
A controlled file list can make the process easier to repeat, but it does not remove the compatibility requirement. Build the list from the exact files you intend to use, and inspect them rather than relying on a playlist’s display information. Metadata tools can report what streams are present; the point is to compare those reports across the set and confirm that each corresponding stream can be joined as expected.
Duration deserves particular attention. The demuxer uses each file’s duration to adjust timestamps for the next one. If a container’s reported duration is inaccurate, the calculated start of a later segment can be wrong. That can make a file appear to join cleanly in a quick glance while producing a gap, overlap or sync issue at a transition.
A useful test is to inspect the output around every join, including audio. Look at the end of one item and the start of the next, then listen across the same boundary. For a large playlist, sample each distinct source type and any known problem files, but do not treat a single successful sample as proof that files with different stream layouts are safe to copy. If the input set changes, repeat the checks.
If stream copy is attractive because your always-on stream already runs on a modest computer, compare it with the decode-and-encode option before choosing hardware or an operating method. The guide to setting YouTube ingest bitrate for an FFmpeg stream addresses a separate part of the live path: output bitrate. It does not remove the need to make the media itself compatible.
Using the concat filter with decoded inputs
The concat filter provides a different route. FFmpeg decodes the inputs, applies any needed filters, and joins the resulting streams. Because the frames are available to filters, you can scale them or otherwise prepare them before concatenation. This is more flexible than stream copy, but it requires the corresponding streams to have compatible parameters for the filter and a suitable output encode.
The filter’s segments need to be synchronised, and corresponding stream parameters must match. Different frame rates can be handled, with variable-frame-rate output as a possible result; that is not the same as making every other difference disappear automatically. Resolution changes that should not remain in the output need to be converted explicitly. Audio also needs attention: the example of two video inputs with audio is not directly reusable if an item has no audio or uses a different channel layout.
A typical planning sequence is to inspect the inputs, decide whether they should share one output canvas, then construct a graph that decodes each video, applies the chosen scaling and padding, and joins the video and audio streams. Finally, map the filtered outputs and encode them into the target file or output. FFmpeg’s command-line documentation explains that applying filters such as resizing requires transcoding, rather than stream copy (FFmpeg command-line documentation).
Do not treat a sample graph as a verified command for an arbitrary YouTube playlist. Real items may differ in whether they contain audio, their frame rates, pixel formats, codecs or stream layouts. Check those properties first and adapt the mappings and filters. If you have many inputs, a generated graph or a staged preparation process may be easier to maintain than manually writing a long command, but the compatibility decisions remain yours.
There is also a distinction between making a joined file and sending the sequence live. A joined file can be reviewed from beginning to end before it becomes the input to a broadcast. A direct live workflow has less room to catch an issue after the transition has gone out. For a channel whose main concern is leaving a prepared programme running, a 24/7 FFmpeg setup on Ubuntu may help frame the operating steps, but the media still needs to be suitable for the chosen workflow.
Normalize resolution before encoding
Normalisation means choosing a target frame size and converting every input to fit that canvas before the final encode. This is useful when a viewer should not see the frame size change between playlist items or when the output path expects a stable frame size. Scaling can preserve the source’s aspect ratio while fitting it within the target dimensions; padding can fill unused space so that the whole output frame remains consistent.
The target is a decision, not a universal setting. Choose it based on the material, intended presentation and encoding capacity. Scaling a lower-resolution clip up does not restore detail that was not in the source. Scaling a larger clip down can reduce the amount of picture information, but may make transitions more consistent and reduce the size of the output frame. Cropping can fill the canvas but removes part of the image, so it is not interchangeable with padding.
Normalise video before the concat filter when the filter needs corresponding video streams to have the same dimensions. A common pattern is to scale each frame so it fits inside a chosen width and height while preserving aspect ratio, then pad the remaining area. The precise filter parameters depend on the target and the source material. If the playlist includes portrait phone footage and landscape artwork, preview the result: padding may leave visible bars, while cropping may cut off text, faces or devotional imagery.
The same care applies to audio and frame rate. A resolution filter does not fix missing audio, differing channel layouts, or all timing issues. Decide whether the output can be variable frame rate or needs a target frame rate, then test the joins. Re-encoding can make the streams more consistent, but it adds processing and can reduce quality, especially if the sources have already been compressed heavily.
For a small station, an overnight test can reveal whether the chosen output is sustainable on the actual computer. Check processor load, playback smoothness, and whether transitions stay in sync over a representative run. If scaling and encoding are too demanding, prepare a normalised file in advance on a suitable machine, or revisit the target canvas and encoding choices rather than assuming stream copy will solve a mismatch.
Test the output before sending it to YouTube
Start with a representative subset of the playlist, including the smallest and largest resolutions and any files that differ in audio or codec. Inspect the streams and durations, then test the workflow you actually intend to use: stream copy, filter-based concatenation, sequential playback or live output. A test should include transitions, not only the middle of a clip where a problem is less likely to show.
Watch and listen through the joins. Check that the frame does not unexpectedly change size, that padding or crop choices look acceptable, and that audio begins and ends where expected. If the output is a file, seek near each boundary and review it. If the output is sent live, first test privately or with a controlled broadcast plan that fits your channel’s needs; confirm the current YouTube requirements rather than relying on an old command or a setting copied from another source.
FFmpeg can read network streams and write mapped outputs, but there is no universal command for every playlist and live destination. The exact approach depends on source formats, stream layouts, whether you are joining or rotating items, and the target. yt-dlp’s selection and output behaviour also needs to be made explicit when you use it to retrieve playlist items. Keep the steps reproducible: retain the selected item list and the settings that produced your tested output.
If your need is to leave a prepared video running while your own computer is off, StreamNeo removes the specific burden of keeping your local machine on to sustain that broadcast: you upload the video and provide the YouTube stream key. It does not change the need to prepare a compatible file, choose the right resolution, or check that the playlist content is suitable for your channel.
Before committing, compare the operating options on the pricing page. When the file and channel are ready, start free — 24-hour trial, no card.
FAQ
Can I concatenate videos with different resolutions using stream copy?
Do not assume so. The concat demuxer requires matching streams, including codecs and time bases, and it relies on duration information when adjusting timestamps. Check the actual inputs and test the joins; if you need a common frame size, decode and scale before encoding.
Does a mixed-resolution playlist have to be converted?
Not simply because the resolutions differ. If you want a consistent output size, scale each item to a common canvas, then encode; if you preserve dimensions, the output may retain changes between items. The choice depends on how the sequence will be played or broadcast.
Is the concat filter the same as the concat demuxer?
No. The demuxer joins compatible compressed streams without decoding when stream copy is used. The filter works on decoded streams, making filtering such as scaling possible, but the corresponding streams still need compatible parameters and suitable timing.
Can I send any YouTube playlist directly from FFmpeg to YouTube Live?
There is no one command that is reliable for every playlist and destination. Retrieval, stream compatibility, output mapping and live ingest all matter. Inspect and test the specific inputs and check YouTube’s current official guidance before broadcasting.