When MP4 clips have different resolutions, first inspect their streams, then normalise the properties that need to match and join them with FFmpeg’s concat filter. This route re-encodes the clips into a consistent output; stream-copy concatenation is not a shortcut for mismatched resolutions.
After checking the joined file from start to finish, you can play it repeatedly as the source for a YouTube live stream. The right output dimensions and frame rate depend on your footage, intended delivery and ability to encode and upload reliably.
Inspect every input before choosing settings
Start by noting what each file actually contains. An MP4 is a container: its name alone does not tell you the video codec, dimensions, frame rate, pixel format, audio format or whether it has an audio stream at all. A media inspector with a stream view is enough for a small batch; FFmpeg’s ffprobe is useful when you want repeatable output you can compare across many files.
For example, run this for each input:
ffprobe -v error -show_entries stream=index,codec_type,codec_name,width,height,r_frame_rate,avg_frame_rate,pix_fmt,sample_rate,channels,channel_layout -of default=noprint_wrappers=1 first.mp4
Record the results in a small table rather than relying on memory. For video, look at width and height, reported frame rates, codec and pixel format. For audio, note whether a stream exists, its sample rate, channel count and layout. Also check the duration and play the clip: metadata will not tell you whether the sound is distorted or the picture is rotated incorrectly.
You might find one 1920 × 1080 landscape clip at 30 frames per second, a 1280 × 720 clip at a different rate, and a vertical phone recording. That does not mean every property must automatically be changed. It means you should decide which differences matter for one continuous output, and how to handle them, before joining.
Inspect for practical trouble as well as technical variation. Does a clip have no sound? Is its audio mono while the others are stereo? Does its aspect ratio differ? Is the footage already low resolution? These answers affect the filter graph and the viewing experience. If one file is suspect, check whether a video file is corrupted before adding it to an OBS playlist before treating a failed join as a settings problem.
Pick one output frame size
Choose a single output canvas for the finished video. Base it on the material you have and the delivery you want, not on the largest number in a folder. Upscaling small footage cannot restore detail that was not recorded; downscaling larger footage can make it easier to combine with smaller sources but also reduces the output dimensions.
Aspect ratio needs a deliberate choice. If you scale a vertical or square clip directly to fill a widescreen frame, you will stretch faces and objects. A more generally useful treatment is to preserve the source shape, fit it within the chosen canvas and pad the unused area. FFmpeg documents force_original_aspect_ratio=decrease for fitting within specified dimensions while retaining the original proportions, and force_divisible_by=2 for keeping dimensions divisible by two. See the FFmpeg scale filter documentation.
For a 1920 × 1080 canvas, for example, a portrait clip can be scaled down to fit and centred against a background. You can choose black padding or make a designed background in an editor; the technical point is that the video remains proportional. If the channel’s visual identity calls for a crop instead, preview it carefully because material near the edges will disappear.
A filter pattern for fitting and padding one input is:
scale=1920:1080:force_original_aspect_ratio=decrease:force_divisible_by=2,
pad=1920:1080:(ow-iw)/2:(oh-ih)/2
This is an example, not a universal target. Substitute the canvas you have chosen, and verify the result with your actual clips. If a source is already smaller than the canvas, avoid assuming that enlarging it will improve its appearance. For a channel built from still or devotional footage, a clean, consistent frame may matter more than matching a high resolution that the sources cannot support.
Decide whether frame rate and pixel format need normalising
The concat filter needs corresponding segments to have compatible properties. In practice, normalising frame rate and pixel format can make the output more predictable, especially when source clips were exported from different apps. But inspection should drive the decision: do not add transformations without a reason, and do not assume that every input requires one fixed setting.
If you want one output frame rate, choose it based on the source material and intended stream. The fps filter can convert each video segment to that rate. Conversion may duplicate or drop frames, so a mismatch is not cost-free. For a mostly static prayer-song loop, the visible effect may be modest; for footage with quick movement, inspect motion around cuts after conversion.
Pixel format is another compatibility consideration. yuv420p is commonly used for broad video compatibility, and the sample workflow below converts to it, but select output settings suitable for the files and encoder you use. setsar=1 sets the sample aspect ratio to square pixels, which is often appropriate for ordinary computer-generated output. Again, check the encoded file rather than trusting that a command completed without an error.
The important distinction is between normalising what differs and blindly applying every possible filter. If all clips already share a frame rate or pixel format, you may not need a conversion for that property. The finished output still needs a consistent stream, but your route there should follow the inspection results.
Align audio only where it is present
Video and audio need their own compatibility checks. If every clip has audio but sample rates or channel layouts differ, you can resample and convert them to one chosen format before concatenating. A stereo output is a practical choice for many music and spoken-word channels, but it is not necessarily right for every source or production.
In FFmpeg, aresample can convert sample rate and aformat can request a sample format or channel layout. In the illustrative command below, both inputs are assumed to have audio, and both are made stereo at 48 kHz. That is a working pattern to adapt, not a claim that 48 kHz stereo is mandatory for every project.
Silent clips require a different plan. If a clip has no audio stream, an audio filter chain referring to [0:a] or [1:a] will fail because that stream does not exist. You can create silence of the appropriate duration for the missing segment, or build a video-only concat output and add a separate audio bed afterwards. For clips with mixed audio presence, handle each missing segment explicitly before connecting a single audio concat chain.
Listen as well as inspect. A technical match does not prevent an abrupt loudness change, a cut-off chant or a pause that sounds unintended. If the loop uses continuous music, consider whether each clip’s audio should remain in sequence or whether a separate continuous track better suits the programme. The guidance on streaming the same song playlist 24/7 on YouTube is also relevant when your loop is built around repeated music; check your rights and the current YouTube rules for your own material.
Concatenate with the FFmpeg filter
For files with differing dimensions or other stream properties, use the concat filter as part of a re-encoding workflow. FFmpeg’s FAQ says the concat filter is recommended when re-encoding is needed. Its concat demuxer documentation describes a different path, whose inputs need matching streams, including codec and time base. A resolution mismatch is a clear reason not to assume that stream-copy concatenation will work.
Here is a representative two-input command for two files that both contain audio. It fits each video into a 1920 × 1080 canvas, converts both to 30 fps and yuv420p, prepares stereo audio at 48 kHz, then joins the segments:
ffmpeg -i first.mp4 -i second.mp4 \
-filter_complex "[0:v]scale=1920:1080:force_original_aspect_ratio=decrease:force_divisible_by=2,pad=1920:1080:(ow-iw)/2:(oh-ih)/2,fps=30,setsar=1,format=yuv420p[v0];[1:v]scale=1920:1080:force_original_aspect_ratio=decrease:force_divisible_by=2,pad=1920:1080:(ow-iw)/2:(oh-ih)/2,fps=30,setsar=1,format=yuv420p[v1];[0:a]aresample=48000,aformat=sample_fmts=fltp:channel_layouts=stereo[a0];[1:a]aresample=48000,aformat=sample_fmts=fltp:channel_layouts=stereo[a1];[v0][a0][v1][a1]concat=n=2:v=1:a=1[outv][outa]" \
-map "[outv]" -map "[outa]" \
-c:v libx264 -c:a aac -b:a 128k -movflags +faststart joined.mp4
This example makes specific assumptions: two inputs, video and audio in each, an available FFmpeg build with the named encoder, and a chosen 1080p 30 fps output. Change the input names, canvas, frame rate and audio handling to match your files. For more than two inputs, add a video and audio filter chain for every segment, then connect them in the desired order and update the concat filter’s input count.
The demuxer with stream copy may be faster and avoids another generation of encoding, but that is not an advantage if the files do not meet its compatibility requirements. The concat filter lets you normalise and join in the same workflow, at the cost of processing time and some potential generation loss from encoding. FFmpeg’s FAQ also notes that simple file-level concatenation only applies to a few containers; joining MP4 files by putting their bytes together is not a general solution.
| Method | Best suited to | Re-encoding | Main caution |
|---|---|---|---|
| Concat filter | Inputs needing normalisation or differing stream properties | Yes, for the workflow described here | Takes processing time and may reduce quality slightly |
| Concat demuxer with stream copy | Inputs whose streams already meet the documented compatibility requirements | No | Mismatched streams are unsuitable; inaccurate durations can cause timestamp artefacts |
| File-level concatenation | Limited container-specific cases | Not necessarily | Not a general way to combine MP4 files |
If you are deciding whether to run the workflow on your own computer or use a different arrangement for a long-running channel, separate the editing problem from the broadcast problem. The free ways to stream pre-recorded videos on YouTube 24/7 in India cover broader operating choices; this article’s immediate task is to prepare one consistent file.
Re-encode, then verify the joined file
Re-encoding writes a new video stream using your selected output properties and joins the filter outputs into one file. It can take time, particularly on an older computer or with a long programme. The trade-off is a predictable output stream that is better suited to the next step than a collection of incompatible clips. Keep the source clips unchanged until you have checked the result.
Open joined.mp4 in a player and watch the transitions, not just the first few seconds. Confirm that the picture stays within the intended canvas, that portrait or square footage is not stretched, and that no clip has been accidentally omitted or reordered. Listen at each cut for missing audio, a sudden change in level, a click, or a moment of silence you did not expect.
You can use ffprobe again to confirm the output’s dimensions, frame rate, pixel format and audio properties. A successful command only means FFmpeg produced a file; it does not mean that your visual choices or audio transitions are good. Check a few points near the beginning, middle and end, then watch the join between the end and beginning as a separate transition if you intend to loop it.
If the output looks soft, revisit the canvas decision rather than blindly increasing dimensions. If movement stutters, reconsider frame-rate conversion and compare against the original. If one input’s audio is too loud, correct that clip or the audio chain and render again. For a longer programme, batch-compressing wedding videos for a YouTube live loop offers a related workflow for managing source files, although compression and joining solve different problems.
Repeat the finished file for a live stream
Once the joined file has been checked, you can use it as a repeating source. FFmpeg documents -stream_loop -1 as an instruction to loop an input indefinitely. In a local FFmpeg setup, a starting pattern for sending a file to a YouTube ingest endpoint might look like this:
ffmpeg -re -stream_loop -1 -i joined.mp4 \
-c:v libx264 -pix_fmt yuv420p -r 30 -g 60 \
-b:v 5M -maxrate 5M -bufsize 10M \
-c:a aac -b:a 128k -ar 44100 -ac 2 \
-f flv 'rtmps://YOUR_INGEST_ENDPOINT/YOUR_STREAM_KEY'
Treat this as a pattern, not a command to paste unchanged. The correct ingest address and protocol are those shown for your YouTube account. Keep the stream key private, including when you share logs or screenshots. The video settings should match the output and your connection, and the bitrate values in the example are not recommendations for every resolution or network.
YouTube’s live encoder settings guidance covers supported protocols and encoder settings. It recommends RTMP or RTMPS for those workflows, constant bitrate and a two-second keyframe interval, with recommendations varying by resolution and frame rate. Its page lists 8 Mbps for 720p at 30 fps and 14 Mbps for 1080p at 30 fps for H.264; these are YouTube’s published recommendations, not a promise that a connection can sustain them. Check the current official guidance and test with your own upload connection before relying on a long broadcast.
There is also a separate choice about where the repeat operation runs. A computer that encodes and uploads continuously must remain on, connected and capable of keeping up; a local power or internet interruption can interrupt the broadcast. StreamNeo removes the need to keep your computer running for the repeat broadcast by letting you upload the prepared video once and connect it to your YouTube channel, which is useful when the recurring problem is a home machine that cannot stay on overnight. It is YouTube-only, so it does not suit a workflow whose destination is another platform.
Whichever arrangement you use, test with movement and sound similar to the intended programme, watch the whole joined file at least once, and check the live control room’s stream health after starting. Confirm that the broadcast survives the file’s end-to-start transition as intended. The YouTube live streaming help page is the place to check current platform settings, rather than relying on an old command or someone else’s bitrate.
Before committing, compare the operating options on the pricing page. When the file and channel are ready, start free — 24-hour trial, no card.
FAQ
Can I join MP4 videos with different resolutions using -c copy?
Do not assume so. Stream copying avoids re-encoding, but the concat demuxer requires compatible input streams, including matching codec and time base; different resolutions are a reason to normalise and use the concat filter instead. Inspect the files first and follow FFmpeg’s current documentation for the method you choose.
Do all clips need to be converted to 1080p and 30 fps?
No. Those are the settings used in the illustrative command, not universal requirements. Choose a canvas and frame rate suited to your source quality and intended delivery, then normalise only the properties that need to agree.
What if one of my clips has no audio?
The sample audio chain assumes every input has an audio stream, so it will not work unchanged for a silent clip. Add an appropriate silent segment or build a video-only join and handle the audio separately; then listen through all transitions to confirm the result.
How do I make the joined video repeat on YouTube Live?
A local FFmpeg stream can use -stream_loop -1 to repeat the input, with output settings adapted to your file, connection and YouTube’s current encoder guidance. Test the file’s loop boundary and monitor stream health; neither the command nor a particular bitrate guarantees an uninterrupted broadcast.