Skip to content
streamneo.
Streaming Settings13 min read

How to Add a Stream Title Card Between Videos with FFmpeg for YouTube Live

Build a title-card segment with FFmpeg, place it between videos, normalise the inputs, and test the finished sequence before YouTube Live.

sn.
StreamNeoPublished 4 October 2026
Worth sharing?

A title card between prerecorded videos is a separate video segment in your FFmpeg sequence. You create or load the card, make its audio and video properties compatible with the surrounding clips, then concatenate everything before sending the finished feed to YouTube Live.

YouTube does not insert the card for you. The reliable part of this workflow is separating sequence construction from live ingest: first prove that the assembled media plays correctly, then connect that output to YouTube with the stream URL and key from Live Control Room.

Choose the type of card you need

There are two practical ways to make the interstitial. You can generate a full-frame background and render text over it with FFmpeg, or you can start with a designed image and turn that still image into a timed video segment. Both can work, but they place different demands on your FFmpeg build and your preparation workflow.

A generated card is useful when the wording changes often. For example, a devotional channel might show the next programme name, a local news loop might display a location and update time, and a study channel might show the next subject. The color source supplies the background and drawtext renders the words. You can control the card's dimensions, duration, font, position and colour within the filtergraph.

An image-based card is better when your logo, typography and layout have already been prepared in a graphics application. The image can be placed over a generated background with overlay, or used as the visual source for the whole card. This keeps branding outside the command, but you still need to convert the still image into video with a defined frame rate, pixel format, duration and, where necessary, audio.

Do not treat either method as a universal command. The correct filtergraph depends on the FFmpeg build, the source codecs, the number of audio channels, the frame rates, the dimensions and the output codec you select. The official FFmpeg documentation describes the individual building blocks, but it does not make unrelated input files identical automatically.

The card should be an editorial decision as well as a technical one. Decide what the viewer needs to read, how long it must remain visible, whether it should have music or silence, and whether every transition should look the same. A simple dark background with a short, high-contrast line is often easier to read than a busy graphic, especially on a small mobile screen.

Check that your FFmpeg build can render text

The drawtext filter is not guaranteed to be present in every FFmpeg build. Its text-rendering capability depends on how FFmpeg was compiled, including support for FreeType. Font selection and shaping behaviour may also depend on optional libraries and the configuration of the build.

Before designing the whole sequence around generated text, inspect the filters available in the copy you intend to use. A command such as ffmpeg -filters can show whether drawtext is listed. You can also inspect the build information and test a very short local render using a known font file. The purpose is not to prove that every possible font feature works, but to find out whether your own executable can create the card you need.

A typical generated-card pattern contains these parts:

color=colour=black:size=WIDTHxHEIGHT:rate=FRAME_RATE,
drawtext=fontfile=/path/to/font.ttf:text='Your title':fontcolor=white:
fontsize=SIZE:x=POSITION_X:y=POSITION_Y

This is a pattern, not a complete command. Replace the dimensions, frame rate, font path, text escaping and positions for your files and build. Text containing apostrophes, colons, backslashes or other special characters may need escaping. A font path that exists on one computer may not exist on another, so use an explicit path and test it in the same environment that will perform the live encoding.

If drawtext is unavailable, an image-backed card avoids that particular dependency. Create the artwork elsewhere, export it at the intended output dimensions, and use FFmpeg to make it a timed video source. That does not remove every compatibility issue: the image still needs a defined pixel format and frame rate, and the resulting card still needs to match the other segments at the concat point.

This is also where a small local test saves a long overnight failure. A command that appears correct in a tutorial may rely on a build with libraries that your package does not include. Check the actual executable, not just the command syntax shown in an example.

Create the card as its own segment

Think of the card as a short piece of media with a beginning and an end. It should have a defined video duration, dimensions, frame rate and pixel format. If your final sequence contains audio, it should also have deliberate audio behaviour rather than an accidental gap or an undefined stream.

For a generated card, the conceptual chain is:

background source -> text rendering -> format and timing -> card video

For an image-backed card, it is usually:

still image -> timed video source or overlay -> format and timing -> card video

The phrase “timed video source” matters. A still image is not automatically equivalent to a video clip. A video sequence has frames advancing at a rate, and FFmpeg must know how long to keep producing them. Set the card duration deliberately rather than relying on the length of a source image or on an implicit filter default.

You can give the card silent audio if the surrounding concat structure expects an audio stream. This is often cleaner than allowing one segment to have audio while another has none, but the silence must have matching sample properties and the intended duration. Alternatively, you can build a video-only sequence if your complete workflow is designed that way. Decide at the sequence level, not after the stream has already reached YouTube.

If the card includes music or a voice announcement, prepare it as part of the segment. Keep its end aligned with the visual card and check that it does not overlap the next clip unexpectedly. A card that ends visually while its audio continues can sound like an unexplained gap or an early start to the next programme.

Make a short card file and inspect it before joining it to anything. Confirm that the text is visible, the margins are safe, the image is not stretched, and the duration is what you intended. This isolates card problems from concat problems.

Put the card between the clips

Once the card exists as a compatible segment, place it in the sequence in the same way as any other clip. The order is conceptually straightforward:

clip A -> title card -> clip B -> title card -> clip C

The FFmpeg concat filter is the flexible route when you are re-encoding the assembled output. FFmpeg's FAQ recommends the concat filter when re-encoding is needed, because it lets you bring streams into a common filtergraph and produce one output. The filter inputs correspond to the streams you supply, so the topology needs to be planned rather than guessed.

The concat demuxer is a different tool. It can be useful when all the files, including an already encoded card, have suitable matching streams and parameters and you want to avoid re-encoding. It is not a method for creating a card that does not yet exist, and it is not a guarantee that clips with different properties will join cleanly.

A simple decision table helps:

Situation More suitable approach Main trade-off
Text must be generated in FFmpeg Filtergraph with color and drawtext Flexible, but depends on build support and normally requires encoding
Artwork is already designed Timed image-backed segment, then concat Easier branding control, but timing and media properties still need checking
All clips and card already match Concat demuxer may be suitable Can avoid re-encoding, but offers less flexibility and strict compatibility matters
Clips differ in size, frame rate or audio Filtergraph with deliberate normalisation More control, with additional processing and testing

The sequence builder and the YouTube connection should remain separate in your thinking. First write an output file or send the assembled output to a local test destination. Only after the transitions, audio and stream properties are correct should you add the RTMP or RTMPS output and the protected stream key.

If you are deciding whether a local computer should stay on for the complete sequence, YouTube 24/7 streaming software versus cloud streaming service covers the operational difference. That decision comes after you understand what media the encoder must produce.

Match timing, dimensions and stream properties

Concat failures often come from differences that are not obvious when you play the files separately. Two clips may both look like landscape video while having different dimensions, frame rates, time bases, pixel formats or audio layouts. A title card created at a different size or with a different frame cadence adds another mismatch.

Before building the filtergraph, inspect every input. Look at the video dimensions and frame rate, then check the audio sample rate, channel layout and sample format. The exact normalisation steps depend on the files and on the output you want. Do not assume that a clip labelled “1080p” has every other property required by the next clip.

At the video side, choose one output width and height, one frame rate and one pixel format. Scale or pad clips deliberately so that a vertical source does not become unexpectedly stretched. If a source has a different aspect ratio, decide whether to crop, pad with a background, or accept a different composition. Apply the same decision to the card.

At the audio side, choose one sample rate and channel layout for the assembled output. A segment with stereo audio and a segment with mono audio may need conversion before concat. If the card is silent, create silence with the same audio characteristics as the surrounding programme rather than leaving the filtergraph to infer what should happen.

The output encoder settings then need to match the platform and your available upload capacity. YouTube's encoder guidance, accessed on 3 October 2026, lists RTMP and RTMPS ingestion, H.264, H.265 and AV1 video options, frame rates up to 60 fps, AAC or MP3 audio, and constant bitrate encoding. It recommends a two-second keyframe interval and says not to exceed four seconds. Check the current YouTube encoder settings and bitrates before choosing your final values.

Use YouTube's bitrate table for the codec, resolution and frame rate you have selected, then check that your upload connection can sustain the feed with room for ordinary variation. For example, the page lists 10 Mbps for H.264 at 1080p30 and 12 Mbps at 1080p60. Those are YouTube's platform recommendations, not a universal answer for every network or content type.

If the output is too demanding for the connection, reducing the chosen resolution or frame rate may be more useful than repeatedly changing the card. A card itself is light visually, but the live encoder still sends the complete output format throughout the broadcast. For more background on reducing file and stream demands, see how to compress long videos for YouTube Live streaming.

Handle transitions and audio deliberately

A hard cut is the simplest transition between a clip and a card. It is also the easiest to reason about when you are building a dependable overnight loop. The last frame of clip A is followed by the first frame of the card, and the card is followed by the first frame of clip B.

A fade can be useful when a hard cut feels abrupt, but it adds timing work. Decide whether the fade belongs to the end of clip A, the beginning of the card, the end of the card, or the start of clip B. Then check the actual output frame by frame or in a player that shows the timeline. A transition that is defined in more than one place can make the card shorter than intended or cause two audio sections to overlap.

Audio deserves the same attention. You have three broad choices:

  • Keep the programme audio and use silence during the card.
  • Add a separate music bed or announcement to the card.
  • Let the card carry a continuation of the surrounding audio, if that is editorially appropriate.

None of these is automatically correct. Silence may be right for a news separator and wrong for a devotional station. A music bed may make a title card feel intentional, but it can also mask an accidental timing problem. Listen to the join with headphones and speakers, because a small click or abrupt level change may not be obvious while watching the picture.

Keep the audio timeline continuous if you want a continuous broadcast. This does not mean every moment must contain music. It means the output should contain intentional samples rather than an unplanned stream disappearance. If your existing playlist has gaps, how to fix audio gaps in a 24/7 YouTube playlist stream explains why the join needs to be treated as a media problem, not only a visual one.

Do not use the card to hide an unverified source clip. If clip B has a damaged timestamp or starts with an audio offset, the card may only move the problem later in the programme. Inspect the source files and correct their timing before you build the final order.

Test the sequence before connecting YouTube Live

Render a local test containing at least one normal clip, the card and the following clip. If your real sequence has multiple source types, include each type in the test. A card that works between two similar MP4 files may still fail when placed between a variable-frame-rate recording and an audio track with a different channel layout.

Watch the complete test from before the first join until after the next clip has started. Check that the card appears for the intended duration, that text remains inside safe margins, and that the background does not flicker. Listen through the join for silence, clicks, sudden volume changes and early or late audio.

Inspect the output with a media-information tool as well as a player. Confirm that the finished file has the expected dimensions, frame rate, codec, pixel format, audio codec, sample rate and channel layout. If you see a mismatch, fix the normalisation step rather than adding another arbitrary filter at the end.

Then test the actual live path in YouTube Live Control Room. YouTube says to test before starting the live stream and to monitor stream health and messages during the event. Use the same kind of movement and audio that the real channel will produce, because a static card alone does not exercise the complete encoder path.

When you connect the encoder, use the stream URL and key shown in Live Control Room. YouTube describes the key as similar to a password, so do not place it in a public command example, a recording, a screenshot or a source repository. Prefer the RTMPS address shown by YouTube when your encoder supports it; YouTube describes RTMPS as the secure extension of RTMP. Its live streaming setup guidance explains where the connection details are provided.

Do not start with an unattended overnight broadcast. Let the test run long enough to pass through the card more than once, observe the health messages, and confirm that the output does not stop when a source clip ends. If the test is stable, save the exact tested command and media order. If you change the font, dimensions, audio layout, codec or card duration later, test again.

For a channel that should continue while your own computer is switched off, uploading the file once and keeping the card in the sequence still requires a reliable way to manage the YouTube feed. StreamNeo removes the need to leave your computer running for this particular uploaded-video workflow by running the prepared channel continuously and restarting it automatically if the feed drops.

Before committing, compare the operating options on the pricing page. When the file and channel are ready, start free — 24-hour trial, no card.

FAQ

Does YouTube add the title card between my videos?

No. YouTube receives the encoded feed that your software sends, so the card must already be part of that feed. Create it as a segment, place it between the clips, and verify the assembled output before connecting the encoder.

Can I use one FFmpeg command for any two videos?

No. The command depends on the inputs, FFmpeg build, available filters, dimensions, frame rates, pixel formats and audio properties. Treat examples as patterns, inspect your files, normalise the streams deliberately and test the result locally.

Is drawtext included in every FFmpeg installation?

No. drawtext depends on how FFmpeg was built, including FreeType support, and related font behaviour may depend on additional libraries. Check the filters in the executable you will use, or create the card as an image and turn that image into a timed video segment.

Should I use the concat filter or the concat demuxer?

Use the concat filter when you need to re-encode or normalise unlike sources, especially when the card is generated inside the filtergraph. The concat demuxer may avoid re-encoding when every file and stream already matches, but it cannot create a missing card and is less forgiving of incompatible media.

YOU’VE REACHED THE END

Keep the ideas coming.

More guides, useful tools and a little help for your next broadcast.

Back to the journal ↗
YOUR NEXT READ

A little more to explore.

More Streaming Settings guides ↗ · All topics ↗