Skip to content
streamneo.
Setup Guides11 min read

What Are WebVTT Captions? A Guide to Adding Captions to Videos

Learn how WebVTT cues work, when to use captions or subtitles, and how to add and check a VTT track in HTML video.

sn.
StreamNeoPublished 4 October 2026
Worth sharing?

WebVTT is a plain-text format for timed text associated with media, not a video format. You can write cues with start and end times, then connect a .vtt file to an HTML video using a <track> element.

The important authoring choice is what the text needs to tell viewers. Captions include dialogue and relevant sounds; subtitles commonly present dialogue, often as a translation. A track’s label alone cannot make dialogue-only text communicate sounds that matter.

What WebVTT is and what it carries

WebVTT stands for Web Video Text Tracks. It carries text cues tied to playback time, and may be used for captions, subtitles, chapters, descriptions or other time-aligned information. The video and the timed text are separate: a VTT file does not contain the picture or audio.

A browser or player reads the file as a track associated with media. In an HTML page, the <track> element supplies context such as the track’s kind and language. The W3C specification says a WebVTT file must contain data of one kind, rather than mixing kinds in the same file. For example, keep captions in one file and chapter markers in another.

That distinction helps when a project grows. If you make an English caption track and a Hindi subtitle track, use separate files and identify each track clearly. Likewise, do not put chapter labels into a caption file merely because both are timed text. The file carries the cues; the surrounding page and track tell the player how those cues should be used.

The W3C WebVTT specification describes the format and its cue structure. If you are preparing video for a live or continuously played channel, captions remain a separate accessibility asset; the playlist itself does not create a caption track. For channel planning, see these faceless 24/7 channel ideas, including formats where clear spoken narration or sound context can matter.

How cues and timestamps work

A cue is a timed block. It has a start time, an arrow, an end time, and text that is displayed during that interval. The basic timestamp form is hours, minutes, seconds and milliseconds, separated by colons and a full stop before the milliseconds. A cue therefore might begin at 00:00:03.200 and finish at 00:00:05.000.

The cue boundaries do not have to match sentence boundaries exactly, but they should make reading comfortable and keep the text close to the corresponding speech or sound. If a cue stays on screen after the speaker has moved to a different idea, it can mislead. If several words appear only briefly, a viewer may not have enough time to read them.

A minimal track looks like this:

WEBVTT

00:00:01.000 --> 00:00:03.000
[Door closes]

00:00:03.200 --> 00:00:05.000
<v Maya>Hello, everyone.

The first line identifies the file. A blank line separates the header and cues, and blank lines separate cue blocks. The first cue tells the viewer about a relevant sound; the second gives a speaker label and dialogue. The example is intentionally short, but the same structure continues for the rest of the recording.

Cue settings can influence where and how text is presented. Start with plain cues unless you have a clear reason to control positioning. A complex layout can be harder to read across players and screen sizes. The W3C documentation is the reference for syntax; do not rely on visual spacing alone to establish timing, because the timestamps are what associate text with playback.

Captions and subtitles are not always the same

In common usage, captions communicate speech and relevant non-speech audio. That can include a door closing, a phone ringing, a meaningful musical cue or a change in background sound. Subtitles often show dialogue only, particularly when they translate speech for someone who can hear the original audio. Terminology varies by region and by platform, so explain what your track contains instead of assuming the name settles the question.

The distinction matters for a viewer who cannot hear the video. If a cooking demonstration says “now” while a timer beeps, dialogue alone may omit a useful cue. If a devotional reading uses a bell or chant response as part of the meaning, a caption may need to describe it. A translated subtitle track can still be useful, but it is not automatically a complete account of the audio.

For HTML, use kind="captions" when the track is intended to provide captions, and kind="subtitles" when it serves as subtitles. W3C WAI notes that what matters to accessibility is the information conveyed; a track marked as subtitles may meet a captions criterion in particular cases if it also contains the relevant audio information. That is not a reason to label every dialogue transcript a caption. Review the actual content against the viewer’s needs.

A useful test is to imagine someone watching with sound off. Would they understand who is speaking, what important sound occurred, and whether the sound changes the meaning? If not, consider what information is missing. For multilingual audiences, separate language tracks can avoid crowding a single cue with several translations. The track’s language and label should make selection straightforward.

Create and structure a VTT file

You can author a VTT file in a plain-text editor. Save it with a .vtt extension and make sure it remains plain text rather than a word-processing document that has acquired formatting. Begin with WEBVTT, add a blank line, then add each timed cue and its text. Keep the syntax consistent so a player can parse it.

A practical workflow is to work from the finished audio or video, not from memory. Draft the text while listening, note the cue boundaries, and replay each section against the picture. A rough transcript is a starting point, not a finished caption file: speech recognition can miss names, devotional terms, local pronunciations and speaker changes. Correct those details before publishing.

You may include speaker information, simple emphasis and voice spans in cue text. These can help identify the speaker or distinguish a sung phrase, but keep them restrained. Avoid adding decoration that competes with the words. If you maintain multiple languages, use one file per language and label them in a way that viewers can recognise without guessing.

Chapters or descriptive text are different kinds of timed data. Keep each kind in its own track file, rather than combining them with captions. The MDN WebVTT API guide covers browser-side text-track concepts as well as authored cues. For most non-technical channel owners, preparing a file and attaching it in the page is easier to review than generating cues in code.

If a collection or archive needs to preserve how a caption file was made, provenance may be useful alongside the cues. FADGI publishes specialised embedded metadata guidance for WebVTT. It addresses collection and administrative context; it is not a universal requirement for every web caption track.

Add the file with an HTML track

Put the .vtt file somewhere the page can load, then add a <track> inside the <video> element. For example:

<video controls>
  <source src="/media/lesson.mp4" type="video/mp4">
  <track
    kind="captions"
    src="/media/lesson.en.vtt"
    srclang="en"
    label="English captions"
    default>
</video>

Here, src points to the VTT file, srclang identifies its language, and label gives the viewer a readable choice in the player. kind indicates the track’s purpose. The optional default marks a preferred track when user preferences do not select a better match; only one track per media element may be marked default. Do not assume that adding the element will make the text visible in every player or embedding context. Test the actual player your audience will use.

For another language, add another <track> with its own source, language and label. Do not mark every track default. The HTML track element documentation on MDN describes the element’s attributes and the WebVTT format it uses. A page must also be able to retrieve the referenced file: check the URL from the page that contains the video rather than only opening the file on your own computer.

A JavaScript-created track is another route. Code can create a text track and add cue objects, which can be appropriate for an application assembling media dynamically. It adds code to maintain and debug, though, and does not remove the need to author accurate text and timing. For a lesson page or a small business product video, a separate reviewed file is often the simpler handoff.

This <track> example applies to HTML video, not to every streaming service or player. A platform may provide its own caption upload workflow, or may not expose external VTT tracks in the way your HTML player does. Check the documentation for the destination player and verify selection in practice. If your channel runs a rotating programme, the background music workflow for a YouTube live stream is a separate concern from attaching a caption file to a web video.

Check timing and text quality

Validation has two parts: the file must be readable by the player, and the content must make sense alongside the media. First check that the file begins with WEBVTT, each cue has an end time later than its start time, and the blocks are separated properly. Check that the file path works from the page where the video is embedded, and that the language and label are accurate.

Then play the video while watching the track. Look at the beginning, the transitions between speakers, pauses, and the ending. A file may parse successfully but still have cues that start late, disappear too early, or cover the wrong speaker. Listen for proper names and local terms that a transcript tool may have guessed incorrectly. Captions are intended to represent what is audible, not simply what the author expects to be said.

Check selection behaviour in the target player. Confirm that a viewer can turn the track on, distinguish it from other language tracks, and see the intended default behaviour where applicable. On a page you control, test the rendered page rather than merely opening the VTT URL. On a hosted platform, use its own current guidance and controls; the HTML technique does not establish that a platform supports an external file.

For a continuous channel, plan around the actual media segments. A caption file for one pre-recorded video does not necessarily match a playlist that changes order or inserts other clips. Keep track versions with the video versions they describe, and recheck cues if an edit changes duration or removes a scene. For operational planning beyond captions, the guide to fixing audio gaps between videos discusses continuity of sound; a caption track should be checked against that final sequence too.

Styling can be adjusted with the ::cue CSS pseudo-element, but keep contrast and readability in mind. MDN notes that ::cue-region is not supported by browsers, so do not build a workflow around region styling that the browser may not implement. The safest practical test remains reading the cues on the actual display and player, not assuming a style will travel unchanged.

Include accessibility-relevant audio

A useful caption track tells viewers about audio that contributes to the meaning, not every incidental noise. Add concise descriptions such as [door closes] or [soft music begins] where the sound helps the viewer understand the scene. Identify a speaker where the identity is not obvious, especially when voices alternate or the screen does not show the speaker.

Use judgement with music and ambience. A continuous background bed may not need to be repeated in every cue, while a musical change, lyric or bell can be significant. For a bird-sound channel, the identity or direction of a bird call may be central to the programme; for a spoken news bulletin, a quiet air-conditioner may add nothing. Captions should preserve relevant information without cluttering the screen.

Sound descriptions belong in the timed text at the point they occur. Keep the description short and distinguish it from spoken dialogue, commonly with brackets. If the audio is unclear, do not invent what happened; describe only what you can establish, or review the source recording. A viewer should be able to tell whether text is speech or a sound cue without the caption becoming a running commentary.

W3C WAI’s H95 technique for captions gives an HTML track example and discusses dialogue alongside relevant sound effects, music cues and other audio information. It is an example technique, not the only way to meet accessibility requirements. Use it to understand the information a caption track may need to carry, then check the applicable standards and platform guidance for your own context.

If the file is ready but keeping a personal computer running is the part that makes a continuous YouTube channel difficult, StreamNeo removes that particular operating burden by turning an uploaded video into a YouTube live stream that can run with your computer switched off. It does not write captions or make a VTT track appropriate for every player; prepare and check the text separately.

When the file and channel are ready, compare the operating options before choosing how to run the broadcast.

Before committing, compare the operating options on the pricing page. When the file and channel are ready, start free — 24-hour trial, no card.

FAQ

Is WebVTT a video format?

No. WebVTT is a text-based format for timed media information, with cues associated to playback times. It does not carry the video picture or audio.

Can one VTT file contain captions and chapters?

A WebVTT file should contain only one kind of data. Use separate files or tracks for captions and chapters, and identify each track’s purpose so the player and viewer can distinguish them.

Does an external VTT track work in every player?

No. The HTML <track> approach applies to compatible HTML video contexts, and other platforms may have different caption workflows or limitations. Check the destination’s current documentation and test the track in the player your viewers will use.

Should I call my track captions or subtitles?

Use the label that best describes its contents and audience. Captions include dialogue and relevant sounds, while subtitles commonly show dialogue, often in translation; terminology can vary, so make the language and purpose clear to viewers.

YOU’VE REACHED THE END

Keep the ideas coming.

More guides, useful tools and a little help for your next broadcast.

Back to the journal ↗
YOUR NEXT READ

A little more to explore.

More Setup Guides guides ↗ · All topics ↗