If audio is out of sync in an FFmpeg YouTube Live product demo, first determine whether the gap stays about the same or grows as the stream runs. A stable offset and accumulating drift point to different timing questions, so do not change offsets or frame-sync settings until you have measured the mismatch.
The title alone does not establish that YouTube caused the problem, or even identify its source. Without the exact command, FFmpeg build, input devices, logs and a sample, you can follow a diagnostic path but cannot name this stream’s root cause or prescribe a guaranteed fix.
Is the mismatch fixed or growing?
Pick an event that is easy to locate in both picture and sound. In a product demo, this might be a presenter saying “now” as they click a button, a visible clap, or a clear notification sound paired with an on-screen action. Note whether the sound is ahead of or behind the matching action, then check again later in the same stream. Avoid judging from memory: write down the approximate gap at the beginning and at a later point.
If the lead or lag is roughly unchanged, you may be looking at a fixed offset. For example, the presenter’s word may consistently arrive after the visible mouth movement by about the same amount. That is a useful observation, not proof of which input or stage is responsible. It tells you an input-timestamp adjustment could be worth investigating once you have the direction and size measured.
If the gap grows, shrinks, or changes direction, do not treat it as a fixed offset problem. Accumulating drift is consistent with ongoing timing differences, timestamp behaviour, or output handling that changes over duration. Adding one fixed delay at the start may make one moment look better while leaving later moments worse.
Also distinguish audio/video sync from delivery delay. A stream can reach viewers several seconds after capture while picture and sound remain aligned with each other. Conversely, a stream can arrive promptly yet have its audio and video misaligned. Keep the question specific: does a known sound match the corresponding visible event, and does that relationship change over time?
Repeat the check at more than one point, ideally during a short controlled test before relying on the demo. If the problem only appears on the viewer-facing live output, note that separately from local capture or playback. The available evidence here does not identify any YouTube-specific sync setting or establish platform-side handling as the cause.
Capture a sample and inspect timestamps
Before editing the command, preserve what you are actually running. Copy the full FFmpeg command, noting which arguments precede each -i and which follow the inputs. Record the FFmpeg version, input types, device names, and whether audio and video come from one source or separate devices. Redact stream keys and other credentials before sharing logs or commands.
Capture a short representative sample that includes a recognisable spoken cue and visible action, and keep a note of where the same event occurs in the live output. A local capture can help reveal what the inputs and encoder produce, but it is not automatically identical to what viewers receive. If the mismatch is visible only on the live stream, retain evidence from that output too, without assuming the ingest platform is the source.
Read FFmpeg’s output rather than relying only on a player’s impression. Look for input start times, packet or frame timestamps, warnings, dropped frames, and whether timestamps jump or progress unevenly. The exact log details depend on the command and build; there is no universal line that proves an audio device is at fault. Keep an unedited log alongside any notes about the sample and the time the mismatch was observed.
FFplay can serve as an observation aid while reproducing the issue. Its documentation describes statistics for audio/video synchronization drift and the use of audio, video or external master clocks; master-clock choice is mainly useful for debugging. Consult the FFplay documentation for the current behaviour, and use the same event over time rather than a single playback instant as your comparison.
A practical record can be simple: event description, whether sound leads or lags, approximate gap at the start and later point, input arrangement, and the command and log file used. This makes it possible to compare one change at a time. For a longer-running channel, the same discipline helps separate timing trouble from unrelated loop or file handling; see the guide to keeping a playlist running when a file is missing for that separate continuity problem.
Understand FFmpeg input offsets
FFmpeg’s -itsoffset and -isync operate on input timing, but they do different things and neither repairs every possible upstream timing issue. The placement of an option matters: input options are associated with an input, while output options affect the output stage. Consult the FFmpeg documentation for the option definitions that apply to your installed version, then check the command’s ordering before experimenting.
With several inputs, first map each stream to its source. A camera may provide video while a USB microphone or mixer provides audio. A screen-capture input and an audio-capture input may have separate timestamp origins. If the picture and sound originate together in one device, that is a different arrangement from two devices whose clocks may not share a reference. Do not assume they are independent or common-clocked until you know the hardware and capture path.
Timestamp adjustment is not the same as changing the audio samples themselves. An input timestamp offset changes when that input’s streams are presented on the timeline; it does not, by itself, prove that the source timestamps were generated correctly. If an upstream device clock or timestamp sequence is unstable, a start-time adjustment may not stop the relationship from changing during the broadcast.
There is also an important boundary between what you observe locally and what a viewer receives. A file or local preview can expose input timing, while the final encoded and muxed output can behave differently. Preserve enough evidence to compare the same audiovisual event at the capture, encoded output and viewer-facing points you can access. Do not make several adjustments at once, or you will not know which one changed the result.
If you need broader context on the role of the command-line workflow, the FFmpeg versus OBS comparison for a prerecorded loop discusses when each approach suits a looping channel. That choice does not diagnose an existing sync error: keep this investigation focused on timestamps and the particular inputs in use.
When -itsoffset may help
FFmpeg documents -itsoffset offset as an input option: the specified duration is added to timestamps of the corresponding input, and a positive offset delays that input’s streams. In plain terms, this can be useful when you have measured a stable mismatch and want to shift one input’s timing relative to another. It does not tell you which input needs moving; your event test must establish which side leads and what adjustment direction is appropriate.
Place the option with the input whose timestamps you intend to adjust, then test the result through the whole demo. Do not copy an offset value from somebody else’s command or infer a precise duration from a vague impression. The required value depends on the measured direction and amount, and the correct adjustment may also depend on how the other inputs are timestamped and combined.
A controlled test changes one thing at a time. Save the original command, apply a measured input offset in a test run, and compare the same spoken cue near the beginning and later in the output. If both points improve similarly, that supports the idea that a fixed adjustment may be useful for this setup. If the first point improves but later points diverge, the remaining issue is not solved by that fixed shift alone.
Keep the expected trade-off in view: moving timestamps can align a stable lead or lag, but it can also make other streams or sections worse if they do not share the same timing relationship. Verify the result against the actual live output, not just a local monitor. If you are also preparing media for a long stream, the video compression guide for YouTube Live covers a separate preparation concern; compression settings should not be substituted for sync evidence.
What -isync requires
-isync adjusts timestamps of a target input based on the start-time difference between that input and a reference input. FFmpeg’s documentation says that, for expected results, source timestamps for the two inputs should derive from the same clock source. It also specifies that no adjustment is made if either input lacks a start timestamp. These are meaningful limits, not optional caveats.
That means -isync is not a generic “make audio follow video” switch. You first need to know which inputs are being compared, whether they have usable start times, and whether their timestamps derive from a shared clock. With separate capture devices, a common wall-clock start does not by itself establish that their media timestamps share a clock source. If you cannot establish the needed relationship, do not expect -isync to correct a growing mismatch.
Check the documentation for the exact syntax and any related input selection options for your FFmpeg version. Preserve the complete command when asking for help: a fragment showing only -isync omits the reference input and the placement that determine what it affects. Then test the output over time. A correction based on start times may be relevant to a start-time difference, but it does not prove that clocks remain aligned as capture continues.
When documenting a multi-device setup, write down which device provides each stream and any known clock relationship from the device documentation. Do not buy replacement hardware on the basis of an untested hypothesis. The source material does not establish that any particular interface or capture device is required for this unspecified demo.
Check output frame synchronisation behaviour
Input timing is only part of the path. FFmpeg’s documented video sync modes describe how output frames are handled: passthrough passes frames with their timestamps, cfr duplicates or drops frames to achieve a requested constant frame rate, vfr passes frames with timestamps or drops them to prevent duplicate timestamps, and auto selects a mode according to the output format. The FFmpeg documentation also notes that timestamps may be further modified by the muxer.
These modes govern video frame handling; they are not interchangeable fixes for bad input timestamps, and they do not directly establish an audio correction. For example, a constant-frame-rate output may repeat or discard pictures while trying to meet the requested rate. That behaviour could matter when inspecting how the final output progresses, but it does not tell you why a microphone’s timing differs from a camera’s.
Read the relevant output options in the full command and record the resulting log around the affected section. If frames are duplicated or dropped, that is evidence to consider alongside the timing measurements, not a complete diagnosis. A muxer’s timestamp handling may also affect what is observed after encoding, so test the actual output rather than relying only on input-side assumptions.
Avoid changing frame rate, sync mode and input offsets together. Make a baseline sample, alter one setting only if the evidence points to that stage, then repeat the same event check at multiple points. If the mismatch is constant, compare offset direction and measured size. If it grows, compare input clocks, timestamp progression, frame timing and output mode over the duration. This keeps a plausible setting from becoming a false explanation.
The viewer-facing stream is the final test for a live demo, but it should be compared with a local sample and logs so you can see where the behaviour first appears. Do not label YouTube as the cause solely because the problem is noticed there. For ongoing operations, StreamNeo removes the need to leave a personal computer running FFmpeg just to keep an uploaded demo file on air, but that does not diagnose or repair a sync problem in an existing FFmpeg capture command.
A measured next step
Before committing, compare the operating options on the pricing page. When the file and channel are ready, start free — 24-hour trial, no card.
FAQ
Should I try -itsoffset first?
Only after you have established that the mismatch is roughly constant, identified which side leads, and measured the gap. It adjusts timestamps for an input; it is not a general remedy for drift that keeps changing over time.
Will -isync keep separate audio and video devices aligned?
Not necessarily. FFmpeg bases it on input start times, and its documentation says expected results depend on source timestamps deriving from the same clock; if either input lacks a start timestamp, no adjustment is made. Establish the input and clock arrangement before testing it.
Is YouTube causing the audio sync problem?
The available information does not establish that. Compare a sample, logs and the viewer-facing output, and identify where the mismatch first appears before attributing it to any stage.
What details are needed for a stream-specific command?
Provide the full redacted FFmpeg command, FFmpeg version, input devices and how audio and video are connected, relevant logs, and a sample or timed event description. Include whether the gap is ahead or behind and whether it stays stable or grows; without those details, a precise command would be guesswork.