Skip to content
streamneo.
Setup Guides11 min read

How to Set Up a systemd Watchdog for a 24/7 YouTube Stream on an Indian Cloud VM

Learn what systemd watchdogs do, how to restart an encoder safely, and how to test recovery without mistaking a VM reboot for stream health.

sn.
StreamNeoPublished 4 October 2026
Worth sharing?

To keep a YouTube stream running on an Indian cloud VM, run the encoder as a systemd service and choose recovery behaviour that matches how it fails. WatchdogSec= only supervises a service that sends systemd watchdog notifications; adding it to a unit does not make systemd monitor an ordinary FFmpeg process.

A hardware or virtual machine watchdog is a separate mechanism: it can reboot a guest that stops responding, if the provider exposes a watchdog device. Neither mechanism proves that YouTube is receiving a healthy stream. You need to check the encoder, the network path and YouTube's stream health separately.

Two mechanisms share the name watchdog

A service watchdog is configured per service with WatchdogSec=. Once the service has completed startup, systemd expects regular WATCHDOG=1 notifications through its notification socket. If a notification-capable process misses the deadline, systemd treats the service as failed. The configured Restart= policy then determines whether it starts again.

An arbitrary encoder does not send these notifications simply because it is started by systemd. A plain FFmpeg command can be restarted when it exits abnormally with Restart=on-failure, but it does not become watchdog-supervised when it hangs or stops producing useful output. For that case, a wrapper or supervisor must send notifications based on meaningful checks of the encoder.

RuntimeWatchdogSec= belongs to the system manager rather than an individual service. It asks systemd to use a watchdog device, commonly represented in the guest as /dev/watchdog or /dev/watchdog0, so that the machine can be rebooted if the manager stops servicing it. A cloud VM may not expose such a device, and the setting has no effect without one. Even when available, a system reboot is a broad recovery action, not a signal that a YouTube broadcast is healthy.

Mechanism Scope Required support What recovery can mean
Restart=on-failure One service The process exits with a failure, or systemd marks it failed Restart that service
WatchdogSec= One service The process or wrapper sends timely systemd notifications Mark service failed on missed notifications, then apply restart policy
RuntimeWatchdogSec= The system manager and VM A watchdog device exposed to the guest Reboot an unresponsive machine

That distinction answers a common question: WatchdogSec does not restart a plain encoder merely because the encoder is stuck. First establish what the encoder supports, then choose a supervisor that can observe the failure you care about.

Check whether the encoder can notify systemd

Check the encoder's documentation and build, not just its name or command line. The relevant question is whether the launched service process implements systemd's notification protocol and sends WATCHDOG=1 at regular intervals. If it does, confirm that systemd can pass it the notification socket and that the watchdog support is actually enabled for this invocation.

A service may advertise Type=notify, which is commonly used for readiness notifications, but that alone is not proof it sends watchdog keep-alives. Likewise, a program that logs progress or prints frames is not necessarily talking to systemd. Consult the installed systemd service documentation and the encoder's documentation for the version on your VM. Directives and behaviour can vary with systemd version and distribution.

For a simple FFmpeg service, assume there is no systemd watchdog support unless you have confirmed otherwise. You can still use systemd to start it at boot, collect its logs, and restart it after a failure exit. If you need to detect a process that remains alive but is no longer encoding or transmitting, add a supervisor that checks for that condition rather than relying on WatchdogSec= alone.

Before changing the unit, identify the service account, working directory, exact command, input files and credential handling. For the YouTube side, create or select a broadcast in the Live Control Room and provide the server URL and stream key to the encoder. Treat the key as a password: avoid putting it in a command visible to other users, world-readable configuration, or logs. If it is exposed, reset it in Studio and update the service configuration.

Choose a wrapper that can report useful state

If the encoder has no notification support, you have two sensible starting points. Keep the setup simple with a normal systemd service and restart on failure, or run it under a wrapper or supervisor that can detect loss of useful work and notify systemd. The right choice depends on whether your actual failure is a process exit, a hang, a broken input, a network interruption, or a failed YouTube ingest.

A notification-capable wrapper should do more than emit WATCHDOG=1 on a timer. It should stop notifying when the child encoder exits, and it should have a health check relevant to your channel. For example, if a wrapper observes that the encoder's output progress has stopped advancing, or that the child process has died, it can stop its heartbeat or terminate the child so systemd can recover it. A wrapper that continues to ping while its child is dead or wedged hides the very failure the watchdog is meant to catch.

Systemd makes the watchdog interval available to a service as WATCHDOG_USEC= when watchdog supervision is configured. The systemd notification API advises that keep-alive messages should usually arrive around halfway through the timeout interval, leaving some margin for scheduling delays. Use the manual for the installed systemd release, and test the wrapper's behaviour before trusting it unattended.

A shell loop that blindly pings systemd every few seconds is not a health check. It can make systemd believe everything is well even when FFmpeg has stopped reading its source, lost its connection, or is producing invalid output. If you do not have a reliable way to assess child health yet, omit service watchdog notifications and use normal restart-on-failure behaviour while you improve monitoring.

Set restart behaviour for the failure you expect

Create a dedicated unit for the encoder, with a clear service name, explicit User=, WorkingDirectory= and paths to configuration and media files. Keep the command reproducible and keep the stream key in a protected configuration file or environment file with limited access. The exact unit depends on the encoder command, distribution and image; treat any example from another machine as a template to adapt, not a tested configuration for your VM.

For a process expected to run continuously, Restart=on-failure is a practical default: systemd restarts it after abnormal termination, but does not repeatedly relaunch it after a deliberate clean exit. Add a modest RestartSec= delay so an immediate failure does not cause a tight restart loop. If a clean exit should also cause the broadcast to resume, consider Restart=always instead. Read the installed systemd.service manual for the effects of restart rate limiting and the directives available on that image.

Enable the unit for boot, start it, and inspect its state before leaving it unattended. With an actual service watchdog configured, a missed notification leads systemd to mark the service failed; it is still the restart policy that determines what happens next. Without notifications, WatchdogSec= is not a substitute for Restart=on-failure and does not provide arbitrary-process liveness detection.

Keep logs persistent enough to answer what happened after an overnight interruption. systemctl status <unit> is useful for the current state, while journalctl -u <unit> can show the command's output, exit status and restarts. Avoid logging the stream key or full credential-bearing URLs. When repeated failures occur, slow or stop the restart cycle until you find the cause; automatic retries do not fix a missing file, bad key or inadequate network path.

If your goal is simply to recover from a crashed encoder, systemd's ordinary service restart can be enough. If your problem is a computer that cannot be left powered on, the guide to restarting a YouTube bhajan stream after disconnecting covers the broader continuity problem, but the same principle applies: identify whether the encoder, host, or broadcast is actually failing before choosing a recovery layer.

Make heartbeats reflect useful health

A heartbeat is evidence that the process issuing it is alive enough to send a message. It says nothing by itself about whether the child encoder is making progress, whether packets are reaching YouTube, or whether the stream is visible to viewers. Design the check around the condition you need to recover from and state what it cannot establish.

For a wrapper around FFmpeg, a basic check might confirm that the child process still exists and that its output progress advances. A stronger check can also observe connection or output errors. Neither necessarily confirms the whole path to a viewer. YouTube's Live Control Room provides stream health messages and a preview, so use those during a controlled test to check ingest as well as local process state.

You also need to distinguish a transient network interruption from a persistent fault. Restarting quickly may restore a dropped encoder connection, but a bad stream key or blocked outbound traffic will usually fail again. Use logs to identify repeated errors and leave enough delay to avoid hammering a failing endpoint. Do not keep sending heartbeats only to suppress systemd's failure response.

For a picture-based channel, confirm that the source file loops or that the encoder has an intentional way to continue after reaching its end. A process can be alive after the input ends without delivering the programme you intended. For a local news loop or devotional channel, check both audio and video in the preview rather than inferring success from a running PID. Advice on YouTube Live audio bitrate and sample rate settings can help when audio is the part that fails while video appears normal.

Know what a hardware watchdog can and cannot do

Before enabling RuntimeWatchdogSec=, inspect the guest and the cloud provider's documentation for a watchdog device and its supported behaviour. A device node alone may not tell you every detail about how the provider handles it; availability, permissions and recovery behaviour are specific to the VM image and provider. If no usable device is exposed, systemd documents that the runtime watchdog setting has no effect.

Where a device is available, its purpose is to recover a machine that has stopped responding to the system manager's watchdog contact. It is not a per-process monitor and does not inspect FFmpeg's output, the stream key, the network route, or YouTube's ingest state. A guest can remain responsive while its encoder has failed, and an encoder can be healthy while a host-level issue interrupts the VM.

Do not treat a reboot as proof that the broadcast recovered. After a host-level restart, confirm that the service was enabled at boot, that credentials and media paths remain available, and that YouTube sees a new or continuing ingest as intended. For a 24/7 channel, decide what should happen if the broadcast itself needs to be recreated rather than merely reconnecting an encoder to an existing ingest.

Test recovery safely on the cloud VM

Start with a test broadcast or a private/unlisted workflow appropriate to your channel, and verify YouTube's current stream-health guidance before changing a production broadcast. YouTube recommends testing before going live and checking its stream health messages. Use the official encoder settings guidance to check your output format: it recommends RTMPS, constant bitrate (CBR), and a two-second keyframe interval, with the interval not exceeding four seconds.

Bandwidth is part of the test, especially when your VM is in an Indian region with a route to YouTube that may differ from your home connection. YouTube recommends 5 Mbps video bitrate for H.264 at 1080p30 and 6 Mbps at 1080p60. These are encoder recommendations, not a promise that a particular VM can sustain them. Check actual sustained upload and leave headroom for variation rather than choosing a bitrate at the limit of measured capacity. If you are weighing resolution against available throughput, the resolution and bandwidth guide explains the trade-off.

Then test one failure mode at a time. First stop the encoder cleanly and see whether the chosen Restart= policy behaves as intended. Next, in a non-production test, terminate the process abnormally and confirm the service restarts. If you have a wrapper and watchdog, test that a missing or stalled child stops the heartbeat and triggers recovery; do not simulate a host reboot on a live channel merely to see what happens. Check systemctl status <unit>, journalctl -u <unit>, and the YouTube preview and health messages after each test.

A restart in systemd is only one step in recovery. Check that the new process has reconnected, that audio and video are moving, and that the viewer-facing broadcast is in the expected state. If the channel relies on a continuing archive, account for YouTube's documented behaviour: streams under 12 hours are automatically archived, so do not assume a 24/7 broadcast becomes one complete VOD. Plan rotation or broadcast recreation around your channel's archive needs, and confirm current Studio behaviour before relying on it.

Once the file, key, unit and test procedure are ready, choose how you want to operate it. StreamNeo removes the need to keep your own computer running for a file-based broadcast: you upload the video and provide the YouTube stream key, while the stream is monitored and restarted if it drops.

Before committing, compare the operating options on the pricing page. When the file and channel are ready, start free — 24-hour trial, no card.

FAQ

Does WatchdogSec restart FFmpeg if it hangs?

Not by itself. FFmpeg must send periodic systemd watchdog notifications, or a wrapper must send them based on checks that reflect the encoder's health. Without that, use normal restart behaviour for process exits and add an appropriate supervisor if you need hang detection.

What is the difference between WatchdogSec and RuntimeWatchdogSec?

WatchdogSec supervises one service that sends watchdog notifications. RuntimeWatchdogSec asks systemd to use a watchdog device to recover an unresponsive machine, and may do nothing if the cloud guest has no such device.

Should I use Restart=on-failure or Restart=always?

Use on-failure when a deliberate clean exit should stay stopped and abnormal failure should restart. Use always only when a clean exit should also relaunch the encoder; either way, check the installed systemd manual and logs for restart limiting.

Does a VM reboot prove my YouTube stream is healthy?

No. A reboot only says that machine-level recovery occurred; it does not show that the encoder reconnected successfully or that YouTube is receiving a healthy stream. Confirm the process, preview and stream-health messages after recovery.

YOU’VE REACHED THE END

Keep the ideas coming.

More guides, useful tools and a little help for your next broadcast.

Back to the journal ↗
YOUR NEXT READ

A little more to explore.

More Setup Guides guides ↗ · All topics ↗