Skip to content
streamneo.
Troubleshooting11 min read

How to Automatically Restart FFmpeg After a 24/7 Lecture Stream Crashes

Use systemd to restart failed FFmpeg processes, then handle input reconnects, output recovery and stalled-stream checks separately.

sn.
StreamNeoPublished 4 October 2026
Worth sharing?

If FFmpeg exits on a Linux host that uses systemd, run it as a service and set Restart=on-failure so systemd can relaunch it after qualifying failures. That does not reconnect every input, repair every output problem, or detect a stream that has frozen while FFmpeg remains alive.

Think of recovery in layers: reconnect a supported input inside FFmpeg, attempt recovery from certain output errors where the muxer permits it, and let systemd restart an exited process. For a process that is still running but no longer making useful progress, you need a separate health check. The examples here are a Linux/systemd pattern; the right details depend on your operating system, FFmpeg build, and input and output protocols.

Identify what failed and where

Before changing restart settings, establish whether FFmpeg exited, whether its input stopped delivering media, whether its output connection failed, or whether the process is alive but the broadcast has stopped progressing. Those symptoms can look similar to a viewer, but they call for different recovery mechanisms.

Start with the service or process status and the most recent FFmpeg and system logs. Look for the exit status, termination signal, timeout, network error, authentication or protocol message, and the last media progress reported. An FFmpeg process that exited gives a process supervisor something to restart. A running process with a stale output does not trigger a restart policy merely by appearing unhealthy to viewers.

Also identify the direction of the failing connection. An HTTP input that drops is not the same as a network output that cannot reach YouTube. Reconnecting the input may restore incoming media while leaving a separate output fault untouched. Conversely, relaunching FFmpeg after it exits starts the command again but does not necessarily cure the condition that caused the exit.

Write down the existing command and test it manually before wrapping it in a service. Check that the file or input is reachable, that credentials are available to the account which will run the service, and that the output target and stream key are correct. If the issue is missing audio rather than a crash, use a diagnosis aimed at the media path, such as this guide to fixing missing audio in an FFmpeg YouTube Live stream.

This first pass saves time later: a service manager can repeat a broken command indefinitely. Repeated restarts are evidence of repeated failure, not proof that the lecture is reaching viewers.

Run FFmpeg under systemd

On a Linux machine managed by systemd, define FFmpeg as a service rather than relying on an interactive terminal session. The unit gives systemd a process to supervise and a place to define restart policy, execution user, working directory, and dependencies. Keep the actual command and access requirements specific to your host.

Here is a configuration pattern, not a tested drop-in file. Replace the executable path, account, directory and placeholder arguments with values from your deployment:

[Unit]
Description=FFmpeg lecture stream
Wants=network-online.target
After=network-online.target

[Service]
Type=simple
User=stream
WorkingDirectory=/srv/lecture-stream
ExecStart=/usr/bin/ffmpeg [your existing FFmpeg arguments]
Restart=on-failure
RestartSec=5s

[Install]
WantedBy=multi-user.target

Use an absolute path for FFmpeg and ensure the service account can read the lecture file, access any required credentials, and write any logs or output needed by the command. The example assumes a persistent Linux host and does not establish that the network is usable simply because the system has reached the online target. Your command may need to tolerate a delayed connection as well.

The RestartSec value shown is an example spacing choice, not a universal recommendation. A very short cycle can produce a rapid series of failures and make the underlying problem harder to see. Choose a delay appropriate to the failure and inspect systemd's rate limiting behaviour on the installed host. If you are deciding where to run a persistent service, this guide to running a continuous YouTube stream on a rented Indian VPS covers the hosting context, not a substitute for a correct service unit.

Follow your distribution's normal systemd workflow to install or reload the unit, enable or start it, and inspect status. Avoid putting a shell wrapper in ExecStart unless you need one and understand its exit behaviour. A wrapper that catches FFmpeg's error and exits successfully can prevent on-failure from acting as expected.

Set Restart=on-failure

Restart=on-failure asks systemd to restart the service when the main process exits with a non-zero status, is terminated by a qualifying signal, times out, or has a watchdog timeout. The systemd service manual documents the restart policy and its triggers. The exact applicable details should be checked against the systemd version on your host.

A stop requested through systemd is different: systemd does not treat its own intentional stop operation as a failure that should be undone immediately. That distinction lets you stop the broadcast for planned maintenance without an automatic restart fighting you. If you use a wrapper or custom stop procedure, verify what signal and exit status systemd actually observes.

Do not treat on-failure as a catch-all. It responds to process outcomes known to the service manager; it cannot infer that viewers see a frozen picture, that the stream key has become invalid, or that the media content is unsuitable. It is the process-recovery layer, not an end-to-end broadcast monitor.

If you deliberately exit FFmpeg with a successful status after detecting a condition, on-failure may not restart it. Conversely, if a persistent bad input causes non-zero exits, systemd may keep relaunching a command that cannot yet succeed. Read the journal, correct the cause, and choose restart spacing and rate limits deliberately rather than trying to force endless immediate retries.

This mechanism is useful when the process genuinely dies from a transient failure, host signal, or timeout. It is less useful when the command itself is invalid or the remote service consistently rejects it. For readers comparing whether to operate a local process or another arrangement, the trade-offs between an OBS spare PC and a VPS are about operating choices; whichever host you choose, the same distinction between an exited process and a stalled broadcast remains.

Use FFmpeg reconnect options for supported input errors

FFmpeg has protocol-specific reconnect controls for some input situations. The FFmpeg protocol documentation describes HTTP options including reconnect, reconnect_at_eof, reconnect_on_network_error, reconnect_on_http_error, and reconnect_streamed, along with retry and delay controls. These are not generic switches that work for every input type or every FFmpeg build.

For example, if your lecture source is an HTTP input and that connection drops before the media is complete, an applicable reconnect option may let FFmpeg retry the input without the process exiting. A live or endless source may need different handling at end-of-file from a finite file. Read the option description for the protocol and build you actually use; do not paste a collection of HTTP flags into a command using a different input protocol and assume it applies.

Retry bounds and delays matter. A retry loop that waits for a service to recover may be appropriate for a transient network interruption, but long or unbounded retries can leave the broadcast without fresh media for an unacceptable period. A short retry window may instead hand control back to systemd sooner by allowing FFmpeg to exit. Decide which layer should own the retry for the failure you have observed, and avoid having multiple layers retry without understanding their combined timing.

Reconnect options concern the input side. They do not restart an FFmpeg process that has already exited, and they do not necessarily recover an output connection. Confirm that the installed executable supports the selected option and test it with the actual protocol and source. If the lecture is a local prerecorded file, a network-input reconnect flag is unlikely to be relevant to a file that simply reaches its end; the playlist or looping behaviour is a separate command design question.

Consider FIFO recovery for supported output errors

Some output failures can be addressed inside FFmpeg with the FIFO muxer. The FFmpeg muxer documentation describes attempt_recovery as an effort to recover after failure, particularly for network output. It also documents drop_pkts_on_overflow, which can allow encoding to continue when a queue fills, at the cost of dropping packets.

This is a trade-off, not a guarantee of uninterrupted viewing. Attempted output recovery may be useful for transient network problems, but the result depends on the output protocol, muxer configuration, and failure. Dropping queued packets may prevent a blocked queue from stopping progress, but dropped media can mean gaps or disruption for viewers. For a lecture, decide whether preserving continuity, avoiding a long freeze, or restarting cleanly is preferable for the specific situation.

FIFO recovery is not the same as relaunching an exited FFmpeg process. It applies to supported output failure cases while FFmpeg is operating; systemd addresses the process after it exits. Neither one proves that YouTube is receiving a progressing broadcast. Before changing output options, confirm what output protocol and muxer the deployed command uses and consult documentation for that configuration.

Keep a simple failure map beside the service definition: input disconnect, output error, process exit, and suspected stall. For each, record the expected signal, the layer that should respond, and what a viewer may experience. This prevents a successful test of one path, such as input reconnection, from being mistaken for proof that all other paths recover.

Add a health check for stalled progress

A systemd restart policy cannot detect a process that remains alive unless some additional mechanism reports a failure, such as a configured watchdog notification or a monitor that deliberately stops or fails the service when its checks fail. A separate health check is therefore needed if the important fault is a frozen or non-progressing stream.

First define a measurable signal. Depending on the command and output, that might be whether FFmpeg's progress output advances, whether frames continue to be encoded, or whether an independent check can observe fresh media at the destination. None is universal: a counter can advance while the wrong content is sent, and an external view can lag. Pick a signal that corresponds to the failure you need to catch, and document its limitations.

Then define how the check acts. A monitor can alert an operator, request a controlled restart, or integrate with a supported systemd watchdog arrangement. An alert-only check is often safer while you are validating the signal. If an automated action is used, make it bounded and observable: log why it acted, avoid rapid restart loops, and ensure a manual stop for maintenance is not mistaken for a fault.

Test false positives as carefully as missed failures. A brief network pause, a quiet section of a lecture, or a deliberately paused input should not automatically be interpreted as a dead stream unless the check has a reliable way to distinguish it. Systemd's watchdog support also requires application or service integration; merely adding a watchdog setting does not make FFmpeg report media health.

If you cannot build or maintain a suitable monitor, be explicit about the limitation and arrange periodic human checks. A continuously running process is not a substitute for a verified stream. Where the primary pain is keeping a prerecorded lecture loop live without leaving a personal computer on, StreamNeo can remove the need to manage a local FFmpeg process and its restart behaviour; it is YouTube-only and does not remove the need to verify the channel and content.

Test recovery and review logs

Test one failure layer at a time during a planned maintenance window. Start with a known-good command and confirm that the service launches under its service account. Then test process exit, an input interruption where relevant, output interruption where safe, and the health check separately. Do not deliberately interrupt a public lecture without warning viewers or choosing an appropriate test stream.

After each test, confirm what happened from more than one perspective: systemd status and journal, FFmpeg messages or progress, and the destination's observed stream. Note whether FFmpeg exited, stayed alive while retrying, or continued with gaps. A service showing as active only confirms a process state from systemd's perspective, not that the broadcast is healthy.

Review logs after a real overnight incident as well. Record the timestamp, symptom, last useful progress, process exit reason if any, retry behaviour, and whether viewers saw recovery. Redact stream keys and credentials before sharing logs. If failures repeat, check the original command, permissions, network path, input availability, output configuration, and restart rate limits before increasing retry complexity.

Keep the tested unit and command under change control, and repeat the test after changing FFmpeg versions, protocols, credentials, or systemd configuration. Option availability and behaviour are version-dependent. If you are checking whether your wider channel plan can support the intended operating arrangement, the comparison of concurrent stream limits addresses that separate question; it does not establish that this particular FFmpeg process will recover.

Before committing, compare the operating options on the pricing page. When the file and channel are ready, start free — 24-hour trial, no card.

FAQ

Will Restart=on-failure restart FFmpeg after every crash?

It asks systemd to restart after documented failure outcomes such as non-zero exit, qualifying signals, timeouts, or triggered watchdog timeouts. It does not restart a service stopped intentionally through systemd, and it does not fix the reason a process failed. Check the unit and journal on your installed systemd version.

Should I use FFmpeg reconnect options or systemd?

Use reconnect options for input errors supported by the input protocol and FFmpeg build; use systemd to relaunch an FFmpeg process that exits under the selected policy. They address different failure layers and can be used together when the failure modes justify it. Neither is a universal recovery mechanism.

What if FFmpeg is still running but YouTube looks frozen?

An exit-triggered restart policy may do nothing because the process has not exited. Check for a meaningful progress signal and use a separately designed monitor or watchdog if you want automatic stall detection. Verify the destination as well as the local process state.

Can FIFO recovery guarantee that viewers will not see an interruption?

No. FIFO muxer options attempt recovery for supported output failures, and queue overflow handling can involve dropping packets. The result depends on protocol, configuration, and fault; test it with a safe stream and decide whether gaps or a restart are acceptable for your lecture.

YOU’VE REACHED THE END

Keep the ideas coming.

More guides, useful tools and a little help for your next broadcast.

Back to the journal ↗
YOUR NEXT READ

A little more to explore.

More Troubleshooting guides ↗ · All topics ↗