Skip to content
streamneo.
Troubleshooting11 min read

How to Set Up CloudWatch Alarms for AWS MediaConvert

Choose the right MediaConvert signal, alarm on queue conditions and use EventBridge with SNS to receive email about ERROR jobs.

sn.
StreamNeoPublished 4 October 2026
Worth sharing?

CloudWatch metric alarms are for conditions that develop in a metric over time, such as a MediaConvert queue building up. For an email about an individual job entering ERROR, use an EventBridge rule with SNS; the two mechanisms solve different notification problems and can be used together.

Start by deciding what you need to know: whether a queue is falling behind, whether MediaConvert API operations are failing, or whether a particular job ended in error. That choice determines the signal, its scope and the notification path. An alarm threshold is not a substitute for matching a job-state event.

Decide what condition needs notification

Write the operational question in plain language before opening a console. “Tell me when the queue has too much work waiting” is a metric question. “Email me when this job fails” is an event question. “Tell me when requests to MediaConvert are returning errors” points towards operation metrics.

A metric alarm evaluates a metric using a configured statistic, period and threshold, then changes state when its evaluation criteria are met. It can notify through a configured alarm action, but it does not report every individual job failure. Metrics describe measurements, not a guaranteed record of each job’s state transition.

EventBridge rules instead match events as they are published. A rule for MediaConvert job state changes can match the ERROR status and send the matching event to an SNS topic. That makes it the relevant path when the requirement is a notification for each matching job event. Email delivery should not be treated as instantaneous or guaranteed, so retain the service console or another operational record as the source for investigation.

Need Signal and scope Notification path
Notice queue pressure or processing trends Queue metric, scoped to the relevant queue CloudWatch metric alarm and its configured action
Notice API-operation errors Operation metric for the relevant operation CloudWatch metric alarm, with threshold and evaluation choices suited to the metric
Receive notice when a job enters ERROR Job-state event for an individual job EventBridge rule targeting SNS, with an email subscription

If you operate a service that processes a steady stream of files, both paths can be useful: the queue alarm can indicate accumulating work, while the event rule can draw attention to a failed job. They are not interchangeable, and configuring one does not imply the other is in place.

Understand MediaConvert metric groupings

AWS documents MediaConvert metrics in operation, queue and job groupings. Operation metrics describe errors from API interactions; queue metrics describe activity across a queue; job metrics relate to output and workflow behaviour. AWS’s MediaConvert CloudWatch metrics reference explains the metric dimensions and scope, while its metrics guide provides further detail on what is published.

The grouping matters because a metric name alone does not tell you which work it represents. A queue metric scoped to one queue answers a different question from a job metric tied to a particular output or workflow. Before creating an alarm, inspect the available dimensions and confirm they identify the queue or activity you actually mean to watch.

MediaConvert publishes measures including Errors, JobErrorsByCode and output-duration measures to CloudWatch. The documentation says these are sent at the end of every job. That publication behaviour is important when choosing a metric for operational alerting: a measure reported at job completion is not necessarily a live indicator of an in-progress queue condition.

For a queue backlog, AWS specifically mentions StandbyTime as a metric that can be used to create an alarm. The service documentation does not set a universal threshold, statistic or period for every workload. Those values depend on what delay is acceptable for your queue and on its normal behaviour; they are not settings to copy without checking your own service objective.

Choose a queue metric and scope

Use a queue-level metric when you need to observe activity across a queue rather than infer the queue’s state from a single job. StandbyTime is AWS’s documented example for detecting a large backlog. Think of it as a signal to investigate whether jobs are waiting longer than your operation can tolerate, not as a magic value that fits every queue.

First identify the queue that matters. If a service separates urgent work from routine conversions, a single alarm on a broad or unintended dimension can obscure which queue is accumulating work. In the CloudWatch metric view, select the MediaConvert namespace, inspect the available metric and dimensions, and verify that the queue dimension corresponds to the intended queue. The precise console labels can change, so rely on the displayed dimensions and the current AWS guidance rather than on a remembered click sequence.

Then define what “too much” means for your operation. A team with a short publishing window may need to act on a smaller delay than a team processing archival files overnight. Compare the metric’s normal pattern with the time at which a human can still intervene usefully. Choose a threshold and evaluation window that reflect that objective, and record the reason so the next operator understands why the alarm exists.

Do not treat a queue alarm as a job-failure detector. A queue may be healthy while one job fails, and a job failure may not cause a queue-level metric to cross the chosen threshold. For a continuous YouTube workflow, this distinction is familiar: stream health, file processing and an individual asset’s failure are separate things to monitor. A practical primer on the separate mechanics of a long-running stream is our guide to setting up a 24/7 YouTube stream with FFmpeg in India.

Create a CloudWatch metric alarm

Once you have chosen a metric and verified its dimensions, create an alarm for that metric in CloudWatch. Select the appropriate statistic and period for the way the measure is published and for the time in which you need to respond. AWS does not prescribe a single statistic, period or threshold that suits every MediaConvert queue, so use your measured baseline and service objective rather than an assumed default.

Set the condition that should move the alarm into an alert state. For a backlog signal, that means deciding how much waiting is unacceptable and how sustained it must be before waking someone or creating a ticket. A very sensitive threshold may create noise during normal variation; a permissive one may leave too little time to act. Use historical observations where available, then review the setting after real workload changes.

Choose an alarm action if a notification is required. This is the CloudWatch alarm’s response to its state, not the EventBridge route for a discrete job event. Make sure the action is directed to a channel someone monitors and that the right people know what the alarm means. A useful notification says which queue and metric crossed the condition, and points to the next place to inspect.

CloudWatch also lets you view metrics and alarms together as part of monitoring. That context can help an operator distinguish a backlog signal from a broader processing change. Keep the alarm narrow enough to be actionable: if the queue is one of several, identify it explicitly in the alarm name and document its purpose. For a small team maintaining both the stream and its media pipeline, separate stream settings from AWS processing signals; our GStreamer latency and bitrate guide addresses a different layer of the workflow.

Create an EventBridge rule for ERROR jobs

For an email when a job enters ERROR, configure an EventBridge rule to match the MediaConvert job-state event. AWS’s MediaConvert EventBridge instructions show the relevant event source, detail type and status. The event pattern is:

{
  "source": ["aws.mediaconvert"],
  "detail-type": ["MediaConvert Job State Change"],
  "detail": {
    "status": ["ERROR"]
  }
}

This pattern matches events with the stated source and detail type when the job status is ERROR. It is not a metric alarm: it does not wait for a queue metric to exceed a threshold. Keep the match specific to the failure status if the operational requirement is failed-job notification; broadening the pattern to other states changes which events generate messages.

Create the rule in the AWS account and Region in which the relevant MediaConvert jobs run, and confirm the event source and target are being configured in the intended place. AWS permissions and available console choices depend on the account setup. The research guidance does not establish a complete IAM policy for every account, so follow the current AWS instructions and your organisation’s access controls rather than copying an assumed policy.

The event contains job details that help identify what failed. Keep the notification useful without treating an email as a replacement for inspecting the job in MediaConvert. If the operator has multiple queues or workflows, include enough context in the process around the alert to find the relevant job and investigate its input, settings and outputs.

Connect the rule to SNS email notifications

Create an SNS topic to receive the EventBridge rule’s matching events, then add an email subscription for the people who should receive them. The AWS tutorial uses a standard SNS topic and an email subscription; the recipient must confirm the subscription before it receives notifications. Use AWS’s current MediaConvert event-rule walkthrough for the supported configuration sequence and permissions.

Set the SNS topic as the EventBridge rule target. Check that the target is the intended topic and that the rule is enabled. A successfully created rule with no usable target, or an unconfirmed email subscription, will not meet the intended notification requirement. Consider who needs to receive alerts when the usual operator is away, but avoid sending operational detail to an unnecessarily broad audience.

SNS email is a delivery route, not a promise about delivery time or receipt. Mail filtering, subscription status and downstream handling all matter. The alert should therefore prompt a check of the job and the EventBridge/SNS configuration where necessary, rather than serve as the only evidence that a job did or did not fail.

These AWS configuration steps address media conversion and job notifications. If your wider workload is a 24/7 YouTube channel, the stream itself has a separate operating concern: your computer and connection must stay available if you run the broadcast locally. StreamNeo removes that specific always-on computer burden by taking an uploaded video and running it as a YouTube live stream with your computer switched off; it does not monitor MediaConvert jobs or replace the AWS alerts described here.

Test the alert path and tune it

Test the EventBridge and SNS route with a job that is expected to fail. AWS’s tutorial gives a missing input location as an example. Submit a controlled test job, then confirm that the subscribed email receives a notification and that its details let you identify the job. Do not use a production asset or workflow whose failure would disrupt a publishing schedule when a safe test is available.

That test checks the event rule, target and subscription as a combined path. If no message arrives, check whether the job entered ERROR, whether the rule pattern matches the event, whether the rule is enabled, whether the target is correct and whether the email subscription has been confirmed. Inspect the relevant AWS event and notification configuration rather than concluding that the job succeeded simply because no email appeared.

Test a metric alarm separately. A failed-job event test does not prove the metric alarm is configured correctly, and an alarm test does not prove that individual job events reach email. Use safe test conditions or representative metric data where your environment permits, and verify the selected dimensions, evaluation settings and alarm action. Avoid forcing production metrics across a threshold if doing so could trigger disruptive downstream actions.

After the system has run under ordinary conditions, review whether the metric threshold catches actionable queue pressure without generating frequent noise. If the normal workload changes, revisit the baseline and objective. Similarly, review the event notification with its recipients: remove stale addresses, confirm the intended people can act on a failure, and keep the path documented so another operator can test it.

Good alerting makes the next step clear. A queue alarm should direct you to the queue and its waiting behaviour; an ERROR email should direct you to the failed job. If an alert is too broad to point to an investigation, refine its scope rather than adding more notifications. For additional context about keeping a live channel’s overall workflow manageable, see our guide to choosing a YouTube live-streaming platform.

Before committing, compare the operating options on the pricing page. When the file and channel are ready, start free — 24-hour trial, no card.

FAQ

How do I get an email when a MediaConvert job fails?

Create an EventBridge rule that matches MediaConvert job-state change events with status ERROR, then set an SNS topic as its target and subscribe the recipient by email. The recipient must confirm the subscription. Test with a controlled job expected to fail, and do not assume email is instantaneous or guaranteed.

Which MediaConvert metric should I alarm on for a queue backlog?

AWS specifically identifies StandbyTime as a metric you can use to detect a large queue backlog. Choose the queue dimension, threshold, statistic and evaluation period to match your own latency objective and observed workload; AWS does not give one universal threshold for every queue.

Can a CloudWatch metric alarm notify me about every individual job failure?

No. A metric alarm evaluates a metric against configured conditions over time, so it should not be treated as a per-job failure feed. Use an EventBridge rule matching the job’s ERROR state when you need notifications about individual failed-job events.

Should I configure both an alarm and an EventBridge rule?

Use both when you need to see queue trends as well as individual job failures. The metric alarm can surface a growing backlog, while the EventBridge rule can match a job entering ERROR; test each path independently because one does not validate the other.

YOU’VE REACHED THE END

Keep the ideas coming.

More guides, useful tools and a little help for your next broadcast.

Back to the journal ↗
YOUR NEXT READ

A little more to explore.

More Troubleshooting guides ↗ · All topics ↗