Skip to content
streamneo.
Tools11 min read

How to Detect Profanity in Live Streams with Amazon Transcribe

Configure Amazon Transcribe vocabulary filters for live captions, choose how matches appear, and understand the limits of exact-word filtering.

sn.
StreamNeoPublished 4 October 2026
Worth sharing?

Amazon Transcribe can filter configured words in streaming transcription. To use it, create a custom vocabulary filter, attach it to the streaming request, and choose whether matched words are masked, removed or tagged.

This is exact-word filtering against a list you supply, not a comprehensive profanity detector. It can help shape live captions or transcript output, but it cannot guarantee that every profane expression will be caught or hidden.

What vocabulary filtering does

Amazon Transcribe converts incoming speech to text and can apply a custom vocabulary filter to that transcript. You provide the words you want handled; the service applies the selected action when it matches a configured term. AWS documents offensive and profane words as one possible use for a custom filter, alongside other words a customer may want to mask, delete or flag. See AWS's guide to custom vocabulary filters for the supported behaviour.

The important distinction is that the filter acts on words in transcription output. It is not a general audio-moderation system deciding whether an utterance is acceptable in context. It does not replace a publisher's moderation policy, a human review process or any separate controls needed to manage the live stream itself.

For a devotional channel, for example, you might choose a short list of terms that should not appear in captions. A news broadcaster could have a different list and workflow, perhaps tagging matches for a producer to review rather than automatically hiding them. The right list and action depend on what you publish and how captions reach viewers.

A custom vocabulary filter is available for streaming as well as batch transcription. The streaming API operation is StartStreamTranscription; its request includes language and audio configuration as well as the vocabulary-filter settings. The feature's value is operationally narrow but useful: you control specified transcript words without having to remove the entire caption line.

Create a custom vocabulary filter

Start by deciding which terms your policy intends to handle and which language or languages your audio uses. Build the list around your actual content and moderation requirements rather than assuming a generic profanity list covers every case. AWS lets you create a filter from words or from a text file, and the filter is associated with a language.

In the AWS Management Console, open Amazon Transcribe and create a vocabulary filter. Give it a name that your streaming application can refer to, select the appropriate language, then enter the terms or upload a text file. The precise console labels can change, so consult the current AWS instructions for creating a vocabulary filter if the screen differs.

Treat the filter as a maintained policy artefact. Keep a record of why each term is included, who can change it, and how you will test changes before they affect a public caption feed. If different shows use different languages or standards, separate lists may be easier to audit than one mixed list. Do not assume that a filter created for one language will work as intended for another.

AWS notes that a mismatched filter language may mean the filter does not apply. Language availability is also feature-specific and can change. Before launching, check the current streaming language support table for both the language and the streaming features you need. If you use automatic language identification, follow the API's rules for specifying vocabulary filters in that configuration rather than assuming a fixed-language example applies unchanged.

A sensible first pass is deliberately small. Include only words your channel has decided to handle, then test recognisable speech containing those words and ordinary speech that should remain unaffected. A small, reviewed list is easier to reason about than a broad imported list whose intended behaviour nobody has checked.

Attach the filter to a streaming request

Your application must pass the filter name and chosen filtering method in the streaming request. It also needs a language configuration, media encoding and sample rate that reflect the audio being sent. The relevant request parameters are documented in the StartStreamTranscription API reference; check that reference for the current syntax and constraints.

The request is bidirectional: the application sends audio while receiving transcription events. AWS supports streaming through its SDKs, HTTP/2 or WebSockets. AWS recommends its SDK route as the simpler, more reliable way to set up a stream. Lower-level HTTP/2 or WebSocket integration can suit a team that needs direct control, but it brings more transport details to implement and test.

Audio format matters. The streaming formats documented by AWS include FLAC, OPUS in an Ogg container, and signed 16-bit little-endian PCM. AWS recommends lossless FLAC or PCM. Do not treat a file extension as proof of encoding: WAV is a container, and the request needs to describe audio in a supported format. The sample rate you declare must correspond to the actual input stream, not an assumed default.

In a typical implementation, the service that captures or reads audio opens a transcription stream, sends properly encoded audio chunks, and consumes result events. The filter settings travel with that stream request. If you change the policy list, verify that the application references the intended filter version or name; otherwise a successful stream may continue using a stale configuration.

Think through where the transcript goes next. A caption renderer, logging system and moderation dashboard may all consume results differently. If the stream is public, confirm that the renderer receives the filtered output you intend, rather than an unfiltered parallel transcript. If you store transcripts for review, note which action was used so an auditor can distinguish an omitted word from a recognition gap.

This transcription setup is separate from a 24/7 broadcast pipeline. If you are deciding how to keep a prerecorded YouTube channel running without a machine at home, the practical trade-offs are covered in using a hosted service instead of a 24/7 streaming PC. That is about broadcast operation, not a substitute for configuring transcript filtering.

Choose mask, remove or tag

The three actions have different consequences for the text your application receives. Choose based on whether the viewer should see an indicator, whether the transcript should omit the match, or whether another stage must decide what happens.

Method What happens to a matched configured word Use it when
mask The word is replaced with ***. The caption should visibly conceal the term but show that something was filtered.
remove The word is deleted from the transcript output. You want the returned text to omit the configured term without a marker.
tag The word remains, with an indicator that it matched the filter. A later review or replacement step will act on the result.

These behaviours are documented in the AWS filter guide and API reference. A tag does not hide the word from viewers by itself. If tagged text is sent directly to a caption renderer, the original word remains visible unless your application interprets the tag and changes what it displays.

Masking is often easier for a viewer to understand because the marker preserves the fact that a word was present. It may, however, make a caption line less natural and can draw attention to the filtered moment. Removing a word produces cleaner text but can alter the meaning or make a sentence hard to follow. Test both against the style of your programme rather than choosing on technical grounds alone.

Tagging is useful when a human or downstream rule should review matches. For instance, a news producer may want to inspect a caption before deciding whether to display a masked form. That extra step adds operational work and may not be suitable where captions are rendered immediately without review. Decide what happens to the tag before enabling it; otherwise you have detected a match but have not defined the viewer-facing result.

Do not confuse vocabulary filtering with PII redaction or identification. Transcribe has separate options for handling personally identifiable information, but those controls are not a replacement for a list of profanity terms. If your use case includes both, treat them as separate requirements and check the API documentation for the relevant settings.

Test configured terms in live captions

Test the complete path, not only the AWS request. Speak or play representative audio into the same capture chain you plan to use, send it through the streaming application, and inspect both the events returned by Transcribe and the final text shown to a viewer. Confirm that the correct filter is attached, that the language is configured appropriately and that the selected action has the expected effect.

Include several kinds of test material: a configured term spoken clearly, an unlisted term, ordinary language that resembles a configured term, and speech with background music or other noise typical of your channel. The point is not to produce a pass rate, but to find where your particular capture and caption workflow behaves differently from your expectations. Keep the test private and controlled if the material would be unsuitable for a public stream.

Transcribe returns partial results incrementally. Those early results can change as more speech context arrives, so do not treat every partial event as final text. For a durable caption record or transcript, use finalized segment results. AWS describes partial-result stabilization as a way to return output faster, with a possible accuracy trade-off; test that setting against your timing and readability needs rather than treating speed as free.

A useful check is to view the live caption output while also logging event types and timestamps in a private test. If a word briefly appears in a partial result and is later revised or filtered in the final result, the viewer-facing behaviour depends on how your renderer handles updates. If you cannot retract text already displayed, you may need to wait for finalized results before publishing captions, accepting the delay that introduces.

Keep a short change log when you add or remove terms. Re-run the representative test after changes to the list, language, audio encoder or caption renderer. This is especially helpful for channels that operate overnight: a filter that works in a desk test may not see the same mix, microphone level or prerecorded material during a full broadcast. For the separate issue of delivering a prerecorded programme to viewers in different time zones, see scheduling a prerecorded YouTube live stream for Indian viewers.

Understand the limits of filtering

The core limit is exact-word matching against your configured list. If a spoken expression is not represented in the list in a form the transcript matches, the filter may not act on it. Different spellings, inflections, deliberate obfuscations, homophones, transcription errors and unlisted words are practical reasons your results can differ from your expectations. They are not evidence that the feature performs semantic analysis.

Recognition itself can also be revised as more context arrives. A partial caption is a provisional result, not a promise that the final transcript will preserve the same words or filtering outcome. Background music, overlapping speakers, pronunciation and the capture chain can affect what text the recogniser returns; the filter can only work on the transcript match it receives.

AWS describes a default filter for certain sensitive terms, but recommends that customers review and supplement coverage for their own language and use case. Do not rely on the existence of a default list as proof that a particular expression will be handled. Review the current documentation and test the terms that matter to your channel.

Language support is not uniform across every streaming feature. Check current support for the intended language and region, especially if you use language identification or a less common language. A filter that is valid for one language does not become a multilingual moderation policy merely because the audio contains more than one language.

Finally, transcript filtering is only one part of moderation. It does not establish whether speech violates YouTube's policies, decide whether a clip should be removed, or guarantee compliance. Review the current official YouTube Community Guidelines and build moderation procedures appropriate to your channel. For a prerecorded 24/7 video channel, reliable playback and audio quality are separate concerns; our guide to setting bitrate for a continuous YouTube VOD stream covers that different part of the workflow.

If the task you need is simply keeping a prerecorded visual stream on YouTube while your own computer is off, StreamNeo removes the need to leave that computer running; it does not replace a transcript policy or make the content moderation decision for you.

Before committing, compare the operating options on the pricing page. When the file and channel are ready, start free — 24-hour trial, no card.

FAQ

Can Amazon Transcribe censor swear words in real time?

It can apply a custom vocabulary filter to streaming transcription and mask, remove or tag configured word matches. The result depends on what is in the list and what the speech recogniser returns; it is not a guarantee of complete censorship.

Does tag hide the word from live captions?

No. Tagging leaves the matched word in the transcript and marks it as a filter match. Your application must interpret that marker and decide whether to mask, remove or otherwise handle the text before it reaches viewers.

Can I use one filter for every language on my channel?

Do not assume so. Filters are associated with language, and AWS says mismatched language settings may prevent the filter from applying; check current streaming support and configure for the language in use.

Should I show partial results as captions?

Only if your renderer can handle revisions and the risk that provisional text changes. If already displayed text cannot be retracted, consider publishing finalized segments instead, while testing whether the added wait suits your live-caption needs.

YOU’VE REACHED THE END

Keep the ideas coming.

More guides, useful tools and a little help for your next broadcast.

Back to the journal ↗
YOUR NEXT READ

A little more to explore.

More Tools guides ↗ · All topics ↗