Building A Swear Filter That Actually Works

I spent about two years maintaining a profanity filter for a community platform with roughly 40,000 daily active users. What follows is the actual system we ended up running, plus the things that broke along the way. The foundation is a word list. There is no way around this. You need a solid dictionary of terms you want to block or flag, and it has to be maintained. The most common starting point is the H-List or the Bad Words API, but both of those are dated in places. I wound up merging multiple open-source lists, removing duplicates, and then manually auditing about 600 entries that had drifted into slang that was either context-neutral or simply not relevant to our moderation team's thresholds. Here is the part beginners get wrong. A flat word list catches about 60 percent of violations in the wild. The rest is leetspeak, intentional misspellings, character spacing, and foreign-language terms. If you only load a basic list into your app and call it a day, you will spend the next six months fielding support tickets from users who cannot figure out why their comment with a space inserted between letters went through.

We used a combination of several techniques:

  • Fuzzy matching with edit distance — This catches "sh!t" and "sh1t" and other common substitutions.
  • Character-level normalization — Strip diacritics, normalize full-width characters, collapse whitespace.
  • Contextual lookup — Some words are fine in certain contexts. "Ass" in "backup" versus "asshole" in a comment thread.
  • Custom blocklist overrides — Our mods could add terms on the fly when a new one surfaced in chat.

The fuzzy matching alone accounts for the biggest chunk of blocked violations we see. I would recommend starting with an edit distance of 2. Anything lower and you are blocking too many legitimate words. Anything higher and your false-positive rate climbs fast enough to make your moderation team hate you. Our stack was Node.js on the ingestion side with a Redis cache for the word trie. The preprocessing pipeline runs before anything hits the database. Here is roughly how the flow works: Input text goes through a normalization layer first. We lowercase everything, strip HTML tags, remove emoji unless they carry specific harmful meaning, and then run a regex pass to tokenize. After tokenization, each token gets compared against the blocked list using a combination of exact match and edit distance scoring. Tokens that score above the threshold get flagged. The flagged tokens are then run through a context checker that looks at surrounding words to determine severity.

Get the Full Details

Funny Swear Words (And Insults) From Around The World: Swear Word Adult Coloring Book : Buy ...
Funny Swear Words (And Insults) From Around The World: Swear Word Adult Coloring Book : Buy ...

Severity matters because not every violation gets the same treatment. An accidental match on a word that happens to share a root with a blocked term might just get a warning. A repeated pattern of evasive spelling triggers a timeout. Hard slurs get immediate removal and a strike. The system I built uses a three-tier model: warn, shadow-flag, and remove. I ran into a specific edge case that took us three weeks to track down. A user was bypassing the filter by inserting zero-width characters between the letters of blocked terms. The filter saw clean strings and let them through. The raw text rendered to humans as a swear word, but the string comparison found nothing blocked. The workaround was a preprocessing step that strips all Unicode characters in the U+200C through U+200F range before any matching runs. Once that was added, the bypass stopped entirely.

Keeping The List Updated

This is the part nobody tells you about. Language changes. New slang surfaces every few months. A word that was irrelevant last quarter becomes a primary vector for harassment this quarter. Our team did a monthly audit where we reviewed the top 50 unblocked terms that appeared in flagged content during the previous month. About 80 percent of those ended up being added to the custom blocklist with a severity level of one or two. There are a few automated sources you can pull from. GitHub repositories like "badwords-lists" and "profanity-filter-words" get updated periodically. The Bad Words npm package ships with its own list and updates when you reinstall. For our purposes, we wrote a cron job that pulled the latest from two open-source repos, merged the lists, ran deduplication, and generated a deployment commit that our CI pipeline picked up automatically. This kept the core list fresh without requiring manual intervention. I should note the main failure mode here. When you merge lists from multiple sources, you accumulate variants of the same word written differently across dialects. You also pick up regional slang that might be completely irrelevant to your user base. The cleanup step is where the real work lives. I would budget at least two to four hours per month for list maintenance if you are running this at scale.

When This Approach Fails

A word-list-based filter is not a complete solution. It misses context entirely. It cannot tell the difference between someone using a slur as a descriptor of behavior and someone quoting it in an educational context, and it cannot understand sarcasm. For high-stakes moderation at scale, you eventually need a model-based layer on top. We ran a small transformer model fine-tuned on labeled moderation data alongside the word filter, and the model caught about 20 percent of violations that the list missed. The tradeoff is latency and cost. The model added roughly 80 milliseconds per request on our setup. If your platform is small and you do not have engineering resources, a well-maintained word list with fuzzy matching is still worth it. It catches the vast majority of obvious violations at near-zero cost. Just do not treat it as the final word on the problem. The code repository we open-sourced for this is available on GitHub under the name mod-filter-core. It includes the preprocessing pipeline, the trie-based matcher, and the custom blocklist management tools we used in production. The README has the setup instructions. Installation is a standard npm install after cloning.

18+ Crazy Dutch Swear Words And Insults
18+ Crazy Dutch Swear Words And Insults