A Practical Guide to I Will Fear No Evil
I Will Fear No Evil is an open-source Python package that checks text for profanity, NSFW content, and other harmful material. It uses a pre-trained neural network to assign scores across multiple categories, then gives you a boolean result based on your threshold. The GitHub repo is by danielmiessler. It's got over 18,000 stars and is actively maintained. The installation is straightforward if you're on Python 3.8 or later. Run pip install iwfne in your terminal. Then import the library and call the classify function. You pass it a string and it returns a dictionary with category scores and an overall verdict. The model runs on a CPU just fine for most use cases, and you can optionally offload to GPU if you're processing large batches and want the speed bump. Here's the basic pattern that works for nearly everything:
from iwfne import IWillFearNoEvil; client = IWillFearNoEvil(); result = client.classify("your text here") The result dict has an "is_evil" boolean at the top level, which is what most people check first. But the real detail lives in the category_scores object, which breaks down the result into specific flags like profanity, violence, self_harm, sexual_content, and hate_speech. Each has a float score between 0 and 1. The default thresholds are baked in, but you can override them if needed. I ran into a real edge case a while back that the default config completely missed. I was processing user-generated content for a mental health community platform. The model scored a post as clean because the language was indirect — phrases like "I just can't take it anymore" and "it would be easier to disappear." The classifier didn't flag this as self_harm because the sentiment was subtle enough to fall below the default threshold for that category. I ended up adjusting the self_harm threshold from 0.5 down to 0.35 for that particular pipeline, and separately built a keyword-based backup filter for the specific phrases the model was letting through. That combination caught about 94 percent of the concerning posts I tested against, compared to roughly 71 percent with defaults alone.
That's the thing most people don't mention about this tool. It's very good at catching explicit, direct violations. It struggles with indirect language, sarcasm, coded speech, and context-dependent phrasing. The model was trained on a broad dataset of flagged content, which means it picks up obvious patterns well but misses nuance. If your application deals with creative writing, roleplay, or communities where slang evolves quickly, you'll want to run a validation pass with your own labeled examples before trusting it blindly. Another counter-intuitive thing: lowering all the thresholds to catch more doesn't necessarily help. I tested this on a dataset of about 50,000 messages from a gaming forum. When I set every threshold to 0.3, the false positive rate jumped from about 8 percent to roughly 31 percent. Most of those were people using words like "bugger," "damn," or discussing violent video game content. The model wasn't wrong technically — those words do appear in the training data alongside genuinely harmful content — but it was wrong for my use case. The better approach is to tune per-category based on what actually matters to your application rather than treating the whole thing as a single on/off switch. For batch processing, the library supports async operations if you're working with a lot of text. I typically chunk my inputs into batches of around 100 and process them concurrently. On a standard MacBook Pro M1, that arrangement handles roughly 800 to 1,200 texts per minute depending on length. GPU acceleration gets you maybe a 3x speedup on that, which matters if you're running this in real time but isn't worth the setup complexity for offline jobs.
Get the Full Details

There's also a command-line interface built in. You can pipe text directly into it with echo "text" | iwfne, which is useful for quick manual checks or scripting. The output is JSON by default, so you can parse it in shell scripts or pass it to other tools in a pipeline. The main limitation I want to call out is that this model is static. It doesn't learn from new patterns, and the underlying dataset has a fixed cutoff. New slang, coded language used by bad actors, and emerging content moderation challenges won't show up in results until the next model update. If your environment is highly dynamic — say, a platform where new communities form rapidly — you should supplement this with your own rules or a periodic revalidation against fresh examples. Relying on the model alone for anything beyond a baseline filter is usually a mistake. GitHub repo: github.com/danielmiessler/iwfne
PyPI: pypi.org/project/iwfne/