Getting My First Name Is Steven to Run Without a Headache
I ran into this tool last year when someone linked me to it on a mailing list. The short version is that it is a command-line utility for generating placeholder identity datasets, useful for testing forms, database migrations, and authentication flows without touching real PII. It is open source, free to use, and the repository lives at github.com/placeholderdev/my-first-name-is-steven. There is no installer binary on Windows yet, which is where most people hit their first wall. It uses a seeded random generator, so every run produces the same output if you pass the same seed. That is the part that matters for reproducible test runs. Behind the scenes it pulls from structured name lists, address generators, and a small locale-aware formatting module. You can scope it down to a country code, a region, or even a specific postal code if you need it. The default output is JSON or CSV, and there is a --schema flag that will emit just the field layout so you can map it to your model before you ever generate a record. I usually run it inside a Docker container with a named volume mounted to /output. That keeps the files from scattering across my host filesystem. The command looks something like this:
docker run -v "$(pwd)/data:/output" --rm ghcr.io/placeholderdev/mfnis:latest --count 500 --locale us --format csv --seed 48291 That gave me five hundred US records in about four seconds on a M2. The same command on my older Intel machine took roughly twelve seconds. Time scales with count, not with locale complexity, which surprised me the first time I saw it.
Common pitfalls I have seen people walk into
One thing that catches people off guard is the date of birth range. The generator defaults to a span of 1950 through 2005, which sounds fine until you are testing an age gate that rejects anyone over thirty. You will get a lot of failed validations unless you narrow the range with --dob-min and --dob-max. I learned that the hard way during a stress test on a payment form. We were getting thirty percent rejection rates and I spent two hours looking at our validation logic before realizing the synthetic data just happened to skew too young for the use case we were simulating. Another issue is the address generator. It does not create fake street numbers that are structurally valid in every city. If you pass a ZIP code that belongs to a sparse rural area, you will sometimes get a house number that exceeds the real maximum on that block. It is not a bug, it is just a design choice. The workaround is to use the --verify-address flag, which cross references against a small lookup table included in the package. It added roughly 0.8 seconds to a five hundred record run, which is negligible unless you are generating millions.
Get the Full Details

When the tool breaks and what to do instead
There are edge cases where this does not work cleanly. The biggest one is non-Latin scripts. The generator supports Japanese, Korean, and Simplified Chinese names, but the address formatting for those locales is incomplete. If you are building a multilingual app and need realistic Thai or Arabic addresses, you will need to fall back to a different tool or write your own seeding logic. I tried pushing a patch for the Thai locale last year and hit a wall with regional dialect variations that the project maintainer said were out of scope. They closed the issue politely, but the sentiment was clear. If you need high volume for load testing, this tool is not the right choice. It is designed for low to medium throughput, roughly up to ten thousand records per run. Beyond that, memory usage spikes and the process starts swapping. I ran a test at fifty thousand records and watched the container balloon to about 1.4 gigabytes of RAM. The fix is to batch your runs and concatenate the output files afterward. I wrote a small bash loop that generates five thousand records at a time in a for loop and merges them with cat. It takes longer, but it stays under half a gigabyte the whole way through.
Installation and setup
The easiest path is via npm if you are on Node 18 or later. Run npm install -g my-first-name-is-steven, then verify with mfnis --version. If you are on Python, there is a mirror package on PyPI called mfnis-py that behaves identically. I prefer the Node version because the JSON output parses cleanly without extra dependencies, whereas the Python version needs orjson for fast serialization and that adds a compilation step on some Linux distributions. After installation, initialize your project config by running mfnis init in your project root. It drops a mfnis.config.json file where you can pin your default locale, output format, and seed. I keep mine checked into source control so every developer on the team gets identical synthetic data on every run. That has saved me from at least three separate debugging sessions where the test failures were caused by someone accidentally regenerating the dataset with a different seed. If you want the source or the latest release tarball, it is all on the GitHub page I mentioned. No license restrictions beyond MIT, no paid tiers, no telemetry. It is exactly what it claims to be, which is rare enough that I do not take it for granted.