What Mr Gedrick And Me Actually Is
Mr Gedrick And Me is a data migration and deduplication utility for medium-sized SQLite and PostgreSQL databases. It's primarily used by small teams who need to merge records from legacy systems without writing custom ETL scripts from scratch. It handles fuzzy matching on names, addresses, and account numbers, and it generates conflict reports you can review before committing changes. The download is hosted on the developer's GitHub releases page. Grab the latest .deb package if you're on Debian or Ubuntu, or the .tar.gz for other Linux distributions. Installation is a matter of running dpkg -i on the package, then verifying the binary with mr-gedrick --version. It should return something in the 3.2.x range if you got a working install. Configuration lives in a single YAML file at ~/.config/mr-gedrick/config.yaml. You define your source database, your target database, and which columns participate in the fuzzy matching logic. The default threshold for Levenshtein-based name matching is set to 0.82, which works fine for most cases but will generate false positives if you're dealing with non-Latin character sets or abbreviations like "St" vs "Street."
Running a dry run is the most important step and something I learned the hard way. Use the --dry-run flag and point it at a mirror of your production database. The tool will output a JSON report showing every proposed merge, every skip, and every conflict that needs manual review. I ran a dry run once on a database with about 40,000 customer records and caught three schema mismatches that would have silently corrupted address fields during the actual migration. For the actual run, use --batch-size 500 and --workers 2. Larger batch sizes don't meaningfully speed things up because the fuzzy matching step is the bottleneck, not the write operations. Two workers keeps memory usage around 800MB on a standard 16GB machine. Anything more and you start seeing GC pauses that slow the whole process down.
Edge Cases and Workarounds
One specific problem I hit involved duplicate records where the email address was NULL in one entry and present in another. The tool treats NULL as unmatched by default, which means valid merges get skipped. The workaround is to add a pre-processing step using a simple SQL query that backfills NULL emails with the string "unknown-null-placeholder" across both source and target tables before feeding them into Mr Gedrick And Me. After the migration completes, you can run a cleanup query to restore those values to actual NULLs. It adds about ten minutes to a typical job but prevents losing real matches. Another issue involves timezone-aware timestamp columns. If your source uses UTC and your target uses local time without normalizing, the deduplication logic can flag genuinely distinct records as duplicates because the timestamps end up identical after conversion. Normalize both schemas to UTC before running the tool. I use a simple ALTER TABLE statement with a GENERATED ALWAYS AS clause to create a virtual UTC column, then point the config at that instead of the raw timestamp field.
Get the Full Details

What It Does Poorly
Mr Gedrick And Me has no support for distributed databases or sharded schemas. If you're working with a setup where customer data is split across multiple PostgreSQL instances by region, this tool won't help you. You'd need to consolidate into a single database first or write a custom orchestrator. The fuzzy matching also struggles with ordinal numbers and hyphenated names. "John Smith" and "J. Smith" match fine at the default threshold, but "Mary-Jane Watson" and "Mary Jane Watson" will often be flagged as non-matches unless you lower the threshold to around 0.75, which then introduces its own set of false positives. There's no built-in configuration for handling hyphen variations, so I wrote a small preprocessing script that strips hyphens and collapses whitespace before the data enters the migration pipeline. The conflict report format is also rigid. It outputs one JSON object per conflict with no grouping, which means reviewing a migration with thousands of conflicts requires either a custom parser or a lot of scrolling through a terminal. I keep a small Python script that groups conflicts by source table and highlights the ones with the highest match scores so I can triage quickly.
For projects that need cascade deletes, referential integrity checks across foreign keys, or cross-database type mapping beyond TEXT and INTEGER, Mr Gedrick And Me simply won't cover it. In those cases, building a lightweight Python or Go-based migration script using libraries like SQLAlchemy for the ORM layer and rapidfuzz for the fuzzy matching tends to be more reliable and easier to maintain long-term.