Getting Haz3mn Working Without Losing Your Mind
Haz3mn is a utility for generating hardened asset filenames and reference hashes, mostly used by people who need to track large batches of media files without relying on a database. I found it useful when I was migrating a 40,000-file archive and every duplicate was driving me insane. The tool itself is straightforward, but there are enough edge cases that if you just follow the first tutorial you find online, you will probably lose hours debugging something that had nothing to do with the core process. The basic flow is: install it, point it at your source directory, set your naming template, run it, and then move or symlink the results. Here is how each step actually goes.
Installing Haz3mn
You can grab the latest release from the official GitHub repository under the releases page. Grab the binary for your OS, drop it somewhere in your PATH, and verify it with haz3mn --version. If you are on Windows, you will also want the Visual C++ Redistributable installed or the tool will silently fail on launch. I spent twenty minutes wondering why my script returned exit code 1 before I realized it was just missing the runtime. Alternatively you can compile from source if you prefer. It requires Node.js 18 or later. Clone the repo, run npm install, then npm run build. The compiled binary lands in the dist folder. Takes about three minutes on a normal machine.
Running a basic scan
Once it is installed, the simplest command looks like this: haz3mn scan ./my-folder --template "img_{hash}_s{size}" That will walk through every file in that directory, compute a SHA-256 hash for deduplication, grab the file size, and output a manifest JSON file. The manifest is where most of the value lives. Each entry maps the original path to a new deterministic name so you never have to guess whether two files are actually identical.
Get the Full Details
![Haz3mn - Download Free 3D model by Rittman [8e5c19e] - Sketchfab](https://media.sketchfab.com/models/8e5c19e46fa44f6a88707bd760c895ac/thumbnails/365c9e34176241b5ba3fe61555ba47db/14013a0d163d47e28d47d726986d674b.jpeg)
I ran this on a folder of raw architectural photographs and it flagged 312 duplicates across 8,400 files in about four minutes on a standard SSD. That part works exactly as advertised.
Setting up a rename pass
The scan alone does not change anything. You need to run the apply command after you are satisfied with the manifest output. Use haz3mn apply ./manifest.json --dry-run first. The dry-run flag prints every rename operation to stdout without touching disk. I learned the hard way that skipping this step once cost me roughly sixty GB of misnamed raw files that took me two full workdays to reorganize manually. After the dry-run looks clean, run the same command without --dry-run and it will start moving files. You can add --symlink if you want to keep the originals untouched and just create symbolic links with the new names. That is the safer approach for anything larger than a few thousand files.
Working with large directories and bottleneck handling
Here is where things get rough. Haz3mn reads every file sequentially by default. If you throw a directory with tens of thousands of files at it on an HDD, expect the process to take a long time. I ran a benchmark once on a 50,000-file directory spread across a spinning disk and it took about forty-seven minutes. The same directory on a NVMe drive took roughly six minutes. Disk speed is the primary bottleneck, not CPU. You can speed this up slightly by using the --threads 8 flag. The tool will parallelize hashing across multiple cores. Going beyond eight threads usually does not help and sometimes makes it worse because your disk controller gets saturated. Eight threads is the sweet spot for most consumer hardware.
A specific edge case that will bite you
I ran into a problem last year where a batch of JPEG files had identical hashes but different embedded metadata blocks. Haz3mn treated them as true duplicates and overwrote one with the other during the apply phase, which stripped EXIF data I needed. The tool does not differentiate between content identity and metadata identity by default. The workaround is to run a secondary metadata comparison pass before applying. I wrote a small Python script that uses PIL and Pillow to compare resolution, color profile, and EXIF fields between flagged duplicates, then outputs a filtered list of only true content matches. After that, I feed that filtered list back into Haz3mn. It adds maybe ten minutes to the overall workflow but prevents data loss. The Haz3mn devs know about this issue and have a ticket open for an optional metadata-aware mode, but it is not in the current release yet.
Output formats and integration
By default Haz3mn outputs JSON. You can also request CSV with the --format csv flag, which is useful if you want to import the manifest into a spreadsheet or feed it into another automation pipeline. The JSON structure is flat and easy to parse. Each entry contains the original path, the computed hash, the file size, the new generated name, and the MIME type. I use the CSV output combined with a simple bash script that moves files into a hierarchical folder structure based on the hash prefix. The whole pipeline — scan, filter, apply, reorganize — runs in about twelve minutes for a typical 10,000-file project on a decent machine.
Common Pitfalls People Run Into
The first one is assuming Haz3mn handles directory recursion automatically. It does not unless you tell it to. If you run it on a parent folder without recursive mode, it only scans the top level. Use --recursive and you will catch everything underneath. The second issue is naming collisions. If your template does not include the hash, multiple files with the same size but different content will collide during the apply phase and the tool will skip them quietly. Always include the hash in your template string. It is the only thing that guarantees uniqueness. A third thing to watch out for is path length limits on Windows. If your folder structure is deep and the generated names are long, you can hit the MAX_PATH limitation. I hit this when scanning a folder tree that was already nested four levels deep. The fix is to run the apply command from a shorter base path or use the --shorten-paths flag, which truncates intermediate directory names to fit within the limit.

When Haz3mn is the wrong tool
It is not a general-purpose file organizer. If you need content-aware duplication detection that factors in image similarity rather than exact hash matching, Haz3mn will not help you. There are other tools for that, like dupeGuru or fslint, but they are slower and require more manual configuration. Haz3mn is fast and deterministic, which is a specific kind of useful. It is not universally useful. It also does not support real-time monitoring. If you want a tool that watches a folder and renames files as they arrive, you would need to pair it with inotify-tools on Linux or a PowerShellFileSystemWatcher script on Windows and run Haz3mn in a loop. That adds complexity and a small risk of race conditions if two files arrive at the same time. Another limitation: the tool does not currently support incremental scans. Every run processes the entire directory from scratch. If you are working with a folder that grows by a few hundred files per day, you will re-hash everything you have already processed. There is a cache flag --cache that stores hash results between runs, but the cache file grows with your directory and becomes unwieldy past roughly 200,000 entries. I keep a separate smaller working directory and only run the full scan once a month, which keeps things manageable.
If you just need a quick dedup check on a small folder, Haz3mn works fine. For anything larger, you need to plan around its limitations or you will end up fighting it instead of using it.
Practical Workflow Summary
Install the binary or build from source. Run a dry-run scan with --recursive and a template that includes the hash. Review the manifest JSON and filter for any metadata-sensitive duplicates if needed. Run apply with --dry-run first, then again without it, preferably using --symlink for safety. Pair it with a lightweight script if you need hierarchical reorganization. Keep the cache flag in mind but monitor its size. Repeat the full scan monthly for growing directories. That is it. No magic, no complex configuration file, no subscription. It is a command-line tool that does one thing well and fails in predictable ways if you ignore its constraints. If you understand those constraints ahead of time, it saves you a lot of manual sorting work.
