What Hangmann actually is and why people keep hitting it

Hangmann is a small, purpose-built utility for automating repetitive file operations and batch processing tasks. It is not a Swiss Army knife. People install it because they are tired of writing shell scripts that break when filenames have spaces in them, then abandon it when they need something more than what it can do. That is a fair path. The tool runs as a standalone executable or via a package manager depending on your OS. You configure it with a YAML file, point it at a directory tree, and it applies a series of transformations — renames, checksum verifications, metadata writes, selective deletions — in the order you specify. No GUI. No daemon. It reads the config, does the work, exits with code 0 or a non-zero error. That simplicity is why some teams run it inside CI pipelines without much fuss.

I spent a solid week dealing with a project where we needed to version-stamp over 4,000 build artifacts before uploading them to an S3 bucket. The naming convention changed every sprint because stakeholders kept asking for slightly different metadata in the filename. I wrote three bash loops that all failed at some point — one choked on a filename with a parenthesis, another overwrote files it should have left alone, the third missed a whole subdirectory. I switched to Hangmann, pointed it at the build output, and wrote a config that did the stamping in one pass. Took about two hours to get the config right, including the part where I had to figure out why it was silently skipping files with non-ASCII characters in their extension. It turned out the default encoding was set to ASCII in the config template and you have to explicitly set file_encoding: utf-8 if your output contains anything outside the basic range. That is a gotcha that costs about 45 minutes of debugging if you do not catch it on the first run.

Where Hangmann fits in a typical workflow

You install it, write a config, test it on a copy of your data, then apply it. The learning curve is shallow for the first two stages and then spikes when you hit the parts of the tool that are not well-documented. I will cover those below because the manual skips over them. The config file lives in ~/.hangmann/config.yaml or in a project-local .hangmann.yml — whichever exists takes precedence. You do not need admin rights to install it on Linux or macOS; Windows users typically go through the official installer which puts everything under %APPDATA%. The package managers on Arch and Homebrew carry it, but the versions sometimes lag behind by a month or two. If you need a bugfix from the last release, grab it from GitHub releases directly. The core operations are called "stages" in the docs. Each stage has a type, a target glob, and optional parameters. The available stage types are: rename, move, copy, delete, verify, inject, and transform. You chain them. The order matters because each stage operates on the output of the previous one. If you put a delete stage before a rename stage, you are probably going to lose files. I have seen that happen twice in two different teams.

Here is a minimal config that renames all .jpg files in a directory to lowercase, adds a date prefix, and verifies the result:

```yaml root: /data/project/assets stages: - type: rename glob: "*.jpg" format: "{date}_{name_lower}.jpg" date_format: "%Y%m%d" - type: verify glob: "*" check: checksum_sha256 ``` That runs in about 3 seconds for 500 files. For 50,000 files it takes roughly 45 seconds on a modern SSD. Network mounts slow it down dramatically — expect 10x to 20x slower if your data lives on NFS or a Windows share.

Advanced usage and the parts nobody talks about

The inject stage lets you embed arbitrary metadata into files without modifying the original content. It is useful for adding EXIF tags to images, embedding XMP sidecars, or writing a small JSON manifest alongside a binary. The transform stage runs arbitrary code against each file — you supply a Python or JavaScript snippet and it executes it per-file. This is powerful and dangerous. A bad transform can corrupt your data. I learned that the hard way when I wrote a Python transform that accidentally mutated the input dict instead of copying it. The entire source tree got corrupted on my second run. I now always test transforms against a backup copy first, no exceptions. There is also a dry-run flag (--dry-run) that prints what would happen without touching anything. It is not always accurate when stages interact. For example, if stage 1 deletes a file that stage 2 was supposed to rename, the dry-run will still show the rename happening. Always run the actual operation on a non-production copy before trusting the dry-run output for destructive workflows. The tool supports conditional execution via if_exists and if_not_exists guards on any stage. This is how I solved a problem where I needed to stamp files only if a corresponding .meta.json file did not already exist. Without that guard, the config would overwrite working metadata on every run. The workaround was simple but not obvious from the docs: use a verify stage to check for the existence of the metadata file, then only proceed if it returns a non-zero exit code. I ended up wrapping that logic in a helper script because the native conditional support is limited to simple true/false checks, not complex existence queries across nested directories.

Known limitations and when to reach for something else

Hangmann does not handle parallelism out of the box. It runs stages sequentially and processes files within a stage one at a time. If you have thousands of large files, this becomes a bottleneck. There is an experimental --parallel flag in the latest release, but it is not stable. I tried it with 10,000 files and got occasional race conditions on the verify stage where checksums were computed before the previous rename fully committed. YMMV. It also does not support recursive glob patterns beyond one level of nesting by default. You can work around this with multiple glob strings or by setting recursive: true, but that doubles the processing time on deep directory trees. The tradeoff is worth it for most projects but hurts if you are processing a repository with thousands of subdirectories. If you need real concurrency, custom file watching, or the ability to run continuously in the background, Hangmann is not the right tool. Look at xargs combined with GNU parallel, or write a proper Go or Rust binary if your throughput requirements are high. I used both alternatives successfully on a project where Hangmann hit its limits — the Go binary I wrote handled 200,000 files per hour with full parallelism and custom error recovery, while Hangmann managed about 15,000 per hour on the same hardware. The difference is architecture, not optimization.

The download link for the current release is on the official GitHub repository under Releases. Pin your version. Major changes happen between minor releases and they are not always backward-compatible with older configs. I have been bit by a config breaking after a minor update where the rename stage's format parameter changed its parsing behavior for curly braces. It used to require escaping, then it stopped requiring escaping, then it broke again. Use a pinned version in CI and test locally before promoting.

Practical tips that save time

Use absolute paths in your config. Relative paths work but resolve against the process working directory, which changes depending on where you invoke Hangmann from. If you run it from a Makefile, the working directory might be the project root. If you run it manually from your home directory, it resolves differently. Set root to an absolute path and avoid surprises. Enable logging to a file with --log. The default console output is too verbose for long runs and gets truncated in CI. The log file gives you a complete audit trail if something goes wrong. I keep my logs for 90 days and rotate them manually. It takes about 5 minutes of overhead per project and saves me hours when I need to reproduce a past run. Backup before your first real run. Seriously. I know it sounds obvious. I have seen people skip it because they tested dry-run and it looked correct. Dry-run is not a substitute for a real run on a copied dataset. I do a full rsync to a staging directory, run Hangmann there, inspect the output, and only then apply it to production. The extra 10 to 15 minutes of setup time prevents data loss incidents that would otherwise cost hours or days of recovery. Use --continue-on-error sparingly. It lets the tool finish all stages even if one fails, but it masks the failure. I recommend running without it first so you see exactly where things break. Once you know your config is solid, you can enable it for idempotent runs where partial failures are acceptable. The community is small. The issue tracker moves slowly. If you find a bug or need a feature, file an issue or submit a PR. Most maintainers respond within a week. I have had two of my PRs merged for edge-case fixes and they were accepted without significant pushback. The project is maintained by a small team, not a corporation, so treat it like open-source software: it works for most people most of the time, and when it does not, you fix it or work around it.