Understanding Hexaunaut

Hexaunaut is a tool I've been working with for about eighteen months now. It started as something for parsing binary data, then expanded into a broader analysis suite. The name comes from hex + navigator, which is exactly what it does — you feed it raw bytes and it walks you through the structure without forcing a schema on top. I should say upfront that Hexaunaut isn't built for everyone. If you're looking for a point-and-click GUI that tells you what your data means, this isn't it. It works best when you already have a hypothesis about the format and need a fast way to verify or disprove it.

Hexaunaut in practice

Here's how I use it on a typical day. I'll drop a file or pipe raw bytes into it, run the auto-detect pass, then manually trace through whatever structural patterns show up. The initial scan usually takes about 10-15 seconds on a file under 50MB on my machine (quad-core i7, 32GB RAM). After that, it depends on complexity. A straightforward fixed-width record layout might take another five minutes to parse cleanly. Something with nested variable-length fields and endian mismatches can go all afternoon. The detection engine tries to identify common structures — fixed offsets, repeating patterns, probable field boundaries. It's not perfect. I've hit cases where it confidently claimed a header was a length-prefixed array when it was actually a bitmask. The workaround I use is to stop trusting the auto-classification after the first pass and instead look at the raw hex dump alongside the parsed view. Side by side makes mismatches obvious within a couple of minutes.

How to get Hexaunaut running

Get the latest release from the official repo: github.com/hexaunaut/hexaunaut. I'd recommend pulling the tagged release rather than main branch unless you need a feature that was merged yesterday. The release builds have regression tests passing. Main branch is where breaking changes happen. Build time is roughly three minutes on a standard machine. Dependencies are minimal — Go 1.21+, a C compiler if you're building the optional native extensions (you don't need them for basic operation). The Go module handles the rest. Installation is just go install ./... followed by adding the binary to your PATH. I skip the full install and run from the source tree during development. Saves me having to rebuild after every config tweak.

First run expectations

On your first launch, you'll see the config generator. It asks about default byte order, whether to treat unknown structures as opaque blobs or attempt heuristic classification, and output format preferences. I set byte order to little-endian by default since that's what 90% of the data I touch uses. The opaque blob fallback is important — without it, the tool will waste cycles trying to force a structure onto random data and fill your terminal with false positives. Config lives at ~/.hexaunaut/config.json. I keep a copy of mine checked into a private repo. The format is straightforward and versioned, so migration between releases has always been painless for me.

A workflow that actually works

Here's the process I've settled on after going back and forth for months: Step one: get a sample file. Even 4KB is enough for initial orientation. Put it in a scratch directory and run hexaunaut analyze sample.bin. The output goes to stdout by default, which is fine for quick checks. For anything longer than a few seconds of analysis, redirect to a file or use the built-in log option. Step two: look at the structural map. Hexaunaut will output a tree of detected fields with confidence scores. Ignore fields below 60% confidence for now. They're usually noise or coincidental patterns that happen to match something in the heuristic library.

Step three: validate with known boundaries. If you know the file format has a magic number or fixed header, check that first. Hexaunaut should flag it. If it doesn't, either the magic number isn't in the heuristic database, or you're looking at something genuinely non-standard. Both are useful findings. Step four: iterate. Take whatever you've confirmed, build a custom parser rule, run it against the same file. This usually cuts manual inspection time from hours down to about twenty minutes for well-structured formats. Poorly structured or obfuscated data? Could still take a while. No tool fixes fundamentally ambiguous input.

Pitfalls I've hit

Endian confusion is the biggest one. Hexaunaut detects endianness heuristically but gets it wrong on files that mix representations — which happens more often than you'd think in legacy formats. When this happens, field values look reasonable but the internal structure collapses on close inspection. The fix is to check cross-field consistency: do derived values make sense together, or do they contradict each other? Another issue is what I call detection fatigue. After running analysis on twenty similar files, the heuristic library starts producing repetitive false positives. The tool doesn't have a learning component that reduces this, so you have to manually tune confidence thresholds or write exclusion rules. I spend maybe ten minutes per project adjusting the threshold config, then move on. The third gotcha: large files. Hexaunaut loads everything into memory during the initial scan. I hit an OOM crash on a 2.3GB capture file. The workaround is to pipe through dd with a smaller block size or use the --slice flag to analyze in chunks. Chunk analysis is slower but avoids the memory spike entirely.

When Hexaunaut isn't the right tool

It struggles with fully encrypted data, obviously. Also with formats that deliberately avoid structure — some DRM implementations and custom game engines do this on purpose. If your input is fundamentally random from a structural perspective, Hexaunaut will find patterns anyway because the heuristics are aggressive. That's a feature and a problem. You get results fast, but you need to verify them carefully. For those cases, I fall back to manual hex inspection or use a different tool like Hexinator or even just xxd with a editor. Hexaunaut excels at semi-structured or partially understood formats. It's mediocre at completely unknown binary with no anchor points.

Advanced configuration

The config file supports custom heuristic rules, field-type definitions, and output formatters. I write custom formatters in Go — they're just interfaces with a single method. Takes about an hour to get something reasonable working if you know Go. Not worth it unless you're doing this repeatedly for the same format family. Custom heuristic rules are more immediately useful. I've added rules for proprietary protocols used in industrial equipment. The syntax is pattern-based with optional constraints. A typical rule looks like a regex with metadata attached: field name, type annotation, confidence weight. I spend maybe fifteen minutes writing a rule for a new format, then it pays for itself on the second file. Rule files live in ~/.hexaunaut/rules/. They're YAML. The tool watches for changes and reloads automatically — no restart needed. This saved me probably twenty hours of downtime over the past year alone.

Bottom line

Hexaunaut is solid for what it does. It's not magic. It won't solve formats that resist structure, and it will occasionally lead you astray with confident but incorrect classifications. Use it as a starting point, not an endpoint. The time savings are real — I'd estimate 60-70% reduction in initial analysis time compared to doing everything by hand — but verification still requires a human who understands the domain. If you're doing binary analysis regularly and your formats have any structure at all, it's worth the thirty minutes it takes to get set up. If you're doing one-off analysis on a never-before-seen format with no reference material, expect to spend more time tuning than you would just reading the hex dump in an editor. Context matters more than the tool.