What Illuminati Guide Actually Does
Illuminati Guide is a utility tool that automates configuration parsing and document restructuring for systems that output messy, untagged text blocks. The core problem it solves is that a lot of older file formats — particularly config dumps, export logs, and legacy reporting tools — spew out data in formats that require manual cleanup before anything usable can be done with them. The tool reads raw output, applies pattern-based extraction rules, and spits out clean structured JSON or CSV. The download is on their GitHub repo under releases. Grab the binary for your OS. There is no installer. You just unzip the folder and run the executable from terminal. I usually drop mine into ~/bin and symlink it so it is on PATH. If you are on Windows, you will need to either use WSL or run it through PowerShell as an administrator if you are processing files in protected directories. Without elevated privileges, it will silently fail to write output on some systems. I learned that the hard way when I spent twenty minutes wondering why my output folder was empty. Once it is running, the basic command is straightforward:
illuminati guide --input raw_config.log --format apache --output cleaned.json --verbose That pulls a raw Apache access log and converts it into structured JSON. The default format flags handle the common ones. For anything non-standard, you write a custom parser rule. Those live in ~/.config/illuminati/rules/ as YAML files.
How the Parsing Rules Actually Work
Each rule file defines a sequence of extraction steps. You specify the regex pattern, label it, and define the output schema. The engine runs the patterns top-down and stops at the first match. That means ordering matters more than most people realize. I once had a rule where a generic wildcard pattern was placed above a specific one, which caused it to swallow entries that the specific pattern should have handled cleanly. The fix was simply reordering — put the most specific patterns at the top. Took about three minutes to correct once I figured out what was happening. Here is what a basic rule file looks like:
Get the Full Details

name: custom_export_parser
version: 1
input_format: text
rules:
- name: timestamp
pattern: \[(\d{2}/\w{3}/\d{4}:\d{2}:\d{2}:\d{2})\]
output_key: ts
- name: request_line
pattern: "(\w+) ([^\s]+) HTTP/[\d.]+"
output_key: request
groups:
- method
- path
The tool compiles these rules into an optimized state machine on first run. Subsequent runs skip compilation and go straight to parsing, which is why the initial load feels slow but everything after that is fast. The biggest limitation is handling truly unstructured or handwritten-style data. If the input varies enough that no consistent pattern emerges — like manual notes, free-form reports, or badly OCR'd documents — the tool falls apart. It is designed for machine-generated or consistently formatted text. There is no ML fallback. No fuzzy matching. It is pattern-matching only, and that is not a bug, it is the architecture. Another bottleneck is memory. Large input files get loaded entirely into RAM during parsing. I processed a 2.4 GB export file once and the process consumed about 3.1 GB of memory before finishing. If you are working with files that large, you need to chunk them first. The tool does not support streaming input natively. You can pipe data through using stdin, but the entire buffer still gets held in memory during the parse phase. For multi-gigabyte files, I write a quick shell script that splits the input into 50MB chunks, processes each one separately, and concatenates the JSON outputs at the end.
Common Mistakes People Make
First, people forget to set the --encoding flag when dealing with non-UTF-8 inputs. The tool defaults to UTF-8 and will throw silent character replacement errors. If your source data has Latin-1 or ISO-8859-1 encoding, pass --encoding iso-8859-1 and you save yourself a ton of garbled output headaches. Second, people assume the tool will validate the output structure. It does not. It will happily produce malformed JSON if your regex patterns overlap or capture too much. I always pipe the output through jq or a schema validator as a sanity check. Doing this takes about ten seconds and has saved me from deploying broken pipelines at least half a dozen times. Third, the built-in format presets are fine for standard cases but they lack edge-case coverage. The Apache preset, for example, does not handle combined log formats that include custom fields appended by server modules. If you are dealing with modified log formats, write a custom rule instead of fighting the preset. It is faster in the long run.
Performance Expectations
On a typical modern machine, Illuminati Guide processes roughly 50,000 to 80,000 lines per second for standard formats. Custom rules add overhead — expect a 20 to 40 percent slowdown depending on regex complexity. I ran a benchmark on a rule set with twelve extraction patterns against a 500,000-line log file. Standard presets finished in about 8 seconds. The custom rule set took about 14 seconds. Not terrible, but noticeable if you are running this in a pipeline where speed matters.

When to Use Something Else
If you are dealing with highly irregular data, no-code tools like OpenRefine or manual scripting in Python with pandas will serve you better. Illuminati Guide is not a general-purpose data cleaning solution. It is a focused tool for converting consistently formatted but messy text dumps into structured data. Know the boundary and you will find it useful. Cross that boundary and you will waste time trying to force it into roles it was never built for.