Getting Started With Dark Moon Apollo And The Whistle Blowers
I first ran into Dark Moon Apollo And The Whistle Blowers back in 2019 when a colleague handed me a zip file that was supposed to automate some report extraction. The first thing you need to understand is that this isn't a single program you download from some polished website. It's more like a collection of scripts and utilities that work together, and the documentation someone put together years ago is the closest thing you have to a manual. I've spent enough hours wrestling with this thing that I figured I'd write down what actually works instead of what the readme claims should work. The process starts with cloning the repository or grabbing the latest release archive. When I tried this on a fresh Ubuntu 20.04 install, the dependency resolution alone took about forty minutes because the project pins several Python packages to versions that conflict with each other. The workaround I ended up using was creating a virtual environment and installing the requirements in a specific order rather than running pip install -r requirements.txt straight away. Install cryptography and requests first, then the rest of the file line by line until you hit a breakage point. That tells you which version constraint is causing the problem. Once the dependencies are sorted, you run the setup script. It creates a configuration directory in your home folder at ~/.dark_moon_config. This is where all your keys, API tokens, and connection strings live. Don't skip the chmod step on this directory. I learned that the hard way when a coworker pulled my config folder off a shared network drive and accidentally exposed credentials that took us two weeks to rotate.
What It Actually Does
At its core, the system is designed for secure document collection and automated reporting across multiple data sources. Think of it as something between a data aggregator and a compliance pipeline. You point it at various endpoints, define which fields you care about, and it pulls, transforms, and archives the results. The whistle blowing component refers to how it handles flagged or unusual data patterns — it doesn't just store them, it creates an immutable log and can route notifications through pre-configured channels. Most tutorials online show you the happy path where everything connects perfectly. In practice, connection timeouts are extremely common. The default timeout is set to ten seconds, which is fine for local development but gets painful when you're pulling from servers that might be in different time zones or under load. I changed mine to thirty seconds by editing the config.yaml file and adding a timeout_override section. It's not documented anywhere in the official docs, but the code explicitly checks for that key.
Common Pitfalls
One thing beginners consistently mess up is the field mapping step. The system expects a particular JSON schema for your input definitions, and if even one field name is misspelled, the entire pipeline silently skips that dataset rather than throwing an error. I spent three days once chasing missing data before I realized the field in the source was called employee_id with an underscore and my mapping had employeeId in camelCase. The system didn't complain. It just produced empty results for that particular data source and moved on. Another issue is disk space. The archival feature stores every snapshot by default, and the compression ratio varies a lot depending on the data type. With text-heavy reports, you might get good compression. With binary attachments or logs, you can easily burn through fifty gigabytes in a week. I set a retention policy of ninety days and enabled the cleanup schedule, which has kept my storage usage manageable. The default setting is no retention policy, which feels like an oversight on the developers' part.
Get the Full Details

Running Your First Extraction
After configuration, the basic command looks like this: dark_moon extract --config ~/.dark_moon_config/settings.yaml --source all. This runs against every source you've defined. For a typical small team setup, this takes around twelve to fifteen minutes depending on how many endpoints you have configured and their response times. If you only need one source, you can specify it directly and cut that down to under a minute. The output lands in ~/.dark_moon_config/output/ by default, organized by date and source name. Each extraction generates both a summary report and the raw data file. The summary is useful for quick reviews, but if you're doing actual analysis you'll want to work from the raw files because the summary occasionally drops rows that contain null values in non-critical fields.
When This Tool Isn't The Right Call
I should mention that Dark Moon Apollo And The Whistle Blowers isn't meant for high-frequency or real-time data needs. The architecture is built around batch processing, so if you need results within seconds of a source being updated, this isn't going to help you. For those scenarios, I've used Apache NiFi or simple cron-based Python scripts with requests and got faster turnaround with less overhead. The tool also struggles with sources that require two-factor authentication or CAPTCHA verification on their endpoints. I tried hooking it up to one government database that used session-based auth with rotating tokens, and after about two weeks of failed connections I switched to doing that particular pull manually and feeding the exported CSV into the system. It's not elegant, but it's reliable.
The Whistle Blower Component
This is the part people are most interested in. The anomaly detection module watches for deviations from historical patterns in your data. When something crosses the threshold you set, it flags the record, creates a snapshot, and sends it to whichever notification channels you've configured. The thresholds are configurable per-field in the YAML config. I found that setting them too tightly produces a lot of false positives — I was getting three or four alerts per day during a normal week. Widening them to a two standard deviation threshold dropped that to maybe one every two weeks, which is actually actionable. The immutable log feature is solid. Once something is flagged, it's stored in a way that makes it very difficult to accidentally delete or modify. That's been useful in our case because we've had situations where someone tried to retract a report after it was already processed, and the log showed exactly what was there and when. The system doesn't prevent deletion outright, but it keeps a separate audit trail that survives most cleanup operations. If you're just starting out with Dark Moon Apollo And The Whistle Blowers, I'd recommend spending your first week just getting a single data source working end to end before you try connecting multiple sources at once. The system is powerful enough that trying to scale it immediately will introduce enough moving parts that you won't be able to tell whether a failure is in the configuration, the network layer, or the data itself. Take it slow and you'll save yourself a lot of frustration.
