Recovery Tools You Actually Need
When Bad Things Happen isn't something you buy on a whim. It's the thing you're grateful exists after your server cluster decides to stop being a cluster and starts being a collection of very expensive paperweights. I still remember 3 AM calls from clients who had skipped their recovery drills because "everything was running fine." It never is. It's a disaster recovery and system restoration framework designed for environments where downtime directly equals revenue loss. Unlike the backup-and-hope approach most IT departments rely on, this gives you automated failover, point-in-time restoration, and cross-region redundancy without requiring a team of five DBAs to operate it. The core workflow takes an image of your production state at configurable intervals, compresses it, and replicates it to a staging zone that can be spun up independently. I set this up for a mid-size e-commerce platform last year. They were using a combination of manual snapshots and whatever their hosting provider offered by default. The host offered daily incremental backups with a retention window of fourteen days. Fourteen days. In that window, ransomware has an entire fortnight to rot your data before anyone even notices the disk latency spike. I moved them to hourly application-consistent snapshots with three-day retention on hot storage and a seven-day archive in cold storage, replicated across two regions. The initial configuration took about four hours. A proper failover test took thirty minutes afterward.
Installation and Initial Configuration
Grab the latest release from the official repository. The current version requires at least 8GB of RAM on the coordinator node and 16GB if you're running the analytics dashboard alongside it. Don't skimp on the dashboard node. I've seen people run everything on a single VM and then wonder why the backup windows kept overlapping with their peak traffic hours. Start by deploying the coordinator on a dedicated machine or container. Extract the archive, run the init script, and it will generate a configuration file at /etc/when-bad-things/config.yaml. The defaults are conservative and will work for a basic single-region setup, but you should modify at least these three fields before you go anywhere near production: snapshot_interval (default is every 6 hours, change it to 1 hour if your data changes significantly between snapshots), retention_days (minimum 3, ideally 7 for hot tier), and replication_target (this is where your secondary region endpoint goes). The initialization process creates a staging bucket in your configured storage backend and seeds the metadata database. This takes roughly ten minutes for a clean install. After that, you add your source nodes by editing the sources section of the config file. Each source needs an identifier, the paths you want to watch, and whether it should use application-consistent snapshots or just file-level copies. Application-consistent is worth the extra overhead for database-heavy workloads. File-level is faster and sufficient for static content or logging partitions.
A Problem I Ran Into
During a deployment for a financial services client, the coordinator kept failing health checks on the replication pipeline. The logs showed intermittent timeouts when pushing snapshots to the secondary region. The issue wasn't network bandwidth or DNS. It was the snapshot compression ratio hitting a wall with their particular dataset. They had large sparse files from old database exports that were being included in the backup path. The compression algorithm would try to process them, burn CPU for minutes at a time, and then time out before finishing. The workaround was adding an exclude pattern for *.bak and *.old in the source configuration, then running a one-time cleanup pass on those files before reinitializing the snapshot pipeline. After that, replication latencies dropped from an average of 45 seconds per snapshot to under 8 seconds. I learned to recommend a pre-flight scan that checks for large, low-value file types before anyone enables replication.
Get the Full Details

Running Your First Failover Test
Don't skip this. I can't stress this enough. A backup you haven't tested is not a backup, it's a wish. Run a dry-run failover with the wbth drill --target=staging command. This spins up a temporary instance from your most recent snapshot in the staging environment, runs a series of connectivity and integrity checks, and tears it down automatically. The whole process usually takes between 5 and 15 minutes depending on your snapshot size. If the drill passes, you have baseline confidence. If it fails, the output will tell you exactly which check broke and often which snapshot is corrupted. Fix it now, not during an actual outage. I once watched a client attempt a real failover and discover their latest snapshot had a metadata mismatch that the automated checks should have caught. They had disabled the pre-flight validation because it was "slowing things down." It slowed things down by three minutes during setup so they wouldn't lose three hours during recovery.
Counter-Intuitive Things Nobody Tells You
First, more frequent snapshots don't always mean faster recovery. There's a tipping point where the overhead of managing thousands of small incremental chains actually increases your restore time because the system has to traverse and validate each chain segment. For most workloads, hourly snapshots with a 72-hour retention window hits the sweet spot. Going beyond that rarely improves your RPO meaningfully and can degrade your RTO. Second, application-consistent snapshots are overrated for read-heavy workloads. If your primary concern is serving static pages or cached API responses, file-level copies with write-order fidelity are faster to restore and easier to debug. Save the application-consistent option for transactional databases and message queues where state integrity matters more than speed.
Known Limitations
When Bad Things Happen doesn't handle hybrid cloud migrations well. If you're trying to move a workload from on-premises to AWS or Azure while keeping it running, you'll hit friction around network latency and credential management between environments. The tool works best when both source and target are in the same cloud provider or both on-prem. Cross-cloud replication is supported but requires manual configuration of the replication endpoints and sometimes custom networking rules. It also struggles with encrypted volumes. If your disks use full-disk encryption, the snapshot engine can't read the blocks to create a consistent image unless you decrypt them first. This means either keeping a decryption key in the coordinator's keyring (which introduces its own security questions) or running snapshots on plaintext volumes and relying on transport-level encryption for the replication channel. Neither option is perfect. The second approach is more common in production and usually adequate. If you need true multi-cloud disaster recovery with automated failover between providers, you might be better served by pairing this with a separate orchestration layer like Terraform or Pulumi to handle the cross-cloud provisioning. When Bad Things Happen covers the snapshot and restore side reliably, but it won't manage the infrastructure-as-code piece for you.

Cost Expectations
The community edition is free and supports up to five source nodes with single-region replication. The professional tier runs about $200 per source node per month and adds cross-region replication, application-consistent snapshots, and the analytics dashboard. For a small team managing ten servers, that's roughly $2,000 per month, which is still cheaper than two hours of downtime for most businesses. The enterprise tier adds SSO, audit logging, and priority support, and pricing scales from there. Download the latest version from whenbadthings.dev/download. Check the compatibility matrix before installing on older Linux distributions. The current release drops support for kernels below 5.4, which catches a lot of people off guard.