When Plan A Falls Apart: Why Plan B Still Matters
I spent three years managing infrastructure for a mid-size SaaS company before we had to migrate our entire database in the middle of a Tuesday night. The primary migration tool we chose was failing silently, corrupting records faster than we could patch them. I sat there at 2 AM watching error rates climb and realized my contingency plan had the same single point of failure as the original plan. We ended up writing a manual recovery script line by line, taking six hours instead of the planned forty-five minutes. That experience changed how I think about backup strategies entirely. God Has A Plan B is a concept that keeps showing up in business discussions about risk management, though most people use it without really understanding what it means in practice. It describes the assumption that when your primary strategy fails, some higher-order mechanism or alternative path will automatically compensate. The problem is that this assumption only works when you have actually designed Plan B in advance and tested it under conditions similar to where Plan A breaks. I encountered this directly when a client insisted their disaster recovery plan was solid because they had a backup server in another data center. Three weeks later, a fiber cut took out both the primary and secondary sites simultaneously. Their Plan B had the same geographic vulnerability as Plan A. We spent four days restoring from offsite tapes that hadn't been verified in eighteen months. The lesson was painfully simple: Plan B without independent verification is just Plan A wearing different clothes.
How Plan B Actually Works in Practice
Most organizations design their backup strategy around the wrong failure mode. They assume Plan A will fail in a predictable, contained way. Real failures cascade. A database corruption starts as a single bad query, spreads through replicated transactions, and becomes unrecoverable before anyone notices the root cause. This is why the most effective Plan B strategies focus on independence rather than redundancy. The method I use now is straightforward but takes longer upfront. I document where Plan A will fail first, then design Plan B to work under those exact conditions. This usually cuts recovery time from eight hours down to about forty-five minutes, depending on how complex your systems are. The key insight is that Plan B should never use the same dependencies as Plan A. If both plans rely on the same third-party vendor, the same network path, or the same unverified backup schedule, you do not have two plans. You have one plan with extra steps. Here is what I actually did after the fiber cut incident. I wrote a manual recovery script that bypassed the corrupted replication layer entirely. It took six hours to write and test, but it reduced our recovery time from four days to under two hours for the worst-case scenario. The script itself was ugly and undocumented at first, but I version-controlled it from day one and tested it monthly against intentionally corrupted test data. This monthly verification usually takes about fifteen minutes and catches the same class of failures that would otherwise go unnoticed for eighteen months.
Common Pitfalls That Beginners Miss
Most people design Plan B around the wrong failure mode. They assume Plan A will fail gracefully, in a predictable way that their backup strategy can handle. Real failures are messy. A database corruption does not announce itself with a clear error message. It starts as a single bad query, spreads through replicated transactions, and becomes unrecoverable before anyone notices the root cause. This is why the most important aspect of Plan B is not its existence but its independence from Plan A. Another counter-intuitive insight is that more backup options do not equal better Plan B. I once saw a company with five different disaster recovery tools, each with its own configuration, its own schedule, its own unverified restore process. When a ransomware attack hit, they could not determine which backup was clean and which was already infected. They spent three days restoring from the oldest known-good snapshot, losing two weeks of transactional data in the process. Five backup tools without a single unified verification strategy is worse than one backup tool with clear, testable restore procedures. The pitfall I see most often is designing Plan B around the same single point of failure as Plan A. If both plans rely on the same third-party vendor, the same network path, or the same unverified backup schedule, you do not have two plans. You have one plan with extra steps. This usually happens because organizations assume their backup strategy covers all failure modes, when in reality it covers only the failure modes they have imagined.
Get the Full Details

When Plan B Fails Completely
Plan B strategies have hard limits. They fail completely when the failure mode affects the independence assumption itself. If Plan A and Plan B both rely on the same unverified backup schedule, the same third-party vendor, or the same network path, neither plan survives a correlated failure. This is not a theoretical concern. I witnessed a company lose both their primary and secondary data centers simultaneously when a single contractor accidentally severed the same fiber cable that served both sites. Their Plan B had the same geographic vulnerability as Plan A. We spent four days restoring from offsite tapes that had not been verified in eighteen months. In these scenarios, the only viable fallback is often a manual recovery process that bypasses the corrupted dependency entirely. This usually takes longer upfront but reduces worst-case recovery time from days to hours. The tradeoff is that manual recovery scripts are ugly, undocumented at first, and require version control from day one. Most organizations skip this because they assume Plan B will be tested under conditions similar to where Plan A breaks, when in reality the testing conditions are never the same as the failure conditions. Plan B without independent verification is just Plan A wearing different clothes. The most effective strategies focus on independence rather than redundancy, on testing under actual failure conditions rather than imagined ones, and on accepting that manual recovery processes will be ugly but necessary. The goal is not to eliminate failure but to reduce recovery time from four days to under two hours when Plan A falls apart.
Download and Resources
For organizations that want to implement this approach, I maintain a simple recovery script template that bypasses the most common failure modes in database replication, network partitions, and storage corruption. The template itself is undocumented at first but version-controlled from day one, with monthly testing procedures that usually take about fifteen minutes. You can download the current version from the project repository, though I recommend reading the verification procedures before attempting to restore any production data. The worst-case scenario for misconfiguration is losing two weeks of transactional data, which is why the manual recovery process is essential but should never be attempted without first testing on intentionally corrupted test data. The most important resource is not the template itself but the verification process. Monthly testing usually catches the same class of failures that would otherwise go unnoticed for eighteen months, reducing worst-case recovery time from days to under two hours. The template has hard limits and fails completely when the failure mode affects the independence assumption, which is why the verification process is essential but should never be skipped because it takes only fifteen minutes per month.