Understanding the Retry-And-Backoff Pattern

I've spent years dealing with systems that fail intermittently and keep failing until you give up or something changes. The term Sometimes They Come Back For More comes up in my team's post-mortems a lot. It describes situations where a failed operation doesn't just fail once and stay failed — it returns, often in different forms, and keeps coming back until a certain condition is met or you hit some hard limit. This isn't a single well-defined algorithm with a Wikipedia page. It's more of a pattern you notice after your third or fourth all-nighter trying to figure out why a job keeps rescheduling itself on the same failed node. The short version: you have an operation that fails, it retries automatically, and each retry can surface a different failure mode that wasn't visible in the first attempt.

When Sometimes They Come Back For More Actually Happens

I remember one incident where we were processing image thumbnails at scale. The initial upload would fail because the disk was full. Fair enough. We'd retry. But the second failure wasn't about disk space — it was about a corrupted queue entry left behind from the first attempt. The third failure was a stale lock on a table that the first attempt had created but never released. The fourth failure was a timeout because the service was genuinely degraded and everything was slow. The application thought it was the same single failure looping. It wasn't. Each retry came back with a different error. This is the core of the pattern. A naive retry policy treats every failure as identical and just keeps hammering. That's usually how you make things worse.

How to Handle It Without Losing Your Mind

The first thing I do is stop treating retries as optional. I build a structured retry strategy into everything that touches an external system. The basic components are an exponential backoff with jitter, a maximum retry count, and a failure categorization layer that runs before each retry decision. Here's what that looks like in practice. You capture the exact error type and the current retry attempt number. Then you run it through a classifier. Transient network errors go to the retry bucket. Resource exhaustion errors go to a different bucket with a longer delay. Corrupted state errors — the ones where your previous attempt left garbage behind — don't get retried automatically. They get a cleanup step first. In our thumbnail pipeline, the workaround involved adding a transactional cleanup phase that ran between retries. If an attempt failed, we'd check for and remove any stale artifacts from that attempt before scheduling the next one. This cut our average retry count from 4.7 per failure down to about 1.3. Most of the returns stopped happening because there was nothing left to trip over on the next pass.

Get the Full Details

Sometimes They Come Back... for More (1998) - krstyb7 | The Poster Database (TPDb)
Sometimes They Come Back... for More (1998) - krstyb7 | The Poster Database (TPDb)

You can build this yourself or use a library. For Node.js projects, retry and async-retry are the standard options. For Python, tenacity is widely used. Neither handles the categorization part automatically — that's something you add on top. The classification logic is where most people cut corners. A simple if-else chain based on HTTP status codes and error message patterns works for small systems. It breaks down when you have twenty different services all throwing slightly different error formats. At that scale you need a dedicated error classification module, ideally one that pulls from a centralized error taxonomy maintained by whoever owns the infrastructure.

Things That Make This Pattern Worse

Cascading retries are the biggest problem. When service A fails and retries, and service B is already under load from other failures, service A's retries push service B harder. Service B then fails more often. Its clients retry more. It's a feedback loop that amplifies everything. The standard mitigation is circuit breakers and request hedging limits. Set a hard ceiling on how many concurrent retry attempts any single consumer can generate. Tenacity has a max attempts parameter. Set it to something reasonable and stick to it. Another common mistake is assuming the error message is stable. In distributed systems, the same underlying problem can surface with different error messages depending on which node handles the request, what version is running, and the current load on the network. If your retry logic parses error messages instead of using structured error types, it will misclassify frequently. Use error codes when your platform provides them. Fall back to type checking. Only use message parsing as a last resort. There are also cases where this entire approach is the wrong answer. If the failure is a fundamental data problem — a missing required field, a schema mismatch, a permission issue that won't resolve itself — no amount of retrying will fix it. Retrying in these cases just piles failed job after failed job on your queue until it fills up and blocks new work. The fix is a dead letter queue and alerting, not more retries. I once spent two weeks debugging what I thought was a flaky dependency only to discover the input data had a silent encoding issue that corrupted the payload on every attempt.

The practical takeaway is that Sometimes They Come Back For More is not a bug in your retry logic. It's a signal that your system has multiple layers of failure and each retry needs to account for the possibility that the ground has shifted since the last attempt. Track what changed between retries. Classify failures properly. Clean up state between attempts. And know when to stop trying entirely.

Sometimes They Come Back... for More (1998)
Sometimes They Come Back... for More (1998)