Why Two Systems Keep Colliding and What to Do About It

I spent three years debugging what I now call The Shark And The Goldfish problem before I realized it wasn't a code issue at all. It was a structural one. The pattern shows up in integration work, microservice architectures, CI/CD pipelines, and even basic data migration scripts. You have two subsystems that are fundamentally designed for different purposes, different failure modes, and different performance characteristics, and someone decided to force them to talk to each other directly without a mediation layer. The name comes from an old office parable, not a formal technical document. The shark represents a system built for speed, aggression, and high throughput. It consumes resources quickly and doesn't care if it makes a mess. The goldfish is the opposite: fragile, slow-moving, easily stressed by sudden changes in environment. In practice, the shark is your high-throughput message queue or your event stream processor. The goldfish is your legacy database, your payment gateway, your user-facing API with strict latency SLAs. When you connect them head-on, something breaks. Usually both of them.

Recognizing The Shark And The Goldfish in Your Stack

Most people don't realize they have this problem until production catches fire. The warning signs are usually subtle. You will notice sporadic timeout errors that only appear under load. Your monitoring will show one service spiking while another crashes for no apparent reason. Latency graphs look normal at 3 AM but degrade significantly during peak hours. The service that is failing often reports errors that make no sense, because the error originated in the other system and propagated in a way that doesn't match its contract. I once had a Kafka producer (the shark) pushing events into a PostgreSQL-backed service (the goldfish) through a direct REST API. We saw 429 errors in bursts. We checked the API rate limits. They were fine. We checked PostgreSQL lock contention. Normal. We checked network paths. Clean. The problem was that the producer would burst 200 requests in under 800 milliseconds, then go idle for three seconds. PostgreSQL can handle a sustained 50 requests per second. It cannot handle 200 in a single wave and then nothing. The database spent most of its time recovering from lock spikes instead of processing actual queries. The service appeared unstable even though neither component was misconfigured on its own. The fix was not to optimize either service. It was to insert a buffer between them. A RabbitMQ queue with a single consumer, a steady fetch rate of 50 queries per second, and a retry mechanism with exponential backoff for failed messages. Once the queue sat between them, the 429s stopped immediately. Latency improved across the board. We added dead letter queue monitoring and a dashboard for queue depth. That system has run for two years without a single unplanned outage related to this pattern.

How to Implement a Mediation Layer

The standard approach to solving The Shark And The Goldfish problem is to introduce a buffer or adapter between the two systems. This is sometimes called a circuit breaker pattern, a bulkhead pattern, or a message queue intermediary depending on what community you ask. The principle is the same: decouple the consumer from the producer so neither one directly experiences the other's failures or load spikes. There are three main strategies, and the right one depends entirely on your constraints. Message queue buffering is the most common solution. You route all traffic from the shark through a persistent queue. The goldfish pulls from the queue at its own pace. This handles burst traffic, provides retry semantics out of the box, and gives you observability into queue depth and processing lag. Tools like RabbitMQ, Apache Kafka, Amazon SQS, and Azure Service Bus all work here. Kafka gives you replayability and partitioning but adds operational complexity. SQS is simpler but has less visibility into message ordering guarantees. If you need exactly-once semantics, you will need to build idempotency into the goldfish anyway because no queue guarantees that purely through configuration.

Get the Full Details

The Shark and the Goldfish | Jon Gordon
The Shark and the Goldfish | Jon Gordon

Adaptive rate limiting works when you cannot introduce a full queue. A middleware proxy sits between the two systems and enforces a dynamic rate limit based on the downstream system's health. If the goldfish starts showing elevated latency or error rates, the proxy slows down the shark automatically. This is less elegant than a queue but avoids the infrastructure cost and introduces roughly 5-10 milliseconds of additional latency per request. Istio rate limiting policies, NGINX rate limiting with custom logic, or a lightweight sidecar proxy like Envoy can handle this. I prefer Envoy for this because it gives you granular per-route rate limits and integrates well with metrics pipelines. Eventual consistency with change data capture is the strategy you use when real-time communication is not required. Instead of the shark calling the goldfish directly, you capture changes from the shark's data layer and propagate them asynchronously. Debezium is the standard tool for this with PostgreSQL and MySQL. It reads the WAL or binlog and publishes changes to a Kafka topic. The goldfish consumes from that topic. This completely removes the synchronous dependency. The tradeoff is that your data will be stale for however long the consumption lag is, usually in the range of 100 milliseconds to a few seconds depending on your topic configuration and consumer capacity. For many business processes, this lag is acceptable. For financial reconciliation or inventory management, it may not be.

Common Pitfalls When Adding a Mediation Layer

The biggest mistake I see is treating the mediation layer as invisible infrastructure. People add a queue and then forget to monitor it. Queue depth grows unbounded during traffic spikes. Memory fills up. The broker crashes. Now you have three broken systems instead of two. Always instrument queue depth, consumer lag, message age, and error rates from day one. Prometheus with Grafana works well. Datadog if you already pay for it. CloudWatch if you are fully AWS and want to minimize tool sprawl. The second mistake is assuming the queue solves everything. It does not solve incorrect data formats, missing fields, or schema mismatches. If the shark and the goldfish disagree on what a message looks like, adding a queue just delays the failure. Define a shared schema contract before you build the buffer. JSON Schema, Avro, or Protobuf all work. Avro is my default choice for internal systems because it supports schema evolution with backward compatibility built in. If you use Protobuf, be aware that removing or renaming fields breaks consumers unless you carefully manage the registry. The third mistake is not planning for queue saturation. I once watched a system fail because nobody configured a maximum queue depth. Under a marketing campaign that tripled normal traffic, the queue grew until the broker ran out of disk space. The broker entered read-only mode. Messages stopped flowing. The shark kept producing. It had no feedback that the messages were being dropped. The only warning was a slow disk usage alert that someone ignored because it was yellow not red. Configure queue size limits. Set up alerts at 60 percent capacity. Add a fallback path that drops non-critical messages when the queue is full instead of letting the broker crash.

When The Shark And The Goldfish Approach Fails Completely

Sometimes the mediation layer is not enough. If the goldfish is fundamentally incapable of handling the workload even at a reduced rate, no amount of buffering will save the integration. This happens most often with legacy systems that have hard architecture limits. A monolithic ERP from the 2000s that locks entire tables during writes. A reporting database that cannot handle concurrent queries above a certain threshold. An API that throttles at the network level regardless of what you configure on your side. In these cases, the real solution is to stop treating the goldfish as a real-time system. Build a read replica specifically for the integration. Replicate data to a purpose-built service that can handle the traffic pattern. Let the shark talk to the replica, not the original. This adds latency to data freshness but eliminates the bottleneck. I have done this with Oracle EBS systems where the source database could not sustain more than 20 concurrent integration queries. We built a CDC pipeline to a PostgreSQL read replica and pointed all integration traffic there. Query throughput went from 20 per second to over 500 per second with no impact on the source system. Another scenario where this breaks down is when the shark and the goldfish have conflicting consistency models. One expects strong consistency. The other provides eventual consistency. No mediation layer resolves this. You have to pick one model and adapt the other system to work within it, or accept that certain operations will have race conditions that need application-level handling.

The Shark And The Goldfish Positive Ways To Thrive During Waves Of Change Book By Jon Gordon ...
The Shark And The Goldfish Positive Ways To Thrive During Waves Of Change Book By Jon Gordon ...

A Real-World Example I Actually Lived Through

Two years ago I was consulting for a mid-size logistics company. Their shipment tracking system (the shark) pushed location updates from GPS devices into their customer notification service (the goldfish) via direct HTTP calls. The tracking system generated about 15,000 events per hour during normal operations and up to 80,000 during peak. The notification service was built on a shared hosting plan with a PostgreSQL backend and had no horizontal scaling capability. It could handle roughly 200 concurrent connections before response times degraded and errors started appearing. The symptoms were classic. Customers reported missing tracking updates. The engineering team blamed the GPS hardware. The hardware team blamed the software. Neither was wrong. Both were right. The GPS devices were sending data correctly. The notification service was dropping or delaying records because it could not keep up with the burst pattern. The direct HTTP calls meant the tracking system had no way to know when the notification service was overwhelmed. It just kept sending. We implemented a three-part solution. First, we moved the tracking events into an Amazon Kinesis stream. The tracking system writes to Kinesis, which buffers the data independently of the consumer. Second, we built a Lambda consumer that processes events from Kinesis at a rate of 500 events per second per shard, with automatic scaling based on stream magnitude. Third, we gave the notification service its own connection pool and query batching logic so it could process incoming events more efficiently. The notification service never had to handle more than 500 concurrent connections even during peak loads because the Kinesis stream absorbed the bursts.

The cost increase was about $340 per month for Kinesis and the additional Lambda compute. The previous outage-related support costs ran approximately $2,400 per month. The system has been stable since deployment. We added a CloudWatch dashboard tracking Kinesis lag, Lambda invocation errors, and notification service response times. There is one alert: if Kinesis lag exceeds 5 minutes, the on-call team gets a PagerDuty incident. We have had three false alarms in eight months, all caused by a Lambda concurrency limit that was set too low during a deployment. We adjusted the limit and the alarms stopped.

The Downside Nobody Mentions

Adding a mediation layer increases system complexity. You now have another component to monitor, deploy, debug, and restore from backup. Debugging a failed message requires checking the queue, the consumer, and the downstream system. There is no single log file that tells the whole story. Your incident response time will initially get worse before it gets better because engineers are learning to navigate the new architecture. Budget for that. The first month of operation with a new mediation layer usually has more operational friction than the problem it is solving. There is also a subtle performance cost. Message queue buffering adds latency. Even a fast in-memory queue adds 10-50 milliseconds of delay compared to a direct call. For real-time systems where every millisecond matters, this might be unacceptable. In those cases, adaptive rate limiting through a sidecar proxy is the better tradeoff because it adds minimal latency while still protecting the downstream system. The core insight is that The Shark And The Goldfish is not a bug. It is a design reality. Any system that connects two components with mismatched throughput, latency, or failure characteristics will eventually hit this problem. The question is not whether it will happen but how prepared you are when it does. Building the buffer early, even when the traffic patterns look manageable, is almost always cheaper than building it after production incidents start destroying customer trust.

The Shark and the Goldfish by Jon Gordon (Audiobook) - Read free for 30 days
The Shark and the Goldfish by Jon Gordon (Audiobook) - Read free for 30 days