Why Your Message Queue Is Losing Data and What A Consumer Actually Does
You set up a RabbitMQ queue. You pushed ten thousand messages. You consumed zero. Not because the broker crashed, not because your network went down, but because nobody actually read the acknowledgment logic on the consumer side. This happens constantly. I watch it every quarter when someone brings me a dead-letter queue that grew to two billion messages over a weekend and asks why their application appears to "lose" data. A consumer is simply a process that connects to a message broker, registers interest in a specific queue or topic, and pulls messages from it for processing. That is the textbook answer. In production, it is much messier.
What Is A Consumer In Practice
The consumer holds a reference to a channel, starts listening, and defines a callback that fires when a message arrives. The callback does work. The work might be writing to a database, calling an external API, or running a computation. When the work completes, the consumer sends an acknowledgment back to the broker to confirm delivery. If the consumer crashes before acknowledging, the broker re-delivers the message, assuming someone else can pick it up. The entire model depends on three things working correctly: delivery semantics, acknowledgment handling, and resource limits. Miss any one of those and your system either drops messages, duplicates work, or chokes on its own queue. I spent three weeks debugging a consumer that appeared to process messages twice. Every order in the system was being fulfilled by two separate payment calls. The messages were not actually duplicated in the queue. The problem was that the consumer was using manual acknowledgments but had prefetch set to zero with no timeout on the acknowledgment call. When the handler threw an unhandled exception during the payment API call, the message was never acked, the broker redelivered it, and a second consumer instance grabbed the same message because the first instance had already started its handler but hadn't explicitly nack'd or recovered gracefully. The fix was straightforward: wrap the handler in a try-catch, ack on success, nack with requeue false on failure, and set prefetch to something reasonable like 10 to prevent a single slow consumer from starving others. That reduced duplicate processing incidents by about ninety-four percent in our environment.
Here is the part nobody mentions in the documentation. Acknowledgment is not optional in manual mode. If you connect with manual ACK and never call ack or nack, the message stays assigned to your consumer forever. The broker will not re-deliver it. It sits there consuming your prefetch slots until your application crashes, at which point the connections drop and only then does the broker re-queue everything. I once watched a consumer backlog grow to four million messages over a month because a junior engineer enabled manual ACK mode for "better control" and forgot to implement the nack path for timeout errors. There are delivery modes you need to understand before you write a single line of consumer code. At-most-once delivery means you acknowledge before processing. If the consumer dies mid-handling, the message is gone. Fast, simple, dangerous for anything involving money or state changes. At-least-once means you process first, then acknowledge. Duplicates are possible on crash but nothing is lost. This is the default for most production systems. Exactly-once is a myth in distributed messaging unless you build idempotency yourself and use transactional outboxes or deduplication tables. Kafka approximates it with consumer offsets and log compaction, but the underlying mechanism is still at-least-once with application-level guarantees layered on top.
Get the Full Details

Prefetch count is another setting that destroys more systems than it helps. Setting prefetch to one sounds like good flow control. It actually serializes your consumers and makes throughput terrible when message processing time varies. Setting it too high floods your application memory with unprocessed messages and makes crash recovery take exponentially longer. The sweet spot is usually between 10 and 100 depending on your average message size and processing latency. Run a quick benchmark with your actual payload and measure where throughput starts dropping. Don't guess. Dead-letter exchanges exist for a reason. When a message fails processing after maximum retries, it should move to a DLQ, not loop forever inside your primary queue. Configure x-dead-letter-exchange and x-max-length on your queue definitions. I set max length to around five thousand per queue and route overflow to a separate DLQ exchange with a dedicated consumer that logs failures and alerts the team. This keeps primary queues from growing indefinitely and makes it obvious when a downstream dependency is broken instead of burying the symptom under millions of retries. Consumer group management matters more than people realize. In Kafka, rebalancing between consumers in the same group causes temporary processing gaps. In RabbitMQ, there is no native consumer group concept, so you manage distribution yourself through queue declarations and connection pooling. If you are using RabbitMQ and have five consumer instances, each must connect to the same queue independently. The broker round-robins deliveries across connected consumers. If one consumer disconnects, the broker reassigns its unacknowledged messages to remaining consumers. This sounds fine until you realize that reassignment happens per-message, not per-batch, and during the transition period you may see spikes in processing latency as the broker redistributes.
Resource monitoring is not a nice-to-have. Track message rate in, message rate out, acknowledgment latency, queue depth, and consumer connection health. Alert on queue depth exceeding two times your average hourly throughput. That gives you a buffer before prefetch starvation becomes visible to users. Log acknowledgment timeouts separately. If a consumer is taking longer than expected to ack, you have a downstream bottleneck, not a messaging problem. The biggest failure mode I encounter is assuming the consumer framework handles everything. Most frameworks give you a basic skeleton. They do not handle partial failures gracefully, they do not implement circuit breakers for downstream services, and they do not warn you when your acknowledgment logic has a silent failure path. You have to build that yourself. Add a timeout around your handler. Add retry logic with exponential backoff for transient errors. Add a dead-letter path for permanent failures. Add health checks that verify your consumer is actually processing messages and not just sitting idle with a full prefetch buffer. When a consumer setup works correctly, you barely notice it. Messages flow in, get processed, and disappear. When it is broken, the symptoms look like application bugs: missing data, duplicate charges, stale dashboards, and angry support tickets. The root cause is almost always something simple that the documentation glossed over or that you skipped because it seemed unnecessary at the time.
Start with automatic acknowledgments if you are prototyping. Switch to manual once you understand your failure modes. Set a reasonable prefetch. Add a dead-letter exchange from day one. Monitor queue depth. Write the nack handler before you write the happy path. These are not suggestions. They are the things that separate a consumer that works in production from one that works until it does not.
