Understanding Information Transmission at Its Core

I ran into a problem a while back that made me appreciate just how much of Shannon's work gets glossed over in intro courses. I was working on a custom radio telemetry link for an industrial monitoring system, and we were hitting bit error rates that seemed impossible given our link budget. The math said we should have had a clean signal with plenty of margin. What we'd missed was that Shannon's formula gives you a hard ceiling, not a guarantee. It tells you the maximum rate at which information can be transmitted with arbitrarily low error, but it doesn't tell you whether your actual code comes close to that limit. Our FEC code was nowhere near capacity-achieving, and that was the entire problem. Claude Shannon published "A Mathematical Theory of Communication" in 1948, splitting it across two Bell System Technical Journal issues. Before that paper, people thought about communication as an engineering problem of amplifiers and filters. Shannon reframed it entirely: communication is about the transmission of messages, and messages carry information measured in bits. He separated the engineering concerns from the mathematical ones. The local area of encoding and decoding becomes a problem of error-correcting codes. The effective area of transmitting those codes becomes a problem of signal processing and channel characterization. Everything sits between those two poles. The channel capacity formula is C = B log(1 + S/N). Bandwidth in hertz times the log of one plus the signal-to-noise ratio. That's it. It looks deceptively simple because it is simple, and simplicity here is what makes it dangerous. People treat it like a design formula when it's really a bound. You cannot exceed C, but you can sit dozens of decibels below it and still have a completely broken system if your coding scheme is inadequate.

I want to emphasize something most textbooks don't make clear enough. Capacity is an asymptotic result. It assumes you're using codes with infinitely long block lengths and performing maximum likelihood decoding. Real systems use finite block lengths. The gap between practical code performance and the Shannon limit is called the coding gap, and it typically runs anywhere from 2 to 4 dB for well-designed modern systems like LDPC or Turbo codes. Older convolutional codes with Viterbi decoding sat closer to 5 to 6 dB away. This gap is not a minor detail. It determines whether your satellite link works or your satellite link requires a larger dish. Another thing that catches people out: the formula assumes an additive white Gaussian noise channel. Real channels are rarely that nice. Multipath fading, interferers, narrowband jammers, quantization noise from your own ADC — these all change the problem fundamentally. In a fading channel, the instantaneous capacity drops dramatically during deep fades even though the average SNR looks fine. You handle this with techniques like interleaving, diversity, and adaptive modulation, but you're no longer operating at the capacity given by that simple formula. You're operating in a regime where Shannon's framework still applies conceptually, but the clean analytical result doesn't directly describe what's happening. Here's how I actually used this in practice. When designing a communication link, the first thing I do is calculate the Shannon capacity for the given bandwidth and SNR. That gives me the absolute ceiling. Then I pick a modulation and coding scheme and check its spectral efficiency against that ceiling. If the spectral efficiency of my chosen scheme is within, say, 3 dB of capacity, I know it's reasonable. If it's 8 dB away, I need a better code or a wider bandwidth. This check usually takes me about five minutes and has saved me from designing systems that would have failed in the field.

Entropy is the other half of this theory and it deserves proper attention. H(X) = - p(x) log p(x). This measures the average information content of a source. Source coding, or compression, operates on this principle. You can compress a source down to its entropy rate, but not below. This is separate from the channel capacity question, though they connect through the source-channel separation theorem, which states that for a sufficiently long block length, you can design the source coder and channel coder independently without losing optimality. That theorem is incredibly useful because it lets you design compression and error correction as separate problems. It breaks down when you're working with very short packets or constrained latency, which is exactly the regime many IoT and real-time systems operate in. I've seen people try to apply Shannon capacity calculations to channels with memory, like DSL lines or power line communication links, without accounting for the correlation structure. The capacity formula for a channel with memory is different. You need to consider the eigenvalue distribution of the channel covariance matrix. In practice, what this means is that treating a colored noise channel as white noise and plugging numbers into C = B log(1 + S/N) will give you an optimistic answer. The real capacity is lower. I learned this the hard way on a power line communication project where our measured throughput was roughly 60 percent of what the white-noise capacity formula predicted. The theory also has hard limitations that nobody likes to talk about. It says nothing about security. A channel can be operating at near-capacity rates and still be completely intercepted by anyone with a receiver on the right frequency. Information theory later addressed this through Wyner's wiretap channel model and related work, but the original 1948 paper does not. It also doesn't handle semantic meaning. Two messages can carry identical information content by Shannon's measure while being completely different in what they actually communicate. The theory treats all content as interchangeable symbols.

Get the Full Details

The Mathematical Theory of Communication de Shannon, Claude E., and Warren WEAVER: [8], [1 ...
The Mathematical Theory of Communication de Shannon, Claude E., and Warren WEAVER: [8], [1 ...

If you're working with this material, the practical takeaway is straightforward. Calculate capacity as your benchmark. Design your coding scheme to approach it within a few dB. Verify your assumptions about the channel model match reality. And don't trust the formula blindly when your channel has structure that the basic model doesn't capture. The theory gives you the boundary of what's possible. Getting close to that boundary is where the actual engineering work lives.