Building Polyphase Filters When the Textbook Diagram Lies to You
I spent three weeks debugging an LTE downlink front-end where the simulated channel response looked perfect and the actual hardware measurement showed 4 dB of ripple at the band edge. The problem wasn't the filter coefficients. It wasn't the ADC sampling clock either. It was the order of operations between the polyphase decomposition and the decimation stage, specifically how the individual E_k(z) branches interact with finite precision arithmetic at the junction where two sampling rates meet. What I learned from that won't be in Oppenheim and Schafer, and it probably won't help you on a certification exam. But if you're building something that actually runs on silicon, it matters.
Practical Multirate Signal Processing For Communication Systems
The core idea is simple enough that you've probably seen the block diagram a hundred times. You have a prototype filter H(z) and you want to change the sampling rate by an integer factor L or D. Instead of running the full-rate filter and then decimating or interpolating, you decompose H(z) into N polyphase components. Each component operates on a sub-sampled version of the signal. The math rearranges but the output is identical in infinite precision. The decomposition itself is straightforward. For a filter with impulse response h[n], the k-th polyphase component is: E_k(z) = sum over n of h[nN + k] * z^(-n), where k ranges from 0 to N-1.
So if your prototype filter has 1024 taps and you're doing a 4x decimation, each polyphase branch gets roughly 256 taps. You run four parallel filters at the lower rate instead of one serial filter at the high rate. That's the part everyone understands. Here is where it gets interesting. The standard textbook shows the polyphase branches with their delays arranged so the outputs align perfectly at the summation node. In an FPGA implementation, those alignment delays become pipeline registers. Each extra register adds latency. A 4x polyphase interpolator with 1024-tap prototype filters will typically introduce somewhere between 256 and 1024 clock cycles of latency depending on how you distribute the pipeline registers across the branches. That latency is real and it affects timing closure in a real receiver chain. I found this out the hard way when the control loop for my automatic gain control started oscillating. The loop was designed assuming 64 cycles of group delay through the front-end. The actual polyphase filter introduced 892 cycles. The compensating delay I added in the digital domain was too crude and created a phase margin problem at the crossover frequency. The fix was to model the exact group delay of the polyphase structure, including the resampling phase, and bake that into the AGC filter design before routing any signals.
Where the Theory Breaks Down in Hardware
Polyphase decomposition does not reduce computational complexity in the way most people assume. This is the most common misunderstanding I see. If your prototype filter has no structural zeros, the total number of multiplications per output sample is identical whether you implement it as a single high-rate filter or as polyphase branches. You are just moving work from one clock domain to another. The complexity savings only appear when your filter has zeros built into its coefficient structure. A half-band filter is the classic example. Every even-indexed coefficient except the center tap is zero. When you polyphase-decompose a half-band filter, approximately half of the branches have only zero coefficients and can be completely eliminated. The remaining branches are themselves half-band filters, and you can recurse. This is how multistage decimation actually saves work. But if you are designing an arbitrary-shape channel selection filter with no symmetry or zero-coefficient structure, you get zero complexity reduction from polyphase decomposition alone. Another thing that trips people up is the interaction between the polyphase phases and the input signal timing. When you downsample by N, the phase of the input sample relative to the polyphase branch selection determines which E_k(z) processes each incoming sample. If your input clock and your decimation clock are not phase-aligned, and you are not accounting for that phase offset in your filter design, you will get unexpected spectral content at the output. This is not aliasing in the classical sense. It is a phase-dependent redistribution of energy across the polyphase branches that manifests as passband ripple.
Implementing a Multistage Decimator That Actually Works
Let me walk through a concrete example. Say you are building a receiver that samples at 61.44 MHz and needs to deliver baseband samples at 3.84 MHz. That is a 16x decimation. You could try to do it in a single stage with a 512-tap FIR filter. You could also split it into four 2x stages, each with a 128-tap half-band filter. The second approach is almost always better for a reason that has nothing to do with complexity and everything to do with stability. When you cascade half-band filters, each stage relaxes the transition bandwidth requirement for the next. The first stage at 16x decimation needs to reject frequencies above 3.072 MHz while passing up to 1.92 MHz. That is a transition band of about 1.15 MHz. As you proceed through stages, the required transition band widens relative to the new sampling rate, which means fewer taps per stage. The total tap count across all stages is lower, and more importantly, the dynamic range requirements on each individual filter are relaxed. But here is the practical issue that nobody warns you about. When you cascade half-band filters in a polyphase structure, the effective filter order of each stage determines the group delay variation across the passband. A 128-tap half-band filter has approximately 64 taps of group delay. Four stages in cascade give you about 256 taps of total group delay variation. If your communication system is sensitive to group delay distortion—and most modern OFDM systems are, because excess delay spread eats into your cyclic prefix—you need to account for this in your equalizer design. The equalizer will see the group delay as part of the channel response and compensate for it, but only if you tell it what the polyphase filter's group delay actually is. Guessing or measuring it empirically will introduce errors that manifest as increased bit error rate at the decoder output.
I dealt with this on a DVB-T2 receiver project. The symbol timing recovery loop was locking correctly in simulation but the post-FEC BER floor was stuck around 1e-4 instead of dropping below 1e-5. The root cause was that the polyphase interpolator in the resampler was introducing a frequency-dependent phase rotation that the timing loop interpreted as a timing error and tried to correct by adjusting the interpolator control word. This created a feedback cycle that added residual phase noise to the recovered symbols. The workaround was to pre-compensate the polyphase filter phases so that the overall group delay was flat across the passband, which required adjusting the tap weights in the odd-phase branches by a small amount. The adjustment was on the order of 0.1 percent of the coefficient values but it eliminated the feedback interaction entirely.
The Aliasing Problem That Only Appears After Deployment
There is a specific failure mode in multirate systems that is nearly impossible to catch in simulation and only shows up when the hardware is running under real signal conditions. It involves image rejection in the interpolator chain when the input data is not perfectly aligned with the polyphase branch boundaries. When you upsample by L and then filter with a polyphase interpolator, each branch E_k(z) receives a different subset of the input samples. If the input sample stream has jitter or if the upsampling clock and the input clock are derived from different sources with untracked phase drift, the mapping between input samples and polyphase branches becomes time-varying. The filter was designed assuming a fixed mapping. The time-varying mapping turns what should be clean image rejection into phase-noise-modulated images that appear as spectral splatter around the desired band. In my experience, this typically shows up as an elevated noise floor about 2*L*f_sample away from the carrier, where f_sample is the input sampling rate. The splatter is not wideband noise. It is discrete in the sense that it tracks the carrier frequency but its amplitude varies with the phase relationship between the two clocks. If you are using a single clock domain and a digital resampler with a fractional delay filter, this problem goes away. If you are crossing clock domains, which is common in software-defined radio architectures where the ADC runs from one PLL and the baseband processor runs from another, you need to either synchronize the clocks at the source or accept that your image rejection will be limited by the phase noise of the clock relationship rather than by your filter design.
The workaround I ended up using was to derive both clock domains from the same reference oscillator and use a phase-locked resampler architecture. The resampler internally interpolates between input samples to account for the fractional phase offset, effectively making the polyphase branch selection continuous rather than discrete. This costs additional computational resources—roughly 30 to 50 percent more DSP slices on a Xilinx FPGA compared to a fixed-ratio polyphase structure—but it eliminates the clock-domain image problem entirely. For a production system where spectral mask compliance is a regulatory requirement, that overhead is usually worth it.
When Polyphase Is the Wrong Choice
I want to be clear about where this approach fails. Polyphase multirate processing assumes integer resampling ratios. If your system requires a non-integer ratio like 3.2x or 7/5, you need a rational resampler that combines an interpolator and a decimator with a gcd-based reduction. The polyphase structure still applies within each integer stage, but the overall architecture is more complex and the latency analysis becomes significantly harder. Another scenario where polyphase is problematic is when your filter coefficients change dynamically. In a software-defined radio that supports multiple modulation schemes, you might need to switch between different channel filters in real time. A fixed polyphase structure with static coefficients is efficient but inflexible. If you need reconfigurable filters, a direct-form FIR implemented with a serial multiplier may be simpler to manage despite the higher computational cost, because you are not committing to a specific decimation ratio at the hardware level. There is also the question of memory bandwidth. In a resource-constrained embedded system, the polyphase decomposition requires you to store all N branches of the filter coefficients in accessible memory simultaneously. For a 64-tap filter with a 16x polyphase decomposition, you are storing 64 coefficients distributed across 16 branches. This is trivial for an FPGA with dedicated BRAM. It is less trivial on a microcontroller with a few kilobytes of SRAM where coefficient storage competes with buffer space for the incoming and outgoing sample streams. In that environment, a commutated direct-form implementation that processes one tap at a time may be more practical even though it runs at the full input sample rate.
A Working Reference
If you want to experiment with this yourself, the GNU Radio framework has a well-tested polyphase filter bank implementation that you can study and modify. The source is available at github.com/gnuradio/gnuradio under the blocks/polyphase_module directory. The C++ implementation handles the polyphase decomposition, the commutator logic for decimation, and the interpolation with fractional delay support. It is not the most optimized implementation for production hardware, but it is correct and well-documented, which makes it useful for verifying your own designs against a known reference. For Python-level prototyping, the scipy.signal.decimate function implements a cascaded IIR+FIR approach that is simpler but less flexible than a full polyphase structure. It is adequate for quick simulations where filter shape is more important than precise group delay characteristics. Do not use it for anything that will ship in a product without verifying the frequency response against your requirements independently. The key takeaway from all of this is that multirate signal processing is not just a mathematical exercise in filter decomposition. The practical challenges come from clock domain boundaries, finite precision effects, group delay management, and the interaction between the resampler and the rest of the signal chain. Get those right and the theory works beautifully. Miss any of them and you will spend weeks debugging symptoms that look nothing like the root cause.