Getting Polyphase Filter Banks to Actually Work
Multirate Systems And Filter Banks are one of those topics that looks straightforward on paper and becomes a headache the moment you try to implement them. I spent three weeks debugging a resampling chain that kept introducing aliasing at just the wrong frequencies. The root cause wasn't the math. It was how I was handling the transition band in my prototype filter design. Once I understood what was actually happening inside the polyphase decomposition, everything clicked. A multirate system processes signals at different sampling rates. That's the simple version. In practice, you'll be upsampling by an integer factor L, downsampling by an integer factor M, or doing both simultaneously. A filter bank splits a signal into multiple subbands, processes each one differently, and then reconstructs the original or a modified version. The classic example is a two-channel quadrature mirror filter (QMF) bank. You split audio into low and high bands, apply gain or crossover filtering, and recombine them. But the real work happens in the polyphase representation, which is where most people get stuck.
Practical Design with Multirate Systems And Filter Banks
Start with your prototype lowpass filter H(z). Instead of working with it directly, decompose it into polyphase components. A Type-1 decomposition splits H(z) into M branches, where each branch is H_r(z) = sum over h[mM + r] z^(-m) for r = 0, 1, ..., M-1. The filter becomes H(z) = sum from r=0 to M-1 z^(-r) H_r(z^M). This decomposition isn't just mathematical gymnastics. It's what lets you push the downsamplers before the filters in a decimator, which reduces computational load by a factor of M. For an audio resampler running on a embedded DSP, that difference between 48 kHz processing and 6 kHz processing is the difference between real-time and not real-time. For implementation, I recommend using the Parks-McClellan algorithm or the window method to design your prototype filter. The choice depends on your constraints. Parks-McClellan gives you an equiripple design with precise control over passband and stopband ripple. The window method is simpler but produces a Butterwick-like response that may require a higher filter order for the same specifications. If you're designing a 64-channel Audio Distribution Network (ADN) multiplexer, the computational savings from a well-designed polyphase decimator can reduce the filter coefficients storage requirement from thousands of taps down to a fraction of that. Here's the thing nobody tells you: the transition bandwidth of your prototype filter determines everything about the filter bank. If your prototype has a narrow transition band, your subband filters will overlap significantly in the frequency domain. That's fine for some applications but catastrophic for others. I once designed a filter bank where I assumed a wider transition band would be acceptable because the reconstruction error would average out. It didn't. The aliasing components from the individual polyphase branches interfered constructively at certain frequencies, creating notches in the overall response that were 12 dB deep. Took me two days to trace it back to the prototype filter's transition band being too narrow relative to the channel spacing.
Understanding the Aliasing Problem
When you downsample a signal, aliasing occurs. That's basic. In a filter bank, each subband channel is filtered and then downsampled. The filter is supposed to remove content above the new Nyquist frequency. But the filter isn't perfect. Residual aliasing components fold back into the passband and corrupt the signal. This is why reconstruction filters are essential. They cancel out the aliasing introduced during the analysis stage. The aliasing cancellation condition for a two-channel filter bank requires that the aliasing transfer function T_1(z) equals zero. In the frequency domain, this means H_0(e^(j)) and H_1(e^(j)) must satisfy specific relationships. For a maximally flat filter bank, you design the filters so that H_0(e^(j)) + H_0(e^(j(-))) = 2 across the passband. This is the allpass condition. Most people skip this step and just design the filters independently, which works okay but leaves significant reconstruction error. I've seen hobbyist projects where the reconstructed signal sounds fine at first listen but exhibits comb filtering when you sweep through the frequency range. That's aliasing leakage. In a practical implementation, I use MATLAB's filter design toolbox to design the prototype filter and then extract the polyphase components programmatically. The command [hp, hs] = firpmord([Fp Fst], [1 0], [Dp Ds]) gives you the filter order and normalized frequencies. Then firpm(N, F, A) designs the equiripple filter. From there, you can extract the even and odd polyphase components using indexing. The key insight is that the polyphase components themselves are lower-rate filters, which means you can implement them more efficiently.
Get the Full Details

Implementation Pitfalls
I ran into a nasty issue with fixed-point implementation once. The polyphase filter coefficients were designed in floating point with 24-bit precision, but the target system used 16-bit fixed-point arithmetic. The quantization noise from the coefficient truncation interacted poorly with the downsampling operation. Specifically, the stopband attenuation dropped from the designed 80 dB to about 45 dB. The reconstruction filter couldn't cancel the increased aliasing, and the output had a noticeable hiss. The workaround was to increase the filter order by about 30% and use a larger coefficient word length for the polyphase branches. It cost more computational resources but the result was clean. Another common mistake is assuming that the upsampler and downsampler are perfect. Real samplers have timing jitter, and the interpolator doesn't always produce exactly zero-valued samples between the input samples. In simulation this isn't an issue, but in hardware it matters. If you're building a software-defined radio that uses multirate filter banks for channelization, the timing accuracy of your sample rate converters directly affects the isolation between channels. I've seen designs that claimed 60 dB channel isolation but measured closer to 35 dB in practice due to sampler imperfections. The computational complexity of a polyphase filter bank decimator is roughly N/M multiplications per output sample, where N is the filter length and M is the downsampling factor. Compared to a direct implementation that does N multiplications per input sample, this is a reduction of M times. For M = 8, that's an 8x speedup. But there's a catch. The polyphase branches need to be stored in memory, and the memory access pattern can become a bottleneck on architectures without prefetch or cache support. On a Cortex-M4 running at 168 MHz, I found that a 256-tap polyphase filter bank consumed about 75% of the available CPU time for the filter operations alone, even with the M-fold reduction. The remaining cycles had to handle the data movement and control logic.
Reconstruction and Perfect Reconstruction Conditions
Perfect reconstruction (PR) means the output signal is a delayed and scaled version of the input. For a two-channel filter bank with analysis filters H_0(z) and H_1(z) and synthesis filters G_0(z) and G_1(z), the PR conditions are: G_0(z)H_0(z) + G_1(z)H_1(z) = 2z^(-d) [aliasing-free condition] and G_0(z)H_0(-z) + G_1(z)H_1(-z) = 0 [aliasing cancellation]. These are the fundamental equations. Most textbook treatments stop here, but the practical question is how to actually find filters that satisfy these conditions. The simplest approach is to design H_0 as a lowpass filter and H_1 as a highpass filter related by H_1(z) = H_0(-z). Then choose G_0(z) = H_0(z) and G_1(z) = -H_1(z) for synthesis. This gives a Hilbert transform pair that cancels aliasing. But the frequency response won't be flat across the entire band. For a more practical solution, I use the Vaidyanathan approach, which constructs PR filter banks from paraunitary polyphase matrices. The polyphase matrix E(z) must satisfy E(z)E(z^(-1))^T = I. This guarantees PR regardless of the specific filter coefficients, as long as the paraunitary condition holds.
In practice, I design the analysis filter bank first using a standard lowpass prototype, then derive the synthesis bank from the polyphase matrix. The MATLAB function firpm with appropriate constraints can handle this, but you need to provide the right boundary conditions. I typically iterate on the passband and stopband edge frequencies until the reconstruction error is below -60 dB across the entire band. For a 4-band audio crossover, this usually means a filter order around 120 to 160 taps depending on the bandwidth allocation.

When Filter Banks Fail
Filter banks don't work well when your signal has spectral content that doesn't align with your channel bandwidths. If you're building a spectrum analyzer and your signal has narrowband interference that falls between two filter channels, that interference will leak into both channels through the sidelobes. The amount of leakage depends entirely on your filter's stopband attenuation. A Chebyshev Type I filter might give you 40 dB of attenuation but with ripple in the stopband, while an elliptic filter could give you 60 dB but with nonlinear phase that distorts transient signals. I use elliptic filters for spectrum analysis because the attenuation matters more than phase linearity. For audio crossover networks, I switch to Butterworth or Linkwitz-Riley because phase alignment at the crossover point affects the summed response. There's also the issue of non-integer resampling ratios. If you need to resample by a ratio that isn't a simple integer fraction, the standard polyphase filter bank structure breaks down. You either need a fractional delay filter or you need to use a rational resampling approach with separate integer upsampling and downsampling stages. The rational approach introduces additional complexity because you need to handle the gcd of the numerator and denominator factors. In one project, I needed to convert from 44.1 kHz to 48 kHz, which requires an L/M ratio of 160/147. The polyphase implementation required 147 branches, each with about 64 taps. That's nearly 10,000 coefficients to store and access, which was impractical on the target hardware. The workaround was to cascade a 16x upsampler followed by a 147x decimator, processing the data in two stages with an intermediate buffer. It added latency but ran comfortably within the CPU budget. Memory bandwidth is another hard limitation I encountered. A 32-channel filter bank with 128-tap filters and 32-bit coefficients requires about 16 KB of coefficient storage. On a system with limited RAM, that's significant. More importantly, if you're processing multiple signal streams simultaneously, the coefficient memory needs multiply. I was designing a radar signal processor that needed eight parallel filter banks, each with 64 channels. The total coefficient storage exceeded 512 KB, which was more than the available on-chip memory. The solution was to use a shared coefficient ROM with time-multiplexed access, trading throughput for memory footprint. The data rate dropped by about 40% but the system remained functional.
The choice between frequency-domain and time-domain implementations is another decision that depends entirely on your constraints. For long filters (N > 256), overlap-add or overlap-save methods in the frequency domain can be faster due to the O(N log N) complexity versus O(N*M) in the time domain. But the FFT itself has overhead, and for moderate filter lengths, the direct polyphase implementation is often faster. I benchmarked both approaches on a Texas Instruments C6678 DSP and found that for N = 128, the time-domain polyphase approach was about 20% faster. For N = 512, the FFT-based approach won by about 15%. The crossover point depends on your specific processor architecture and cache configuration.
Real-World Applications
Digital audio workstations use filter banks extensively for multiband compression and equalization. Each band is processed independently, and the results are recombined. The filter bank design here needs to be transparent. Any coloration from the filter bank itself will be audible, so the reconstruction error needs to be well below the noise floor of the system. I've worked on plug-ins where the filter bank introduced measurable distortion that was inaudible in blind tests but showed up on spectrum analysis. The fix was to increase the filter order and use a linear-phase design, which eliminated the phase distortion but doubled the computational cost. In telecommunications, OFDM systems rely on filter banks for subcarrier generation and detection. The DFT-based approach is the standard, but polyphase filter banks offer better out-of-band rejection and can reduce the cyclic prefix overhead. I designed a filter bank-based OFDM transmitter for a proprietary protocol and achieved 20 dB better spectral regrowth compared to a standard DFT approach. The tradeoff was increased complexity and latency. The polyphase implementation required about 3x the processing time of the FFT-based approach, which was acceptable for our off-line encoding application but not suitable for real-time transmission. For hardware implementation, FPGA designs benefit greatly from polyphase decomposition because the parallel structure maps naturally to pipelined architectures. Each polyphase branch can be implemented as a separate FIR filter with its own clock domain, allowing you to balance the pipeline depth across branches. I've synthesized designs on Xilinx Artix-7 FPGAs that achieve over 200 MSPS throughput with 64-channel filter banks. The key is using distributed arithmetic for the multiply-accumulate operations, which avoids the need for dedicated DSP slices and frees up resources for other functions.
