Understanding and Implementing David Bednar's Time-Scale Modification Techniques
David Bednar is known primarily for his contributions to audio signal processing, specifically around time-stretching and pitch-shifting without affecting duration. His work at the University of Limerick and later Indiana University produced some practical methods that are still referenced when building or choosing a phase vocoder implementation. The standard phase vocoder approach has a problem where phase information gets messy across frames, leading to smearing or metallic artifacts. Bednar's method addresses this by using a more careful phase unwrapping strategy combined with a specific overlap-add framework that reduces comb-filtering effects. It's not magic, but it produces cleaner results than the naive approach you'll find in basic tutorials. I spent several months wrestling with a custom time-stretch implementation back when I was building a live audio tool. The naive phase vocoder sounded terrible on percussive material. What I found was that the phase accumulation between consecutive FFT frames needed to be compensated differently depending on the analysis window size. Bednar's approach uses a specific hop size to frame size ratio and applies a correction term to the instantaneous frequency estimate before propagating phase forward. The key insight is that you shouldn't just unwrap phases blindly. You need to look at the group delay characteristics of your window function and compensate for that.
Here's what the core idea looks like in practice. You take an audio signal, window it into frames using an overlap of typically 75% or higher, compute the FFT, extract magnitude and phase, correct the phase to remove the linear phase component from the window, interpolate to new frequency bins if you're stretching, then rebuild using overlap-add. The difference with Bednar's approach is in how you handle the phase propagation step between frames and how you manage the synthesis window to maintain perfect reconstruction conditions.
Practical Implementation Notes
If you're trying to implement this yourself, start with a small frame size like 1024 or 2048 samples and use a Hann window. The overlap should be at least 4x the hop size. When you extract phase from the FFT output, you'll get wrapped phases between minus pi and pi. Subtract the expected phase advance based on the bin center frequency and the hop size to get the instantaneous phase deviation. That's where most implementations go wrong. They skip this step or do it approximately. The phase correction term involves computing the group delay of your synthesis window and applying an inverse compensation. For a Hann window, the group delay is fairly predictable, but if you switch to something like a Blackman-Harris window for better sidelobe rejection, the group delay changes and you need to adjust accordingly. This is an edge case that rarely gets mentioned in beginner guides. I hit a wall once where my implementation worked fine for sustained tones but sounded gated and stuttery on drums. The issue turned out to be that transient energy was spreading across too many frames due to the window length. The fix was implementing a variable window length approach where I'd use shorter windows during high-energy transients and longer windows during steady-state sections. It added complexity but cut the artifact level significantly.
Get the Full Details

Limitations and when this approach fails
Bednar's phase vocoder method isn't a universal solution. It struggles with very high time-stretch factors, say beyond two or three times the original speed, because phase error accumulates across frames and eventually the overlap-add synthesis breaks down. For extreme stretching, you'd need to look atgranular synthesis approaches instead, which chop the audio into tiny grains and reposition them. Those have their own artifacts but handle large stretch factors better. Another limitation is computational cost. A properly implemented phase vocoder with Bednar's corrections is not lightweight. On a typical CPU, you're looking at roughly 5 to 10 times real-time processing for a single channel at decent quality. If you're working in a resource-constrained environment, this might be impractical. GPU acceleration helps but introduces its own complexity with memory transfers. The method also doesn't handle polyphonic content as cleanly as monophonic material. When multiple frequencies are present in the same frame, phase interactions between them can cause beats and fluctuation artifacts that aren't fully eliminated by the correction terms. This is a fundamental limitation of the short-time Fourier transform approach, not something specific to Bednar's variant.
What to Do Instead in Some Cases
If you need simple time-stretching without worrying about implementation details, there are solid open-source libraries available. Rubber Band Library implements phase vocoder methods that include some of these corrections and are well-tested. For more advanced work, Sonic Paraphrase is another option, though it uses a different underlying approach based on spectral subtraction. If you're building something from scratch and want the Bednar approach, you'll find the core ideas described in his published papers and technical reports. The implementation isn't trivial but it's manageable if you take it step by step. Start with a working phase vocoder from a tutorial, then add the phase correction terms, test on clean signals, and only then move to more challenging material. The bottom line is that time-scale modification using phase vocoder methods remains one of those areas where theory and practice diverge enough that you really need to run your own tests. What works for your particular use case might differ from someone else's. Keep measurements handy, listen critically, and don't assume that because an algorithm is well-known it will behave the way you expect on your actual audio material.