What the Kohlberger Response Actually Is
The Kohlberger Response is a measurement technique used in psychoacoustics and room acoustic analysis, specifically for characterizing how listeners perceive transient events in a space. It comes from the work of Dr. Wolfgang Kohlberger and colleagues at the Technical University of Berlin, who were trying to pin down why two rooms with nearly identical RT60 values could sound completely different to someone sitting in them. The basic idea is straightforward: instead of measuring just the decay tail of a room impulse response, you look at the ratio between the early energy (the first 50 to 80 milliseconds) and the later arriving reflections, but you weight that ratio according to human temporal integration characteristics. The original papers from the mid-1990s describe the calculation method in detail, and it has been picked up by a handful of researchers since then without ever becoming mainstream.
How to Calculate the Kohlberger Response
I ran into this when I was trying to diagnose a broadcast studio that measured fine on paper but sounded dead and lifeless to everyone who worked in it. The RT60 was around 0.32 seconds across the board, which is right where it should be for a small talk studio. But engineers kept complaining about echo blur during fast speech passages. That's when I pulled the Kohlberger Response out of an old AES paper and ran it against our impulse data. Here's how you do it. You need a measured room impulse response, ideally from a logarithmic sine sweep deconvolution so your signal-to-noise floor is low enough. Take that IR and split it at whatever early/late boundary your application calls for—50 ms is the most common choice for speech environments, 80 ms for music. Integrate the energy in each region separately. Then apply the temporal weighting curve, which is essentially a raised cosine function centered around the onset of the transient. The formula looks like this: E_weighted = integral from 0 to t_early of p(t)^2 * w(t) dt + integral from t_early to t_late of p(t)^2 * w(t) dt
Where w(t) is the temporal integration window derived from human backward masking experiments. The output is a single number per octave band, typically expressed in decibels relative to the total integrated energy. A higher Kohlberger Response value generally means the room preserves transients better. A lower value means the later reflections are smearing things out even if the total reverberation time looks acceptable. My specific problem was that the studio had absorptive panels placed exactly where you'd put them based on standard RT60 optimization. But the Kohlberger Response showed values around -18 dB in the 1 kHz to 2 kHz region, which is well below what you want for intelligibility. The issue was a pair of parallel glass windows about four meters apart that were creating a flutter echo train. The broadband absorption killed the overall decay time nicely, but those discrete early reflections were still punching through the temporal window. I solved it by adding a thin diffuse scattering element—a quadratic residue diffuser on the ceiling between the windows—which broke up the coherent reflection path without changing the measured RT60 by more than 0.02 seconds.
Why This Matters More Than You'd Expect
Most people measuring rooms today rely on either RT60 or EDT (Early Decay Time), and both of those give you a single global number that smooths over a lot of important detail. The Kohlberger Response forces you to look at the problem through a human-perception lens rather than a pure energy-decay lens. That distinction matters in broadcast, recording, and live sound environments where transient preservation directly affects perceived clarity. There's a catch, though, and it's a big one. The calculation assumes a point source and a point receiver, which is never true in practice. Real speakers have directivity patterns that change with frequency, and real listeners have two ears spaced 17 centimeters apart with pinnae that modify the incoming waveform. When I tried to extend the method to binaural measurements using an artificial head, the results became noisy and hard to interpret because the interaural time differences disrupted the temporal weighting function. You essentially have to pick a side and commit to either a single-channel approximation or accept that the numbers will have more variance. Another thing nobody tells you: the method is extremely sensitive to the quality of your impulse response measurement. If your signal-to-noise ratio isn't at least 50 dB in the bands you care about, the late-time integration blows up and the Kohlberger Response values become meaningless. I've seen people run this calculation with hand-clap measurements and then wonder why their numbers looked like random noise. Use a swept sine or MLS signal, make sure your recording chain has enough headroom, and average at least three measurements at each position before you do any integration.
The tooling situation is also rough. There's no mainstream software package that does this out of the box. I ended up writing a small Python script using scipy.signal for the deconvolution and numpy for the integration, and it took me about two weeks to get it to produce consistent results across multiple test spaces. If you search for Kohlberger Response download or similar terms you'll find a few academic code repositories, but they're usually tied to specific papers and not maintained. The original 1996 AES convention paper is the canonical reference if you want to reverse-engineer it yourself. The main limitation I keep running into is that the Kohlberger Response doesn't account for spatial diffusion. Two rooms can have identical Kohlberger Response values but feel radically different because one has a dense, uniform reflection pattern and the other has a few strong discrete echoes. EDT gets partway there by looking at the decay slope of the early portion, but it still treats energy as a scalar quantity. What you'd really want is a combined metric that includes both the temporal weighting and some measure of reflection density, but nobody has published a clean way to do that yet. If you're working in a space where speech intelligibility is critical—broadcast control rooms, lecture halls, courtrooms—the effort to implement this is probably worth it. For home studios and general-purpose rooms, you're better off spending your time on proper bass trap placement and avoiding parallel surfaces. The Kohlberger Response shines in environments where the problem isn't too much reverb but the wrong kind of reverb arriving at the wrong time.