What Scenes Speech Therapy Actually Is
I ran into this term several years ago when I was dealing with some spatial audio projects, and it kept coming up in discussions about scene-based speech processing. At its core, it is a framework for representing and manipulating speech within a scene context rather than as an isolated audio signal. The main idea is that speech does not exist in a vacuum. When someone speaks in a room, in a car, or outdoors, the sound carries with it environmental characteristics. Scenes Speech Therapy treats those environmental characteristics as data that can be modeled, analyzed, and manipulated. The terminology comes from a combination of spatial audio research and speech processing work. People working in scene representation, ambisonics, and acoustic modeling eventually started applying those same principles to speech signals. Instead of just isolating the voice, you treat the voice and the scene together as a unified acoustic object. This changes how you do noise reduction, how you do source separation, and how you reconstruct speech for different listening conditions.
Scenes Speech Therapy: The Basics
Here is how the process generally works. You start with a recording that contains both speech and scene information. Scene information includes reflections, reverb tails, ambient noise, directional cues, and whatever else is present in the acoustic environment. A standard speech enhancement pipeline would try to strip all of that away and leave only the dry voice. Scenes Speech Therapy takes a different approach. It separates the speech component from the scene component, but it keeps both available for manipulation. The separation is typically done using deep learning models trained on large datasets of speech in various environments. Models like DeepFilterNet, Demucs, or custom scene-separation architectures are commonly used. The output is usually a set of stems: the clean speech, the scene reverb and reflections, and the ambient noise floor. From there you can do things like transplant the speech into a different scene, adjust the perceived distance, or modify the room characteristics without making the result sound artificial. I spent a few months building a pipeline around this for a project where we needed to generate realistic speech in simulated environments for training ASR systems. The basic workflow looked like this. First, run the input through a scene-separation model to get the three stems. Second, process the clean speech stem with whatever enhancement you need. Third, synthesize a new scene stem matching your target environment. Fourth, recombine the stems. The whole thing took about forty-five seconds per minute of audio on a decent GPU, which was fast enough for our production needs.
Why Standard Speech Enhancement Fails Here
This is the part most people miss when they first encounter the concept. Standard speech enhancement models are trained to remove everything that is not speech. That works fine if your goal is a dry, studio-quality vocal track. It fails completely if your goal is natural-sounding speech in a realistic environment. When you strip all scene information out of a recording and then try to add reverb back in with a simple convolution, the result sounds wrong. The reverb does not match the original vocal characteristics. The spectral envelope of the voice and the spectral characteristics of the room are coupled in ways that a simple add-reverb step does not preserve. Scenes Speech Therapy addresses this by treating the scene as an inseparable part of the original recording. The model learns the joint distribution of speech and scene during training, which means the recombined output maintains the natural coupling between voice and environment. The difference is subtle but noticeable. A listener might not be able to point to exactly what is different, but a recording processed with standard enhancement plus artificial reverb will sound like it was recorded in a studio and then dropped into a virtual room. The Scenes Speech Therapy approach sounds like the room was there from the beginning.
Get the Full Details

Common Pitfalls I Have Run Into
The first problem is model selection. Not all scene-separation models handle speech well. Some are trained primarily on music or general audio and will artifact the speech stem in ways that are hard to fix afterward. I ended up going with models specifically trained on speech-in-scene datasets because the difference in speech quality was immediately obvious. The tradeoff is that those models are usually more computationally expensive. The second problem is what happens when the scene is very loud relative to the speech. If ambient noise or competing sounds dominate the recording, the separation model may misattribute speech components to the scene stem or vice versa. I had a case where a recording with heavy traffic noise resulted in the model removing parts of the speech consonants and putting them into the noise stem. The workaround was to run the input through a directional preprocessing step first to boost the speech signal before separation, which gave the model a better starting point. It added about ten seconds of processing time per file but saved me from having to manually fix the artifacts. The third problem is that the approach does not work well with recordings that have very little scene information to begin with. A dry vocal recording in a treated booth has almost no scene content for the model to learn from. In those cases you are essentially just adding complexity with no benefit. If your source material is already clean, standard enhancement is faster and more predictable.
When This Approach Is Actually Worth Using
I would recommend this workflow when you need speech that sounds naturally integrated into a specific environment. That covers a few common scenarios. If you are building synthetic training data for speech recognition systems and need varied acoustic conditions. If you are working on audio for virtual reality or spatial media where speech needs to move naturally through a 3D space. If you are doing forensic audio analysis where understanding the original scene characteristics matters for evaluating the authenticity of a recording. If you are working in audio post-production and need to match dialogue to a different environment without it sounding processed. It is not worth using when you simply want to clean up a podcast recording. It is not worth using when you need real-time processing on limited hardware. It is not worth using when your source material is already high-quality and close to your desired end state. The processing overhead and the complexity of getting good results do not justify the approach in those cases.
Scenes Speech Therapy in Practice
Here is a concrete example from my own work. I had a client who needed a series of spoken instructions recorded in a noisy warehouse environment for a safety training module. We did not have access to an actual warehouse, so the approach was to record the voice in a quiet room and then use Scenes Speech Therapy techniques to place it into a synthesized warehouse scene. The process took roughly three minutes per minute of final audio. We ran the dry recording through a scene generation model conditioned on warehouse impulse responses, separated the speech from any artifacts introduced during generation, and then blended the results. The final output passed a blind listening test against actual warehouse recordings for most listeners, which was the requirement. The most important detail from that project was the conditioning step. Simply generating a warehouse scene and layering the voice over it did not produce convincing results. The voice needed to be processed through the same acoustic transformation that defined the scene. That meant using a neural vocoder or spectral transformation model that could apply the room characteristics directly to the speech signal rather than just adding reverb afterward. This step alone accounted for about sixty percent of the processing time in the pipeline. If you are looking to try this yourself, the main open-source tools you would need are a scene-separation model, a scene generation or conditioning model, and a recombination pipeline. Python-based workflows with PyTorch are the standard. There are implementations available on GitHub that cover most of the individual components, though you will likely need to stitch them together yourself depending on your specific use case. The learning curve is steeper than a standard audio processing tutorial because you are dealing with multiple models in sequence rather than a single filter or effect.
