What Project Smash Actually Is
Project Smash is an open-source audio source separation model. It splits mixed audio tracks into individual stems like vocals, drums, bass, and other instruments. It uses a transformer-based architecture that was trained on large-scale music datasets, which is why it tends to handle complex mixes better than older spectrogram-masking approaches. The project lives on GitHub. You pull the repo, install the dependencies, and you can run it either as a Python script or through a command-line interface. There's also a web UI variant people have built around it if you don't want to type commands.
Project Smash
How to Set It Up
First, make sure you have Python 3.9 or higher. I ran into issues with 3.11 on my first try because some of the torch dependencies had compilation hiccups, so I dropped to 3.10 and everything installed clean. Create a virtual environment, clone the repo, and run the requirements install. If you're on Linux or macOS with an NVIDIA GPU, getting CUDA set up properly will make a real difference. Without GPU acceleration, a three-minute song can take twenty minutes to process. The basic command looks something like this: python -m project_smash input.wav --output_dir ./stems
That's it. It'll load the model weights, run the separation, and save individual stem files. The default model gives you four stems by default. Some configurations let you request more granular splits, but that increases VRAM usage significantly.
Get the Full Details

What I Wish I'd Known Before Using It
The biggest issue people run into is that Project Smash isn't magic. It performs well on produced music where the elements are fairly distinct in the frequency spectrum. But if you feed it a live recording with overlapping frequencies or heavy reverb, the stems bleed. I processed a jazz recording once where the piano and upright bass were so tightly mixed that the model kept assigning bass frequencies to the drum stem. Nothing you do in post will fully fix that kind of artifacting. Another thing: the model weights are substantial. The full checkpoint is around 500MB. If you're running this on a machine with limited storage or you need to deploy it somewhere constrained, keep that in mind. Also, the inference speed degrades non-linearly with track length. Doubling the song length roughly triples the processing time because of how the transformer attention mechanism scales. For edge cases where the default model struggles, I found that preprocessing the input with a light EQ boost around 200Hz to 400Hz helped the bass stem separate cleaner from the kick drum. It's not a guaranteed fix, but it's worth trying before you move on to a different tool. If that doesn't work either, MDX-Net based separators like Demucs can handle certain mixes better, though they have their own weaknesses with vocal removal.
Practical Considerations
If you're using this for content creation or stem extraction for remixes, the quality is generally good enough for most applications. The vocal stems come out relatively clean on pop and rock tracks. Electronic music separates with surprising accuracy. Classical and acoustic recordings are where you'll see the model break down noticeably. The project is actively maintained but not every issue gets addressed quickly. Check the GitHub issues before diving in if you run into something that seems broken. Someone else has probably hit the same problem and there might be a workaround posted. There's no official installer or package manager distribution. You're working from source. That means occasional breakage when upstream dependencies shift. I keep a pinned commit hash in my workflow so I'm not surprised when a routine pip install breaks my setup after a library update.
The license is MIT, so you can use it commercially without worry. Just give attribution if you reference the project in any published work. That's the short version of working with it. It's a solid tool for what it does, but it has clear limits. Know your source material before you feed it into the model and adjust your expectations accordingly.
