Getting Started With I Hear A Symphony

I used to work with this a lot back when I was doing mix work for a small label out of Richmond. It came up again recently when someone sent me a project that needed stems separated out, and I realized I had no clear idea of where the current version of I Hear A Symphony stood or how people were actually using it now. It is a stem separation and source separation platform. You feed it a mixed audio file and it outputs isolated tracks -- vocals, drums, bass, other instruments. The approach is rooted in deep learning models trained on large datasets of multi-track recordings. It is not magic. It is pattern matching at scale with some decent post-processing on top. You upload your audio file, pick a model or let it auto-select, and wait. The longer the track, the longer the render. I usually run it through the web interface first, then drop the files into a script for batch processing if I need more control. File size limits used to be tight. They opened up a bit in recent updates, but I still send anything over about 100 megabytes through the desktop client when one exists, because the browser tab tends to hang halfway through.

The output comes as individual WAV stems. Sampling rate and bit depth match your input by default, which saves you a step. You can choose different model flavors depending on whether you care more about vocal isolation versus drum transients. The defaults are fine for most things. The speciality models matter when your source material is messy.

A Specific Problem I Hit

Once I fed it a live concert recording where the crowd noise was roughly as loud as the snare drum. The model kept pulling the clap track from audience hands into the drum stem and leaving the actual kit pieces scattered across two or three outputs. What worked for me was running it twice -- once with the vocal-focused model and once with the instrumental-focused one -- then manual phase alignment and subtraction between the two result sets. It is not elegant. It gets the job done in about twenty minutes per track instead of the five minutes the software promises on clean studio material.

Get the Full Details

The Supremes - I Hear A Symphony
The Supremes - I Hear A Symphony

Counter-Intuitive Things Beginners Miss

First, source separation is a lossy process. You will never get back the original stems perfectly, no matter what marketing copy says. Second, feeding it a heavily compressed or brick-wall-limited master makes everything worse. The model struggles with transients that have been flattened. If your source is a club mix with -6 LUFS integrated loudness, run a light normalization pass first or feed it the pre-master if you have it. Third, the quality varies wildly by genre. Electronic music with isolated synth lines separates cleanly. Jazz fusion with overlapping acoustic guitars and brushed drums? Expect smeared artifacts in the midrange.

When It Fails Completely

It fails on mono source material where all instruments occupy the same frequency space with no stereo separation to exploit. It also struggles with recordings that have heavy reverb tails bleeding across channels, because the reverb is already mixed together and the model has no spatial cue to pull from. If your project is a vintage FM broadcast recording or a phone memo, just accept that it will sound terrible and move on to manual editing or acceptance.

Alternatives Worth Knowing

If I Hear A Symphony does not give you clean enough results, splitting to RX by iZotope for manual stem extraction or using Demucs from Facebook Research locally is the next step. Demucs runs on your own hardware and gives you more control over the model version, but it requires a GPU for anything reasonable. For quick web-based jobs where you do not want to install Python dependencies, I Hear A Symphony remains one of the faster options available, even if the results need some cleanup afterward.

I Hear A Symphony - The Supremes - Vintage vinyl album cover Stock Photo - Alamy
I Hear A Symphony - The Supremes - Vintage vinyl album cover Stock Photo - Alamy