How Tiny Jukebox Actually Works and How to Use It Without Losing Your Mind

Tiny Jukebox is a browser-based stem splitter. You drop an audio file in, and it uses a small deep learning model to separate it into roughly four tracks: vocals, drums, bass, and other instruments. It runs locally in your browser using TensorFlow.js and the Web Audio API. No server uploads, no accounts, no processing queues. The whole thing is built on Google's Demucs-inspired architecture, but scaled down to fit in a client-side JavaScript bundle.

The interface is embarrassingly simple. There's a drag-and-drop zone, a few slider controls for each stem, and a play button. That's pretty much it. You can also download individual stems or the full separated mix as WAV files. The free version supports files up to about 25MB, which covers most mp3s and lossless tracks you'd reasonably throw at it. Anything larger and you'll get a timeout before the model even starts processing. You don't actually download anything. The entire application lives at tinyjukebox.com and loads from there. Sometimes people ask about a desktop version or a GitHub repo they can self-host, and I just tell them it doesn't exist in that form. There is a self-hostable version on GitHub by the original developer, but it requires some node setup and isn't really meant for casual use. For 99% of people, just going to the website is the correct approach. Here's what I wish people understood before they start using it: this tool is not precise. It's a "tiny" jukebox for a reason. The separation quality is decent for quick reference tracks, rough remixes, or practice purposes, but if you're trying to isolate a clean vocal for professional use, you're going to be disappointed. The model trades accuracy for speed and browser compatibility. You'll get artifacts around complex passages, the bass line often bleeds into the drum track, and anything with heavy reverb on the vocals tends to leak into every stem.

I spent about three weeks last year trying to use it for a personal project where I needed isolated guitar tracks from a live recording. The model kept splitting the guitar between the "other" and "drums" stems depending on the frequency range. What I ended up doing was running the track through Tiny Jukebox first to get a rough separation, then feeding the vocal stem into a separate dedicated vocal isolation tool for cleanup. It's not ideal, but it cuts the problem down significantly compared to starting from scratch. One edge case that took me forever to figure out: Tiny Jukebox handles stereo files fine, but if your source audio has phase issues between channels — common with some live recordings or poorly mixed tracks — the separation gets weird. The model essentially treats left and right channels independently during preprocessing, and phase cancellation throws off the frequency estimates. My workaround was to just mono-sum the track first before feeding it in. You can do this in any DAW or even with a free tool like Audacity. Select both channels, route to mono, and save. Then load that into Tiny Jukebox and the results are noticeably cleaner.

What You Actually Get Out of It

The four stem output is the core of it. Vocals, drums, bass, and other. That's it. The "other" category is a catch-all that picks up everything that doesn't fit neatly into the first three — guitars, synths, vocals with heavy effects, percussion that doesn't hit on the main beats, etc. The model wasn't trained to distinguish between a snare and a hi-hat, or between a piano and a string section. It groups by broad spectral characteristics. Processing time depends heavily on your hardware. On a mid-range laptop from 2021, a three-minute song takes roughly 45 seconds to a minute to separate. On an older machine or a Chromebook, it can take several minutes or just hang entirely. The browser will sometimes ask if you want to close the page because a script is running unresponsive. This is normal. Don't close it. Just wait. If you're processing multiple tracks in a row, clear your browser cache between uses. The model gets loaded into your browser's memory each time, and Chrome tends to hoard it. I noticed my tab memory usage climbing to over 600MB after three or four songs in a row, and the processing started getting slower and slower. Reloading the page resets everything.

Get the Full Details

Jive Rock Sixty Retro Mini Jukebox Dark Wood by Steepletone – Unusual Designer Gifts
Jive Rock Sixty Retro Mini Jukebox Dark Wood by Steepletone – Unusual Designer Gifts

When to Use It and When to Walk Away

Tiny Jukebox is good for: quick stem extraction for practice sessions, rough mashups, DJ training tracks, sampling projects where you don't need surgical precision, and anything where you need a fast answer without signing up for a service or paying a subscription. It's bad for: professional remix work where you need clean stems, podcast voice isolation when you need broadcast quality, removing vocals from a track for a karaoke version where the instrumental needs to sound natural, and anything with dense, layered productions where the model's limited capacity can't disentangle overlapping frequencies. For those cases, you'd want to look at something like Demucs directly, RipX, or a dedicated service like Moises or Lalal.ai. Those use larger models, often server-side, and give you better separation at the cost of speed, money, or privacy since your audio goes somewhere else.

The one thing Tiny Jukebox does better than most paid tools is that it never touches your audio. Everything happens on your machine. If you're working with unreleased music or anything you don't want uploading to a cloud service, that's genuinely useful. I've had producers tell me that's the only reason they use it despite the quality limitations.

Practical Tips That Actually Matter

First, always normalize your input audio before feeding it in. Tiny Jukebox's model expects a certain amplitude range, and if your track is quietly mixed or dynamically compressed differently than the training data, the stem separation skews. A quick pass through a free compressor/limiter or even just the normalize function in Audacity makes a real difference. I usually aim for peaks around -3dB to -1dB. Second, the sliders on the output are actually mixed together, not independent volume controls for each stem in the way you might expect. Turning down the vocal slider doesn't just lower the vocal stem — it also reduces how much the vocal model influences the other outputs. If you're getting weird ghost artifacts in the drum track after lowering the vocal slider, turn the vocal back up and adjust differently. Third, try exporting as WAV instead of MP3 if the tool lets you. Some versions offer both formats, and the WAV export preserves more of the separation quality. MP3 compression adds its own artifacts on top of the model's artifacts, and they compound in unpleasant ways.

Jive Rock Sixty Retro Mini Jukebox Light Wood by Steepletone - Theperfectlivingroom
Jive Rock Sixty Retro Mini Jukebox Light Wood by Steepletone - Theperfectlivingroom

Finally, don't expect the same quality from every genre. Pop and electronic music with clear separation between elements tend to work well. Jazz, classical, and heavily layered rock are where the model shows its limitations most clearly. I once ran a full orchestral recording through it and got something that sounded like everyone was playing in the same room but nobody could hear each other. Not useful, but interesting.