How a song gets split
Four stems from one mix
The separator is HTDemucs, a model published by Meta Platforms, Inc. and licensed under the MIT License. It reads a stereo mix and writes four stems: vocals, drums, bass, and other. Other is every instrument that is not one of the first three. That is why a guitar and a piano come out together. The weights used here are the HTDemucs v4 hybrid transformer, converted to ONNX so a browser can run them.
The default file stores those weights in 16-bit floats and does the math in 32-bit floats. It is about 92 MB. An optional high-quality file is the original 32-bit weights, about 181 MB. In listening tests the two were nearly the same. Help & settings is where you pick. Each file is cached on its own, so choosing one does not delete the other.
Inference runs in a worker so the page can still scroll and play. When the browser is allowed to use several processor threads, it uses up to four. When it is not, it uses one thread and finishes more slowly. The result is the same four stems either way.
Where the model file comes from
The page itself does not contain the weights. The first time you separate a song, the browser downloads them from models.stemsplitterlab.com and stores them in this browser under a versioned address. Later separations read that copy. The song you opened is not part of that request. Only the model file moves across the network, and only in that direction.
A small WebAssembly runtime is served with the site so the worker can run on the CPU. A larger graphics runtime, used only if this browser can reach the graphics chip, is fetched from the same model host, because it is too big to sit on the page host. If that fetch fails, the worker continues on the CPU runtime.
Tempo, key, and the player
BPM and key are estimated with ordinary signal processing on the original mix, before any stem exists. The BPM and key finder explains half-tempo readings and relative keys. They are estimates, and the player says so when a reading is weak.
Playback can change speed without changing pitch, and pitch without changing speed, using SoundTouchJS (LGPL-2.1). Exports to MP3 use @breezystack/lamejs (LGPL). WAV export is uncompressed audio written in the browser. None of these steps upload the mix.
Credits and the sample clip
ONNX Runtime Web runs the model (MIT). The segment loop and STFT helpers come from demucs-web (MIT). The practice quartet in the library is an original 20-second synthetic recording, dedicated to the public domain under CC0, so the tool pages can show a spectrum of a file this site is allowed to publish. It is not a commercial track.
A fuller notice of those licenses is in the project’s NOTICES file for anyone shipping a copy of the software. The model authors’ MIT license requires the copyright line and the permission notice to travel with the weights. This page is the public form of that credit.
Last updated 28 September 2026