Split long audio files into short, overlapping chunks for AI vocal/instrumental separation services with per-request duration limits — then merge the processed chunks back together with a real equal-power crossfade instead of a naive concatenation.
Most browser-based AI stem-separation tools (vocal/instrumental splitters) cap how long a single upload can be (commonly around 60 seconds). Splitting a full song into chunks and running each one separately is the only option — but that introduces two problems:
- Where you cut matters. Cutting at the single loudest-drop sample instead of a stable, sustained quiet passage puts the cut where the AI model has the least stable context on either side, so its output is more likely to diverge right at that edge.
- How you rejoin matters. Naively concatenating the processed chunks
(e.g.
ffmpeg -f concat) just glues PCM data together — it does not blend anything. Because each chunk was processed independently by the AI, even small differences in level or phase at the boundary produce an audible click, pop, or seam.
StemSplit solves both:
stemsplit splitfinds cut points inside a smoothed loudness curve, favoring sustained quiet passages over a single noisy minimum, and exports each chunk with a configurable overlap into its neighbors.stemsplit mergere-assembles the AI-processed chunks using an equal-power crossfade across each overlap region, so the seam is blended rather than glued.
song.mp3
│
▼ stemsplit split
chunks/part01.mp3, chunks/part02.mp3, ... + chunks/manifest.json
│
▼ (run each chunk through your AI vocal/instrumental separator, e.g. a
web tool with a ~60s per-file limit — keep the same order, renaming
is fine)
processed/part01_instrumental.mp3, processed/part02_instrumental.mp3, ...
│
▼ stemsplit merge
song_instrumental_full.mp3
Requires Python 3.9+ and ffmpeg on your PATH
(used by pydub to read/write compressed audio formats such as MP3).
git clone https://github.com/ndhn27/stemsplit.git
cd stemsplit
pip install -e .stemsplit split "song.mp3" --target 52 --overlap 4 --out chunks/Key options:
| Option | Default | Description |
|---|---|---|
--target |
52 |
Target length of each chunk in seconds, before overlap is added. Tune so target + after-slack + overlap stays under your AI service's limit. |
--before-slack |
20 |
How much shorter than --target a chunk may end up, while searching for a quiet cut point. |
--after-slack |
3 |
How much longer than --target a chunk may end up. |
--overlap |
4 |
Seconds shared between two neighboring chunks, for crossfading later. 0 disables overlap (equivalent to a plain split). |
--smooth-ms |
250 |
Moving-average window (ms) applied to the loudness curve before searching for the quietest point, to favor sustained silence over a single noisy sample. |
--fade-ms |
8 |
Fade in/out (ms) applied only at the very start and end of the whole track — internal chunk edges are left raw since they'll be crossfaded during merge. |
--min-final |
5 |
If the last chunk would be shorter than this (seconds), it's merged into the previous one. |
--out |
chunks |
Output directory. |
This produces chunks/<name>_part01.mp3, chunks/<name>_part02.mp3, ... and
chunks/manifest.json, which records the source file, the overlap length,
and the part order — stemsplit merge needs it later.
Process every file in chunks/ through your vocal/instrumental separation
service of choice and save the results into a new folder (e.g.
processed/). Keep the same order the parts were split in — files may
be renamed (AI services often append suffixes like _instrumental or
_vocals), but don't drop or reorder any of them.
stemsplit merge processed/ --manifest chunks/manifest.json --out "song_instrumental_full.mp3"If a single source chunk produced more than one output file (e.g. both a
vocals and an instrumental version), narrow it down with --match:
stemsplit merge processed/ --manifest chunks/manifest.json \
--match instrumental --out "song_instrumental_full.mp3"Key options:
| Option | Default | Description |
|---|---|---|
--manifest |
(required) | Path to the manifest.json produced by stemsplit split. |
--overlap |
(from manifest) | Override the overlap length used for crossfading. |
--match |
– | Filter which processed file to use when a part maps to several files. |
--method |
equalpower |
equalpower (recommended, keeps perceived loudness constant across the transition) or linear (pydub's built-in crossfade). |
--bitrate |
320k |
Output MP3 bitrate. |
--samplerate |
44100 |
Output sample rate. |
Across each overlap region, the outgoing chunk's tail is multiplied by
sqrt(1 - t) and the incoming chunk's head by sqrt(t), with t going
from 0 to 1. Because sqrt(1-t)² + sqrt(t)² = 1, the combined power
stays constant through the transition — unlike a plain linear fade, which
dips in perceived loudness at the midpoint.
pip install -e ".[dev]"
pytest
flake8 src testsMIT — see LICENSE.