Areas We Cover
Categories
TWO REQUESTS, ELEVEN DAYS, AND THE REHEARSAL ROOM WHERE THEY STOPPED BEING THE SAME PROBLEM
by Jim Allen | August 1, 2026
in Technology
“Can we just get the song without the singer on it?”
Every music director has fielded that sentence, usually from a producer, usually on a Tuesday, usually about a number that opens in eleven days. What makes it interesting is that it is almost never one request. Buried inside it are two, and they have completely different answers.
The first is about audio: point a vocal remover at it and give me the track with the voice taken off. The second, which surfaces about a day later, is about notation: nobody can tell me what the second trumpet is actually playing in bar 34.
For years both got the same shrug. They no longer do, and the tools that answer them are not the same tool.
Request one: the track without the voice
A browser-based vocal remover takes a finished stereo recording and hands back two files — the voice on its own, and everything else on its own. It runs in a tab. A four-minute number comes back in roughly a minute.
What a rehearsal room actually gets from a vocal remover is narrower than the marketing suggests, and more useful than the skepticism suggests.
Cast members stop unconsciously copying the original performer’s phrasing, because that performer is no longer in the room. A director can loop sixteen bars at half speed without a vocal line fighting the tempo change. And a singer who has been rehearsing along with the commercial recording usually discovers, within about ninety seconds of the vocal disappearing, that the original singer had been carrying the pulse for them the entire time.
The limits are physical rather than fixable, and worth knowing before you burn an afternoon. A voice recorded in a large hall arrives at the microphone several times over — once directly, then repeatedly as the room returns it. Strip out the direct signal and those returns remain, sitting under the band like a singer who left but did not stop. Dense harmony stacks give the software no clean edge to cut along. Saxophone lives close enough to a human voice, in both register and vibrato, that it sometimes leaves with the vocal.
And the file you start from sets a ceiling nothing downstream can lift. A vocal remover cannot restore detail that compression already discarded, so the CD beats the phone download every time.
Request two: what are the notes
This is the one that used to end the conversation, and it is a genuinely different question. Removing a voice rearranges audio. Getting a chart out of a recording means turning audio into symbols — pitch, start time, duration — which is a translation, not a subtraction.
Pitch-detection systems read the performance and propose what was played. Browser tools that convert mp3 to midi follow the fundamental frequency through time and decide where one note stops and the next begins, then let you look at the result before exporting it.
That review step is the part worth caring about. The output is a first draft, not a transcription, and it behaves accordingly: clean on an exposed solo line, progressively less certain as the texture thickens. Octaves are the honest illustration of why. The second harmonic of any note sits exactly where the note an octave above it would sit, so an algorithm listening to a single sustained pitch genuinely cannot always tell whether it heard one note or two.
A music director can work around that, because you read music and the software does not. You are not asking it to be right. You are asking it to save you from transcribing forty bars by ear at eleven at night, and then you fix bar 34 yourself in ninety seconds.
Where nobody can read the output, the same tool becomes a liability — it produces confident-looking files that a musician will later rebuild from scratch.
The order of operations that saves the most time
Answer three questions before opening anything.
Rehearsal or performance? Everything described here belongs in the first category. Teaching a harmony, studying an arrangement, working out a horn line — that is ordinary preparatory use. A file running through the house PA on opening night is a different matter entirely, governed by the venue’s licences and the producer’s clearances, and settled by the people who handle rights rather than by anybody’s browser.
What is the source? Take the CD over the download. This decision costs nothing and determines the ceiling for both tasks.
Who checks the output? A vocal remover needs an ear on it. Note detection needs someone who reads music. Where neither person is in the room, the tool is not saving labour, it is deferring it.
One rule for the callboard
Use it to learn, not to perform.
Both tools produce reconstructions. What a vocal remover returns is an estimate assembled from a mix; a converted MIDI file is an estimate assembled from a waveform. Estimates are excellent for rehearsal and wrong for delivery.
If the original session files exist in someone’s archive, use those instead. They beat an estimate every time, and they always will.