Background Noise Remover

The recording

Nothing is uploaded: the recording is decoded, cleaned and written back inside this tab. Any container the browser can open, up to 20 minutes. A video keeps its picture exactly as it is — only the sound is replaced.

Engine

Strength

The engines hear up to 8 kHz, and above that this follows what they heard below it. Switch it on for a recording whose interest lives up there — music, a room with air in it, a cymbal — and that band comes back untouched.

Volume

Brings the speech to a set level and, if you want, follows the speaker as they drift — leaning in, turning away, trailing off at the end of a sentence. Runs after the noise is removed.

Free Background Noise Remover for Video and Audio

Remove background noise from a recording and keep the voice: fans, traffic, keyboards, air conditioning, room hiss, a second conversation across the room. A speech model decides, moment by moment and band by band, what is a voice and what is not — the same idea behind the noise suppression in Microsoft Teams — and it all runs inside your browser, so the recording never leaves your computer.

Works with any video or audio file the browser can open. A video keeps its picture untouched and gets its sound replaced; audio comes back as WAV, M4A, MP3 or OGG.

Instructions

Choose the engine

Voice model is the one to use. It is GTCRN, published at ICASSP 2024: a network of forty-eight thousand parameters that scores alongside models thirty times its size, which is why half a megabyte of weights can be served with this page. It looks at the whole spectrum at once and recognises speech, so it removes the noise that a simple filter cannot — a keyboard, a door, someone else talking. The first run fetches the runtime it needs, about thirteen megabytes, and your browser then keeps it.

Classic is RNNoise, from 2018. It is already inside the page, so there is nothing to fetch and it is roughly twice as fast. On constant noise — a fan, hiss, mains hum — the two are close. On anything that starts and stops it is clearly behind.

Choose the strength

The strength is not how much of the noise is removed but how far any part of the recording may be turned down. A limit rather than a switch, because a gate that closes completely sounds like a gate: the room falls silent between words and every breath cuts off. Leaving a defined floor of the original underneath is what makes the result sound like a good recording instead of a processed one.

Balanced suits almost everything. Gentle is for a recording that was nearly clean already. Maximum is for a difficult one, and is the setting to listen to carefully before you save.

Evening out the volume

Separate from the noise removal, and off unless you ask for it. It measures how loud the speech actually is and brings it to a target — around −20 dBFS by default — so a recording made too quietly comes back at a normal level instead of needing the volume turned up. The same levelling is in the Video Editor, and it is literally the same code, so a file cleaned here and a clip levelled there arrive at the same place.

Only the overall level applies one correction to the whole recording and leaves everything inside it alone. Also the drift while they speak adds a slow curve that follows the speaker, lifting the sentences they trailed off on and holding back the ones they leaned into. The curve is deliberately slow — it changes over seconds, not syllables — because a fast one is heard as the volume breathing.

It never boosts silence. Room tone, breath and microphone hiss all sit under the noise floor, and lifting them to speech level is the most recognisable way to make a recording sound processed. It also runs after the noise is removed, which is the only order that measures the voice rather than the voice plus the background.

How the voice survives

The engines are trained at 16 kHz, and using their audio directly would hand back a voice with a telephone's bandwidth. This tool never does that. It takes only the decision they made — how much to keep, in each band, at each moment — and applies it to your recording at its own sample rate, with its own phase. The band above the model's reach, where sibilance and breath live, follows the top of the band it can see. Nothing is resampled, and a passage the engine leaves alone comes back sample for sample as it went in.

Listen before you save

The two buttons under the player switch between the recording as it arrived and the cleaned one at the same instant. Judge the voice, not the silence: any setting makes the pauses quieter, and the question that matters is whether the words still sound like the person who said them. If consonants have gone dull, or the voice flutters at the start of a word, use a gentler strength.

What it cannot do

It removes noise, not echo: a room with hard walls still sounds like a room with hard walls. It cannot separate two people talking at once — both are voices, and it keeps voices. Music behind speech is treated as noise and will be damaged. And nothing here recovers detail that the noise destroyed; the quieter the voice was against the background, the less there is to work with.

Time and memory

Recordings up to 20 minutes are accepted. The whole soundtrack is held uncompressed while it is worked on, which is what sets that limit; longer files should be split first. Expect the voice model to take somewhere near a fifth of the recording's length on a current laptop, and the classic engine about half of that. Stop cancels immediately and keeps nothing.