Your file stays on your device. Nothing is uploaded to our servers — the media is read, analyzed and cut entirely inside this browser tab.
Click or drop a video or audio file here
MP4, MOV, M4V, WebM, MKV, MP3, WAV, M4A, AAC, OGG and FLAC — codec support depends on your browser
Free Silence Cutter — Remove Silence from Video and Audio
Automatically find and remove silent sections from video or audio. Fine-tune the silence threshold, preview every cut on the waveform before anything is written, and process your media directly in your browser. It is free, your file never leaves your computer, and you see exactly what will be removed before you cut.
The tool works with recordings and with music: lectures, interviews, podcasts, screen recordings, voice memos and any file where long pauses make the result longer than it needs to be.
How to remove silence from video or audio
1. Choose your media
Drop a video or audio file onto the upload area, or click it to pick one. The file is inspected by reading its header, not its extension, so a renamed or remuxed file is still identified correctly. Duration, container, codecs, resolution and channel layout appear as soon as they are known; anything that could not be determined with certainty is simply not shown.
2. Configure the detection
Start from a preset — Balanced, Aggressive, Natural speech or Long pauses only — and adjust from there. The four settings that matter most are on the main panel; the rest live under "Advanced settings". Changing any of them marks the analysis as out of date rather than silently re-running it, which keeps the sliders responsive on long files and saves your battery.
3. Analyze the media
Click "Analyze Media". The audio is decoded in small pieces and reduced to loudness statistics as it goes, so a multi-gigabyte recording is analysed without ever being loaded into memory. The interface keeps responding throughout, progress is reported per stage, and you can cancel at any point without losing the selected file.
4. Review the cuts
The waveform appears with every region that will be removed marked in a hatched overlay, alongside a timeline whose granularity adapts to the length of the file. The summary tells you how much will be removed, how many sections there are, and how long the result will be. Click anywhere on the waveform to move the player there, and turn on "Preview cuts" to hear the edit — the player skips the removed regions live, with nothing rendered yet.
5. Export the result
Click "Remove Silence", confirm the summary, and the file is written. Where the browser supports it you choose the destination up front and the output is streamed straight to disk instead of being assembled in memory. Your original file is never modified — the result is always a new file.
How silence detection works
Silence is not the absence of signal; it is the absence of anything loud enough to matter. The tool splits the audio into short windows — 20 milliseconds by default — and measures the RMS level of each one, which is the root mean square of its samples and a good stand-in for perceived loudness. That level is converted to decibels relative to full scale (dBFS), where 0 dB is the loudest a digital signal can be and every value below it is negative.
A window whose level falls below the silence threshold is a candidate. Consecutive candidates are grouped, and a group only becomes a real cut once it lasts at least the minimum silence duration. That second condition is what separates a pause from a gap between two words: without it, every syllable boundary would be a cut.
Finally the margins are applied. Keep before speech and keep after speech shrink each silence from its edges, so the cut never lands on the attack of a word or clips the tail of a sentence. The ranges that survive are the ones you see hatched on the waveform.
What silence threshold should I use?
−40 dB is a sensible default for a clean recording. If the room is noisy, air conditioning, traffic or computer fans sit somewhere around −50 to −35 dB, and a threshold below that noise floor will never detect anything — raise it until the pauses light up. If quiet speech is being cut, lower it instead.
A practical way to find the right value: analyze once at −40 dB and look at the waveform. If the hatched regions cover parts where someone is clearly talking, the threshold is too high. If long, obviously empty stretches stay untouched, it is too low. Two or three passes are usually enough, and each one costs only the analysis, never a render.
Why keep padding around speech?
Speech does not start at full volume. A word begins with a consonant that rises out of the noise floor over a few tens of milliseconds, and a sentence ends with a tail that decays just as gradually. A detector working purely on level will place the cut inside those transitions, which is what makes automatic edits sound clipped and abrupt.
The two padding values buy that time back. 120 ms before and 180 ms after is a good starting point: enough to keep the onset of a word intact, short enough that the pause still feels removed. The asymmetry is deliberate — trailing decay lasts longer than the onset that precedes speech.
Silence cutter settings
Silence threshold (dB) — the level below which audio may be considered silent.
Minimum silence duration — how long the level must stay below the threshold before the stretch counts as a removable pause.
Keep before / keep after speech — margins preserved at each end of a detected silence.
Minimum kept segment — a fragment of audio shorter than this, trapped between two silences, is absorbed into the cut. Useful for isolated clicks and false starts; aggressive if set too high.
Detection window — how much audio is measured at a time. Shorter windows react faster to sharp transitions, longer ones are steadier on noisy material.
Channel analysis — how a stereo or multichannel signal is reduced to a single decision. "Combined" measures every channel together and is the safe default; "Any channel" is stricter, "All channels" more conservative.
Automatic zoom after long pauses — when a pause is removed, the two halves it joined were filmed with the same framing, so the join shows up as a jump. Tick this and the picture pushes in slightly on what follows the longest pauses, which hides the jump and gives the edit a rhythm. Only pauses that stand out are used: the tool takes the average length of everything it removed and zooms only after the ones at least 30% longer than that, so an ordinary gap between two sentences is left alone. The gear beside the checkbox opens the settings — how far above average a pause has to be, how far the picture pushes in (10% to 20% by default), whether the amount is picked at random inside that range or in proportion to how long the pause was, how long the push-in takes and how long it is held. The random choice is seeded from the file itself, so the same recording always gets the same zooms and a second export never differs from the first. Nothing is stretched: the picture is enlarged and the edges are cropped. Like the crossfade, this is applied while writing the file, so changing it never invalidates an analysis.
Cut smoothing (crossfade) — a few milliseconds of fade at each cut point, which removes the click a hard splice would leave. It is applied while writing the file, so changing it does not invalidate the analysis.
Does the video get uploaded?
No. There is no upload, no server-side processing and no third-party service involved. The file is read from your disk by the browser, decoded by the browser's own codecs, and written back out by the browser. Nothing about the media — not its name, not its content, not the waveform — leaves the tab. You can verify this by opening your browser's network panel and watching it stay silent while a file is analysed.
Can I use large video files?
The tool is designed for them. Media is never loaded into memory as a whole: the file is read in ranges on demand, audio is decoded in small pieces and folded into statistics that are kilobytes in size, and the output is streamed to disk as it is produced where the browser supports it. The waveform stores a few thousand aggregated peaks rather than millions of samples.
Real limits still depend on your browser, the codec, available memory, CPU and GPU, and free disk space. A desktop browser will handle multi-gigabyte recordings; a phone may not, and the tool says so rather than failing halfway through. Analysis is far cheaper than export — if a device can do only one of the two, it will be the analysis.
Supported formats
Video: MP4, MOV, M4V, WebM and MKV. Audio: MP3, WAV, M4A, AAC, OGG and FLAC. What actually works depends on the codecs inside the container and on your browser: H.264 and AAC in MP4 are supported almost everywhere, VP9 and Opus in WebM nearly as widely, and HEVC or AV1 depend on the platform. The file is validated before any work starts, and an unsupported codec produces a clear message rather than a failure partway through.
The output keeps the input's container whenever the browser can encode into it. When it cannot — no browser ships an MP3 encoder, for instance — the audio is written as uncompressed WAV, which never loses a second generation of quality.
How the tool picks a processing pipeline
The rule is to avoid work before accelerating work. When nothing needs to be removed, the encoded streams are copied across untouched, which beats any encoder and costs nothing in quality. When a cut forces a re-encode, the video pipeline validates a hardware-accelerated configuration first and falls back to a compatible one when the browser declines. Loudness statistics go to the GPU only when the media is long enough for the transfer to pay for itself; below that, a Web Worker with typed arrays is faster.
Hardware acceleration is a preference the browser may or may not honour, and no web API reports which GPU or media block actually ran. The tool therefore says "hardware acceleration preferred" and never claims a specific device is in use.
Frequently asked questions
What is a silence cutter?
A silence cutter is a tool that finds the quiet stretches in a recording and removes them automatically, producing a shorter, tighter version of the same material. It is the mechanical part of an edit that would otherwise be done by hand, one pause at a time.
Can I remove silence from a video for free?
Yes. This tool is free, requires no account and has no export limits. It runs entirely in your browser, so there is no per-minute cost on our side to pass on to you.
Can I remove silence from audio too?
Yes. Audio-only files are handled the same way as video, and the output preserves the sample rate and channel layout of the original whenever the browser can encode them.
What dB level should I use for silence detection?
Start at −40 dB. Raise it towards −30 dB if a noisy room means nothing is being detected, and lower it towards −50 dB if quiet speech is being cut. The waveform tells you which of the two is happening.
Will removing silence reduce video quality?
Cutting requires re-encoding the parts that are kept, and re-encoding always costs a little quality. The tool encodes at a high quality setting to keep that loss small, and skips re-encoding entirely when there is nothing to remove. Audio exported as WAV is not re-compressed at all.
Why do I need to reanalyze after changing the settings?
Because the waveform and the cut regions on screen are the result of the previous settings, and showing them next to different values would invite you to cut using numbers that no longer apply. Rather than re-running a heavy analysis on every slider movement, the tool marks the result as out of date and waits for you to ask.
What happens to a video with no audio track?
Silence detection needs sound to measure, so a video with no audio track is reported as such and no analysis is started.
Is my original file modified?
Never. The tool only reads the file you select and writes a separate new one, named after the original with -silence-cut appended.