The recording
Nothing is uploaded: the file is decoded and listened to inside this tab. Up to 60 minutes. You can import audio or video recordings exported from Zoom, Teams, Meet or another meeting app.
Record a meeting
Choose the meeting tab and enable Share audio in the browser dialog. Only audio is saved. Transcription starts after you stop recording; this is not live captioning. For desktop meeting apps, audio sharing depends on your browser and operating system.
Use headphones when including your microphone to reduce echo. Let participants know you are recording. Recording stops automatically after 60 minutes or when sharing ends. Keep this page open.
Meeting audio capture is unavailable in this browser. Import an exported recording above.
Recognition
Model files are downloaded from Hugging Face and can be cached by your browser. Your recording stays on this device. Download sizes are approximate; loading and processing still take time.
Caption shape
These change the file, not the recognition — move them and the transcript is rebuilt at once, without listening again.
Video Transcription and Subtitle Generator
Transcribe speech locally with Whisper and export SRT, SBV, WebVTT or plain and timed text. Select the spoken language when you know it. Small is the default quality option; Base and Tiny reduce download size and processing time. Large v3 Turbo requires much more memory.
Recording a meeting
Capture the meeting's audio without leaving the page: choose the tab or window in the browser dialog and enable Share audio. The video is discarded and only the sound is kept, optionally mixed with your microphone. Recording stops when you press stop, when sharing ends, after 60 minutes, or when the recording grows too large; transcription then starts on its own. Whether the audio of a desktop meeting app can be shared at all depends on your browser and operating system — where it cannot, export the recording from the meeting app and import it above. Nothing is live: the transcript appears after the recording stops.
Accuracy and review
Recognition uses overlapping audio context and word alignment to assemble captions. Automatic timestamps and words can still be wrong, particularly with music, multiple speakers, names or numbers. Use the player and editable captions to check your recording. Assign participant names manually to each caption; names are included in exports. Automatic speaker detection is not provided.
Caption shape changes reformat existing text. Timing takes priority when a minimum display time would overlap the next caption. Notices flag fast reading or captions outside your preferred shape. Manual edits remain in the transcript when you change the shape.
Processing and export
Recordings up to 60 minutes are decoded in bounded blocks. Recognition runs in a worker; Stop cancels it and keeps completed captions. Model downloads use the network, but your audio is not uploaded. Speed depends on your processor and model; long recordings can take a substantial time.
Use SRT or SBV for caption uploads, WebVTT for web players, or text for a readable transcript. Downloaded files include your edits. Review the result before publishing.