Drop an audio or video file with English speech, up to 20 minutes.
Tiny is fastest. Base is more accurate. The model is downloaded once.
Edit the text, then copy it or download TXT, SRT or VTT.
Interviews, meetings and voice notes never leave your device.
Get SRT and VTT with timestamps for editors and players.
Fix names and terms in the text before you export.
The model is fetched when you press Transcribe, then cached.
Whisper is an open-source speech recognition model. Here it runs in your browser: your audio is converted to 16 kHz mono, split into 30-second windows and turned into text with timestamps. English only.
Tiny is the fastest and smallest, good for clear speech and quick drafts. Base is a step up in accuracy and handles accents and background noise better, at roughly twice the size and time. Neither replaces a human review for anything important.
Use a clear recording with one speaker at a time, close to the microphone. If the audio is noisy, clean it first with the Noise Reduction tool. Names, brands and technical terms are the most common errors, so check them.
SRT and VTT files include timestamps for each line, ready for video players and editors. Lines follow the model’s own segmentation, so you may want to re-time them in an editor for polished captions.
Recordings are limited to 20 minutes in the browser, and a phone will be much slower than a computer. Very long files are better split into parts. There is no speaker labelling.
Your file is processed on your device and is never uploaded to any server. Only the open-source model files are downloaded the first time.
Resize, crop, convert and clean up images, audio and video — all in your browser.
Browse all tools →