All tools Audio Speech to text
AUDIO Free · no sign-up In your browser · AI model

Transcribe speech to text

Drop an English recording or video, get a transcript with timestamps, and download text or subtitles — all processed on your device.

EnglishTXT · SRT · VTTFile stays on your deviceNo sign-up
Your audio is transcribed in your browser and is never uploaded to any server. The first time, an open-source Whisper model (40–75 MB) is downloaded and cached.
AI model. This tool downloads an open-source AI speech model (≈40–75 MB) to your browser the first time you use it, which uses data and memory, and it is then kept in your browser cache. Everything runs on your device and your audio is not uploaded to any server, but transcripts can contain mistakes — always read and correct them before you use them.
How it works

Transcribe in three steps

1

Add a recording

Drop an audio or video file with English speech, up to 20 minutes.

2

Pick a model

Tiny is fastest. Base is more accurate. The model is downloaded once.

3

Transcribe and export

Edit the text, then copy it or download TXT, SRT or VTT.

Why use it

Transcription that stays private

Private

Interviews, meetings and voice notes never leave your device.

Subtitle files

Get SRT and VTT with timestamps for editors and players.

Editable

Fix names and terms in the text before you export.

Loads only when used

The model is fetched when you press Transcribe, then cached.

Good to know

Getting a clean transcript

What it does

Whisper is an open-source speech recognition model. Here it runs in your browser: your audio is converted to 16 kHz mono, split into 30-second windows and turned into text with timestamps. English only.

Tiny or Base?

Tiny is the fastest and smallest, good for clear speech and quick drafts. Base is a step up in accuracy and handles accents and background noise better, at roughly twice the size and time. Neither replaces a human review for anything important.

Improve accuracy

Use a clear recording with one speaker at a time, close to the microphone. If the audio is noisy, clean it first with the Noise Reduction tool. Names, brands and technical terms are the most common errors, so check them.

Subtitles

SRT and VTT files include timestamps for each line, ready for video players and editors. Lines follow the model’s own segmentation, so you may want to re-time them in an editor for polished captions.

Limits

Recordings are limited to 20 minutes in the browser, and a phone will be much slower than a computer. Very long files are better split into parts. There is no speaker labelling.

Privacy

Your file is processed on your device and is never uploaded to any server. Only the open-source model files are downloaded the first time.

FAQ

Frequently asked questions

Is the transcription tool free?
Yes — completely free, with no account and no watermarks. Part of Mokivo.
Is my audio uploaded to a server?
No. Whisper runs in your browser. Only the open-source model files are downloaded, once.
Which languages work?
English only for now.
How accurate is it?
Tiny is quick but makes more mistakes, especially with accents, names and noise. Base is better. Always read the result before you use it.
How long can the recording be?
Up to 20 minutes in the browser. Processing runs on your device, so long files take a while, and a phone will be slower than a computer.
Can I get subtitles for a video?
Yes. Drop the video, export an SRT or VTT, and use it in your video player or editor.
More audio tools

Keep going

Every creator tool, one place

Resize, crop, convert and clean up images, audio and video — all in your browser.

Browse all tools →