Runs on your device · WebGPU
Turn video into text
Add a video or audio file of up to 20 minutes and get a written transcript, with or without timestamps. The speech is recognised by Whisper running in your browser, so the recording never leaves this page. Copy the text, or download it as TXT for notes or SRT and VTT for subtitles.
Downloads once: Whisper tiny (42–117 MB) or base (76–199 MB); browsers with WebGPU fetch the larger files
How it works
1Add the recording
Videos (MP4, MOV, WebM) and audio (MP3, M4A, WAV and more) up to 20 minutes. Only the sound is used, decoded by your browser and turned into 16 kHz mono, which is what Whisper listens to.
2Whisper transcribes it
Long recordings are cut into pieces of up to 28 seconds that start at a pause where they can and overlap by at least 2 seconds. Each gets a time for every word, and the pieces are joined at a pause so each word is kept once.
3Copy or save
Read it as paragraphs, or switch on timestamps for one line per sentence. Copy it, or save a TXT, an SRT or a VTT file named after your recording.
What a transcript is good for
- Writing the caption for a Reel from what you actually said in it, instead of retyping it from memory.
- Pulling quotes from a live stream you recorded, a podcast clip or an interview to post as text or carousel slides.
- Checking the exact wording of a sponsor’s script, a giveaway’s rules or a price you mentioned on camera.
- Making notes from a long voice note or a recorded call, and finding the minute where something was said.
- Getting an SRT file for a video ad: Meta Ads Manager lets you add captions to a Facebook or Instagram video ad by uploading one.
Switch on timestamps when you need to find your way back into the recording: each line starts with the time it was said, such as [4:32]. Leave them off when you want clean paragraphs to paste into a caption, an email or a document.
How accurate it is
Whisper is an open speech recognition model from OpenAI, released under the MIT licence. This page uses its two smallest multilingual versions: tiny (39 million parameters) as Fast and base (74 million) as Better. On a clear voice recorded close to the microphone, both get most words right: in our check with a 10-minute computer-generated English recording, Fast differed from the script on about 7 words in 100, mostly spelling variants such as color for colour, numbers and technical terms. Accuracy drops with music under the voice, several people talking at once, a phone held at arm’s length in a noisy street, and words the model rarely heard, like brand names and usernames.
| Speech model | Fast | Better |
|---|---|---|
| Model | Whisper tiny | Whisper base |
| One-time download | 42 MB, or 117 MB with WebGPU | 76 MB, or 199 MB with WebGPU |
| Speed | Quickest | Slower: a bigger model to run |
| Choose it for | Clear speech, quick drafts, phones | Accents, background noise, languages other than English |
Whisper knows 99 languages, and OpenAI’s model card says it performs unevenly across them, with lower accuracy where it had less training data. Auto-detect picks the language from the first 28 seconds with speech; if a recording mixes languages, or the detection is wrong, choose the language yourself and transcribe again. Hindi is where these small models struggled most in our check with computer-generated voices: Fast turned it into rough English, and Better managed a few words in Urdu and Latin letters before getting stuck. In Spanish, Better got all 28 words of a short sentence right and Fast got about a third wrong.
Tip: Numbers are the easiest thing to get wrong in a transcript. Before you post a price, a date or a discount code taken from one, play that moment back and check it.
Files, limits and privacy
- TXT is the transcript as you see it, with or without timestamps. SRT and VTT hold numbered subtitle lines of up to two rows of about 42 characters and no more than 6 seconds each.
- Files can be up to 20 minutes long, 600 MB for video and 100 MB for audio. The whole file is read into your browser’s memory to get at the sound, so on a phone keep videos to a few hundred MB.
- Speech recognition runs on your graphics chip through WebGPU where the browser offers it, and on the processor otherwise. If WebGPU gives output that can’t be speech, the tool redoes the job on the processor rather than show it.
- Your recording stays on your device. The speech model is downloaded from Hugging Face the first time and kept by your browser, so later transcripts start straight away.
Questions people ask
Can it transcribe a Reel from a link?
No, it works on files, not links. Use the video file on your phone or computer, such as the version you edited before posting it.
How long can the recording be?
Up to 20 minutes. Split a longer recording into parts in any video or audio editor and transcribe each one. On a recent Mac (an Apple M4 Pro), Fast transcribed a 10-minute recording in about 27 seconds with WebGPU and about 46 seconds without it, once the model was downloaded. Phones and older computers take longer.
Is my recording uploaded anywhere?
No. The file is read by your browser and the speech model runs on your own device. Nothing from the recording or the transcript is sent to DMFast or to anyone else.
Why does the transcript say the wrong language?
Auto-detect guesses from the first 28 seconds with speech, and short or mixed-language clips can fool it. Open “Transcribe again”, choose the language from the list and run it again.
What is the difference between SRT and VTT?
Both are subtitle files: numbered lines, each with a start and end time. SRT writes times with a comma before the milliseconds and is the format Meta Ads Manager takes for video ad captions, under a file name such as my_video.en_US.srt. VTT uses a dot, starts with the word WEBVTT, and is the format web video players read.
More free tools
- Video caption generatorAuto captions for Reels: transcribe in your browser, fix the words, then burn them in or save an SRT.
- Image to textGet the words out of a screenshot or photo as editable text, read on your device.
- Caption generatorYour photo matched to caption topics, with ready-to-copy captions for each, on your device.
- Alt text generatorA short alt text and a longer description for each photo, written on your device.
- Character counterCount caption, bio, DM and Facebook ad text against limits that say where they come from.
- Auto reframeLandscape to 9:16, 4:5 or 1:1, with the crop following the person on screen.
Related guides
Answer every comment and DM, automatically
When someone comments a word like GUIDE, DMFast replies and sends them your link in a DM. Free plan, no card.
Start free