Audio to Text
Convert audio or video to text online - free, private, no upload, no sign-up.
Turn a recording into text without installing anything or sending your file anywhere. This audio to text tool runs Whisper speech recognition locally in your browser by default - free, with no upload and no sign-up - and gives you an editable transcript with timestamps you can export as plain text or subtitles.
Drop an audio or video file here, or click to choose one
MP3, WAV, M4A, MP4, MOV and more - up to 30 minutes on the free local engine
Recording... 0:00
The free engine never uploads your file. Recording and transcription both run in this browser tab.
- File
- -
- Duration
- -
- Size
- -
- Type
- -
This is a video file - only its audio track will be transcribed.
Speaker labels (who said what) aren't supported yet - the transcript comes back as one continuous stream of text, split into timed segments.
Downloading the model...
Local transcription runs at roughly real-time speed on a modern laptop - a 10-minute recording takes about 10 minutes. Older machines and phones are slower.
Why use this transcriber
Runs in your browser
The free option processes your audio locally, in a Web Worker. Nothing is uploaded, and there's no server in the loop.
Subtitles included
Export as plain text, Markdown, or SRT and VTT subtitle files with timestamps, ready to drop into a video editor.
Edit before you export
Fix names and technical terms directly in the transcript, segment by segment, instead of cleaning up a finished export.
How to convert audio to text
- Drop an audio or video file above, or click Record audio to capture something directly in your browser.
- Choose a model size and language, or leave both on the defaults.
- Click Transcribe. The first time you do this, your browser downloads the speech model - this is a one-time download, cached afterwards.
- Watch the progress bar while your file is transcribed, in your browser, on your device.
- Read the transcript, click any timestamp to jump to that part of the audio, and fix any names or terms directly in the text.
- Export as .txt, .srt, .vtt or .md, or copy the text straight to your clipboard.
Common audio to text problems and fixes
- The first run is slow: that's the one-time model download, not a bug. It's cached afterwards, so every later run on this device starts immediately.
- Accuracy is poor on accents, crosstalk or background noise: this is a known limit of speech models generally, not just this one. Try the Best quality model size, or edit the errors directly in the transcript.
- Names and technical terms come out wrong: click into the transcript and fix them in place, or wait for the high-accuracy server option, which handles jargon better.
- Long files fail or crash the tab: split the recording into parts under the 30-minute cap and transcribe each one separately.
- No text comes back at all: check the recording actually has sound - a silent track, the wrong audio track in a video file, or a muted microphone during recording are the usual causes. The tool reports "no speech detected" rather than returning an empty box.
- Timestamps drift in the exported SRT: this can happen if the source file's own duration metadata is wrong (common in some screen recordings). Re-export the source file with a standard tool and try again.
- It won't run in this browser: the local engine needs WebAssembly, and runs faster with WebGPU. Update to the latest Chrome, Edge, Firefox or Safari.
- Transcribing a meeting recording specifically: if you're recording meetings regularly, doing this one file at a time isn't the long-term answer - see how Meetrix handles recording and notes together below.
Check your microphone before recording audio directly on this page. Have a large video file to send instead of a transcript? Try the video compressor - same local-first approach, nothing uploaded. Want timed captions on a video instead of a plain transcript? Use the subtitle generator, which previews them live over your video. Need to fix the timing on the SRT this exports, or turn it into VTT? Open it in the subtitle editor next. Once you have a transcript, paste it into AI meeting notes to get a summary, decisions and action items out of it automatically. Browse all Meetrix features for the rest of the pre-call tools, or see the background on how Meetrix records meetings in the first place on the Meetrix blog.
Frequently Asked Questions
Is this audio to text converter free?
Yes. The local engine is free to use, with no sign-up, no watermark, and no limit on the number of files - just the per-file length cap listed above.
Is my audio uploaded anywhere?
Not on the free engine. Your file is decoded and transcribed entirely in your browser, in a background worker, and never leaves your device. The only path that uploads anything is the optional high-accuracy server engine, which discloses the upload before it runs and is not live yet.
What's the difference between the free and high-accuracy options?
The free option runs a smaller Whisper model locally in your browser: no upload, no cost, but slower and somewhat weaker on accents, crosstalk and noisy audio. The high-accuracy option (coming soon) sends your file to Meetrix's servers for a larger model, is faster and more accurate, but requires an upload and an email address.
How long can my file be?
Up to 30 minutes on the free local engine - that's the point at which a browser tab reliably starts running low on memory for this kind of processing. Longer files: split them into parts and transcribe each one.
Which languages are supported?
Whisper supports dozens of languages. Leave the language on Auto-detect, or pick one explicitly from the dropdown if auto-detection guesses wrong on a short or noisy clip.
Can it tell speakers apart?
Not yet. Speaker labels aren't in this version - the transcript comes back as one continuous stream of text, split into timed segments rather than by who's talking.
What formats can I export?
Plain text (.txt), subtitle files (.srt and .vtt) with timestamps, Markdown (.md), or copy the text straight to your clipboard.
Which browsers are supported?
Any recent Chrome, Edge, Firefox or Safari. Chrome and Edge can use your GPU (WebGPU) and run faster; other browsers fall back to WebAssembly, which works everywhere but is slower.
Does this work on a phone?
Yes, but expect it to be slow. Both the model download and the transcription itself take longer on a phone than on a laptop - it works, it just needs patience.
Why is the first run slow?
The first time you transcribe anything, your browser downloads the speech model (40 MB to a few hundred MB, depending on the size you pick). That download is cached, so every run after that starts immediately.
How accurate is it?
On clear, single-speaker audio it does well. Accuracy drops on strong accents, overlapping speakers, background noise, and technical jargon or names - that's true of any speech model, not just this one. Edit the transcript in place to fix what it gets wrong, or try the high-accuracy option once it ships.
Can I edit the transcript before exporting?
Yes. Click into any segment and type - your edit is saved in the transcript immediately, and every export (copy, .txt, .srt, .vtt, .md) uses the edited text.
Transcribing one file at a time? There's a faster way
If you're transcribing meeting recordings regularly, Meetrix can record, transcribe and summarize them automatically on your own infrastructure - see Meetrix's self-hosted AWS products, including the recording and transcription AMIs, to run this at scale instead of one file at a time.
Talk to Us