On this device · No upload · No account

Audio to Text Online

Convert a local audio file to text in this browser. The English speech model runs on your device and the file is not uploaded.

Drop an audio file or click to select

MP3, WAV, M4A, OGG, FLAC, or WebM. One file. Nothing is uploaded. A computer can take about 200 MB or 20 minutes; phones keep a shorter clip.

Audio to text

Speech is converted on this device. A computer can prepare a larger model in the background after the page appears. Phones use a lighter model. The speech pass runs in this tab through WASM.

  • Processed on this device
  • Audio is not uploaded
  • No account

On this page

Transcript

Features

What this audio to text page actually does

The homepage is the working tool. You convert one local file to text, see a progress bar while the speech model caches, then copy or download the transcript without leaving the tab.

One file, one Transcribe button

Drop an MP3, WAV, M4A, or similar clip onto the card and press Transcribe. The transcript stays on this page so you can read it, copy it, or download a txt without opening another tab.

On-device Whisper

The English speech model runs in this tab through Transformers.js. WebGPU is used when the browser exposes it; Safari and iPhone fall back to WASM and just take longer on the first visit.

No upload and no account

Audio stays in this browser. Nothing is posted to our server, you do not create a login, and the transcript is not stored after you close the tab.

Copy or download txt

After the pass finishes you can copy the text, download a .txt file, or open View for a larger dialog. Speaker labels and subtitle files are not offered on this page.

A compact audio file dropping onto a browser card and becoming a page of written lines

What it is

What is Audio to Text Online?

Audio to Text Online is a browser page that turns one local recording into written lines with an on-device English speech model. The model files come from our own host and run in this tab through Transformers.js. WebGPU is used when the browser exposes it; Safari and iPhone fall back to WASM. The file is not posted to our server and is not written to a log we keep after you leave.

This first English homepage is only for a clip you already have. It does not fetch a remote video URL, label speakers, export SRT, summarize a meeting, or take dictation from the microphone. If those jobs are what you need, this page is the wrong tool — we say so instead of pretending a queue exists behind the button.

Engine
On-device English speech in the browser
Upload
None — the file stays local
Account
Not required
Fallback
WASM when WebGPU is missing

How to use

How to use this tool

Three on-page steps: add a file, press Transcribe, then copy or download the txt. Step 1 is choosing the clip. There is no paste-a-link path on this first English homepage.

Step 1

Choose a local audio file

Drop one clip or click the card. Keep it within the limit shown on the card so the tab can finish. There is no paste-a-link field on this first English homepage.

A hand choosing one compact audio clip on a simple drop card inside a browser

Step 2

Press Transcribe

The first visit downloads a small English model and caches it. A progress bar shows that wait. Later files in the same browser skip the download unless you cleared site data.

A laptop running a speech model with a progress bar while a transcript fills the page

Step 3

Copy or download the txt

Confirm the filename, read the transcript here, then copy it or save a text file. Nothing is stored after you leave.

A text sheet leaving a browser window toward a download tray and a clipboard

Why this page

Why convert audio on this device

Some converters upload the recording to a cluster. This page keeps the file in the tab so a voice memo or interview clip does not become a hosted copy we can read later. The increment is privacy and honesty, not a longer feature list.

Private by default

Some converters upload the recording to a cluster. This page keeps the file in the tab so a voice memo or interview clip does not become a hosted copy we can read later.

Works without WebGPU

Safari and iPhone still run the same English model through WASM. The first visit is slower because the model downloads; later files reuse the cache in this browser.

Honest limits

One file, the limit shown on the card, English on this device. We do not claim realtime captions, speaker labels, SRT export, or server-grade accuracy on this homepage. Those extras belong to hosted tools, not this tab.

One job, one button

The homepage is the working tool. Extra modes we do not ship stay off the page instead of appearing as grey tiles you cannot click.

Tips

Tips and limits

Practical notes so the first model download and the transcript stay predictable. These are limits, not marketing bullets: phones use a lighter model, the card shows the size cap, and a loud room will still read poorly.

Who uses this

Who uses this in a real job

People who already have a recording and need text they can edit. Each card is a real job with an input, a setting, an output, and a caveat — not a slogan. This page is only for that local file, not a remote video URL or a meeting bot.

Phone memo into editable notes

You already recorded a thought on your phone and need lines you can paste into a note app without sending the clip to a host.

Input: One phone memo saved as M4A or MP3, 25 MB or smaller, already on the device.

Setting: Default Transcribe. English Whisper-tiny. No extra language pack and no speaker mode.

Output: A plain transcript on this page plus a txt you can paste into notes.

Caveat: Quiet speech works. Music beds and overlapping talk will come out messy.

Short interview first pass

You recorded a conversation yourself and want a first-pass transcript so you can clean names and quotes by hand.

Input: An interview you recorded as WAV or MP3, one file, under the 25 MB cap.

Setting: Press Transcribe once. WebGPU if the laptop has it; WASM on Safari.

Output: One block of English text on this page. Copy or download txt. No speaker columns.

Caveat: This page does not label speakers. You split turns yourself after the pass.

Lecture excerpt already downloaded

You already have a class clip on disk and want searchable words in the same tab, not a bot that joins the call.

Input: A lecture excerpt already saved locally as MP3, WAV, or M4A.

Setting: 25 MB cap. English-only model. No remote course URL field.

Output: Searchable text in this tab that you can copy into study notes.

Caveat: We do not fetch a course page or split a remote video. Save the audio first.

FAQ

Questions about audio to text

How do I convert audio to text?

Use File, Link, or Record. Drop a local clip, import a direct audio or video file URL, or record in this tab. The speech model runs on this device and the transcript appears on the same page so you can copy it or save a txt file.

Does audio stay on my device?

Yes. The browser reads the file or recording in this tab and feeds it only to the in-browser Whisper model. AudioToText.im does not post the audio to our server or keep a copy after you leave.

Do I need an account?

No login and no signup. Add audio and transcribe. The first visit downloads an on-device speech model and caches it in this browser so later clips skip that wait.

What audio formats can I use?

MP3, WAV, M4A, AAC, OGG, FLAC, and WebM work when the browser can decode them. Keep the clip within the limit shown on the card. Link imports a direct file URL. YouTube and other page links are not pulled here.

Does this work on iPhone or Safari?

Yes. Phones use a lighter on-device model and automatically write the language you spoke. There is no Auto / English faster switch on a phone. The first model download can take a minute; later files reuse the cache. A computer can run a stronger pass if the phone transcript is not clean enough.

Is there a file size limit?

One file. A computer can take about 200 MB or 20 minutes — that pair is roughly a 20-minute uncompressed WAV. Phones stay around 80 MB or 10 minutes so the tab is less likely to lock up. Longer jobs belong to a hosted API, not this page.

Is the model downloaded every time?

The first visit downloads an on-device speech model from our own files and caches it in this browser. A computer can start that download after the page appears. Later files skip the wait unless you cleared site data. Switching Auto and English faster on a computer can fetch a second model once.

Can I download the transcript?

After Transcribe finishes, Download TXT saves a .txt file and Copy puts the same text on the clipboard. View opens a dialog on this page. If you recorded in this tab, you can also download the audio clip. Nothing is stored after you close the tab.

What languages does this page handle?

Auto recognizes the language you spoke and writes matching text — speak French, get French. Coverage is about 99 languages from the official Whisper set. We do not claim 100+ languages, and this page does not translate speech into English.

Does it detect my language automatically?

Yes on Auto. The tab listens to the clip, picks the spoken language, and writes that language. A computer remembers Auto or English faster on this device. Phones always use Auto with the lighter model.

What is Auto vs English faster?

On a computer, Auto is the multilingual pass. English faster is only for English speech; it is a smaller download and usually quicker. File, Link, and Record all use the choice you last picked. Phones hide this switch.

Can I record audio in the browser?

Yes. Open Record, allow the microphone, then stop when you are done. The clip stays in this tab until you download it. After you stop, the same tab transcribes it. Mic permission is required.

How accurate is the transcript?

It is an on-device speech pass, not a studio caption desk. Quiet speech works better than music or overlapping talk. Auto keeps the spoken language; English faster is only for English. We do not claim server-grade or realtime accuracy.

Can I paste a YouTube or WhatsApp link?

No. Link only imports a direct audio or video file URL. It does not fetch a YouTube page, a chat export, or a cloud album. Save the audio first, then use File — or pick Dropbox if you already have the file there.

Can I use this offline after the first visit?

After the model is cached, the tab can transcribe without sending audio out. You still need the site files themselves if the tab is fresh or the cache was cleared.

Do I need an app or browser extension?

No install and no extension. Open this page in a current browser, add a file, a direct link, or a recording, and transcribe. Inputs the decoder cannot read are skipped rather than guessed.

Does this label speakers or export SRT?

Speaker labels, SRT or VTT, summaries, and mind maps are not on this page. You get one plain transcript you can copy or download as txt.

Can I transcribe a video file?

If the browser can decode the soundtrack, a local video file or a direct video-file URL can work. We do not fetch or split a YouTube page or other site that is not a file.

Does this work inside CapCut or Canva?

Those editors have their own caption tools. AudioToText.im is a separate browser page: add audio here, then paste the txt wherever you already edit.

Back to the tool

Drop a file in the first viewport and press Transcribe. This link only scrolls there.

Convert audio to text