Drop an MP3 above and it transcribes into editable, timestamped text on your own device. No upload, no minute counter, no account.
Search "MP3 to text" and every result hands you the same deal: create an account, upload your file to their servers, wait, and then watch a free-minute meter count down until it asks for a card. This page is the opposite arrangement. The MP3 is read straight off your disk and the words are produced inside this browser tab, and there is no server on the other end metering minutes, because there is no server involved at all. That single structural difference is why the transcription here is genuinely unlimited: nothing is counting your audio, so nothing can run out.
The practical upshot is that a four-minute jingle and a four-hour recorded conference cost exactly the same to transcribe, which is to say nothing, and neither one is queued behind a stranger's upload pipeline. You are only ever limited by your own machine's patience.
MP3 throws away data to shrink the file (it is a lossy format), which understandably worries people who assume a smaller file means a worse transcript. In practice speech is astonishingly robust to that compression. The parts of the sound MP3 discards are mostly the very high and very low frequencies your ears barely register, while the mid-range that carries vowels and consonants is exactly what the encoder tries hardest to preserve. So an old dictaphone export saved at 64 kbps, a podcast archived at 96 kbps, or a radio rip at 128 kbps all transcribe about as well as a pristine studio file. What actually degrades a transcript is the recording itself: a phone left across the room, two people talking over each other, or a fan roaring next to the microphone.
One MP3-specific behaviour worth knowing: long silent stretches are skipped automatically rather than fed to the recogniser. Pushing pure silence through speech recognition is a known way to make it hallucinate words that were never spoken, so those gaps are detected and passed over. That is a claim about silence only: a heavy music bed under someone talking is a different matter and can still confuse recognition, so a track that is mostly song with occasional narration is not this tool's strong suit.
Whether your MP3 is mono or joint-stereo makes no difference to you either. Stereo is folded down to a single channel before recognition, so a two-channel interview where each person sits in one ear is merged into one clean transcript rather than being processed twice.
MP3 is the format of things that were meant to be listened to and kept. Published podcast episodes are almost always MP3: a feed full of them is a natural transcript project, whether you are turning your own show into show-notes or mining a favourite series for quotes. Downloaded lecture and sermon archives arrive as MP3 too, often as a folder of dozens of files that you can queue all at once. Handheld voice recorders (the little devices journalists and students still carry) export MP3 by default, as do the interview libraries that pile up on a reporter's drive over the years.
Take one concrete case: a ninety-minute podcast episode you want a searchable transcript of. You drop the single MP3 in, the transcript starts streaming almost immediately, and while the back half of the episode is still being read you can already select and copy the opening segment. When it finishes you copy the whole thing to your clipboard, timestamps and all, and paste it wherever your show-notes live. Nothing about that episode ever left the laptop.
.mp3 files onto the drop zone (or click to pick them). Select a whole folder of episodes if you like; they queue up and transcribe one after another.Working with a different file type? Use the full audio transcriber, or jump to a format-specific guide: