Drop a video and the sound track is read straight out of it on your device, with no separate audio step, no upload, and subtitles included.
Almost every "how to transcribe a video" guide opens with the same chore: first extract the audio into a separate file, then feed that to a transcriber. Skip it entirely. Drop your MP4 or MOV straight onto this page and the app opens the video, finds the sound track inside it, and transcribes that. The picture is never even looked at, let alone needed. And because the whole thing runs on your own machine, none of the video is uploaded: not the audio, and certainly not the frames, which for video is often the more sensitive half.
That local-only design is worth dwelling on for a second, because video carries more than a voice. Footage shows faces, whiteboards, screens full of unreleased work, the inside of someone's home. Sending all of that to a transcription server just to get the words out is a poor trade. Here the trade does not exist: only you ever see the file.
The supported formats are MP4 and MOV: between them, that covers what phones shoot, what screen recorders save, and what a local Zoom or Teams recording writes to disk. Inside those files the sound is almost always AAC, which streams in at any size, so a long recording is no obstacle. There are three honest edges worth stating. First, a video with no sound at all is told apart from a genuine failure: instead of a confusing error you get a plain "No audio was found in this file", so you know the problem is the file, not the tool. Second, a copy-protected video will fail with an error rather than a transcript; it stops cleanly instead of stalling, and it simply cannot be read. Third, WebM is not supported yet; a WebM download needs converting to MP4 first before it will go in.
One quieter boundary sits under the marketing: on the rare video whose audio uses a codec the browser cannot decode, the file falls back to a whole-file rescue path that only works below a couple hundred megabytes. Standard AAC audio (the overwhelming majority of MP4 and MOV files) never touches that path and streams at any size.
The most common case is a meeting you recorded locally, a Zoom or Teams call saved as an MP4 on your own drive, which you turn into written minutes without that recording ever going back to the cloud. Close behind are lecture and webinar recordings, where a long session becomes a searchable transcript, and screen-recorded product demos that need captions before they go public. Interview footage rounds it out: a videographer's raw clips turned into a text the editor can scan.
What ties these together, and sets this page apart from its siblings, is that the timestamps are usually the deliverable, not a side effect. A transcript of a video is very often really a subtitle file in waiting, which is the whole next section.
On an audio-only file the timestamps are a bonus; on a video they are usually the reason you came. Because the transcript is timed at the word level, it exports as an SRT or VTT caption file that drops directly into YouTube, Premiere Pro, or DaVinci Resolve. The caption files follow the conventions professional captioners use by default (a capped line length, at most two lines per cue, and a reading speed slow enough that a cue never flashes past faster than a viewer can read it), so you are not hand-tidying the output before it is usable.
Captioning a demo, end to end:
.mov product demo in.Working with a different file type? Use the full audio transcriber, or jump to a format-specific guide: