Choose a file
Select a video or audio file from your device. It never leaves your browser.
Upload a file or paste a direct media URL, then generate SRT, VTT, and text transcripts in your browser. Your media never passes through our servers.
The AI model is downloaded from Hugging Face on first use and cached by your browser. Processing speed depends on your device.
YouTube and protected Vimeo or Zoom links cannot be imported directly. Our companion browser extension will capture the audio only after you open the video and start it yourself, then generate captions on your device. It will not bypass sign-ins, passwords, or viewing permissions.
Select a video or audio file from your device. It never leaves your browser.
A compact Whisper model transcribes the audio locally on your computer.
Save the result as SRT, WebVTT, or plain text—free and without an account.
Most transcription sites upload your media to a remote server. This tool uses your browser's WebGPU or WebAssembly support instead. That protects private recordings and lets us offer the core tool without usage credits.
Yes. There is no account, watermark, or usage credit. Advertising supports the site.
No. Decoding and transcription happen inside your browser. We do not receive or store your media.
Your browser must download the speech-recognition model once. It is normally cached for later visits.
MP4, WebM, MP3, WAV, M4A, and other formats your browser can decode, up to 500 MB. Shorter files are faster and use less memory. Chrome or Edge provides the best compatibility.
Yes. The tool processes direct media links and authorized Vimeo or Zoom download links. YouTube, sign-in pages, and links whose owners disabled access are routed to the companion browser extension.