How private browser transcription works
CaptionKite performs speech recognition inside your browser. That changes where your media travels, what gets downloaded, and which device resources do the work.

What happens after you choose a file
The video or audio is not uploaded to CaptionKite. Temporary audio data and the generated transcript exist in the page while you work.
What still uses the internet
The website itself must load, and the speech-recognition model is downloaded from its model host on first use. Your browser normally caches that model for later visits. If you paste a direct media URL, your browser also requests that file from the website hosting it.
Those network requests are different from uploading a local media file to CaptionKite. The local file remains on your device.
The privacy and cost benefits
- Private meetings and drafts do not need to be stored on a transcription server.
- No account is required to connect a transcript to your identity.
- Processing does not consume paid server transcription minutes.
- The cached model can make later sessions quicker to start.
The tradeoffs of local processing
Your computer supplies the memory and computing power. Older devices may run slowly, long audio can increase memory pressure, and browser media support varies. Closing the tab interrupts the job, and clearing browser storage may require the model to download again.
Use a local file rather than a remote URL, close unrelated tabs, review your organization's data policy, and remove downloaded captions from shared computers when finished.
Analytics and advertising are separate
CaptionKite may use privacy-conscious traffic measurement and clearly identified advertising to support the free tool. Those services can load on the page, but the transcription pipeline does not send them the contents of your selected local media or generated transcript. See the privacy policy for the current details.