Troubleshooting slow, failed, or inaccurate transcription
Most problems fall into one of four stages: loading the model, opening the media, processing the audio, or reviewing the output. Start with the symptom you can see.

The model will not download
- Confirm that the browser is online and refresh once.
- Try a current version of Chrome or Edge.
- Temporarily disable a content blocker for CaptionKite if it blocks the model host.
- Make sure the browser has free storage for the cached model.
- Try a normal window if private-browsing storage rules interrupt the download.
The first run is the slowest. Once the model is cached, later visits should avoid most of that setup.
The file cannot be decoded
First confirm the media plays outside CaptionKite. Then try another browser. If it still fails, convert a copy to a common format or extract the audio. A familiar extension does not guarantee a familiar codec; see supported media formats for the explanation.
Transcription is very slow or stops
- Connect the computer to power and close heavy apps or tabs.
- Keep CaptionKite visible; some browsers throttle background tabs.
- Use extracted audio instead of a high-resolution video.
- Split a long recording into smaller sections.
- Restart the browser if the device has been under memory pressure.
Speed depends on audio duration, your processor or WebGPU support, and the amount of available memory—not merely the video's file size.
The transcript is inaccurate
Check the selected language first. Then listen for low volume, echo, music, cross-talk, strong accents, or specialized names. Automatic transcription needs human review, particularly for legal, medical, financial, or safety-critical material.
- Use the cleanest audio source available.
- Correct names, numbers, acronyms, and technical terms.
- Compare difficult passages at a slower playback speed.
- Do not treat an automatic transcript as a certified record.
A pasted URL is rejected
Open the URL in a new tab. If it shows a web page, login, or streaming player instead of the media file, it is not a direct media URL. YouTube and protected Vimeo or Zoom pages cannot be imported this way. Cross-origin restrictions can also prevent a browser from fetching an otherwise public file.
Download media you are authorized to use and upload the local copy. The planned extension will support permitted playback in an open tab without bypassing access controls.
Still stuck?
Use the contact form and include your browser, device type, file format, approximate duration, and the exact status message. Do not attach confidential media or paste private access links. Those details are usually enough to reproduce a bug safely.
Start with a file you know works
When something fails, the first useful move is to find out whether the problem is your file or your setup. Those need completely different fixes and it is easy to spend an hour on the wrong one.
Take a short, ordinary MP4 — thirty seconds of clearly recorded speech — and run it through. If that works, your browser and machine are fine and the original file is the problem: a codec it cannot decode, a length it cannot hold, or audio it cannot hear. If even the short file fails, the file is not the issue and you should be looking at the browser.
It costs a minute and it eliminates half the possibilities, which is more than any amount of re-trying the file that failed.
Is it stuck, or is it working?
Transcription of a long recording takes a long time, and a progress bar that has not moved is not automatically a frozen page.
Watch the status text rather than the bar. It should name what is happening — reading the file, downloading the model, transcribing with a time counter that advances. On a long recording those steps can each take minutes, and the transcript now streams in as it is produced, so words appearing is proof the work is genuinely progressing.
Real signs of trouble are different: the status stops changing for many minutes, the tab stops responding entirely, or an error appears. Waiting is uncomfortable but usually correct — cancelling a job at eighty per cent and restarting it means paying the whole cost again.
Bad recognition or bad audio?
An inaccurate transcript has two very different causes, and the fix depends on which you have.
Listen to the recording at the points where the text is wrong. If you can hear the word clearly and the transcript still missed it, that is a recognition limit — an unusual name, an accent, technical jargon — and the answer is correcting it afterwards, or using the larger model where accuracy matters more than speed.
If you listen and cannot make it out yourself, no model will do better. Background music under speech, two voices at once, and echo in a hard-walled room are the usual culprits, and the only real fix is a better recording or a better source track. Re-running the same audio will produce the same result.
One more cause worth ruling out first: the wrong language selected. It produces confidently wrong text that reads as gibberish rather than as errors.
Reporting a problem without sending us your video
We cannot ask for your file, and you should not send it. So a useful report describes the conditions rather than the content.
- Which browser and version, and whether it is desktop or mobile.
- The container and codecs if you know them, plus roughly how long the recording is.
- What the status line said at the moment it failed, word for word.
- Whether the same thing happens with a short test file.
That last point is the one that makes a report actionable, because it tells us whether we are looking for a bug in the tool or an unusual property of one file. If you can reproduce it with a video you are happy to share — a clip of something public — that is the most useful thing you can send.