Free AI subtitles vs Descript/CapCut: what's actually free
June 17, 2026
Search "free AI subtitles" and you'll land on tools that are free the way a free trial is free — free to start, free until you hit a wall. This isn't a takedown of Descript or CapCut, both are genuinely capable editing tools with subtitle generation as one feature among many. It's a straight look at what "free" actually means on each, so you can pick based on what you actually need instead of what a pricing page implies.
How the free tiers typically work
Descript and CapCut are full video/audio editors first, with AI transcription and captioning built in as one feature of a much larger toolset. Their free tiers are structured the way most freemium editing software is: enough to try the feature and see that it works, with a limit designed to convert regular users to a paid plan — commonly a monthly cap on transcription minutes, watermarks or resolution limits on exports, or editing features gated behind a subscription. None of that is unreasonable for tools building a full editing suite with a real cost structure behind AI transcription running on their servers.
That server-side cost is exactly the structural difference: WavyVid's subtitle generator runs the AI transcription model (OpenAI's Whisper) directly in your browser via WebAssembly/WebGPU, instead of sending your audio to a cloud API metered by the minute. There's no per-minute cost on our end because there's no server-side inference happening at all — which is also why there's no per-minute cap on your end.
What you give up running it in-browser
To be specific rather than just favorable to ourselves: running transcription locally means it's bound by your device's own processing power, not a data center's. On a modern laptop this is fast — often close to real-time. On an older or lower-powered device, it's slower, and the first use requires downloading the speech model (cached afterward, so it's a one-time cost per device). Cloud-based tools don't have this constraint since the heavy lifting happens on their servers regardless of your device. If you're transcribing on a low-end machine constantly, that's a real tradeoff worth knowing about upfront, not glossing over.
The other honest limitation: WavyVid is a focused set of single-purpose tools, not a full non-linear video editor. If you need multi-track editing, transitions, a full timeline — Descript and CapCut are built for that in a way a browser-based single-purpose tool isn't trying to compete with.
What you actually get free, side by side
- Transcription limits: WavyVid — none, since it's not metered by a server. Descript and CapCut — a free tier with real limits, typically a monthly transcription cap or feature restrictions that push regular users toward a paid plan.
- Export format: WavyVid exports standard .srt and .vtt files plus burned-in captions, no watermark. Free tiers on editing suites vary — check current terms before assuming an export is unrestricted.
- Account required: WavyVid — no. Descript and CapCut — yes, both require an account to use transcription features at all.
- Where processing happens: WavyVid — entirely on your device. Descript and CapCut — uploaded to their servers for processing, which is standard for cloud-based AI tools but worth knowing if privacy or upload time matters to your workflow.
Accuracy: the part nobody advertises honestly
All three tools (WavyVid, Descript, and CapCut) lean on modern speech-recognition models, and accuracy on clear, single-speaker audio in a quiet room is generally strong across the board — this isn't really a differentiator between them. Where accuracy drops for all of them is the same: heavy accents, overlapping speakers talking at once, background noise, and fast or mumbled speech. None of them get this perfect, which is exactly why an editable transcript step before export matters more than which specific model is doing the transcribing — you're going to proofread regardless, and the tool that makes that proofreading step fast and frictionless is doing the more useful job.
If your source audio is genuinely difficult — multiple overlapping speakers, heavy background noise, a strong regional accent — expect to spend real time correcting any AI-generated transcript, on any of these platforms. Running noisy audio through a cleanup pass first measurably improves transcription accuracy downstream, since the model has an easier time separating speech from noise when there's less noise to begin with.
Which one should you actually use
If you need captions and a transcript for a video — clips, interviews, talking-head content, repurposing long-form into short-form — and don't need a full editing suite around it, a focused, unlimited, no-account tool is the more direct path. If you're already living inside Descript or CapCut for the rest of your editing workflow, their built-in captioning is a reasonable default purely for convenience, with the tradeoffs above being the actual cost of that convenience.
Either way, know what you're actually trading before you hit a paywall mid-project — that's the whole point of writing this out plainly instead of just saying "free" and leaving the fine print for later.
Frequently asked questions
Is WavyVid's subtitle generator really unlimited and free?
Yes — there's no per-minute cap and no subscription, because transcription runs on your own device instead of a metered cloud API.
What do Descript and CapCut actually charge for?
Both offer a free tier with real limits — typically a monthly transcription-minute cap or export watermarks/restrictions — that push regular users toward a paid plan.