Auto captions, no watermark

Whisper transcribes in your browser and the captions burn in on your machine. No upload, no account, no watermark. Click any line to fix a word before burning in.

Drop a video here or click to choose — it never leaves your browser

Long videos or 4K? The CoAnimator desktop app does styled, animated captions as part of full video projects.

how it works

How to add captions to a video

Drop your clip above, let the speech-recognition model transcribe it right in your browser, fix any words it misheard, pick a lower-third style, and export with the captions burned in. Free, no upload, no account, no watermark.

  1. Drop the video — transcription runs locally in your browser.
  2. Edit the caption lines; timing stays attached to each line.
  3. Choose one of the lower-third styles.
  4. Burn in and download the captioned video.

what to check

Where speech recognition slips

Whisper is strong on ordinary speech and weak in a few predictable places. Reading the transcript once with these in mind catches nearly every error, and each line stays clickable so a correction takes seconds.

Product and company names

Anything not in ordinary English gets mapped to the nearest common word. Your product name, your competitors, and any internal jargon are the first things to scan for; they are also the words your viewers most need to see spelled correctly.

Numbers and units

Spoken figures land inconsistently: "fourteen ninety-nine" may arrive as words rather than 14.99, and version numbers often lose their punctuation. Prices, dates and version strings are worth checking one by one.

Homophones with no context

Short technical clips give the model little surrounding sentence to disambiguate from, so their/there, to/two and site/cite slip through. Longer sentences come out correct, because the context resolves them.

Cross-talk and overlapping speech

When two people talk at once the model transcribes one of them and drops the other. Interview and podcast footage needs a closer read than a single-narrator screen recording.

Acronyms and spelled-out letters

API, SDK and URL usually land, but anything read letter by letter tends to arrive as a word or with spaces in odd places. Check product SKUs, coupon codes and short commands character by character.

Sentence breaks in fast speech

Timing is accurate but punctuation is inferred, so a quick run of short sentences can merge into one long caption line. Splitting those improves readability more than any styling choice, because a viewer reads a caption in about two seconds.

terminology

Captions, subtitles and open captions

The three terms describe different jobs. Captions carry the audio in the same language for viewers who cannot hear it, sometimes including sound cues like [door closes]. Subtitles translate dialogue for viewers who do not speak the language and assume the viewer can hear. Open captions are burned into the pixels rather than delivered as a separate file, which is what this tool produces.

Burned-in captions cannot be turned off, which suits social feeds where video plays muted by default and caption settings vary by viewer. The trade-off is that they cannot be translated later, resized for a different aspect ratio, or read by search engines the way a caption file can. For a video you will re-cut or localise, keep the caption text as part of the project rather than baked into a finished export.

readability

Captions people can actually read

A caption has to be read in the time it is on screen, by someone holding a phone at arm’s length, often in daylight. That makes readability a timing problem as much as a design one. Comfortable reading runs at roughly 15 to 20 characters per second, so a 40-character line needs two to three seconds on screen. A line that flashes past faster than that is decoration, not communication.

Two lines is the working maximum, at about 32 to 42 characters each. A third line starts covering the picture, and on a vertical video it pushes into the region where the platform draws its own interface. Break lines at natural phrase boundaries instead of filling to the character limit. “We cut the render time / from forty minutes to four” reads cleanly; a break after “from” forces the viewer to hold an incomplete thought across a line change.

Contrast decides whether any of that matters. White text on light footage is unreadable outdoors regardless of size, so every usable caption style carries either a solid backing bar or a heavy shadow instead of relying on text colour alone. Size scales with the smallest screen you expect. Text that looks generous on a laptop preview is often marginal on a phone, and phones are where most of this footage gets watched.

placement

Where the platform UI will cover your captions

Vertical feeds draw their own controls over your video, and they draw them in the same places every time: a caption sitting at the bottom of the frame lands under the account name, the description and the sound label, while the right edge carries the like, comment and share column. Anything you place there is competing with interface you do not control.

Vertical feeds

Keep captions inside the middle band of the frame, clear of roughly the bottom fifth and the right-hand column. Centring them slightly above the midpoint is safest across Reels, Shorts and TikTok, because all three overlay the same regions.

Landscape video

The lower third is genuinely available, which is where the name comes from. The exception is anywhere a player draws a scrub bar on hover. Leave a margin so a caption is never half-covered by a control that appears when someone reaches for pause.

Square and 4:5

The most forgiving shapes for captions, and the reason 4:5 performs well as an in-feed format. There is vertical room above the interface without the extreme crop of full 9:16, so a two-line caption fits without crowding the subject.

Embedded on your own site

You control the chrome, so captions can sit wherever the design wants. This is the one case where a caption file beats burned-in text, because a track can be toggled, restyled with CSS, and read by search engines as text.

Autoplay previews

Feed previews often start muted and small, sometimes cropped to a square thumbnail. A caption that only works at full size is invisible at the moment someone decides whether to stop scrolling. Size the first line for the preview, not the player.

Repurposed cuts

A landscape video re-cropped to vertical takes its burned-in captions with it, usually straight out of frame. If a video will be cut for more than one aspect ratio, keep captions as a project element and re-render, instead of baking them once and cropping.

why bother

What captions are actually for

The usual argument is that feed video plays muted, and that is true enough to matter. But captions do three other things that get less attention. They make the video usable by people who are deaf or hard of hearing, which is a requirement, not a nicety. The WCAG guidelines list captions for prerecorded audio as a Level A criterion, the baseline tier, and public-sector procurement in several markets checks for it.

They also help people who can hear perfectly well. Anyone watching in a second language, in a noisy office, on a bad connection where audio drops before video, or following a technical term they have never seen written down. A spoken product name is a guess until it appears spelled out. For software demos in particular, captions are often the only place a viewer learns how a command or a flag is actually written.

The third is retention. Captions keep a viewer anchored during the pause between spoken sentences, which is exactly where people scroll away. The format spread from accessibility tooling into mainstream editing because it helped the median viewer, not only the ones who needed it. Burned-in captions of the kind this tool produces cover all three cases at once, with the trade-off that they cannot be switched off or translated later. For that, keep the caption text in the project and re-render instead of baking it into a finished file.

The common mistake is treating captions as a transcript. A transcript records every word, including the false starts, the repeated phrases and the “so, basically” at the head of a sentence. Captions are read, not heard, and reading is less forgiving than listening. Filler that disappears in speech becomes clutter on screen, and it eats the seconds a viewer needs for the words that matter. Trimming a spoken line to its meaning is the same edit a subtitler would make, and it is the difference between captions that get read and captions that get ignored. The exception is anything quoted or legally significant, where the exact words are the point.

Speaker labels are the other judgement call. On a single-narrator screen recording they are noise. On an interview or a podcast clip, where two voices alternate and the speaker is often off camera, a short label at the start of each turn is the difference between a viewer following the conversation and giving up on it. Keep them to a first name or a role, and only where the change of speaker is not obvious from the picture. A caption announcing who is talking while that person is plainly on screen spends characters the viewer needs elsewhere. The same rule applies to sound cues: [laughter] earns its place when it explains a reaction the viewer can see but not hear, and clutters the frame when it does not.

CoAnimator
This tool is the free taste.

CoAnimator is the full studio, driven by your AI agent.

  • Real-footage effects
  • Motion graphics
  • 3D scenes
  • Voiceover
  • Unlimited local renders

FAQ

How do I add captions to a video?

Drop your video into the tool above: Whisper transcribes the speech in your browser, you click any line to fix a word, and the captions burn in on a local canvas — no upload, no account, no watermark. Pick one of three lower-third styles before exporting.

Is this auto caption generator really free, with no watermark?

Yes — completely free and unwatermarked. It can be, because there’s no server doing the work: the Whisper model runs in your own browser and the burn-in renders locally, so there’s no compute bill to recover with watermarks or sign-ups.

What’s the difference between captions and subtitles?

Captions transcribe the audio for viewers who can’t hear it — same language, sometimes with sound cues — while subtitles translate dialogue for viewers who don’t speak it. Social feeds mostly need captions, since most feed video plays muted. This tool generates same-language captions from your audio.

What video formats work, and what do I get out?

Any video your browser can decode works — MP4 and WebM cover most recordings. The output is your video with captions burned into the pixels, which plays everywhere without needing a separate subtitle file or player support.

How do I get styled, animated captions on a full video project?

Burned-in lower thirds cover social clips; for animated caption styles inside produced videos, the CoAnimator desktop app treats captions as timeline elements an AI agent can style and time to the narration. The app is a free download for macOS and Windows.

Does my video get uploaded?

No. Whisper runs in your browser (WebGPU when available, WebAssembly otherwise) and the burn-in renders on a local canvas. The file never leaves your machine — which is also why there is no watermark.

How fast is transcription?

On WebGPU we measured 169 seconds of speech transcribing in about 4.5 seconds; the no-GPU fallback did the same track in 21.5 seconds. The first run downloads the model once and caches it.

Can it handle long videos?

Transcription handles well past two minutes comfortably. For long-form or 4K exports, the CoAnimator desktop app renders styled caption tracks as part of a full project.