Copy the caption style from any video

You have seen the look you want. It is in somebody else’s reel, and every part of the answer is already on screen — the colour, the size, where it sits, whether the spoken word lights up — just baked into pixels instead of written down anywhere you can copy. This reads it back off those pixels and puts it on your captions. Your words stay yours; only the look comes from the video.

Caption my video — freeNo sign-up · nothing uploaded

How to do it

Open the editor with your own video and generate the captions first — a style needs captions to land on, and matching onto an empty timeline changes a setting you cannot see. Then open the Media panel, choose Match style, and pick the video whose captions you want to copy. It plays through it, measures the lettering frame by frame, and applies what it finds.

The reference video is opened, measured and dropped. It is never uploaded, never becomes a clip in your project, and nothing about it is kept — the same promise the rest of this product makes, applied to somebody else’s footage as well as your own.

What comes back is a list of what was found and how confident the match is on each part, so you can keep what landed and change what did not. It is a starting point that is already most of the way there, not a black box.

What it reads

Colour, including the two-tone case where the spoken word is a different colour from the rest of the line — which is most of what makes a karaoke caption look like a karaoke caption. Font size relative to the frame, and weight, measured from how thick the strokes are rather than guessed from a name.

Casing, because a style set in capitals reads completely differently from the same style in sentence case. Position within the frame, so captions that sit low over a speaker’s chest stay there rather than jumping to the middle. Outline and shadow, measured as a ratio of the letterform, which is what keeps text legible over a bright background.

And the entrance: whether each caption fades, pops, slides or types in, classified from how it arrives over its first few frames.

What it cannot do, said before you try it

It matches the font by measurement, not by identification. There is no font-recognition model here — every face in the library is rendered once and measured, the video’s lettering is measured the same way, and the closest one wins. That reliably picks a font that looks like the one on screen, and it cannot pick a font we do not have. If the video uses a licensed brand typeface, you get the nearest thing rather than the thing.

Two similar easings are not distinguishable. Fade, pop, slide and typewriter are; two flavours of ease-out are one answer.

It reads the captions the video actually shows. A style appearing on a single caption in a two-minute video will probably not be sampled, and captions burned in at a tiny size give it less to measure. Clean, large, front-and-centre captions match best — which is to say the styles people usually want to copy are the ones it reads most reliably.

Every field comes back with a confidence figure, and the ones it is unsure about are shown as unsure rather than presented as fact. A tool that guesses silently is worse than one that says it does not know.

Why nothing else does this

The obvious way to build it is to send frames to a vision API and ask what the captions look like. That is one API call and an afternoon, and it would mean the reference video leaves the device — which on a product whose entire pitch is that nothing is uploaded would cost more than the feature returns.

So it is arithmetic instead: masks, blob labelling, stroke ratios and letter measurement over raw pixel data, in the same style as the automatic reframing. It runs on your machine, it runs offline, and it costs nothing per use — which is also why there is no credit meter attached to it.

Questions

Does the video I am copying from get uploaded?

No. It is opened and read frame by frame on your own device and then dropped. It never reaches a server, it does not become part of your project, and nothing about it is stored.

Will it tell me the name of the font?

It picks the closest match from the fonts available here, by measuring letterforms rather than recognising them. That gives you a font that looks like the one in the video. It is not a font identifier and it cannot name a typeface it does not have.

Can I copy the style from a video I did not make?

The look of captions — a colour, a size, a position, a highlight on the spoken word — is not something anyone owns, and copying it is what every style preset in every editor already does. Your words stay yours. What you should not do is republish somebody else’s footage, and this never gives you their footage: it reads the reference and discards it.

What if the match is wrong?

Every field it sets comes with a confidence, so you can see which parts it was sure about. Anything it got wrong is a normal control in the style panel — change it and carry on. It is a starting point, not a lock.

Does it cost anything?

No. It runs on your device, so it costs nothing to run and there is no credit attached to it. AI Caption Cut is free to use, with unlimited exports at 480p carrying a small watermark. Removing the watermark is a one-off payment for that one video — 720p, or full 1080p at 60fps — or an Unlimited Pass that covers every export for thirty days. Subtitle files are free either way. It is charged in your own currency at checkout, and there is no subscription and no account.

Related free tools

Try it on your own video

Nothing to install, no account, and your video never leaves your device. Unlimited free exports, and a one-off payment when you want the watermark gone.

Caption my video — free