Word-by-word

Word-by-word captions — one word, exactly on cue.

A single word, or a tight cluster, appearing as it’s spoken. The default short-form look, and the reason it works.

Start captioning freeFree plan · no credit card
Your video never uploads

Word-by-word subtitles reveal one word at a time, each appearing at the instant it is spoken rather than a whole line sitting on screen for several seconds. LumaCaption times every word from the audio itself, and 82 of its 340 caption presets are built to reveal word by word.

Somewhere in the last few years, short-form captions stopped being subtitles. Instead of a line sitting at the bottom of the frame for three seconds, you get one word — or a tight two- or three-word cluster — arriving at the instant it’s said and leaving when the next one lands. Almost every high-retention talking-head account uses some version of it now.

The reason isn’t decorative. When only the current word is on screen, reading speed is forced to match speaking speed: the viewer can’t read ahead and get bored waiting for your voice to catch up, and can’t fall behind and miss the next sentence trying to finish the last one. There’s also no block of text to take in at a glance, judge and dismiss — the caption is never “done”, so there’s always a reason to stay another beat. LumaCaption times every word from your audio and gives you presets built around exactly that rhythm.

The viewer can’t read ahead

One word at a time removes the gap between reading and listening. Your pacing, pauses and emphasis land the way you delivered them instead of being flattened into a paragraph.

Nothing to skim and dismiss

A full line of text can be absorbed in a glance and written off. A word that keeps changing holds attention through the hook, which is where short-form videos are won or lost.

One word or a tight cluster

Different presets show different amounts. Typist and Baseline Rise run strictly one at a time; Rise Blur, Rise Focus, Cascade and Editorial group two or three words so fast speakers stay readable.

Steady, not jittery

The failure mode of single-word captions is a block that resizes and re-centres on every word. Equal Width holds the caption to a consistent footprint so the eye stays in one place while the words change.

Word-by-word reveal is not the same as word-by-word highlight

Two different looks get the same name, and picking the wrong one is the most common mistake here. In a **reveal**, only the words already spoken exist on screen — the line builds as you talk. In a **highlight**, the whole line is on screen from the start and the current word changes colour as it is reached. Both are word-level; they do opposite things to the viewer.

A reveal removes the ability to read ahead, which is what makes it hold attention on fast, punchy delivery. A highlight preserves context, which is what makes it right for anything a viewer needs to follow rather than just absorb — lyrics, instructions, a language they are still learning. Neither is more advanced; they suit different footage.

Reveal (82 of 340 presets)
Words appear as spoken. Best for talking heads, hooks and high-energy short-form.
Cue (192 of 340 presets)
A short chunk of up to three or four words arrives together. The default, and the calmest to read.
Highlight
Full line visible, active word recoloured. Covered on the karaoke captions page.

How the timings are produced

A word-by-word caption is only as good as its timestamps, and this is where most tools quietly fail: they time the *line*, then divide it evenly across the words in it. That looks fine on a steady sentence and falls apart the moment you pause, stumble or stress a word — the caption drifts away from the voice and the effect inverts, becoming distracting instead of magnetic.

LumaCaption asks the speech model for word-level timestamps directly, so a half-second pause before a punchline is a half-second gap in the captions too. That is also what makes the emphasis features possible: the renderer knows which word is being spoken on any given frame, so it can scale it, recolour it or move a highlight onto it without any manual keyframing.

When it is the wrong choice

One word on screen means no context, and there is footage that needs context. Dense technical explanation, anything with numbers a viewer wants to compare, and long-form talking video all read better in cues — the viewer is trying to hold two ideas at once and a single word denies them the chance. Accessibility guidance also assumes a readable line rather than a single flashing word.

The practical rule: reveal word by word when the video is carried by delivery, and use cues when it is carried by information. Switching between them is a preset change, not a re-edit, so it costs nothing to try both against the same transcript.

Reveal modes across the 340 presets

ModeWhat appearsPresetsBest for
WordOne word at a time, as spoken82Hooks, talking heads, high-energy delivery
CueA chunk of up to 3–4 words together192Explanation, long-form, anything with numbers
SceneFull-frame kinetic typography44Titles, openers, hard emphasis beats
Focus / stackSpecialist single-preset behaviours22One-off looks

How it works

01

Upload your video

Nothing to install, on desktop or phone, and the video file is never uploaded — only the extracted audio is sent, and only to transcribe it.

02

Pick a word-by-word preset

Rise Blur, Rise Focus, Baseline Rise, Typist, Equal Width and Cascade all reveal in time with speech. If you want something bigger, 47 of the 290+ styles are full-frame kinetic typography scenes.

03

Adjust and export

Retime a word, correct a spelling, or star the word that should be highlighted on its line. Exports are free and unlimited, so trying a second look costs nothing.

Common questions

Word-by-word or line-at-a-time — which should I use?

Word-by-word for short-form: Reels, Shorts, TikTok, ads, hooks — anywhere you’re competing with a thumb. Line-at-a-time for long-form: podcasts, interviews, tutorials, anything over a few minutes, where a viewer settling in for twenty minutes wants to read at their own pace rather than yours. Line-based is also the right answer for accessibility and for subtitle files, since SRT and VTT are line-based formats by design.

Doesn’t one word at a time get hard to follow?

Not when the timing is genuinely per-word — the eye is being fed at the same rate the ear is, so it reads as natural. It does get uncomfortable when someone speaks very fast, because single words start flickering. In that case use a preset that groups two or three words: you keep the locked-to-speech feel and give the eye something to settle on.

Can I still get a subtitle file from a word-by-word video?

Yes. The same transcript exports as SRT, VTT or TXT on every plan including free — those formats group words back into readable lines, which is what platforms and screen readers expect. Plenty of creators burn in word-by-word captions for the feed and upload an SRT alongside the long-form version.

What makes word-by-word captions possible?

A timestamp for every individual word, produced during transcription. Tools that only record when a caption line starts and ends can’t place words individually, which is why their “word-by-word” option usually just splits lines into shorter lines. LumaCaption times at the word level, so a word appears when it’s said — and the same timings drive the karaoke highlight styles.

Getyourvideoswatched.

Every style unlocked. No credit card.