Voiceover scripting for short-form: when to write tight, when to ramble
A spectrum, two word-count budgets, and where to leave the silence editors keep cutting out.
"Stop scrolling. This $19 tool replaced my $400 mic. Link in bio."
"So I'm in this Lyft at 2am and the driver looks at me and goes... 'you're the guy from the videos, right?'"
Both are 60-second scripts. One should be 150 words. The other should be 90 and feel longer.
Most short-form voiceover advice gets this collapsed into a single rule — "be punchy, every word counts" — and that's how you end up with storytime videos that feel like ad copy and ads that feel like nothing at all.
The right question isn't how tight. It's which direction on the spectrum, and how committed are you to that direction.
The spectrum
On one end: the ad read. Hormozi's school of writing — every word earns rent, cut anything that survives deletion, write so a fifth-grader can visualize it (Hormozi via Practicing the Write Stuff). Donald Miller's StoryBrand sits on this end too: "if you confuse, you lose" (StoryBrand). Confusion is a tax. Cut it.
On the other end: the breathing storytime. NPR's house style for narrative audio: conversational, contractions, sentence fragments, dependent clauses are toxic (NPR Training). Casey Neistat's vlog voiceover lives here — first-person, pace dictated by music, the famous mid-sentence cut to a new scene (Artlist on Neistat).
The mistake is treating these as opposite qualities. They're opposite budgets.
Word-count budgets per format
Read these as ceilings, not targets.
- 60-second hard ad / DR: ≤ 150 words. About 2.5 words per second. Dense. No throat-clearing.
- 60-second hook-and-CTA (UGC, soft pitch): ≤ 130 words. Leaves room for one beat of personality.
- 60-second storytime / narrative: ≤ 110 words. Roughly 1.8 words per second. The gaps are the content.
- 30-second platform ad: ≤ 75 words. Yes, that's brutal. That's the format.
- 15-second hook: ≤ 35 words. One idea. One.
If your storytime hits 150 words at 60 seconds, you didn't write a story — you wrote an ad with anecdotes. The reason a Neistat-style cut works is that the script left air for it. Music and visuals fill the calorie deficit (In Depth Cine on Neistat).
Cadence and breath placement
Mark your breaths in the script. Castos uses ellipses as breath markers; an old radio trick is // for a hard pause and / for a soft one (Castos).
Three rules that hold across the spectrum:
- 01Breathe at thought boundaries, not comma boundaries. A comma is a grammar artifact. A breath is a meaning artifact. They're not the same.
- 02No more than ~12 words between breaths in tight reads, ~18 in ramble mode. Past that, even pros sound winded, and the listener clenches with them.
- 03Plant the breath before the punchline, not after. "I get to the door //, and the keys are in my hand the whole time." The pause cocks the line.
Read it out loud, with delivery notes, before you record (Castos). If your jaw is tired after one take, the script is too dense. Cut, don't push through.
Where to leave silence (and how editors cut it wrong)
Modern silence-removers are very good and slightly dangerous. Descript's own guidance is to flag punchlines and emotional beats before bulk-removing silence — those pauses are doing the work (Descript). Riverside's guidance is similar: leave at least ~0.3s of room around cuts or speech glues together and sounds robotic (Riverside via Cotovan).
The way silence dies in short-form is predictable:
- A "Remove Silence" pass with default thresholds eats the 0.4s beat before a reveal because it looks identical to a 0.4s lull between sentences. It isn't.
- Aggressive breath-deletion on a storytime makes the host sound like an AI voice clone. That's not a compliment.
- Cutting flush to the consonant — no buffer — produces the "glued words" artifact every editor hears once and can never unhear (Riverside via Cotovan).
Practical rule: in tight ad reads, you can compress aggressively but protect the silence before the offer and after the CTA. In storytime, protect the silence before the turn ("...and that's when I realized—") and after the line that lands. Mark these in the script with [HOLD] so the editor — future you, or someone else — doesn't auto-trim them out.
Script-to-screen translation
Captions are not a transcript. Captions are a second script written for a different organ.
Voice scripts are written in clauses; captions are written in chunks. Voice tolerates a 14-word sentence; a caption past 7 words breaks the eye-line. NPR's "writing for the ear" rule — no dependent clauses, sentence fragments are fine (NPR Training) — is even more brutal in captions.
Two practical moves:
- Don't burn the punchline twice. If the VO says "and the keys were in my hand the whole time," the caption can read "...the whole time." The eye finishes the sentence the ear is still landing.
- Strip the connective tissue from the caption, not the VO. "So", "and", "but", "you know" — those are breath words. They make the voice sound human. They make the caption look amateur. Different scripts.
The contrarian
Here it is: most short-form scripts should be 20% longer than they are, not shorter.
Everyone's been told to cut. The default failure mode in 2026 isn't bloat — it's compression sickness. Scripts that read like LinkedIn posts, voices that sound like they're being chased, no air for the visual to land. The "every word earns rent" doctrine is correct for the ad-read end of the spectrum. Apply it to a storytime and you sand off the thing that made it a story.
Hormozi-tight is a register, not a universal law. If you're writing storytime, UGC narrative, or anything with a turn — write it long, then read it aloud, then cut only the words that fight your breath. Leave the ones that ride it.
One takeaway
Before you record, mark every breath in your script with a /, count the words between marks, and rewrite any run longer than 12 (tight) or 18 (ramble). That single pass will fix the pacing of more videos than any AI editor ever will.