The difference between a script and an article is that the reader cannot skim. They also cannot go back easily, and they leave at any moment with one thumb. So a script carries obligations prose does not: tell people what is coming, tell them where they are, and give them a reason to stay before every section break. Most generated scripts miss all three, because they are written as continuous prose.
Word counts and timing
Spoken English runs at roughly 130 to 150 words per minute for a conversational delivery, faster for high-energy channels and slower for teaching ones. That makes a ten-minute script about 1,300 to 1,500 words, and a six-minute script around 800 to 900. Decide the runtime first, then write to that number — a script that runs long does not get better, it gets trimmed at the end when you are tired.
The eight templates
1. Tutorial — for doing something specific
- Hook (15 seconds, ~35 words): the outcome, then the cost of not knowing it.
- Preview (20 seconds, ~45 words): the steps named in order, so nobody is guessing what comes next.
- Prerequisites (30 seconds): what to have open before starting.
- Steps (60% of runtime): one per section, with the common failure named after each.
- Verification (30 seconds): how the viewer knows it worked.
- Close and next step (20 seconds): one action, one related video.
2. Listicle — for breadth
- Hook (15 seconds): the number and the selection rule — seven tools we actually paid for.
- Why this list (20 seconds): your basis for choosing. This is what separates it from a scrape.
- Items (75%): equal airtime each, each with the catch stated out loud.
- How to choose (45 seconds): a decision rule, because twelve options is a problem.
- Close (15 seconds): which one you would start with.
3. Review — for a verdict on one thing
- Verdict first (20 seconds): who should buy it and who should not. Do not make them wait.
- What it is and the price (30 seconds): including what the plan actually includes.
- The test (40%): what you ran through it, with real numbers.
- Where it struggles (25%): the section that makes the rest credible.
- Alternatives (30 seconds): two, named honestly.
4. Comparison — for two named options
- The decision (20 seconds): who picks which, stated before the evidence.
- Where they overlap (30 seconds): establishes that the difference is real.
- Dimension by dimension (60%): price, limits, fit, support, lock-in — each clearly signposted on screen.
- Switching cost (45 seconds): what it takes to move later.
- Third option (30 seconds): for people neither suits.
5. Explainer — for a concept
- The question (15 seconds): phrased as the viewer would ask it.
- The wrong mental model (45 seconds): what most people assume, and why it misleads.
- The real mechanism (50%): built up in layers, with one concrete example per layer.
- Where it breaks down (60 seconds): the limits of the explanation.
- Close (20 seconds): what to do with it.
6. Story — for a narrative
- The moment (20 seconds): start in the middle of it.
- The stakes (45 seconds): what was on the line.
- The attempt (50%): what happened, in sequence, with specifics.
- The reversal (20%): the thing that changed the outcome.
- The takeaway (45 seconds): what transfers to someone else's situation.
7. Short — for vertical, under 60 seconds
- Line one (2 seconds): the claim, with no greeting.
- The proof (35 seconds): one idea, one example.
- The turn (15 seconds): the counter-intuitive part.
- Close (5 seconds): one line, no "like and subscribe" ramble.
8. Channel trailer — for the landing experience
- Who this is for (10 seconds): named precisely.
- What they get (25 seconds): the recurring formats.
- Proof (20 seconds): one concrete result or credential.
- Subscribe prompt with a reason (10 seconds): the cadence you actually publish on.
Where generated scripts go wrong
Four patterns showed up consistently. The hook is written as an introduction — "hey everyone, welcome back to the channel" costs you the first ten seconds where retention is decided. No signposting between sections, so the viewer cannot tell whether the boring part will end soon. Numbers delivered without a source or a scale, which sounds authoritative and means nothing. And the close asks for three things — subscribe, comment, watch the next one — which reliably gets none of them. Ask for one.
What this costs, from our 30-day benchmark
Our 30-day test measured roughly 22 minutes of editing per 1,000 words on the low-cost dedicated writer against about 8 minutes on the general assistant — near $7 per piece of your time at $30 per hour. Scripts sit closer to the general assistant's end of the range, because holding the thread across fifteen hundred words is exactly where cheaper tools lose it.
A workflow that holds up
Fix the runtime first and write to a word count. Give the tool the raw material — your actual test results, the real story — rather than asking it to invent the substance. Then read the whole script aloud once before recording; anything you stumble over will lose viewers faster than any hook can win them. Mark the retention checkpoints physically in the script so you can see them while recording: the preview before minute one, the "here is why this matters" before every long explanation, and the payoff promised early and delivered late.
Write to a runtime, not a word count
Pick the structure, give it real material, then read it aloud before you record.
Try Rytr →Frequently Asked Questions
How many words is a ten-minute YouTube script?
Around 1,300 to 1,500 for conversational delivery, based on roughly 130 to 150 spoken words per minute. Faster channels and denser topics vary.
Can AI write a full video script?
It can structure one well, particularly for tutorials and listicles. Supply the actual substance — real test numbers, real events — because invented specifics are obvious on camera.
What is the biggest mistake in generated scripts?
The opening. "Hey everyone, welcome back" spends the first ten seconds, which is exactly where viewers decide whether to stay.
Should a script differ from a blog post on the same topic?
Yes. More signposting, shorter sentences, more repetition of where you are, and a payoff structure that rewards listening rather than skimming.