← Back to blog

September 28, 2026 · 9 min read

How to make an AI script sound human when it's read out loud

Every guide to this problem sells you a voice setting. The research says the giveaway is the script. Here's the six-pass edit that makes spoken copy land.

How to make an AI script sound human when it's read out loud

An AI script sounds synthetic because of the writing underneath the voice, not the voice itself. Spoken copy fails when every sentence runs the same length, nothing in the draft hesitates, and every paragraph closes on a tidy summary. Fix the rhythm in the writing before you touch a single setting in the voice tool.

Why an AI script sounds worse out loud than on the page

Reading gives you a second pass. Listening doesn't. Your eye can jump back a clause, skim a line and rebuild the meaning from the shape on the page. A listener gets the words once, in real time, and the meaning has to survive the trip through someone's ears. Drafts written for the page are built to survive the first pass, not the second one.

There's a bigger difference, and it's structural. Punctuation is invisible to a listener. When you read, commas and paragraph breaks silently tell you where one idea ends and the next begins. Out loud, the only thing carrying that structure is the reader's voice, which means pacing, stress and pause length are suddenly doing work that formatting used to do for free.

That gap explains the specific disappointment most people hit. A draft that scored well in an editor falls apart in a voiceover, and the words are identical. What changed is that the page was holding the piece up. Take the page away and you can hear how little internal rhythm the sentences actually have.

What listeners notice that readers never do

Listeners don't hear your vocabulary. They hear rhythm. By the time someone listens instead of reads, the word choices have gone by, and what stays is the shape of each sentence: how long it runs, where the stress lands, and whether the next line arrives at the same tempo as the last one.

Most of these are the same tells we catalogued in our breakdown of common AI writing patterns, and they cost more in audio. On the page one of them is a small irritation. Out loud, two of them together is usually enough for a listener to decide the narrator isn't real.

How many fillers should you actually write in?

Fewer than you fear, and more than zero. Once you learn that AI scripts sound too smooth, the instinct is to sprinkle um and you know through the draft, and that overshoots badly. Spoken hesitation has a measurable range, and the range is narrower than most people guess.

Start with the baseline from natural conversation. Researchers studying task-oriented dialogue counted every filler, repeat and restart that speakers produced, and the average came out at just under six per hundred words.

Bortfeld et al., Language and Speech, 2001: speakers in task-oriented conversation produced 5.97 disfluencies per 100 words, of which fillers alone accounted for 2.38 to 2.74 per 100 words.

A second study gives you the ceiling. A 2024 parametric study in the Journal of Applied Behavior Analysis tested filler rates of 0, 2, 5 and 12 per minute and asked listeners to rate the speaker. Five per minute did not hurt perceived effectiveness. Twelve clearly did. The authors concluded that a rate of five or fewer disfluencies per minute may be acceptable.

Translate that into something usable while writing. One hesitation per forty to sixty spoken words, placed where a real speaker would stall: before a number, after a self-correction, and immediately before the sentence that carries your actual point. Three or four in a ten-minute script, not thirty.

The counterintuitive part is why the range sits so low. Fillers buy a listener time to catch up with a speaker who's thinking. When nothing in the script sounds like thinking, the fillers have nothing to signal, and they read as decoration.

The three tells you can only hear

Some failures survive on the page and only become obvious in audio. These three are the ones that reliably give a spoken script away.

  1. The flat open. People begin a thought before they've finished planning it, so the first words of a real sentence are slightly improvised. AI scripts open with a fully formed sentence, and a perfectly formed opening line is the loudest tell in the format.
  2. The missing breath. Written sentences are sized for comprehension. Spoken sentences are sized for lung capacity. Anything past roughly twenty words forces whoever reads it, human or synthetic, to leak air somewhere in the middle, and the leak lands on a random word.
  3. The uniform landing. When every third sentence closes on a neat summary beat, the listener stops tracking meaning and starts tracking a pattern. It's the audio equivalent of a paragraph that keeps ending on the same cadence.

The quickest way to hear this for yourself is to strip the connectors. Take a paragraph you've written, delete every transition word, and read both versions aloud. Our before and after on deleting transition words makes the effect obvious, because the rhythm has nowhere left to hide.

Here's the difference in one line. Written for the page: creating spoken content at scale requires a consistent process, and the most reliable process is the one that begins with a written draft before any audio is generated. Twenty-six words, one breathless clause, and an argument that arrives fully assembled. Nobody says that out loud, and nobody can without sounding like a memo.

Written for the ear: write the draft first. Then generate the audio. That order is the whole method. Three sentences, three different lengths, and a landing point on the final three words. The second version survives being read aloud because the breaks sit where a speaker would actually breathe.

Why the script beats the voice model

Search for a fix and you'll find hundreds of pages about stability sliders, similarity settings, pause tags and loudness targets measured in LUFS. Almost every one of them treats this as an audio problem. That framing made sense in 2022. It doesn't match what the perception research now shows.

The largest listening study on synthetic speech to date collected 35,532 judgments from 1,768 participants across 138 voice systems. Accuracy at spotting fake audio barely moved from a 2021 baseline, sliding from 72.9% to 71.2%. Accuracy at correctly accepting real human speech fell much further, from 72.7% to 64.1%. Listeners aren't getting better at hearing synthesis. They're getting worse at trusting each other.

arXiv 2605.26136, Eroding Trust in Real Speech, 2026: across 35,532 judgments from 1,768 participants, accuracy on fake samples barely moved from 72.9% to 71.2%, while accuracy on real human speech fell from 72.7% to 64.1%.

A 2026 study in Computers in Human Behavior points at what's left. Listeners rate AI speech lower on humanlikeness, and prosodic variation is what drives the separation between the two. Prosody means the rises and falls, the pace changes, and the pauses that happen mid-sentence. All of those are written into a script. A voice model performs the rhythm it was handed and adds nothing of its own.

That's the argument in one line. A better voice reading a flat script is just a better reading of a flat script, and the money you spent on the model gets spent twice.

How do you edit a script for the ear?

Six passes, in this order. On a ten-minute script it takes about twenty minutes, and it fixes more than any settings change will.

  1. Break every sentence past twenty words. Read it aloud. If you run out of breath, the narrator will too, and the break will land in the wrong place.
  2. Vary the length on purpose. After two long sentences, write a three-word one. The short line is where emphasis goes, and it costs you nothing but a period.
  3. Move the payoff to the end of the line. Stress falls on the last stressed word, so put the word that matters there instead of burying it inside a clause.
  4. Add one hesitation per forty to sixty words, on the beats where a person would actually stop.
  5. Cut the wrap-up. If a paragraph ends by summarising itself, delete the summary and end on the last concrete detail instead.
  6. Read the whole thing aloud again, standing up, at performance pace. Every stumble you hit is a place the narrator stumbles too, and this pass catches more than the five before it.

If the result still reads stiff, the problem has usually moved up a level, from sentence rhythm to register. Making written copy sound conversational covers that layer, and the two passes work well back to back.

What to do when the voice itself is the problem

Some failures genuinely belong to the audio, and it helps to know which ones, so you don't rewrite a good script chasing a settings problem.

Notice that only two of those three are audio fixes, and neither one changes the writing. Nothing in the settings can invent structure the script doesn't have.

The pass that fixes most of it

If you only do one thing, do the read-aloud pass and fix what you stumble on. In most scripts that single pass removes the flat open, the sentences too long for one breath, and the summary that closes the piece.

Then add hesitation on the beats where a person would stop, and shorten the lines that carry the point. Everything past that point is emphasis, which is a writing job. The voice model won't do it for you, and the vendor pages promising otherwise are selling you a slider.

This matters more in audio than on the page because listeners are already primed to suspect. What makes readers decide a piece is AI generated runs through that judgment, and most of it happens in the first few sentences, before any of your content arrives.

One number worth carrying into the edit. Conversational English lands somewhere between 150 and 170 words a minute, which means a single twenty-word sentence takes about eight seconds to say and a ten-minute script runs roughly 1,600 words. Counting seconds instead of words turns sentence length from a style question into a timing problem, and timing problems are the ones a listener can hear.

The judgment itself arrives early. Listeners decide whether a voice is worth trusting long before they've followed an argument, which is why the opening line carries more weight in audio than in any other format. Watch all three tells everywhere in the script, and watch them most carefully in the first fifteen seconds.

Fix the script, then choose the voice. In that order, and you'll spend less time regenerating lines and more time on the part a listener actually notices.

Frequently asked questions

Why does my AI script sound robotic when the words are fine?

Because the words were never the problem. Spoken copy fails on rhythm: sentences of matched length, no hesitation, and a tidy close every few paragraphs. A voice model performs the rhythm you wrote, so a flat draft gives you a flat read no matter which voice you pick.

How many filler words should a script have?

About one per forty to sixty spoken words. Conversation averages 5.97 disfluencies per 100 words (Bortfeld et al., 2001), and a 2024 study found that five fillers per minute did not hurt how listeners rated a speaker while twelve did. Three or four in a ten-minute script is enough.

Should I rewrite the script or switch to a better voice model?

Rewrite first. In the largest listening study to date, listeners were no better at spotting synthetic audio than they were in 2021, and accuracy on real human speech actually fell. Voice quality has stopped being the giveaway, which leaves the script as the thing you can still fix.

How do I know when a sentence is too long to speak?

Read it out loud. If you run out of breath, or the emphasis lands somewhere you did not intend, it's too long. Twenty words is a practical ceiling, and the fix is usually a period rather than a rewrite.