The good synthetic voices no longer fail on the words. They fail on the third minute, when you notice that every sentence has the same shape and nothing is ever said as if it mattered more than the sentence before it. That is the current state of text to speech: good enough that most people cannot pick out a single sentence, and still recognisable over a paragraph if you listen for it.
This post sorts the 2026 options by the questions I would ask before paying for one: what a voice sounds like on your own text rather than the demo, what an hour of audio costs, how to write so the voice has a chance, and what the law requires now that the EU AI Act's labelling rule is in force.
I do not produce audiobooks or podcasts. What I do use text to speech for is reading my own reports back to me, because the ear catches the doubled word and the sentence that never ends where the eye slid past them. That is a small use, and it is the one I can speak for.
Where a synthetic voice goes wrong, and why it is the same three places
Every text to speech system does three jobs, and each one has its own way of failing.
First it reads the text and decides what the symbols mean. "1,500" is one thousand five hundred; "6/12" is either June the twelfth or the sixth of December depending on who wrote it; "NASA" is a word and "FBI" is three letters. This is where you hear a postcode read as a year, or a Dutch surname pronounced as if it were English.
Then it decides the prosody: where to pause, which word carries the stress, how the pitch falls at a full stop and rises at a question. This is the job that separates the 2026 models from the 2022 ones, and it is also where the third-minute problem lives. A model can only guess emphasis from the text, and text does not say which word matters.
Last it generates the waveform. This part is solved. Nobody hears the vocoder any more.
The Text to Speech tool on this site uses the voices built into your operating system and browser: pick a voice, set speed and pitch, and it reads the text aloud. It does not produce a file. It is what I use for hearing a paragraph back, and it is a fair way to test how your text behaves before you pay a cloud service to read it. On a Mac or an iPhone the built-in voices are good; on Windows and in Chrome they vary.

Cloning a voice takes a minute of audio, which is the problem
A recognisable clone of a voice now takes about a minute of clean recording. A clone good enough to narrate a book takes half an hour or more of varied speech, and the better services ask for that much and verify that the person recording is the person being cloned.
The uses that hold up:
- Someone with a long list of written pieces reads them aloud in their own voice without booking a studio. This is the use the services are built for.
- Someone losing their voice to illness banks it while they still can. ElevenLabs has offered this free to people with ALS since 2024, and it is the use I find hardest to argue with.
- One voice across a training course or a phone system that stays the same after the original speaker has moved on.
- A course recorded in English, then generated in German and Dutch in the same voice. The result is the same person speaking languages they do not speak, which is useful and slightly uncanny.
The same minute of audio is all a scammer needs, and the phone call from a relative in trouble is now a standard fraud. The US Federal Communications Commission ruled in February 2024 that AI-generated voices in robocalls are illegal under existing law. Tennessee's ELVIS Act, in force since 2024, made a person's voice a protected likeness. The serious services now watermark their output and require consent for professional clones, and none of that helps the person who receives the call.
A recognisable clone of a voice now takes about a minute of clean recording.
What it costs, per hour of audio and per book
A human narrator costs a few hundred dollars per finished hour at the low end and well over a thousand at the top, and a finished hour takes several hours of work. A 50,000-word book is about five hours of audio.
The cloud APIs charge per character. OpenAI, Google and Microsoft all land in the range of fifteen to thirty dollars per million characters for their better voices. That same book is roughly 300,000 characters, so it costs somewhere between five and ten dollars to generate. ElevenLabs sells subscriptions rather than raw characters; the same book runs to tens of dollars there, and the voices are the ones most people rate highest. Kokoro, an open model small enough to run on a laptop, costs nothing and sounds better than anything you could buy in 2022.
So the cost of the audio is no longer the decision. The decisions are whether the platform will take it and whether the listener will finish it. Amazon's KDP has let authors publish AI-narrated audiobooks since 2023, labelled as such on Audible, and Audible opened AI narration to publishers in 2025. Reviews on those titles are where you find out how the third-minute problem plays out over five hours.
The uses where the arithmetic is simplest: a training course whose narration must change every time a slide changes, a phone system reading back an account balance, a website offering an audio version of every article. In each case the text is already there, the voice is a delivery mechanism, and nobody is choosing the audio for the performance.
Before you convert anything long, run it through the Word Counter to see what it will cost and the Readability Checker to see how it will sound. Long sentences with three subordinate clauses read badly aloud in any voice, human or otherwise.
Write for the ear before you convert anything
The voice can only work with what the text gives it, and text written for the eye gives it very little. Six changes that make more difference than switching provider:
Short sentences. A sentence with three subordinate clauses gives the model three places to put the wrong pause. One idea per sentence and the pauses land themselves.
Punctuation is the control surface. A full stop drops the pitch and stops. A comma is a short breath. A paragraph break is a longer one. Use them for the rhythm you want, not for the grammar you were taught.
Spell out anything with two readings. "Dr." is doctor or drive. "Jan" is a month or a Dutch first name. "St." is street or saint. The model will pick one, and it will pick with confidence.
Formatting does not exist in audio. A bulleted list becomes one long sentence with no boundaries. Headings vanish. Rewrite lists as sentences with "first", "second", "and last" before converting.
For names and technical terms, use the pronunciation controls the platform gives you. Google, Microsoft and Amazon accept SSML, a markup with tags for pronunciation, pauses and emphasis. OpenAI's current model takes a plain-language instruction instead ("read this like a news bulletin, pronounce Schoenmaekers as SKOON-mah-kers"). ElevenLabs accepts a subset, such as pause tags.
And read the text aloud yourself, or have the browser do it, before spending money on a cloud voice. Anything that trips your own tongue trips the model.

Test on your own text, and read the licence before the price list
The 2026 line-up: ElevenLabs, whose voices most listeners rate highest and whose pricing is a subscription; OpenAI's gpt-4o-mini-tts, steered by instructions rather than markup; Google's Chirp 3 HD voices; Microsoft's Azure neural and HD voices; Amazon Polly's generative voices; and the open models, of which Kokoro is the one small enough to run on your own machine. Every one of them has a demo page that sounds wonderful, because the demo text was chosen to.
So do not listen to the demo. Paste two pages of your own text, the kind with your own names and numbers in it, into each candidate and listen to all of it. A voice that is convincing on a news paragraph can fall apart on dialogue or on a page of technical terms, and the only way to find out is with the text you are going to ship.
Then check the things the demo cannot show:
- Every language you need, at the same standard. English is where the effort goes. Dutch, in my experience, is a tier below on most platforms, and smaller languages further still.
- The pronunciation controls from the previous section. If you cannot fix how it says a name, you will be living with the mistake.
- The pricing unit. Per character, per minute of audio, or per subscription credit. Work out your monthly volume in that unit before comparing, because the cheapest option at ten thousand characters is rarely the cheapest at ten million.
- Latency, if the voice talks back in real time. The cloud services answer in a few hundred milliseconds; the fastest of them under a hundred. A phone system notices the difference, a podcast does not.
- The licence. ElevenLabs' free tier is non-commercial and requires attribution. Several providers forbid using the output where a listener could believe it is a human without being told. Read the terms before you build a product on the voice, not after.
What the law requires, as of August 2026
The rules caught up this year, and for anyone publishing in Europe the date that matters is 2 August 2026. Since then Article 50 of the EU AI Act has been in force: AI-generated audio that resembles a real person has to be labelled as artificial, and providers of the generating systems have to mark their output so that it can be detected by machine. The earlier version of this post said the Act treated voice cloning as high-risk. It does not; it treats it as something that must be disclosed, which in practice is the stricter rule because it applies to everyone.
In the United States there is no single law. The Federal Communications Commission ruled in February 2024 that AI voices in robocalls fall under the existing robocall ban. Tennessee's ELVIS Act made a person's voice a protected likeness the same year, and a growing list of states has followed with their own versions. The actors' union struck for nearly a year over AI voice terms in video games and settled in 2025 with consent and payment provisions, which tells you where the working practice is heading: the voice is licensed, per use, with a record of consent.
What that means if you are the one pressing generate:
- Say the voice is synthetic wherever a listener could reasonably think otherwise. A line in the show notes, a sentence in the intro.
- Clone nobody without written consent, and keep the consent with the recording.
- If you licence your own voice, licence it per project with a term, not forever.
None of this stops the fraud calls. It does mean that the honest uses now have a clear rule to follow, and a year ago they did not.
The rules caught up this year, and for anyone publishing in Europe the date that matters is 2 August 2026.
FAQ
Is a synthetic voice good enough for an audiobook?
For non-fiction, yes, and the platforms accept it: Amazon's KDP has published AI-narrated titles since 2023 and labels them on Audible. For fiction with several characters and scenes that turn on tone, listeners still notice, and the reviews on AI-narrated novels say so. If the book is the kind people finish in one sitting, generate it. If it is the kind they live with for a week, listen to the third chapter before deciding.
What does it cost to convert a book?
About 300,000 characters for 50,000 words. On the per-character cloud APIs that is five to ten dollars. On ElevenLabs it is a subscription month plus extra credits, so tens of dollars. On an open model running on your own machine it is electricity and patience.
Can one voice read text that switches languages mid-sentence?
Better than two years ago. The current multilingual models from ElevenLabs and OpenAI handle a French phrase inside an English sentence without switching accent entirely, and Google's newer voices do too. It still fails on a full paragraph in the second language. Split the text by language when you can and mark the parts you cannot split.
Will this replace voice actors?
It is replacing the work that was never about performance: the training module, the phone menu, the article read aloud. It has not replaced the audiobook narrator whose name sells copies or the actor who gives a character a voice, and the 2025 union settlement is built on the assumption that those jobs continue, with consent and payment when a voice is cloned.
Can I tell by ear whether audio is synthetic?
Over a sentence, mostly not any more. Over a few minutes, sometimes: every sentence the same length and shape, pauses that are all exactly as long as each other, no breath before a long clause. The reliable test is not the ear but the watermark, which the larger providers now embed and which their detection tools can read. Treat a confident answer from anyone listening as a guess.
AI Image Prompts in 2026: Write a Brief, Not a Keyword Pile
What works in image prompts on the 2026 models (GPT Image, Nano Banana, Midjourney v7, Flux): how to structure one, which old habits now hurt, how to iterate without starting over, and what the EU AI Act asks of you when you publish the result.
AI Image Generators in 2026: Which One I Would Pick for Which Job
Midjourney, ChatGPT's image model, Flux, Stable Diffusion, Ideogram, Leonardo, Google and Adobe compared on the things that decide the choice: text in the image, control, licence terms and price. Plus what the EU AI Act asks of you since August 2026.
AI Bedtime Story Generators: What to Ask For, and What to Check Before You Read It Aloud
A generator writes a story with the child's name in it in thirty seconds. Whether it is a bedtime story, short, calm and ending asleep, depends on the prompt. The tools I could verify, the prompt that gets a usable story, word counts per age, and the one-minute check before you read it out.
AI Language Apps: The Three Features That Work, the Apps That Have Them, and What No App Teaches
Spaced repetition, pronunciation scoring and reading pitched at your level are the AI features that move a learner. Which apps do them well, a routine that fits in twenty minutes a day, and the list of things you will only learn from people.
