IndexTTS Online Guides
How to Write Better TTS Scripts: Punctuation, Numbers, Acronyms, and Natural Speech
A practical writing guide for turning written copy into scripts that sound clearer and more natural when generated as speech.
Text-to-speech quality depends on more than the voice model. The script itself tells the model where ideas begin and end, how numbers should be read, which abbreviations are ambiguous, and where a human speaker would pause.
A useful TTS script is not simply an article pasted into a generator. It is a spoken-language edition of the content.
Write for one listen
Readers can pause, scan backward, and inspect punctuation. Listeners usually hear a sentence once while doing something else. That means spoken writing should reveal structure early.
Prefer direct sentences, clear transitions, and one main idea at a time. If a sentence contains several nested clauses, split it before generation.
Read the script aloud before generating
This is the fastest quality check available. If you stumble, run out of breath, or naturally change the wording while reading, the text is probably not ready for TTS.
Reading aloud reveals awkward rhythm, repeated words, hidden ambiguity, and punctuation that looked acceptable on screen but does not match speech.
Use punctuation as delivery guidance
Punctuation is not a precise prosody language, but it provides useful structure.
Periods create strong boundaries. Commas can mark short pauses. Colons can introduce lists or explanations. Question marks help distinguish questions from statements.
Do not add punctuation randomly in an attempt to force dramatic delivery. The text should still be grammatically understandable to a human reviewer.
Rewrite numbers in the form you want spoken
A compact number can have several readings. For important narration, reduce ambiguity.
Instead of leaving “2026/08/23” for the model to interpret, write the date in words appropriate to the target language. Decide whether “3.5” means “three point five,” “three and a half,” or something else in context.
For currencies, specify the currency when the symbol could be ambiguous. For percentages, decide whether the symbol should be spoken as “percent.”
Expand abbreviations when the spoken form matters
Some acronyms are spoken letter by letter, some as words, and some expanded into full phrases. A model cannot always infer which reading your project expects.
If an abbreviation is critical, write it in a speech-friendly form. Maintain the same treatment across the project so repeated terms stay consistent.
This is especially useful for company names, technical products, medical terminology, and internal abbreviations.
Handle URLs and email addresses deliberately
A visible URL may not need to be spoken at all. If it does, decide whether the audience needs the full address or a simpler instruction such as “visit our website.”
Email addresses can sound awkward when read character by character. In many voiceovers, it is better to display the address on screen and say a simpler call to action.
Do not automatically feed every visible string into narration.
Break lists into spoken structure
Written bullet lists often become monotonous when converted directly to speech. Add transitions such as “first,” “next,” and “finally” when they help the listener understand order.
For long lists, group related items rather than reading many disconnected terms in sequence.
If the audience needs to remember exact items, consider pairing narration with on-screen text.
Introduce names before repeating them
A difficult proper noun is easier to review when it appears in a short ordinary sentence rather than as an isolated word. Test names early and reuse the successful spelling treatment throughout the project.
For international names, decide whether you want the source-language pronunciation or a localized pronunciation. Consistency matters more than leaving the decision to chance on every generation.
Avoid excessive parenthetical writing
Parentheses are common in articles and reports but can create awkward spoken interruptions. Convert important parenthetical information into a normal sentence or remove it if the listener does not need it.
The same principle applies to footnotes, citations, and references. Decide what belongs in the spoken experience.
Use shorter paragraphs for generation
Even when the final content is long, smaller coherent sections are easier to generate and review. A paragraph-sized section gives you enough context for natural rhythm while keeping corrections manageable.
Split at meaning boundaries, not arbitrary character counts. Keep complete sentences together.
Write transitions explicitly
Visual documents rely on headings and layout to signal structure. Audio needs spoken transitions.
A heading such as “Security” may need a line like “Next, let's look at security.” A new chapter may need a short orientation sentence so the listener knows the topic changed.
Do not assume a silent visual heading will communicate through audio.
Keep tone consistent
A script assembled from marketing copy, technical documentation, and support text may switch tone abruptly. Edit the pieces into one speaking style before generation.
Choose a level of formality, typical sentence length, and preferred terminology. This often improves perceived voice consistency even without changing model settings.
Use pronunciation testing as part of writing
When a word repeatedly fails, treat it as a script problem to solve. Create a short test sentence, experiment with a clearer written form, and record the solution in a pronunciation guide.
Do not wait until the final batch to discover that a product name appears incorrectly in every section.
Review the generated audio against the script
After generation, confirm that the audio matches the approved text. Listen for omissions, repeated words, unexpected pauses, and places where the spoken interpretation changed meaning.
If the same issue happens repeatedly, fix the script convention rather than correcting each output individually.
A speech-first editing checklist
Before generation, check that sentences are comfortable to read aloud; punctuation reflects natural pauses; important numbers are unambiguous; acronyms have an intended spoken form; names and brands have been tested; URLs and visual-only text are handled deliberately; headings have spoken transitions where needed; long paragraphs are divided at meaning boundaries; and tone is consistent across sections.
A good TTS model can only work with the structure it receives. Writing for speech reduces regeneration, makes pronunciation easier to review, and produces audio that feels designed for listeners rather than converted from a document.