IndexTTS Online Guides
A Long-Form Voiceover Workflow for Podcasts, Videos, Courses, and Audiobooks
A structured method for turning long scripts into consistent AI voiceover without losing pronunciation, pacing, or editability.
Long-form AI narration fails when it is treated like one giant Generate button. The practical way to produce ten, twenty, or sixty minutes of usable audio is to divide the work into stable sections, test difficult material early, and keep the reference voice and generation settings consistent.
This workflow is designed for podcasts, video essays, online courses, audiobooks, product walkthroughs, and other projects where the final audio must feel continuous even though it is generated in smaller pieces.
Clean the script before generating
Remove material that is meant only for editors: camera notes, URLs that should not be spoken, formatting markers, duplicated headings, and comments. Rewrite abbreviations, dates, percentages, and numbers in the form you want the listener to hear.
A narration script should be easier to read aloud than the original article or slide deck. If a sentence is difficult for a human to say in one breath, it is probably too dense for synthetic narration as well.
Divide by meaning, not arbitrary character counts
Break the script at paragraph or topic boundaries. A section should contain one coherent idea and end at a natural pause. Avoid splitting in the middle of a sentence merely to fit a technical limit.
Meaningful sections make corrections cheaper. If a product name is mispronounced in one paragraph, you can regenerate that paragraph instead of recreating an entire chapter.
Establish a project baseline
Before producing the full script, choose one representative section and use it to lock the setup. Test the reference audio, model, pronunciation, pace, and any emotion setting.
The baseline should include the kinds of material that appear in the real project: names, numbers, normal exposition, and at least one longer sentence. Once it sounds right, record the exact setup and reuse it.
Maintain a pronunciation list
Long-form projects repeat important words. Create a list for names, brands, technical terms, abbreviations, and unusual spellings. Decide how each should be written in the narration text.
When a term is difficult, solve it once and reuse the same text treatment everywhere. Inconsistent spelling leads to inconsistent audio.
Generate in reviewable batches
Do not generate an entire hour before listening. Work in small batches, such as a few sections at a time. Review each batch for identity, pronunciation, pacing, and clipping.
This catches systematic problems early. If the reference or style is wrong, discovering it after three minutes is far cheaper than after thirty minutes.
Keep file names predictable
A simple naming scheme prevents editing confusion. For example:
01-intro.wav
02-problem.wav
03-example-a.wav
04-conclusion.wav
If a section needs a revision, keep the section number and add a revision suffix until the final version is chosen. Consistent file names matter even more when several people are editing the project.
Normalize writing style across sections
Different writers often use punctuation differently. One person may write short direct sentences while another uses long clauses. Synthetic delivery will reflect those differences.
Before generation, edit the script into a shared voice. Use similar sentence length, punctuation habits, and terminology across sections. This creates more consistent rhythm before any audio setting is touched.
Use emotion selectively
Long narration usually needs a stable center. If every paragraph uses a different emotional style, the voice can feel like a collection of demos instead of one speaker.
Choose a default neutral or restrained style. Use stronger emotion only where the editorial structure calls for it, such as an opening hook, quoted dialogue, or closing message.
Manage timing at the section level
For video or slide-based content, note the available time for each section. Compare generated duration with the target. If the gap is large, rewrite. If the gap is small, pace adjustment may help.
This is better than forcing the entire project to one global speaking speed. Different sections have different information density.
Listen to joins, not just isolated clips
A clip can sound acceptable alone but awkward next to the previous one. When reviewing, listen across boundaries. Pay attention to sudden changes in loudness, speed, emotional energy, or silence.
Regenerate mismatched sections when necessary. Trying to hide every mismatch with editing can take longer than producing a better source clip.
Build checkpoints into the project
For a long project, approve the work in stages:
- reference voice approved;
- pronunciation list approved;
- first two minutes approved;
- first full section approved;
- remaining batches generated;
- final continuous listening review.
These checkpoints reduce the risk of discovering a fundamental problem at the end.
Do a final real-time listening pass
Waveforms do not reveal every problem. Listen through the final assembled audio at normal speed. Check transitions, missing sections, duplicate lines, pronunciation, emotional consistency, and whether the overall pace feels comfortable.
If the project includes factual or instructional material, compare the audio with the final script rather than relying on memory.
Keep the right to use the voice documented
Long-form production often has greater commercial value than a short demo. Make sure the permission to use the voice covers the full project, distribution channels, and intended duration of use.
If the project is updated later, confirm that the existing permission still applies.
A repeatable long-form system
The core system is straightforward:
clean script → divide by meaning → approve a baseline → maintain pronunciation rules → generate in batches → review joins → assemble → listen end to end.
The value of long-form TTS comes from reducing repeated recording work without giving up editorial control. Consistent inputs and staged review are what turn a cloned voice from a novelty into a production tool.