IndexTTS Online Guides
How to Test Reference Audio Length for Voice Cloning
A controlled method for comparing short and longer voice references without inventing a universal “best length,” including a reusable test matrix and listening criteria.
“How long should my voice-cloning reference be?” sounds like a question with one numeric answer. In practice, length is only one property of the recording.
A clean ten-second clip can be more useful than a longer clip containing background music, multiple speakers, changing microphone distance, or several emotional styles. A longer reference can contain more speech evidence, but it also creates more opportunities for conflicting information.
IndexTTS Online accepts short references in an approximately 3–30 second range. Within that range, the most useful approach is to test candidate lengths under controlled conditions instead of assuming that the maximum duration is automatically best.
This guide shows how.
Do not compare different recordings first
If you want to understand the effect of duration, start from one clean master recording.
For example, record 30 seconds of one speaker in the same room, on the same microphone, with stable delivery. Then create several excerpts from that master.
Possible test set:
- 5 seconds;
- 10 seconds;
- 20 seconds;
- 30 seconds.
Each excerpt should begin and end cleanly and contain complete speech. Avoid creating a five-second sample that happens to contain an awkward pause while the ten-second sample contains a perfect sentence.
The closer the excerpts are in recording conditions, the more meaningful the length comparison becomes.
Keep the target text identical
Use the same generation text for every reference.
A useful target contains:
- ordinary connected speech;
- a mix of short and longer words;
- at least two sentences;
- no obscure names or difficult numbers;
- punctuation that you already know works.
Do not use a target designed to stress pronunciation while you are testing reference length. That adds another variable.
The goal is to make differences in identity and stability easier to hear.
Keep model and controls fixed
If the model changes between runs, the test no longer isolates reference length.
Keep constant:
- model;
- target language;
- target text;
- emotion mode;
- pace;
- account/product path.
If you later want to compare IndexTTS2 and IndexTTS 2.5, repeat the whole length test separately for each model or use the same chosen reference in Voice Lab.
One experiment should answer one question.
What to listen for
Do not reduce the result to “I like this one.”
Use specific criteria.
Speaker identity
Does the generated voice consistently resemble the authorized reference speaker? Listen across the entire output, not only the first few words.
Pronunciation clarity
Are consonants and word boundaries clean? Reference duration may not be the real cause of pronunciation errors, but a bad reference can make the whole output less stable.
Prosody
Does the rhythm feel natural? Is the delivery unexpectedly rushed, flat, or strained?
Artifacts
Listen for buzzing, metallic texture, duplicated sounds, strange breaths, or background characteristics that seem to come from the source recording.
Stability across repeats
Generate more than once when a result matters. Generative systems can vary. A reference that produces one excellent take and several unstable takes may be less useful than a reference that performs consistently.
Use a simple scoring sheet without pretending it is MOS
You can keep private production notes such as:
| Reference | Identity | Clarity | Stability | Notes |
|---|---|---|---|---|
| 5 s | good | good | mixed | misses some voice character |
| 10 s | strong | strong | strong | baseline |
| 20 s | strong | strong | strong | little difference |
| 30 s | mixed | good | mixed | includes room change |
These are subjective project notes, not a standardized Mean Opinion Score.
Do not publish them as scientific MOS unless you actually run an appropriate listener study with a documented protocol. A small internal comparison is useful for production, but it should be labeled honestly.
Why the shortest usable clip can be attractive
A shorter reference is faster to inspect, easier to trim, easier to store, and less likely to contain accidental noise or another speaker.
If a clean short reference already gives stable identity and natural output, adding more duration may not create meaningful value.
This is especially relevant when users repeatedly upload references or when a project has many speakers.
Shorter is not automatically better. It is simply worth testing.
Why a longer clip can help
A longer clean reference can contain more examples of the speaker's vowels, consonants, pitch range, and normal rhythm.
If a five-second clip contains only one short phrase, it may underrepresent the speaker. A longer sample can provide a more representative baseline.
The keyword is clean. Additional duration that introduces music, laughter, another speaker, or a microphone change is not free information.
Content diversity matters more than silence
Ten seconds of actual clear speech is not the same as a ten-second file containing five seconds of silence.
When comparing lengths, track effective speech content. Remove excessive silence and do not count it as useful reference duration.
Likewise, a longer clip made of repeated words may provide less variety than a shorter natural sentence.
Think in terms of representative speech, not file length alone.
Keep delivery consistent within the reference
A clip that starts as calm narration and ends with shouting creates mixed style evidence.
If your intended output is neutral narration, choose a reference that stays reasonably neutral. If you are testing expressive delivery, make that a separate experiment.
Reference length and emotional variety should not be changed at the same time.
Avoid combining clips from different environments
It can be tempting to concatenate several short recordings to create a longer reference.
That can introduce:
- different rooms;
- different microphones;
- different gain levels;
- different compression;
- different speaking styles.
For a duration test, use one continuous recording where possible. If you must combine clips, normalize the production conditions first and treat the combined result as a separate reference strategy rather than a pure length test.
A practical four-run experiment
Here is a compact workflow:
- Record one clean 30-second master.
- Export 5-, 10-, 20-, and 30-second versions.
- Use one fixed target paragraph.
- Keep the model and all controls unchanged.
- Generate one output for each reference.
- Listen in random order if possible.
- Mark identity, clarity, artifacts, and overall stability.
- Repeat the top two candidates once more.
- Save the shortest candidate that remains reliably good for your project.
This creates a decision you can explain and reproduce.
Test the chosen reference on a second script
A reference can perform well on one target sentence by coincidence.
Before making it the project baseline, test it on a second script with different vocabulary and sentence rhythm.
If the same reference remains stable, confidence increases. If performance collapses, the problem may involve text pronunciation or model behavior rather than duration alone.
Do not keep extending the reference automatically. Diagnose the failure.
When to stop testing
Stop when additional duration no longer gives a meaningful production improvement.
The purpose of the experiment is not to find a mathematically perfect duration. It is to find a reliable reference that is easy to reuse.
Once a candidate works across several ordinary scripts, save it and keep the source file stable. Constantly changing the reference makes every later comparison harder.
How this relates to Voice Lab
Voice Lab is designed to compare IndexTTS2 and IndexTTS 2.5 with controlled inputs.
For a fair model comparison, choose the reference first. Then use exactly the same reference and target text on both sides.
If you change reference length at the same time as the model, you cannot tell whether the audible difference came from the model or the source.
Reference-length testing should therefore happen before, or separately from, model A/B testing.
Final principle
There is no useful universal claim that “30 seconds is always best” or “5 seconds is enough for everyone.”
The better rule is:
Use the cleanest, most representative reference that produces stable results for the intended project, and prove that choice with a controlled test.
That approach is slower than repeating a single magic number, but it produces information you can actually trust.