IndexTTS Online Guides
How to Choose Better Reference Audio for AI Voice Cloning
A quality checklist for reference recordings, with practical fixes for noise, echo, multiple speakers, clipping, and unnatural delivery.
Reference audio is the foundation of a voice-cloning workflow. A clean sample makes the model's job easier; a confused sample forces the model to guess which characteristics belong to the target voice and which belong to the room, microphone, background music, or other speakers.
The most useful rule is simple: optimize for clarity and consistency, not for dramatic production value.
What a strong reference sounds like
A strong reference contains one clearly audible speaker at a stable volume. The voice should be easy to understand without headphones, but the recording should also survive a closer headphone check without revealing clipping, heavy echo, or another speaker in the background.
Useful traits include:
- one speaker only;
- normal conversational delivery;
- stable microphone distance;
- low room noise;
- no music under the voice;
- no hard clipping;
- no telephone or radio effect;
- a few seconds of uninterrupted speech.
The target is not “studio perfect.” The target is “unambiguous.”
Why longer is not always better
A longer clip can contain more examples of the speaker's voice, but it also creates more opportunities for problems. A thirty-second clip with music, laughter, interruptions, and changing microphone distance may be less useful than a clean ten-second sentence.
When comparing samples, ask which clip gives the clearest evidence of the speaker's ordinary voice. If the answer is the shorter clip, use it.
Avoid mixed speakers
Multiple speakers are one of the most damaging reference problems because the model may receive contradictory identity cues. This can happen in obvious ways, such as an interview clip, or subtle ways, such as another person saying “yeah” in the background.
Listen from start to finish before uploading. If any second voice appears, trim around it or choose another recording.
Watch for room echo
Echo changes the apparent character of a voice. A speaker recorded in a tiled kitchen, stairwell, or large empty room can sound very different from the same person recorded close to a microphone in a quiet room.
Mild natural room tone is usually less problematic than obvious reverberation. If syllables seem to trail behind the speaker or consonants become smeared, record again in a softer environment. Curtains, carpets, furniture, and closer microphone placement often help more than aggressive software cleanup.
Do not over-process the sample
Noise reduction can help a truly noisy recording, but too much processing creates metallic or watery artifacts. Voice isolation tools can also remove parts of consonants or breaths that contribute to natural identity.
If you have the option to make a fresh recording, a clean original is usually preferable to a heavily repaired one. Use processing as a rescue tool, not as the default workflow.
Check clipping and loudness
A clipped waveform occurs when the recording level is too high and peaks are cut off. Audible clipping often sounds like harsh crackling on loud syllables. Once the original signal is clipped, reducing the volume later does not restore the missing shape.
Record at a comfortable level with some headroom. The speaker should not need to whisper to avoid clipping, and the file should not be so quiet that background noise becomes dominant when amplified.
Use natural speech rather than a performance
If the intended output is ordinary narration, use ordinary speech as the reference. A sample captured while the speaker is shouting, whispering, singing, crying, or doing a character voice can bias the perceived identity and style.
This does not mean emotion is forbidden. It means the reference should match the role you want the synthetic voice to perform. For a neutral narrator, use a neutral reference. For a character project, test whether a more expressive reference actually improves consistency rather than assuming it will.
Record with simple equipment
You do not need an expensive microphone. A modern phone or laptop microphone can produce a usable reference when the environment is quiet and the speaker is positioned consistently.
A practical setup is:
- move away from fans, traffic, and open windows;
- place the device at a stable distance;
- speak across the microphone rather than directly into it if plosives are strong;
- record two or three takes;
- select the cleanest take instead of the most energetic one.
Consistency matters more than brand names on hardware.
What should the speaker say?
Use a sentence that contains normal vowels and consonants and feels comfortable to speak. Avoid a reference made entirely of names, numbers, acronyms, or repeated short words. The model benefits from natural connected speech.
If you plan to generate English and Chinese content, you can test references in the language most representative of the speaker's natural delivery. For multilingual production, verify the final generated language separately rather than assuming the reference alone guarantees pronunciation quality.
A five-minute reference-audio test
Before using a sample for a large project, run a short controlled test. Generate the same two-sentence script with two candidate references. Keep the model and text unchanged. Compare:
- identity similarity;
- consonant clarity;
- stability across the sentence;
- background artifacts;
- naturalness of pauses.
If one reference consistently wins, save it as the project's baseline. This is much more informative than changing several settings at once.
When to replace the reference
Replace the reference when voice identity is unstable across otherwise simple sentences, when obvious environmental artifacts carry into the output, or when the clip contains mixed speakers. Do not replace it merely because one difficult word was pronounced badly; text formatting or model choice may be the better fix.
Reference-audio checklist
Before uploading, confirm: one authorized speaker, no music, little echo, no clipping, natural delivery, sensible volume, no accidental interruptions, and a clean start and end. Then test on a short sentence before committing to long-form generation.
A good reference reduces uncertainty. That is its real purpose. The cleaner and more representative the sample, the easier it becomes to diagnose every later part of the voice-cloning process.