Zero-shot reference cloning
Upload or record a short speaker reference and generate new speech with that voice identity.
IndexTTS2
Clone a voice from a short reference and generate English or Chinese speech online with the established IndexTTS2 path, including emotion controls, Saved Voices, History and WAV downloads.
IndexTTS Online Voice Cloning
OnlineAdd a short reference voice and the text you want it to speak.
Choose the IndexTTS model for this generation.
Add 3–30 seconds of clear, single-speaker audio.
Add a reference voice and text to get started.
WAV output · Signed-in generations can be viewed in History.
By generating, you confirm that you are 18 or older and that you own this reference voice or have permission to use it. Use is subject to our Terms and Acceptable Use Policy.
IndexTTS model samples
Use official source samples to compare IndexTTS2 with the latest multilingual IndexTTS 2.5 model before choosing a model for generation.
IndexTTS2 remains available for the established English and Chinese generation path.
Choose a sample
IndexTTS2
IndexTTS2 remains a first-class model in IndexTTS Online. It uses a short reference voice to guide speaker identity and delivery without requiring a separate training workflow.
Upload or record a short speaker reference and generate new speech with that voice identity.
The current hosted IndexTTS2 path supports English, Chinese and mixed English/Chinese text.
Follow the reference delivery or use the supported Happy, Sad, Angry and Calm presets.
Use Saved Voices, History and WAV downloads without tying a reference asset to one model version.
Reference guidance
A clean reference helps the model focus on speaker identity. Use WAV or MP3 with one speaker and minimal background noise; the current product accepts references around 3–30 seconds.
Model comparison
IndexTTS2 is the established English/Chinese generation path. IndexTTS 2.5 adds explicit five-language and pace controls. Both remain selectable inside one product.
IndexTTS 2.5Generate repeatable English or Chinese speech from a reusable reference voice.
Use the reference voice and emotion presets for narration, dialogue and creator content.
Save frequently used references and return to recent outputs from the same account.
Upload or record a short, clear single-speaker reference, or choose a saved voice after signing in.
Write the script and choose Follow reference or one of the available emotion presets.
Generate the speech, listen to the result, retry if needed and download WAV output.
IndexTTS Online is an independent third-party hosted service. The official project remains the source for upstream model research, code and technical details.
Yes. IndexTTS2 remains an available model in IndexTTS Online and is not removed when IndexTTS 2.5 is introduced.
The current hosted IndexTTS2 path supports English and Chinese text, including mixed English/Chinese input.
No. Saved Voices are shared reference assets and are not permanently tied to IndexTTS2 or IndexTTS 2.5.
Yes. You can follow the reference delivery or use the supported Happy, Sad, Angry and Calm presets.
Choose 2.5 when you need Japanese, Spanish or Arabic generation, an explicit language selector, or speaking-pace presets.