IndexTTS Online
Voice CloningExamplesPricing
Sign inGenerate speech

IndexTTS Online Guides

A Reproducible A/B Testing Method for AI Voice Cloning

Compare voice-cloning models, references, scripts, emotion, or pace with controlled inputs, useful listening criteria, repeat runs, and honest reporting without fabricated MOS.

August 23, 2026·IndexTTS Online Editorial Team

A voice-cloning comparison is difficult to interpret when several inputs change at the same time.

If version A uses one reference speaker, a short sentence, neutral emotion, and one model while version B uses a different reference, a longer script, slower pace, and another model, you can hear a difference but you cannot explain it.

A useful A/B test is designed around one question.

IndexTTS Online includes Voice Lab to make controlled model comparisons easier. The same method also works when testing reference recordings, text formatting, emotion presets, or speaking pace.

This guide describes a reproducible workflow and its limits.

Write the question before generating audio

Good test questions are narrow:

  • Which of these two references gives more stable speaker identity?
  • Does a period improve the pause compared with a comma?
  • Which model is more suitable for this English/Chinese sentence under the same input?
  • Does the slower pace preset improve comprehension for this paragraph?

Bad test questions are vague:

  • Which setup is best?
  • Which model sounds more human?
  • What settings make perfect audio?

A narrow question tells you which variable may change and what should remain fixed.

Define the control variables

Create a simple test record.

For a model A/B comparison, record:

  • reference file;
  • target text;
  • target language;
  • emotion mode;
  • pace;
  • model A;
  • model B;
  • date;
  • number of repeats.

Keep reference and target text identical on both sides.

If one model does not support a control or language that the other model supports, do not pretend the comparison is symmetric. Label the capability difference instead.

For example, the hosted IndexTTS2 path focuses on English and Chinese, while IndexTTS 2.5 exposes Chinese, English, Japanese, Spanish, and Arabic plus pace controls. A Japanese test is useful for 2.5 capability evaluation, but it is not a fair direct A/B against a path that does not expose Japanese.

Use a clean authorized reference

The reference should have one speaker, little background noise, no music, and a stable recording environment.

Do not start a model benchmark with a questionable reference. Otherwise you are measuring the interaction between model and poor source quality.

For public examples, document provenance and rights. Do not publish customer audio or a private user's reference as a benchmark without explicit permission.

Official upstream references can be useful as test inputs when their provenance is clear and their use is appropriate, but outputs generated by your own hosted workflow should be labeled separately from upstream samples.

Choose target text that answers the question

A test sentence should be difficult enough to reveal the property you care about but not overloaded with unrelated edge cases.

For a general identity/naturalness check, use ordinary connected speech.

For punctuation, use a sentence with one meaningful pause.

For numbers, use a short line containing the numeric form you are testing.

For multilingual generation, use a natural sentence written by or reviewed by someone competent in the target language.

Do not put names, acronyms, dates, emotional dialogue, and code tokens into the same “general” test unless those are the actual target domain.

Run more than once

Generative output can vary.

One take can be unusually good or unusually bad. A single sample should not be treated as a universal result.

For a meaningful production decision, repeat the finalists. You do not need a huge research study for every project, but two or three runs can reveal instability that one run hides.

Save the outputs with clear names:

  • case01-modelA-run1.wav
  • case01-modelA-run2.wav
  • case01-modelB-run1.wav
  • case01-modelB-run2.wav

This makes review and later auditing easier.

Listen in a controlled order

If possible, hide which model produced which clip during the first listening pass.

Labels can bias perception.

A simple blind review can rename files A and B temporarily. After notes are written, reveal the mapping.

For team reviews, ask people to write their choice before discussing it together. Group discussion can cause later listeners to follow the first confident opinion.

Separate listening dimensions

“Sounds better” hides several different judgments.

Use dimensions such as:

Speaker identity

Does the output resemble the authorized reference speaker throughout the clip?

Intelligibility

Are the words understandable without reading the text?

Pronunciation

Are names, numbers, abbreviations, and target-language words pronounced as intended?

Prosody

Are stress, rhythm, pauses, and sentence endings natural for the content?

Stability

Does the voice remain consistent from beginning to end and across repeat generations?

Artifacts

Are there metallic sounds, glitches, repeated phonemes, clipped words, odd breaths, or background-like noise?

Task fit

Even if two outputs are both high quality, which better fits the actual task: narration, dialogue, localization, support content, or another role?

These categories produce actionable notes.

Use rating scales carefully

A five-point internal scale can help a production team compare several clips. It does not automatically become a scientific MOS.

Mean Opinion Score normally implies a defined listening-test methodology and aggregation across listeners. Do not publish a small internal rating as “MOS” simply because it uses numbers from one to five.

If you have not run a proper listener study, call the result what it is:

  • internal review score;
  • reviewer preference;
  • production note;
  • sample comparison.

Honest labeling protects users from false precision.

Do not invent latency benchmarks

Latency depends on more than the model:

  • queue state;
  • provider infrastructure;
  • network path;
  • reference upload;
  • text length;
  • cold starts;
  • retries;
  • geographic location.

If you publish latency, define exactly what you measured: start point, end point, sample size, region, text length, date, and whether failures were excluded.

Without that protocol, “Model X is 30% faster” is not a trustworthy claim.

Upstream technical reports can be cited as upstream evidence, but they should not be presented as measurements from your hosted product.

Keep upstream evidence and first-party measurements separate

A benchmark page can contain three different kinds of information:

  1. hosted product facts: which controls and languages the site exposes;
  2. upstream published findings: results reported by the model authors;
  3. first-party measurements: tests actually run by your hosted service.

Label them separately.

This prevents a common credibility problem where an upstream number appears next to a hosted product and users assume the site reproduced it.

IndexTTS Online's benchmark approach intentionally keeps those categories distinct.

Record the exact test date

Models, providers, infrastructure, and product controls can change.

A reproducible result needs a date and enough configuration information to repeat it.

For public tests, include:

  • test date;
  • model identifiers used;
  • reference provenance;
  • target text;
  • language;
  • controls;
  • number of runs;
  • evaluation method.

If a result is later outdated, the page can be updated without pretending the old test represented all future versions.

Test the actual production task

A benchmark sentence is useful, but a project decision should include a representative task.

If you are producing an audiobook, test a paragraph with realistic sentence length.

If you are producing short product UI prompts, test short instructions.

If you are localizing a video into Spanish, test real Spanish narration rather than an English demo sentence.

The best model in a generic sample may not be the best setup for a particular workflow.

Example: model A/B

A controlled model comparison might look like:

Question: Which model gives the more stable result for this English narrator reference?

Fixed: same voice_01.wav reference, same target sentence, neutral/follow-reference emotion, normal applicable pace.

Changed: model only.

Runs: three per model.

Review: identity, pronunciation, prosody, artifacts, preference.

This can support a practical recommendation for that case. It does not prove that one model is universally superior.

Example: reference A/B

Question: Which reference recording is cleaner for repeated narration?

Fixed: model, target text, language, emotion, pace.

Changed: reference file only.

Compare the outputs, then repeat the winner on a second target sentence.

If the same reference remains better, save it as the project baseline.

A reusable benchmark template

For each case, save:

Case ID
Question
Reference source and rights/provenance
Target text
Language
Model/settings
Variable under test
Number of runs
Listening criteria
Reviewer notes
Decision
Test date

This can live in a spreadsheet, Markdown file, or internal database.

The format matters less than consistent documentation.

When a test is inconclusive

Sometimes A and B are both good, or repeat runs reverse the initial preference.

That is a result.

Do not force a winner. Record:

No consistent audible advantage under this test.

Then choose based on product fit, supported languages, controls, cost, or workflow simplicity.

An honest inconclusive result is more useful than a fabricated certainty.

Final testing rules

A trustworthy voice A/B test follows a few principles:

  1. ask one narrow question;
  2. change one variable;
  3. keep reference and text controlled;
  4. repeat important runs;
  5. judge several listening dimensions;
  6. distinguish preference notes from standardized MOS;
  7. document provenance and dates;
  8. separate upstream findings from hosted measurements;
  9. do not force unsupported language/model comparisons;
  10. publish limitations with the result.

A/B testing is valuable because it reduces guessing. Its credibility comes from restraint: control what you can, label what you measured, and do not claim more than the experiment supports.

IndexTTS Online

A browser-based IndexTTS voice product for voice cloning, Saved Voices, History and hosted text-to-speech.

IndexTTS Online is an independent third-party service and is not affiliated with or endorsed by Bilibili or the official IndexTTS team.

support@indextts.online

Product

  • Voice Cloning
  • Examples
  • How It Works
  • Pricing

Resources

  • Voice Lab
  • Benchmark
  • Samples
  • Guides
  • About
  • Contact

Models

  • IndexTTS 2.5
  • IndexTTS2

Legal

  • Privacy Policy
  • Terms of Service
  • Acceptable Use
  • Cookie Policy
© 2026 IndexTTS Online