IndexTTS Online
Voice CloningExamplesPricing
Sign inGenerate speech

IndexTTS Online Guides

How to Record Clean Voice-Cloning Reference Audio at Home

A room-by-room recording workflow for capturing clear 3–30 second reference audio with a phone, laptop, or simple microphone, without over-processing the voice.

August 23, 2026·IndexTTS Online Editorial Team

A useful voice-cloning reference does not require a studio. It requires a recording that makes one speaker easy to identify.

That distinction matters. Expensive equipment cannot rescue a room full of echo, a fan next to the microphone, or a second person talking in the background. A phone in a quiet, soft room can produce a better reference than a professional microphone used badly.

IndexTTS Online accepts short reference recordings, so the objective is to capture a few clean seconds rather than a polished podcast episode.

Choose the room before choosing the microphone

Room sound is usually the first problem to solve.

Hard, empty surfaces reflect speech. Kitchens, bathrooms, stairwells, and nearly empty rooms can create obvious reverberation. Rooms with curtains, rugs, upholstered furniture, bedding, and books usually absorb more reflections.

You do not need to build a vocal booth. Walk into two or three possible rooms, clap once, and listen to the decay. Then speak a normal sentence and record it on your phone.

Choose the room where consonants sound clear and the voice does not leave a long tail.

Remove continuous noise at the source

Noise reduction software is not the first step. Turn off the noise if you can.

Common home noise sources include:

  • fans and air conditioners;
  • open windows near traffic;
  • computer fans under heavy load;
  • refrigerators or other appliances;
  • desk vibrations;
  • distant television or music;
  • other people speaking.

Listen for ten seconds before recording. A noise that your brain ignores in daily life may become obvious in headphones.

If the room becomes uncomfortable with everything switched off, record a short take, then restore the environment. You only need a clean reference, not an hour of silence.

A phone is often good enough

Modern phones can capture clear speech when the microphone is not too far away.

Place the phone at a stable distance, roughly in the range you would use for a voice memo. Avoid holding it in your hand if movement creates handling noise. Put it on a stable surface with the microphone unobstructed.

Do not place it so close that every breath and plosive becomes dominant. Do not place it across the room where the recording becomes mostly room sound.

Record two or three distances and compare them. Keep the closest version that still sounds natural.

Laptop microphones require more attention to distance

A laptop microphone can work, but the normal laptop position may be farther from the mouth than ideal. It can also capture keyboard noise, fan noise, and desk reflections.

If using a laptop:

  1. stop typing during the take;
  2. close heavy tasks if the fan is loud;
  3. place the laptop in a quiet position;
  4. speak consistently toward the microphone area;
  5. compare the result against a phone recording.

Use the cleaner source. There is no prize for using the more complicated device.

An external microphone does not fix a bad room

A USB or XLR microphone gives you more control, but greater sensitivity can reveal more room sound.

Use the microphone according to its pickup pattern and intended distance. Keep the speaker in a consistent position. Monitor loud syllables for clipping.

If the microphone sounds “professional” but the room echo is stronger than the phone version, the phone may still be the better voice-cloning reference.

The model needs a clear identity signal more than it needs expensive frequency response.

Speak naturally

Use the voice you want the generated audio to resemble.

If the final project is neutral narration, record neutral conversational speech. If the reference is a shout, whisper, exaggerated character voice, or highly emotional performance, that style can influence what the model receives as evidence about the speaker.

A useful test sentence is two or three natural clauses with varied sounds and comfortable pacing.

Avoid a reference made entirely of:

  • numbers;
  • proper names;
  • abbreviations;
  • repeated short words;
  • a memorized slogan delivered unnaturally.

Natural connected speech gives the model a more representative sample.

Keep one speaker only

A second voice can be obvious, like an interviewer, or tiny, like someone saying “okay” from another room.

Listen to the full reference before upload. If another speaker appears, record again or trim the clip so only the authorized target speaker remains.

This is both a quality and a rights issue. Only use a voice you own or have permission to use.

Record several short takes

Do not try to make the first take perfect.

Record three versions:

  • normal delivery;
  • slightly slower but still natural;
  • another normal take after a short pause.

Then compare them on headphones.

You are listening for clarity and consistency, not charisma. Choose the take with the least echo, least background noise, stable volume, and most representative voice.

Watch for clipping

Clipping happens when the input level exceeds what the recording system can capture. Loud peaks become flattened and can sound harsh or distorted.

If a loud word crackles, move slightly farther from the microphone or reduce input gain when that control is available.

Lowering the volume after a clipped recording does not reconstruct the lost waveform. A fresh take is better.

At the other extreme, a very quiet recording can require amplification that also raises background noise.

Aim for a healthy, comfortable level with headroom.

Leave clean edges

A reference should not begin halfway through a consonant or end abruptly during a word.

Leave a small amount of room tone before the first word and after the final word, then trim excessive silence if necessary.

Avoid loud clicks, keyboard taps, chair movement, or handling noise at the start and end.

Clean boundaries make the file easier to inspect and reuse.

Do not add background music

A music bed may sound good in a finished video, but it is poor reference material.

The reference is an input to a voice model. Background music creates another signal that the model does not need.

Record the dry voice first. Add music later in the editing stage after the synthetic voice has been generated.

The same applies to sound effects, reverb, radio filters, and strong compression.

Be conservative with noise reduction

If the recording has a small amount of steady noise, a light cleanup may help. Aggressive voice isolation can create metallic edges, missing consonants, or watery artifacts.

Whenever possible, compare:

  • the original cleanest take;
  • a lightly processed version.

Use the one that preserves the natural voice better.

If both are poor, record again. Ten minutes spent improving the source can save far more time than repeatedly generating from a damaged reference.

Export a simple compatible file

Keep the workflow simple. WAV or MP3 are practical formats for a short reference when supported by the product.

Do not repeatedly transcode the same clip through several lossy formats. Each conversion can add unnecessary degradation.

Keep a master copy of the chosen reference so future generations use the same source. Renaming it clearly helps:

narrator-reference-clean-01.wav

A stable reference makes later A/B tests meaningful.

Run a two-sentence validation

Before saving the reference as your baseline, generate a short script with ordinary speech.

Listen for:

  • whether the speaker identity is stable;
  • whether consonants remain clear;
  • whether room echo appears in the output;
  • whether the voice becomes buzzy or metallic;
  • whether the delivery is unusually strained.

If the result is poor, test a second reference with the same text and model. Do not change several variables.

This tells you whether the recording itself is the likely problem.

A home recording checklist

Use this before uploading:

  1. quiet room with soft surfaces where possible;
  2. no fan, music, television, or nearby conversation;
  3. one authorized speaker only;
  4. stable microphone distance;
  5. natural conversational delivery;
  6. no obvious clipping;
  7. clear start and end;
  8. no unnecessary effects;
  9. several takes recorded;
  10. best take tested on a short generation.

The goal is not studio perfection. The goal is a reference that gives the model one clear answer to the question: “What does this speaker normally sound like?”

That is achievable at home with simple equipment and careful recording choices.

IndexTTS Online

A browser-based IndexTTS voice product for voice cloning, Saved Voices, History and hosted text-to-speech.

IndexTTS Online is an independent third-party service and is not affiliated with or endorsed by Bilibili or the official IndexTTS team.

support@indextts.online

Product

  • Voice Cloning
  • Examples
  • How It Works
  • Pricing

Resources

  • Voice Lab
  • Benchmark
  • Samples
  • Guides
  • About
  • Contact

Models

  • IndexTTS 2.5
  • IndexTTS2

Legal

  • Privacy Policy
  • Terms of Service
  • Acceptable Use
  • Cookie Policy
© 2026 IndexTTS Online