How much audio do you need to clone your voice?
About ten seconds. The interesting question is which ten seconds, and why most bad clones are a recording problem rather than a model problem.
Rauno Oidram · 3 August 2026
Short answer: modern voice cloning needs roughly ten seconds of clean speech. Not an hour, not the twenty minutes older tools asked for. dub44 asks you to record thirty to sixty seconds, and then uses about ten of them.
So why record longer than the model needs?
Because not every ten seconds is equally useful. A stretch where you pause, tail off, or say three words teaches the model less than one where you speak continuously with normal energy and range.
Recording longer is not about volume of data. It is about giving the system something to choose from. dub44 scans the recording and picks the liveliest continuous window automatically, so a minute of speech usually contains one genuinely good ten seconds — and you never have to identify it yourself.
Record in the language you will publish in
This is the mistake that surprises people most. A clone copies pronunciation, not just timbre. If you record in Estonian and then generate English, you get your voice borrowing someone else's English — because the model never heard you say an English word.
Record in the language you intend to publish in. An English sample is what teaches the system your English: your vowels, your rhythm, your accent.
The mistake that ruins most clones
Not the microphone. Not the room. Level.
When I was building this, my own clones sounded generic and I could not work out why. The recordings were arriving at about −39 LUFS, roughly a twentieth of the level they should have been. The model was being asked to learn a voice from something close to silence, and it did what you would expect: it filled the gaps with an average voice.
Normalising the sample before analysis fixed it completely. dub44 now does this for you, but the principle holds whatever tool you use: a quiet recording is worse than a short one. If your clone sounds like nobody in particular, check your input level before you blame the model.
What a good sample looks like
- Thirty to sixty seconds, one take, no edits.
- A quiet room. Soft furnishings beat a bare kitchen; no fan, no traffic.
- Speak the way you actually present — same energy, same pace.
- Real sentences, not a word list. Read something you have written.
- Close enough to the microphone to be loud, far enough to avoid popping.
- No music, nobody else talking, no background TV.
You only do this once. Two minutes of care here decides how everything you generate afterwards sounds.
What ten seconds cannot do
It cannot change your accent — nor should it, if people follow you rather than your topic. It cannot invent emotion your sample never contained: record flat and you will generate flat. And it cannot rescue a recording made in a bathroom, since the model learns the room along with the voice.
If you want to hear what a single one-take sample produces, the demo on the front page is exactly that: one recording, one generated clip, no cleanup between them.