Why I built dub44
I kept not making videos. Not because I had nothing to say — because I did not want to hear myself say it in English.
Rauno Oidram · 3 August 2026
English is not my first language. I write it fine. I read it all day. But recording it is a different thing entirely: you hear the vowel you got wrong, start again, get it wrong differently, and after the fifth take you are not thinking about the content at all. You are thinking about your mouth.
So the video does not get made. Not consciously — you just find something more urgent to do, week after week, and the channel stays where it is.
The obvious fix is the wrong one
The AI voices everyone reaches for solve this by replacing you. You pick a confident American narrator, paste your script, and get something that sounds professional and belongs to nobody. It is competent and it is invisible.
That is a fine solution if your voice is not part of what you are selling. But if people follow you rather than your topic, swapping yourself for a stock voice removes the reason they turned up. I did not want to sound like a documentary. I wanted to sound like me, without the forty minutes of takes.
Keeping the accent is the whole point
dub44 clones your voice and keeps everything about it — including the accent. That surprises people. Most tools in this space quietly promise the opposite: neutralise, smooth, sound native.
I think that framing is wrong, and not only for sentimental reasons. Your accent is not a defect in your English; it is evidence of a second language. The problem was never how you sound. It was that recording is slow, tiring and self-conscious, and that is a workflow problem, not a personal one.
So the product fixes the workflow and leaves the voice alone. You record once — thirty to sixty seconds, one take, in a quiet room. After that you type, and it is your voice reading it back.
What I learned building it
That the hard part is not the model. The hard part is everything around it: making a browser recording usable, choosing which ten seconds of a sample to learn from, and — the thing that nearly sank it — discovering that my recordings were arriving twenty times too quiet, so the model was learning a voice from something close to silence.
I also learned that I would not trust this product if someone else had built it. Handing a company a recording of your voice deserves suspicion. So: your audio sits encrypted in Frankfurt, it is never used to train anything, generated clips are deleted after 90 days, and every file carries an inaudible watermark marking it as synthetic. Cloning someone else's voice requires their permission, and you confirm that every single time.
Who it is for
People who write English well and publish in it — YouTubers, podcasters, course makers, anyone narrating their own work — whose first language is something else. Estonians, Latvians, Lithuanians, Finns, Poles, Germans, Dutch. People who have a script ready and keep not recording it.
If that is you, the demo on the front page is me: one clip recorded, one generated. Decide from that whether it sounds like a person.