
Three weeks ago a client emailed me at 11pm asking if we could change one sentence in a 90 second product explainer that had already gone through two rounds of approval. The voiceover actor was on a flight to Lisbon and unreachable for two days. The old answer was “we wait” or “we re-record the whole thing and eat the cost.” Instead I opened ElevenLabs, cloned the line from the actor’s existing 90 seconds of audio, dropped it into the timeline, and had a revised cut in the client’s inbox before midnight. Nobody on their side noticed the swap. That is the moment AI voice cloning stopped being a party trick for me and became a real part of how I deliver video work.
I have been producing UGC ads and product explainer videos for tech clients for a while now, and voiceover has always been the most fragile part of the pipeline. Actors get sick, scripts change after a client sees the first cut, and re-recording a single line for a 30 second ad can cost more than the original session. So over the past two months I actually put four AI voice cloning tools through real, paid client deliverables instead of just demoing them. Here is what happened, with real numbers, not vendor marketing.
The four tools I actually shipped client work with
I tested ElevenLabs, Descript’s Overdub, HeyGen’s voice cloning, and WellSaid Labs. All four can clone a voice from a short sample. Only two of them were good enough that I would put the output in front of a paying client without a disclaimer.
- ElevenLabs won on raw quality. Their cloning model handles breath placement and mid-sentence pacing changes better than anything else I tried, and it took under 60 seconds of source audio to get a usable clone. Creator tier runs about $22 a month for roughly 100,000 credits, which covers around 100 minutes of generated audio. For a shop doing weekly 30 to 90 second ads, that tier is genuinely enough.
- Descript Overdub is the most convenient because it lives inside the same editor where I am already cutting the video and cleaning transcripts with Whisper. The tradeoff is voice quality. It is noticeably flatter on longer sentences and struggles with product names that have unusual pronunciation, which is exactly what tech client scripts are full of.
- HeyGen’s voice cloning is built to pair with their avatar generator, and honestly the voice alone is decent but not best in class. Where it earns its subscription is when a client wants a talking-head explainer without booking an on-camera presenter at all.
- WellSaid Labs produced the most consistent studio-grade tone across a full script, but the onboarding is built for enterprise brand voice libraries, not a solo shop cloning one actor’s voice for one project. Pricing is quote-based and slower to get access to than the other three.
What “good enough to ship” actually means
The bar is not “does it sound human.” Almost all four tools clear that bar on a single clean sentence. The bar is whether it survives being cut next to the actor’s real recorded audio in the same video, because that is what actually happens. Nine times out of ten I am not generating a whole voiceover from scratch, I am patching one or two lines into an otherwise human-recorded track. That means the clone has to match breath timing, energy, and mic tone closely enough that a client watching on their phone does not clock the seam. ElevenLabs is the only one of the four that consistently passed that specific test for me. The other three were fine standing alone but created an audible texture shift when spliced next to real audio.
The real cost math
Here is what actually changed for my business. A standard re-record session with a voice actor for a script fix, even a short one, has historically cost me between $75 and $150 depending on the actor’s minimum booking and how fast I need the turnaround. A cloned line inside my existing ElevenLabs subscription costs functionally nothing beyond the $22 monthly credit allowance I am already paying for. Over the six client jobs I used cloning on in the past two months, I estimate I saved close to $500 in re-record fees and, more importantly, saved two to five days of turnaround time each time, because I was not waiting on an actor’s schedule.
Where it still falls apart
I want to be honest about the limits because most content about this tool category oversells it. Cloning still struggles with genuine emotional range. Excitement, sarcasm, and a real pause for comedic timing all come out flatter than a human take. I would not clone an entire spot that leans on delivery and performance, only ones that are informational and product-focused, which happens to be most of the tech UGC and explainer work I do. Accents outside the training sample also degrade fast. And there is a consent issue that does not get talked about enough: I only clone voices I have an explicit agreement with the actor to use this way, in writing, and I never clone a client’s own voice or a public figure’s voice without the same written sign-off. That is not optional, it is the difference between a useful production shortcut and a legal problem.
My actual workflow now
I run the same Whisper-based transcript pipeline I built for captioning to generate a clean text version of the script first, then feed the exact line that needs fixing into ElevenLabs rather than regenerating a whole paragraph. Smaller generations are more controllable and easier to match to the surrounding audio’s volume and pacing. I always render the cloned line at a slightly lower processing setting than default and add a touch of room tone underneath it so it does not sound artificially clean next to audio that was recorded in a real space. That one trick fixes most of the “too perfect” tell that gives AI voice away.
The takeaways
- ElevenLabs is the one I actually kept using on paid client work, mainly for splicing single lines into human-recorded voiceovers
- Descript Overdub wins on convenience if you are already editing there, but quality drops on longer or unusual sentences
- HeyGen makes more sense when you need an avatar and voice together, not voice alone
- Get written consent before cloning anyone’s voice, every time, no exceptions
- Use cloning for informational lines and fixes, not performance-heavy or emotional delivery
- Add slight room tone under a cloned line so it blends with real recorded audio instead of sounding pasted in
Frequently asked questions
Is AI voice cloning good enough to replace a voice actor entirely?
Not yet, for anything that leans on emotional performance or comedic timing. It is genuinely good enough to replace a voice actor for short, informational fixes and patches inside an otherwise human-recorded voiceover, which is the most common real-world use case in ad and explainer production.
How much audio do I need to clone a voice well?
With ElevenLabs I got usable results from under 60 seconds of clean source audio, though quality improves with two to three minutes of varied sentence structure and pacing from the original recording.
Is it legal to clone someone’s voice for commercial video work?
You need explicit written consent from the person whose voice you are cloning before using it commercially. Voice rights laws are tightening in 2026, and platforms increasingly require proof of consent before allowing commercial use of a cloned voice, so treat this as a contract step, not an afterthought.
Related reading
- Why I Ditched CapCut’s Auto-Captions and Built My Own With Whisper
- I Tested AI Video Generators for UGC Ads So You Don’t Have To
- The Real Monthly Cost of an AI Subscription Stack (And How I Cut Mine by $768 a Year)
If you are building a product and need it turned into video people actually understand and act on, see how we do that.
Written by Osato Umweni, a designer and tech creator based in Lagos. More about me.


