Understanding and Working Around Voice Cloning Artifacts in Production
The Con Oneill Voice Problem
When you're working with AI voice synthesis or digital voice conversion in real projects, there's a persistent class of issues that comes up again and again. People in the industry sometimes refer to this cluster of artifacts as the Con Oneill Voice Problem, named after a vocal engineer who documented it extensively in early voice-cloning forums. The core of it is straightforward: synthetic or semi-synthetic voice systems struggle with certain phonetic transitions, especially around plosives, fricatives, and breath sounds. The output sounds technically correct but carries a subtle mechanical quality that listeners notice even if they can't articulate why. I've been dealing with this in post-production for years, mostly in indie game localization work where we use hybrid pipelines — a base TTS model refined with neural voice conversion. The problem shows up most obviously when the cloned voice handles long consonant chains or emotional inflection shifts. A sentence that should sound tired or frustrated comes out flat because the model smooths over the micro-timing variations that convey human affect.
What's Actually Happening Under the Hood
Voice cloning models typically work in two stages. First, a base model generates audio from text using a learned phoneme-to-spectrogram mapping. Second, a voice conversion layer applies the target speaker's vocal characteristics — timbre, pitch contour, formant structure — to that generated output. The Con Oneill Voice Problem arises primarily at the boundary between these stages. The base model doesn't truly understand prosody the way a human does, so it produces mechanically even timing. The voice conversion layer then paints the target voice on top of that robotic skeleton. You end up with a voice that sounds right in tone but moves wrong in time. The worst offenders are sentences containing clusters like "strengths," "twelfths," or "the sixth street." The model tends to either truncate the final consonant, add a micro-pause that wasn't in the source, or flatten the vowel quality. This is especially noticeable in Irish and Scottish English dialect work, which is likely why the original case studies used those as test cases. Those dialects carry a lot of glottal stops and consonant mutations that standard models aren't trained to handle gracefully.
Practical Workaround: The Three-Pass Method
Here's what actually works in practice. Don't run your text through the pipeline once and accept the result. Use a three-pass approach: Pass one: Run the full script through your base TTS model with a neutral voice. Export the phoneme timestamps. This gives you the timing skeleton. Pass two: Have your voice actor record the same lines. Don't aim for perfection — aim for natural rhythm. You just need reference audio that captures the emotional pacing and breath patterns. This takes about 45 minutes to an hour for a typical five-minute scene.
Get the Full Details

Pass three: Feed both the TTS output and the reference recording into your voice conversion model. Most modern tools (RVC, So-VITS-SVC, or commercial alternatives like ElevenLabs) allow you to provide a reference for prosody transfer alongside the timbre reference. The model then applies the actor's timing and breath patterns to the cloned voice output, rather than just their tonal quality. This usually cuts the rework time from several hours of manual editing down to maybe 20 or 30 minutes of fine-tuning. The tradeoff is that you need a clean reference recording, which means you need either a voice actor or a high-quality source track to begin with.
A Specific Edge Case I Ran Into
Last year I was working on a project where the cloned character needed to deliver a monologue that started very quiet and built to a yell. The model handled the volume shift fine, but it completely lost the gravelly texture of the voice during the loud passages. It sounded clean and bright instead of strained. I spent about two hours trying different model weights and conditioning parameters before I found the actual fix: the issue was that the vocoder's spectral masking was too aggressive during high-energy segments. I had to disable the built-in denoising in the vocoder stage and apply a separate noise gate afterward with a much higher threshold. That preserved the distorted vocal fry without introducing the model's characteristic cleaning artifact. It's a niche problem but it comes up whenever you need dynamic range in a single clip. Most tutorials don't cover it because most people are just generating calm dialogue.
Counter-Intuitive Things Beginners Miss
One thing that consistently surprises people: more training data doesn't always solve the Con Oneill Voice Problem. I've seen projects with 40 hours of reference audio still produce the same mechanical phrasing. What actually matters is the diversity of the reference material. A model trained on three hours of emotionally varied speech — quiet moments, shouting, whispering, laughing, crying — will generally outperform one trained on 20 hours of evenly-toned narration. The model needs to learn the mapping between emotional state and vocal production, not just the mapping between text and sound. Another thing: punctuation in your input text matters more than most people realize. Most TTS systems parse punctuation for pause placement, but they do it inconsistently. A comma might register as a 150ms pause or a 400ms pause depending on the model and the language. If you're getting unnatural phrasing, try replacing commas with ellipses or breaking long sentences into shorter ones with periods. It sounds counterproductive but it often produces noticeably better results than fighting the parser.

Where This Approach Completely Fails
Be honest about the limitations. This three-pass method works well for monologue-heavy content and dialogue with clear emotional arcs. It breaks down for rapid-fire overlapping conversation, where timing precision matters at the millisecond level. It also doesn't help much with languages the base model wasn't trained on well. If you're working in a low-resource language, the phoneme alignment will be garbage regardless of how good your voice conversion is, and you'll spend more time fixing artifacts than you would have saving by using the pipeline in the first place. For those cases, your best option is still full human voice acting with post-processing to match the desired character. No current pipeline handles it reliably, and pretending otherwise just costs you time and money down the line.
Tools Worth Looking At
RVC (Retrieval-based Voice Conversion) is the most common free option. It's well-documented and has active development. The main downside is that inference can be slow on CPU, and the default model settings need tweaking for the edge cases I mentioned above. So-VITS-SVC is another free alternative with slightly better prosody handling out of the box, though it's less actively maintained. On the commercial side, ElevenLabs offers voice conversion with prosody transfer, and it's noticeably better at handling emotional variation, but you pay per character and the pricing scales poorly for long-form projects. If you're doing this professionally, budget for a voice actor reference session even if you're not planning to use their performance directly. The reference audio is the single biggest factor in whether your output sounds human or artificial, and it takes far less time than you'd expect to produce usable material.