What This Is Actually About

People keep asking this because they heard a voice comparison clip go around social media. Someone ran a voice synthesis test and it sounded like Morgan Freeman narrating a Hendrix track, or vice versa. The question itself is a bit of a loop, but the technology behind it is real enough that I thought I should explain how it actually works rather than just shrugging it off. This isn't a real identity question. It's about voice cloning and deepfake audio tools that have been around since roughly 2018 when papers like WaveGlow and DeepVoice came out, then really took off when open-source models like RVC (Retrieval-based Voice Conversion) showed up on GitHub around 2023. The idea is simple enough: you take a clean reference recording of one person speaking or singing, train a model on it for about 20 to 40 minutes on a decent GPU, and then feed it audio from a completely different source. The output maps the timbre and delivery characteristics of the reference onto the target audio. I spent a few weekends messing around with this after seeing the Freeman-Hendrix comparisons pop up. The short answer to whether Morgan Freeman is Jimi Hendrix is no, obviously. But the longer answer involves understanding why the comparison keeps resurfacing and what it actually takes to produce that kind of result consistently.

How the Process Actually Works

Start with a clean reference dataset. This is where most people mess up. You need maybe 10 to 30 minutes of clear audio with minimal background noise, music, or reverb. For a voice like Morgan Freeman's, you'd pull from interviews or narration work where he's speaking plainly. For Jimi Hendrix, you'd need isolated vocal tracks, which is harder since most of his recordings have guitars and production layered in. That isolation step alone can add hours to the process. I found that extracting stems using Demucs or similar source separation tools helped a lot. Getting a clean vocal stem from a Hendrix recording usually means running it through multiple passes and then manually cleaning up the artifacts around the transients. Guitar sounds bleed into the vocal frequency range in a way that's annoying to filter out cleanly. If you skip this step and just use a full mix as your reference, the model learns the guitar tones too, and your output sounds muddy regardless of what you feed it later.

Training and Conversion

RVC v2 is the model most people use now. It's fast, relatively easy to set up, and produces results that are surprisingly good for casual use. You point it at your cleaned reference files, set the index ratio somewhere around 0.7 to 0.9 depending on how clean your source is, and let it train. On an RTX 3090, a standard model trains in about 20 to 40 minutes. Older GPUs take significantly longer and sometimes run out of memory on larger datasets. The conversion step is where the actual magic happens. You run your target audio through the model, and it outputs a new waveform that carries the reference voice's characteristics. Pitch correction matters here. If you're converting speech to speech, you generally leave the pitch alone. If you're doing something like making a spoken voice sing, you need to preprocess the target audio to get the notes roughly right before conversion, or the output will soundstrained and unrealistic. I hit a specific problem early on where the converted audio had this weird robotic artifact on sustained notes. Turns out it was the index threshold being too low. Bumping it from 0.5 up to 0.85 fixed it almost immediately. The index file controls how much the model relies on the reference versus the original audio content, and getting that balance right depends heavily on your source quality. Noisy references need a higher index value to compensate, but too high and you lose all the original delivery nuances.

Get the Full Details

Why Morgan Freeman wore glove on his left hand at the Oscars
Why Morgan Freeman wore glove on his left hand at the Oscars

Download Links and Tools

The main RVC repository is on GitHub under the name RVC-Project. There's a well-maintained fork called ap497's RVC-WebUI that bundles everything into a single installable package with a browser interface. It's the most straightforward option for people who don't want to deal with command-line setup. You can also find pre-trained voice models on Hugging Face and dedicated community servers, though I'd caution against using unverified models for anything beyond personal experimentation. There have been cases where models were trained on copyrighted material and then redistributed without permission. Beyond RVC, there's VoiceRack and OpenVoice, which take a slightly different approach. OpenVoice uses a reference clip to clone a voice without any training, which is faster but generally less accurate than a proper training pass. It works well for quick demos but falls apart if you need consistent quality across a longer piece of audio.

What This Can and Can't Do

The technology is good enough that casual listeners can be genuinely fooled by a short clip. I've had friends listen to converted audio and swear they recognized the voice before I told them what was actually going on. But it breaks down under scrutiny pretty quickly. Long-form content reveals inconsistencies in breathing patterns, consonant articulation, and emotional delivery that don't quite match the reference material. The model copies the timbre well but struggles with the prosody and natural variation that comes from actual speaking or singing. There's also a legal and ethical dimension that can't be ignored. Using someone's voice without their consent, especially for public distribution, is a growing legal gray area. The likeness rights framework in the US doesn't clearly cover synthetic voice replication yet, and different countries handle it differently. Even within a single country, the rules can shift depending on whether the content is commercial or non-commercial. I stick to personal experiments and don't distribute converted audio publicly, partly because it's the right thing to do and partly because the legal exposure isn't worth the hassle. If you're just curious about the technology and want to try it yourself, the learning curve is shallow enough that you can get a working model up and running in an afternoon. The results won't win any awards, but they're convincing enough to understand why these kinds of comparisons keep circulating online. The underlying principle is straightforward, even if the execution requires patience and decent hardware.