Of Shadows Vc Andrews: A Practical Look
I've spent more hours than I care to count tweaking voice conversion models, and Of Shadows Vc Andrews is one of the less documented but capable ones out there. It's not the flashiest tool in the ecosystem, but it does what it says and does it reasonably well if you know what you're doing. The basic premise is straightforward. You feed it a source vocal recording and a reference voice, and it converts the timbre of the source to match the reference while preserving the original speech content. The model architecture is built on a modified RVC (Retrieval-based Voice Conversion) foundation, which means you get decent results with relatively little training data compared to other approaches.
Of Shadows Vc Andrews
Getting started isn't complicated, but there are a few gotchas that trip up most people. First, you need a clean source audio file. Ideally no background noise, no reverb, no music underneath. The model will attempt to remove some of that stuff automatically, but it does a better job when the input is actually clean. I learned that the hard way on a project where I had to convert a recording that was done in a cheap home studio with obvious room echo. The output sounded like someone talking through a curtain, no matter what settings I used. You'll also want your source and reference audio to be roughly the same quality level. If your reference is a pristine studio recording and your source is phone audio, the model still tries its best but you're going to get artifacts in the conversion. The index retrieval step that RVC uses can sometimes compensate a little by pulling similar spectral features, but it has limits. Here's the workflow. You train an index on your reference voice using the standard RVC training pipeline. The default settings work fine for most cases, though I usually bump the feature extraction sampling rate to 48000Hz if the reference audio is high quality. It doesn't change the training time noticeably but it captures slightly more detail in the spectral representation. Then you run inference on your source file with whatever conversion parameters you want.
The key parameter most beginners miss is the pitch shift value. Getting that wrong is the fastest way to make a convincing voice conversion sound obviously synthetic. Set it to zero if the source and reference are the same gender and vocal range. Shift it by +12 if the source is male and the reference is female, or -12 in the reverse direction. Anything more extreme than that and you start getting that robotic warbling effect that gives away the conversion immediately. I ran into a specific edge case last month that took me a few hours to work around. I was converting a vocal recording where the performer was deliberately using heavy vibrato. The model treated each vibrato cycle as a separate pitch event and the index retrieval got confused because the pitch contours in the source didn't match what the model expected from the reference. The output sounded choppy and the timbre kept shifting in and out of focus. The workaround was to apply a mild pitch auto-correction pass to the source audio before running the conversion. Not full autopitch, just enough to smooth out the vibrato cycles into a more steady pitch contour. Then I set the pitch shift to exactly zero and used a lower pitch detection threshold in the model settings. This let the model focus on the timbral conversion rather than trying to reconcile mismatched pitch contours. The result was noticeably smoother.
Get the Full Details

There are some real limitations worth being honest about. The model struggles with non-speech content like singing with strong consonants or sibilance. If you're converting a vocal track that has prominent "s" and "t" sounds, you'll likely get water-like artifacts on those frequencies. It's a fundamental limitation of the spectral folding approach the model uses, not something you can fix with settings. For singing conversion specifically, you might want to look at parallel tools like So-VITS-SVC or the newer Fuse-LLM variants, which handle phonetic content better at the cost of requiring more training data. Another common pitfall is the reverb problem. If your source audio has any natural or added reverb, the converted output will carry that same reverb character but now layered on top of a completely different voice timbre. This creates a weird spatial mismatch that sounds off even when the voice conversion itself is technically good. My solution is to run the source through a dereverb pass first. RNNoise alone doesn't cut it for this, but newer dereverb models like Demucs or some of the specialized voice separation tools do a reasonable job of isolating the dry vocal signal before conversion. If you want to download the model, the official distribution is typically found on the creator's GitHub page or the associated Hugging Face repository. Make sure you're getting the latest release because the training recipes have been updated a few times to address some of the stability issues in early versions. The model files themselves are fairly compact, usually under a gigabyte depending on the included indices, and they run on a standard consumer GPU without issues.
The inference speed is decent. On a 3090, a typical 3-minute vocal track processes in roughly two to three minutes. That's not groundbreaking but it's usable for most projects. If you're batch-processing a lot of audio, the CPU fallback mode exists but it's painfully slow, so don't bother with it unless you have no GPU available. Overall, Of Shadows Vc Andrews is a solid option if you need reliable voice conversion and don't want to spend weeks training a custom model from scratch. It won't solve every problem you throw at it, but for standard speech and moderate singing conversion tasks it holds its own against more popular alternatives. Just make sure your input audio is as clean as possible and pay attention to the pitch settings, because those are where most people go wrong.