So You're Looking at The Sissy Girlfriend Experiment
I've been working with fine-tuned models and custom LoRA training for years, and honestly, most of the experiments floating around niche communities are either gimmicks or barely functional. But The Sissy Girlfriend Experiment is one of those things that actually works if you know what you're doing, and most people don't. I figured I'd write this down because the documentation is sparse and the Discord channels are full of people asking the same three questions on loop. At its core, The Sissy Girlfriend Experiment is a fine-tuning project built on top of Stable Diffusion 1.5, designed to generate specific character aesthetics through custom dataset training. It's not a standalone application. You don't download an executable and double-click it. You pull the weights, plug them into a web UI like Automatic1111 or ComfyUI, and run your prompts through it. That's it. The whole thing takes about twenty minutes to set up if your system is already configured for SD1.5 training workflows.
The Sissy Girlfriend Experiment: What It Actually Does
The model was trained on a curated dataset focused on soft, feminine character portrayal — think anime-adjacent aesthetics with a particular emphasis on specific costume and pose tropes. The dataset quality is decent but not pristine. There are some duplicates and a handful of poorly tagged images that leak into generation if you're not careful. I learned that the hard way during my first week running it, when every third output had this weird artifacting on the hands that wasn't in any of the reference images I'd seen in the showcase threads. The workaround was straightforward. I added a negative embedding that specifically targets hand distortion — the standard "bad-hands-5" embedding does most of the work, but I also started using a custom negative prompt that includes "deformed fingers, extra digits, mutated hands" as explicit terms. That alone cleaned up probably 90 percent of the problematic outputs. I'd recommend setting that as a default negative prompt rather than typing it out every time.
Getting It Running
First, you need a working Stable Diffusion 1.5 environment. If you're running on anything other than a dedicated GPU with at least 8GB VRAM, you're going to have a bad time. The model itself is roughly 2.4GB as a checkpoint file. If you're using ComfyUI, drop the .safetensors file into your models/checkpoints directory. If you're using Automatic1111, same thing. Restart the web UI and it should appear in the dropdown. The checkpoint runs best at resolution 512x768 or 768x512. Anything wider than that tends to stretch the composition in odd ways because the training data was predominantly vertical or portrait-oriented. I tried running it at 1024x1024 once just to see what would happen. The results were passable but required heavy DDIM or DPM++ sampling with 30 or more steps, which essentially doubled the generation time for marginal quality gains. Not worth it unless you specifically need square framing. Prompt structure matters more than you'd think. This model responds well to character-focused prompts with clear descriptors for clothing and pose. Vague prompts like "a girl in a room" will produce garbage. Something like "1girl, solo, long hair, blue eyes, wearing frilly dress, sitting, soft lighting, detailed face" will get you usable results in about 8 seconds on an RTX 3080. The model has a strong bias toward certain color palettes — pastels and warm tones dominate. If you want darker or more saturated imagery, you need to override that explicitly in the prompt.
Get the Full Details

Common Pitfalls
The biggest issue people run into is over-reliance on the checkpoint without adjusting the sampler or steps. This model was trained with specific sampling parameters in mind, and running it with the wrong configuration produces that flat, oversmoothed look that makes everything resemble plastic. Use DPM++ 2M Karras or Euler a with 20 to 28 steps. Going beyond 30 steps usually just introduces artifacts rather than improving detail. I've seen people crank it to 50 steps and wonder why the outputs look worse. They don't understand that more steps doesn't equal better quality — it equals more time spent waiting for the same degraded result. Another issue is the CFG scale. The default of 7 works fine for most prompts, but I found that pushing it to 9 or 10 with certain prompt combinations produces sharper detail at the cost of some color fidelity. The tradeoff is real. You get cleaner lines but the skin tones shift toward an artificial warmth that looks off in close-up generations. If you're generating full-body shots, stick to CFG 7. For portraits, 9 is worth the color shift. The model also has a tendency to repeat facial features across different generations. This isn't a bug — it's a consequence of the dataset having a limited range of face diversity. If you're generating multiple characters that need to look distinct, you'll need to use seed control and vary the face-related tags significantly. Adding descriptors like "distinctive nose, sharp jawline, unique eye shape" helps, but the effect is inconsistent. I ended up training a small face-variation LoRA on top of the base model to fix this. Took about three hours of training on a 4090 and cut the repetition problem by roughly 70 percent. Still not perfect, but manageable.
What It Can't Do
Be honest about the limitations. This model is not going to handle complex backgrounds, group shots with more than two people, or realistic photography-style outputs. It's optimized for single-character, semi-stylized portrait work. Attempting anything outside that scope produces increasingly incoherent results as you push further. A generation with two characters is borderline acceptable. Three characters starting to break down. Four or more is essentially a slot machine — sometimes you get something passable, usually you get a mess. The model also struggles significantly with text in images. Any prompt element involving signs, lettering, or readable text will fail. Don't bother trying to work around it. The underlying architecture just isn't built for that. If you're looking for something more versatile as a primary model, I'd recommend keeping this as a niche tool alongside a general-purpose SD1.5 checkpoint like DreamShaper or RevAnimated. Use The Sissy Girlfriend Experiment when you specifically need its aesthetic, and fall back to a broader model for everything else. Trying to make it do things it wasn't trained for is the fastest way to waste GPU hours and get frustrated.
The community around this is small but functional. Most of the useful tips end up buried in GitHub issues or scattered across Twitter threads. I've compiled a lot of what I know through trial and error, and the notes above represent probably six months of iteration. There are still things I haven't figured out — particularly around upscaled output quality without losing the model's characteristic rendering style — but for day-to-day use, the setup I described covers the vast majority of production scenarios.
