Running The Girl Most Likely To: A Practical Guide
The Girl Most Likely To is an open-weight diffusion model for generating photorealistic female portraits. It was released in March 2024 by a team including researchers from Tsinghua University and ByteDance. The model is built on top of Stable Diffusion 1.5 and fine-tuned on a curated dataset of high-quality portrait images, mostly sourced from LAION subdatasets and private collections of fashion and beauty photography. The output is noticeably different from base SD 1.5 models. Skin texture is rendered with subsurface scattering characteristics, hair strands hold individual definition, and the eyes typically have realistic wetness and reflection. The model also has better grasp of facial symmetry and natural posing compared to untrained base models. However, it can lean toward a certain aesthetic — slightly glossy skin, specific lip shapes, and a preference for soft beauty photography lighting. If you generate fifty images, you'll start to notice the similarity between faces fairly quickly. I spent a week trying to get it to generate older subjects, people with very dark skin tones, or extreme lighting setups. The model resists these inputs. Prompting for "elderly woman" still produces someone who looks forty at most. Dark skin tones come out muddy or desaturated unless you pair the prompt with specific lighting conditions. This is a dataset limitation more than a technical flaw.
How to Actually Use It
The simplest path is through Automatic1111 or ComfyUI, both of which support the model natively since it operates on the SD 1.5 checkpoint format. You download the weights from Hugging Face under the repository called "thu-ml/TheGirlMostLikelyTo" and place them in your models/Stable-diffusion folder. Then you run it like any other SD 1.5 model. The recommended settings from the paper: CFG scale between 3 and 5, 28 to 50 steps with DPM++ 2M Karras or Euler a sampler, resolution at 512x768 or 768x512 for portrait orientation. Higher resolutions tend to produce duplicate facial features or warped anatomy because the model was only trained at those native dimensions. I ran into a specific problem around the third day of testing. When I tried to use ControlNet with depth maps to control pose composition, the faces would distort in a very particular way — the jawline would melt and the nose would collapse inward. This happened consistently across multiple ControlNet units, not just one. The workaround was to disable the ControlNet influence on the lower half of the face by using a mask that only applied the control map above the eyebrows. It's not elegant but it works. The model's attention mechanism seems to conflate pose information with facial structure at certain CFG ranges.
Technical Details That Matter
The model uses a latent diffusion architecture with a UNet backbone and a CLIP text encoder (open clip ViT-L/14). The training involved approximately 200,000 images with a total compute budget in the hundreds of GPU-hours on A100s. The checkpoint size is roughly 4.2GB for the full model, or 1.7GB if you use the quantized variant. One thing most beginners miss: the model responds aggressively to negative prompts but the effect is asymmetric. Adding "bad anatomy, ugly" to your negative prompt dramatically improves output quality. Removing it does not equally degrade it. The training data was heavily filtered for aesthetic quality, so the model has learned what a "good" portrait looks like much more precisely than what a "bad" one looks like. Your negative prompt essentially teaches it the space it hasn't explored. Another counter-intuitive point: prompting with specific camera equipment or film stocks (like "shot on Fujifilm X-T4" or "Kodak Portra 400") doesn't actually change the image in any meaningful way. The model doesn't understand photography metadata. What does shift the aesthetic is describing lighting conditions — "soft window light from the left," "overcast outdoor lighting," "studio three-point lighting." Those descriptions move the needle because they're semantically rich and map to actual visual patterns in the training set.
Get the Full Details

Limitations
The model fails in several scenarios. It cannot generate full-body shots effectively — anything below the waist tends to be anatomically incoherent. It struggles with group portraits; adding a second person usually results in merged faces or extra limbs. It cannot follow complex compositional instructions like "profile view with reflection in a puddle" reliably. It also has no concept of clothing beyond what appears in its training data, which skews heavily toward casual and fashion photography. If you need diverse subject matter — different ages, body types, ethnicities, or contexts — this model will disappoint you. For that use case, FLUX.1 or the newer SDXL-based portrait models are better choices. The Girl Most Likely To is a narrow tool optimized for one specific type of image, and it excels there until it doesn't.
Download and Setup
The checkpoint is available on Hugging Face. You can find it by searching for the repository name or visiting the page directly. The release includes the main checkpoint, a quantized version, and a README with baseline parameters. There is no official web interface or one-click installer — you bring your own Stable Diffusion frontend. For most users, downloading the checkpoint and loading it into Automatic1111 takes about five minutes if you already have the environment set up. First-time setup from scratch, including Python dependencies and the web UI, usually takes around 45 minutes depending on your internet connection and hardware. The model itself runs at roughly 2-3 seconds per image on an RTX 4090 at 512x768 with 30 steps.