Training a train model is mostly about not wasting GPU hours on bad data
The first thing most people get wrong is assuming the model is the problem. It almost never is. The problem is always the data. I spent three weeks debugging what I thought was a learning rate issue on a LoRA fine-tune before I realized the dataset had 847 duplicate images scattered across the training dirs. The model wasn't failing to learn. It was learning the same reference image 12 times and calling it progress. Here's the basic flow. You need a captioned dataset, a base model, and enough VRAM to not give up after epoch one. Most people use Kohya_ss these days because it wraps everything into something manageable. The alternative is running the diffusers script directly, which gives you more control but requires you to understand the argument flags without a GUI holding your hand.
How To Train A Train
Start by collecting or generating your dataset. If you're training a character or a style, you want at least 20 to 40 high-quality images. More is fine but diminishing returns hit hard after that unless you're doing full fine-tuning on a massive compute setup. Caption every image. Not with an auto-captioner that spits out garbage like "a person standing in front of a background." Use something likeWD-1.4 Tagger or BLIP and then manually fix the captions. A bad caption is worse than no caption because the model learns the wrong association and you spend hours wondering why it keeps adding unrelated elements. Set your resolution. The standard is 512 or 768 depending on your base model. SDXL wants 1024. Don't mix resolutions in the same bucketing run unless you know what you're doing with the aspect ratio bucketing settings. I once ran a mixed-resolution job without enabling proper bucketing and the model's convergence looked fine in TensorBoard but the outputs were blurry at certain aspect ratios. The loss curve lied to me. It looked healthy the whole time. For LoRA training, which is what 90 percent of people should be using, here's a starting point that actually works: network dim 32, alpha 16, learning rate 1e-3 for the text encoder and 1e-4 for the UNet. Epochs around 10 to 20. Batch size whatever your VRAM allows, usually 1 or 2 on a 24GB card. Gradient accumulation to fake a larger batch if needed. Use a cosine or constant learning rate scheduler. Warmup steps around 10 percent of total steps.
The trick nobody tells you is that you should validate early and often. Save a LoRA every 200 to 500 steps and test it on a fixed prompt set. Most people train for 10,000 steps and then discover on epoch ten that the model started overfitting around step 3,000. The validation images from step 3,000 would have been better. Keep a log. A simple spreadsheet with step number, validation prompts, and a quick screenshot takes two minutes and saves you from training past the point of no return. Another thing that isn't obvious: regularization images matter more than most tutorials admit. If you're training a specific character, you need regularization data that represents the base model's understanding of that concept without the character's features. Without it, the model will collapse toward your dataset and lose its ability to generate the subject in different poses or contexts. I used ~100 regularization images per concept and it made a noticeable difference in generalization. Skip it and your LoRA will overfit like aIdiot. When to switch from LoRA to full fine-tune. If you're trying to replicate a specific artistic style with hundreds of reference images and you have access to multiple A100s or H100s, full fine-tuning can capture nuances that LoRA misses. But it's expensive, slow, and the resulting model is stuck to one checkpoint. LoRA is modular. You can layer them. Full fine-tune is a hostage situation with your GPU.
Get the Full Details

One edge case I ran into recently: training on anime-style data with SDXL. The base model is trained on photographic and natural image distributions. Feeding it anime outputs without adjusting the VAE or using an anime-specific SDXL checkpoint produced muddy results. The fix was switching to an anime-compatible base like AnythingV5 or using a dedicated SDXL anime checkpoint as the starting point. The architecture doesn't change. The pretraining distribution does. Using the wrong base model is like trying to tune a violin with a cello bow. It works in principle. The results are terrible. Check your output after training. Load the LoRA into a local inference pipeline or webui, run your validation prompts, and compare against the pre-training baseline. If the outputs look identical to the base model, your rank was too low or your learning rate was too conservative. If they look like garbage, you overfit or your captions were wrong. There's no middle ground where bad data produces good results through enough training steps. More steps just amplifies whatever signal you fed it. If you want to download the tools, Kohya_ss is on GitHub. The scripts are free. WD-1.4 Tagger is also on GitHub. Stable Diffusion models come from Civitai or Hugging Face. Nothing here costs money. The cost is your time and your GPU's patience.
The real bottleneck isn't the training. It's the preparation. Captioning, deduplication, resolution bucketing, regularization data, validation setup. That's where the actual work lives. The training loop runs itself. Most of the decisions that determine whether your output is usable or trash happen before you hit start.