Training a Diffusion Model on an Artist's Style: What Actually Works

I spent about three weeks debugging why my LoRA kept producing muddy, over-smoothed outputs that looked nothing like the reference paintings. The images had the right color palette but zero brushstroke character. After burning through a bunch of failed checkpoints, I figured out the main issue was how I was handling the training data preprocessing and the learning rate schedule. Here is what I learned doing this, including one specific edge case that nearly made me abandon the whole project. The process involves taking a base diffusion model, typically SDXL or Stable Diffusion 1.5, and fine-tuning it on a curated set of images that share a consistent artistic style. The goal is not to teach the model a subject, but to teach it how a particular hand draws things. A LoRA, which stands for Low-Rank Adaptation, is the usual vehicle for this kind of work. It modifies only a small subset of the model's weights, keeping the base model intact while learning style-specific features like line weight, texture application, and color handling. The core mechanism is straightforward. You run your images through a training loop where the model attempts to predict noise added to a noisy version of each image. The difference between the prediction and the actual noise is the loss, and backpropagation adjusts the LoRA weights to minimize that loss. Over enough steps, the LoRA captures patterns specific to the style you fed it. The challenge is making sure it captures only the style, not random artifacts from the dataset.

I worked on a project to replicate a traditional watercolor artist known for loose washes and visible paper texture showing through. The first few runs at standard settings produced generic painterly results. Nothing resembled the artist. The issue was that the model was averaging out the very details I wanted to preserve. Once I switched to using a higher rank dimension and dropped the resolution to 512 instead of the default 1024, the fidelity improved noticeably. The lower resolution forced the model to focus on broader compositional patterns rather than getting lost in pixel-level noise.

The Practical Setup

For anyone starting out, Koala Diffusion or Kohya_ss are the two most common training frameworks. Kohya_ss is more widely documented but has a steeper initial configuration curve. I used Kohya with SDXL as the base. SDXL needs more VRAM, roughly 16GB minimum for comfortable training at reasonable resolution, but the quality ceiling is higher. If you are on a consumer card with 8GB, stick to SD 1.5 or use gradient checkpointing to reduce memory pressure. Dataset preparation is where most people go wrong. I found that using 15 to 30 images is usually sufficient for a clean LoRA. More than that introduces redundancy without adding signal, and it increases training time disproportionately. Each image should be as high quality as possible. Blur, compression artifacts, and inconsistent lighting in the source material all get absorbed into the LoRA weights. I trained a dataset of 22 images of a single illustrator's work and got a usable result in about 1,200 steps. Before training, I crop each image to a square aspect ratio and resize to either 512 or 768 pixels depending on the base model. SDXL handles 768 reasonably well on a 24GB card. I also normalize the images to float16 precision, which cuts memory usage roughly in half compared to float32 without any visible quality loss in the output. The captioning step matters less than most tutorials suggest. For style training, generic captions like photo of artwork or even empty captions work fine because the model is supposed to learn visual patterns, not textual associations. Over-captioning can actually harm style transfer by forcing the model to bind irrelevant semantic concepts to the visual features.

Get the Full Details

Multi-Source Training-Free Controllable Style Transfer via Diffusion Models
Multi-Source Training-Free Controllable Style Transfer via Diffusion Models

Training Parameters That Matter

The learning rate is the most sensitive parameter. A rate that is too high causes the LoRA to overfit quickly, memorizing individual images rather than learning generalizable style traits. A rate that is too low means the training never converges and you waste compute. I settled on a range of 1e-4 to 5e-5 for the UNet layers and 2e-5 for the text encoder, using a cosine decay schedule with a warmup period of about 10 percent of total steps. The rank or dimension of the LoRA controls its capacity. A rank of 16 is a safe starting point for SDXL. Going to 32 or 64 can capture more nuance but risks overfitting, especially with smaller datasets. I discovered this the hard way when I ran a trial with rank 64 on just 15 images. The loss dropped to near zero by step 500, but the generated outputs were brittle and failed to generalize to new prompts. The model had essentially memorized the training set. Batch size should be as large as your VRAM allows. A batch size of 4 at 768 resolution gave me stable gradients without OOM errors on my 24GB card. Smaller batches introduce noisier gradient estimates, which can destabilize training. I also enabled network alpha set to half the rank value, which acts as a scaling factor that prevents the LoRA from diverging too far from the base model weights during inference.

A Specific Edge Case That Almost Ruined My Training

During a run training on an artist whose work featured heavy black ink outlines over watercolor washes, the model began producing outputs where every generated line became unnaturally thick and dark, regardless of the prompt. This happened around step 800 and persisted through the final checkpoint. I initially thought it was a captioning problem, but the captions were completely blank. The issue was actually a form of mode collapse where the model latched onto the high-contrast black strokes as the dominant feature and started amplifying them in every output. The workaround was to add a small number of regularized images to the dataset. I pulled in about five reference images from a different artistic tradition that used similar media but with lighter linework. This broke the monolithic association the model was forming. I also reduced the text encoder learning rate to near zero, effectively freezing it, since the problem was purely visual and had nothing to do with textual conditioning. After those changes, the line weight in the outputs normalized within 200 steps. This is a niche but real problem that does not appear in most beginner guides.

Common Pitfalls to Avoid

The biggest mistake I see people make is treating style training like subject training. They feed in hundreds of images expecting better results. More data does not always mean better style capture. Diminishing returns hit hard after 30 images for most styles, and beyond 50 you are usually just reinforcing biases present in your dataset. If all your reference images share a particular composition habit, the LoRA will learn that habit as part of the style even if it is not intrinsic to the artist's work. Another pitfall is ignoring validation during training. I stopped checking intermediate outputs for the first few runs and ended up with checkpoints that looked great on paper but produced garbage in practice. Running a validation image every 100 steps and keeping track of loss curves let me catch overfitting early. A loss that drops below 0.01 on a 20-image dataset is a strong signal that the model is memorizing rather than learning. There is also the question of base model choice. SDXL generally produces cleaner style transfers than SD 1.5 for most contemporary art styles, but it is significantly slower to train. For traditional media styles like sumi-e ink painting or charcoal sketches, SD 1.5 can actually outperform SDXL because its lower resolution architecture preserves fine texture detail better. I have no general rule for this, just a recommendation to test both on a small subset before committing to a full run.

Style Injection in Diffusion: A Training-free Approach for Adapting Large-scale Diffusion Models ...
Style Injection in Diffusion: A Training-free Approach for Adapting Large-scale Diffusion Models ...

Exporting and Using the LoRA

Once training completes, export the LoRA as a .safetensors file. The file size for a rank 16 SDXL LoRA is usually between 200MB and 400MB. Loading it in ComfyUI or A1111 requires setting the strength to around 0.8 for most applications. Going to 1.0 can sometimes produce harsh, unnatural results because the style weights override the base model's natural priors too aggressively. I usually find 0.7 to 0.85 is the sweet spot. Prompt engineering with a style LoRA is different from normal usage. Since the LoRA already encodes the visual style, adding excessive descriptive text can conflict with what the model has learned. I keep prompts short and let the LoRA handle the aesthetic. A typical working prompt looks like a portrait in the style of the artist, soft lighting followed by the LoRA trigger at 0.8 strength. Anything beyond that tends to muddle the output. The whole pipeline from dataset preparation to a working LoRA takes roughly 3 to 4 hours on a modern GPU, with dataset curation being the largest time investment. Most of that time is spent selecting and cleaning images, not running the training itself. If you have a good dataset ready, the actual training can finish in under an hour for a modest number of steps.