The Reality of AI-Assisted Art Production
I spent three years working with Stable Diffusion and ControlNet pipelines before I stopped treating AI as a magic button and started treating it like a really opinionated junior designer who never misses a deadline but also never understands why you're asking for something weird. The process is mundane once you strip away the hype. You write prompts. You tweak denoising strength. You iterate. Sometimes it works, sometimes you get a render that looks like a fever dream of hands with seven fingers each. The whole concept gets discussed a lot in empty philosophical circles. People talk about authorship and creativity while completely sidestepping the actual technical workflow that makes this work. Let me just explain what I do when I need to produce output that passes as legitimate artwork rather than obvious AI slop. Start with Stable Diffusion XL or one of its finetunes. Run it through ComfyUI because the node-based interface actually lets you control the pipeline instead of clicking a generate button and hoping for the best. I use ControlNet heavily — depth maps, canny edges, and openpose pass through — because raw generation without constraints produces garbage about sixty percent of the time. The ControlNet layer is what separates controlled output from random seed lottery tickets.
Here is where most people get stuck and walk away frustrated: they prompt for something specific and then complain when the model ignores their instructions. The workaround is straightforward. Use img2img mode with a reference image. Set your denoising strength between 0.35 and 0.55 depending on how much you want to deviate from the source. Anything above 0.6 and you are basically gambling. Below 0.3 and the output becomes an indistinguishable copy of your input. I had a client project last year where they needed a series of architectural visualizations in a specific brutalist style. I spent two days just building a custom ControlNet depth map pipeline using Blender renders as input. The final workflow took about eight minutes per image after setup, compared to maybe forty-five minutes of manual prompting and inpainting without it. The quality difference was also night and day. The raw model outputs looked uncanny and plastic. The controlled pipeline produced results that required only minor touchup in Photoshop. The tools themselves are not the hard part. Understanding which parameters interact with each other is what takes time. Denoising strength and CFG scale fight each other if you set both too high. A CFG above 8 with a denoising strength above 0.5 will usually produce burnt, oversaturated nonsense. Keep CFG between 5 and 7 for most workflows. That is the range where the model follows your prompt without breaking the image.
Another thing nobody mentions enough: your training data matters more than your prompt. If you are working in a specific aesthetic style, you need a LoRA or a finetuned checkpoint built for that style. Running a generic SDXL checkpoint and hoping it nails a very specific illustrative style is a waste of compute. I train my own LoRAs using Kohya_ss when a project demands consistency across a batch. Takes about forty minutes on a decent GPU and produces a file you can apply to any image in the pipeline. The downsides are real and worth acknowledging. Hardware requirements are steep if you want anything approaching reasonable speeds. A 4090 gets you there, but anything less and you are waiting around. Licensing is another mess. Some finetunes and LoRAs have questionable origin stories. You need to check where models come from before using them commercially. Commercial rights vary wildly between hubs like Civitai and Hugging Face. The biggest limitation though is that AI does not understand physics, biology, or basic cause and effect. It generates patterns that resemble those things. This means every output needs a human eye checking for structural problems. Hands, teeth, text, reflections, shadows — these are the usual failure points. Inpainting fixes most of them but costs additional time and compute. Expect to spend 20 to 40 percent of your total project time on post-processing fixes regardless of how good your initial pipeline is.
Get the Full Details

If you are just starting out, do not buy into the narrative that AI replaces artists. It replaces tedious tasks within an artistic workflow. The person who understands composition, color theory, and visual storytelling still has the advantage. The tool just makes iteration faster. That speed advantage means nothing if you cannot direct it toward something coherent. Download ComfyUI from their GitHub releases page. Grab SDXL base and refiner checkpoints from Hugging Face. Learn ControlNet architecture before you try to combine multiple conditions. The learning curve is about two weeks of frustration before things start clicking. After that it is just engineering.