Running Your Own Image Generation Stack Outside the Cloud
Most people don't realize that Leonardo is just a UI sitting on top of several open-source diffusion models. The actual generation happens on whatever hardware Leonardo spins up in their datacenter, and you're paying for the privilege plus waiting in a queue. If you've got a machine with a decent GPU, you can bypass the whole platform and run these models yourself with direct API access, local inference, or both. The idea of getting Leonardo to the internet means pulling the models off the platform and running them on your own infrastructure so you're not dependent on their uptime, pricing changes, or rate limits. It's not especially complicated, but there are a few traps that catch people who assume this will just work out of the box. Here's how I actually did it. I was generating roughly two hundred images per day for a client project and Leonardo's queue times had started creeping up to forty-five minutes during peak hours. Their API credits were burning through my budget at a rate that didn't scale. So I pulled the weights and ran them locally.
What You Actually Need
The hardware question comes first because it determines everything else. A single NVIDIA GPU with at least 12GB of VRAM will get you through most standard generation workflows. I ran a 4060 Ti with 16GB initially and it handled SDXL fine at around eight seconds per image with default settings. When I needed higher resolution or denser batching, I moved to a 4090 with 24GB and the math changed significantly. The software stack breaks down into three parts: a model runner, a UI if you want one, and a way to expose it. For the runner, ComfyUI is the most flexible option and the one I settled on. It's node-based, which makes it feel like overkill until you're trying to chain LoRAs, control nets, and upscaling in a single workflow. Then it's exactly what you need. Automatic1111 is simpler but heavier and less modular. I wouldn't recommend it for anything beyond casual use. For exposing the thing to the internet, you have two real options. Either run a reverse proxy with authentication in front of your local UI, or package it as a REST API and call it from whatever frontend you build. The reverse proxy route is faster to set up. The API route gives you control.
Getting the Models
Leonardo's models aren't officially downloadable, but the underlying architectures are mostly based on Stable Diffusion checkpoints that exist publicly. SDXL 1.0, SD 1.5, and the Pony Diffusion variants are all available on Hugging Face and Civitai. Some of Leonardo's fine-tunes are derived from these, so you can get very close results by starting with the base checkpoint and applying the appropriate LoRA or text encoder settings. I spent about two weeks mapping which Leonardo prompts produced which kinds of outputs and then finding the closest public checkpoint plus LoRA combination. It's tedious work but necessary because prompt compatibility varies wildly between models. A prompt that works perfectly on Leonardo's fine-tune of SDXL will look completely different on raw SDXL. Document everything as you go.
Get the Full Details

Setting Up the Local Runner
ComfyUI installs with a single Python command and a git clone. The tricky part isn't installation, it's managing your workflow files and making sure your custom nodes don't break when you update. I keep a separate portable installation in a virtual environment and only update when I actually need new features. Left alone, it works for months without issues. Load your checkpoint into the model loader, connect a CLIP text encode node for your prompt, a KSampler for the actual denoising, and a VAEDecode into an image save node. That's the minimum viable pipeline. Add a ControlNet unit node if you need pose or depth conditioning. Add a latent upscaler if you're going past 1024x1024. The node graph is explicit about what's happening at each step, which is why I prefer it over black-box interfaces.
Exposing It to the Internet
This is where most guides skip ahead. If you just want to access your local ComfyUI from other machines on your network or over the internet, you can use ngrok for a quick tunnel or Caddy for something more permanent. I set up a Caddy reverse proxy with basic auth pointing at ComfyUI's localhost port. That gave me a URL I could access from any browser with credentials, and Caddy handled the TLS certificate automatically. For programmatic access, ComfyUI has a built-in WebSocket API. You can send workflows as JSON and receive generated images back. I wrote a small Python script around this that batches ten prompts, queues them, and pulls the results as they complete. The script runs headless on a Raspberry Pi I had lying around, and it communicates with the main generation machine over the LAN. This setup cut my image production time from about three hours down to roughly forty minutes for the same output volume.
What Actually Breaks
The first thing that goes wrong is memory management. OOM errors are the default state when you're not careful. I learned this the hard way when I tried to batch six SDXL images at once on a 4060 Ti and the system froze hard. The workaround is enabling vram_optimization flags and keeping batch sizes at two or three unless you're on 24GB or more. There's also an offloading option in ComfyUI that moves unused tensors to system RAM. It's slower but prevents crashes. The second thing is model file management. Checkpoints are large, usually between 6GB and 14GB each. LoRAs are smaller but add up. I have about 40GB of model files spread across three directories organized by architecture and use case. Running df -h on your drive before you start is a good habit. The third thing is prompt drift. Without Leonardo's curated prompt templates and negative prompt defaults baked in, you have to manage those yourself. I keep a spreadsheet mapping desired output styles to specific prompt structures, model versions, and sampler settings. It took me about a week to build but it's saved me countless hours of trial and error since.

When This Isn't Worth It
If you're generating fewer than twenty images a day, you're probably better off paying for Leonardo. The setup time, maintenance overhead, and troubleshooting cost more than the subscription in most cases. If you don't have a GPU with at least 8GB of VRAM, you're not saving anything either, since cloud API access will be faster than running on CPU. And if you need access to Leonardo-specific fine-tunes that aren't available as public checkpoints, you're stuck using their platform regardless. The real use case is steady high-volume generation where you value reliability and cost predictability over convenience. Once it's running, it runs. I haven't had a queue timeout in fourteen months.