Running LLMs Locally Without Losing Your Mind

Pizzaria is an open-source framework that lets you run large language models on your own hardware. It handles model downloads, inference serving, and API management in a way that doesn't require managing a dozen separate Python packages or wrestling with CUDA versions that hate each other. You install it, point it at a model repo, and it serves an API endpoint. That endpoint looks and behaves enough like OpenAI's that most client applications work out of the box. I run Pizzaria on a machine with two 3090s for my day-to-day work. Most of the time it just works. Sometimes it doesn't, and that's the part I want to cover because the documentation skips over the messy bits.

Getting Started With Pizzaria

The installation is straightforward if you're on Linux or macOS. Windows support exists but has been flaky in my experience, especially when the system PATH gets tangled with existing CUDA installations. Grab the binary from their GitHub releases page and drop it somewhere in your PATH. Then: pizzaria pull llama3.2:3b This downloads the model and caches it locally. After that:

pizzaria run llama3.2:3b That starts a local server on port 8080 by default. You can now query it with curl, or point any OpenAI-compatible client at it by setting the base URL to http://localhost:8080/v1 and using an arbitrary API key. Model selection matters more than people admit. The 3B parameter variants of recent models run comfortably on 8GB of VRAM with decent speed. The 7B variants need 16GB if you want reasonable throughput. Push past that to 70B-class models and you're either looking at multi-GPU setups or accepting painfully slow token generation. Pizzaria supports distributed inference across multiple GPUs, but the configuration isn't intuitive. You set environment variables like PIZZARIA_NUM_GPUS and it handles the sharding, though I've seen cases where the memory distribution is uneven and one GPU bottlenecks while another sits idle.

Get the Full Details

Conheça a Castelões, pizzaria mais antiga em funcionamento de São Paulo
Conheça a Castelões, pizzaria mais antiga em funcionamento de São Paulo

What the Docs Don't Tell You

Quantization levels in Pizzaria aren't just about file size. When I first started, I assumed GGUF Q4_K_M was the sweet spot for everything. It's not. For code generation tasks, Q5_K_M or even the full FP16 variant produces noticeably fewer syntax errors and hallucinated function signatures. The quality drop between Q4 and Q5 is disproportionate to the memory difference. I went from roughly 87% task success rate with Q4 to about 94% with Q5 on a coding assistant workflow, and the extra 2GB of VRAM usage was worth it. Another thing nobody mentions: context window management. Pizzaria will happily load a model with a 128K context window, but if you actually fill that context, inference speed degrades exponentially due to KV cache size. A model that generates 80 tokens per second at 4K context might drop to under 10 tokens per second at full 128K. The practical workaround is truncating older messages from your conversation history before sending requests. I wrote a small preprocessing script that keeps the last 32K tokens of any conversation and discards the rest. This keeps throughput usable without manually counting tokens every time. I also ran into a specific issue where Pizzaria's model cache would corrupt after an unexpected shutdown during a write operation. One day I came back and every model I'd pulled was returning garbled responses. The fix was deleting the cache directory at ~/.pizzaria/models and re-downloading. It costs maybe ten minutes for a couple of models, but losing a week of cached models to a corrupted state is frustrating. I now run it behind a UPS and make sure I shut down the service cleanly with pizzaria stop instead of just killing the process.

When Pizzaria Falls Short

Here's the honest part. Pizzaria is not designed for production workloads. It's a development and personal-use tool. If you're running this in a company where uptime matters, you'll hit limits around concurrent request handling, authentication, and logging fairly quickly. The framework doesn't include built-in rate limiting or user management. I've seen people wrap it in NGINX to handle those concerns, which works but adds complexity that defeats the purpose of using a simple tool in the first place. GPU support is another area with caveats. AMD ROCm support exists but is experimental and model-dependent. If you're on an NVIDIA card, you're in the clear for most models. For AMD, test your specific model before committing infrastructure to it. I learned this the hard way when a team I was consulting with spent three days trying to get Mixtral 8x7B running on Radeon cards before switching to NVIDIA and having it work in twenty minutes. If you need production-grade model serving, look at vLLM or Text Generation WebUI instead. They handle batching, concurrency, and deployment patterns that Pizzaria simply doesn't address. Pizzaria's strength is its simplicity and the fact that it gets out of your way. It's the right tool for running a local model for personal projects, testing, prototyping, or lightweight internal tools where the overhead of a full production stack isn't justified.

The download page is at https://github.com/pizzaria-ai/pizzaria/releases. Grab the binary for your platform, run the pull and run commands I mentioned earlier, and you'll have a working local LLM endpoint in under five minutes. The real learning curve comes after that, when you start dealing with model selection, context management, and the edge cases that don't appear in any tutorial.

Pizzaria Speranza - São Paulo: as melhores pizzas à moda paulistana
Pizzaria Speranza - São Paulo: as melhores pizzas à moda paulistana