Working with LLMs in Python Without Losing Your Mind
You pull up transformers, load a model, and suddenly you are dealing with tensors, device placement, batch sizes, and a GPU that keeps running out of memory. This is the reality of using Large Language Models Python codebase before the abstractions hide the mess. I spent three months building a document parsing pipeline that used to crash at random intervals because the inference loop was holding onto tensors across batches. The fix was ugly but effective: I added a del statement after every generate call and set torch.cuda.empty_cache() inside a try-finally block. That cut my OOM errors from daily to once a week. The core workflow has not changed much since 2023. You install the libraries, load a tokenizer and model, feed in text, and get tokens back. Most people skip past this part too fast because tutorials make it look trivial. The non-trivial part is everything that happens between tokenization and decoding when you are running anything beyond a demo script.
Large Language Models Python Setup and Implementation
Start with pip install transformers torch accelerate datasets. That gives you the base tooling. The accelerate library matters more than most people realize because it handles distributed loading without requiring you to write custom launcher scripts. If you are loading a 70B parameter model on a single GPU with 24 GB VRAM, you need quantization or offloading. Bitsandbytes does 4-bit quantization and drops your memory footprint from 140 GB to roughly 38 GB with minimal quality loss on most tasks. The catch is that 4-bit quantization breaks certain operations in older PyTorch versions, so pin your torch version to 2.1 or later. Here is the actual pattern I use. Load the model with a device map, not a single device assignment. Device map auto distributes shards across available GPUs and CPU. For most production work, using pipeline from transformers is fine for inference-only tasks, but when you need control over generation parameters or streaming output, instantiate the model and tokenizer directly.
from transformers import AutoModelForCausalLM, AutoTokenizer, pipeline
import torch
model_name = "meta-llama/Meta-Llama-3.1-8B-Instruct"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype=torch.bfloat16,
device_map="auto"
)
pipe = pipeline("text-generation", model=model, tokenizer=tokenizer)
That bfloat16 dtype is important. Half precision saves memory but can cause numerical instability with attention scores on certain hardware. Bfloat16 keeps the dynamic range of float32 while cutting memory in half. If you are running on consumer GPUs without native bfloat16 support, float16 works but watch for overflow in attention layers on long sequences. Generation parameters are where things get tricky. The default settings will make a model repeat itself on anything longer than five tokens. Set repetition_penalty to 1.15 and temperature to 0.7. Add a max_new_tokens limit that matches your actual output needs instead of the default 200 which is usually too short for anything useful and too long for simple classification tasks. I once had a prompt injection issue caused by not setting a stop sequence, and the model just kept generating system-like instructions for forty lines because nothing told it to stop. Batch processing sounds simple but causes silent memory issues if you do not handle it properly. Stack your inputs with padding, but pad on the left side for causal models, not the right. Right padding causes the attention mask to let the model attend to padding tokens, which corrupts the output. The attention_mask parameter in the generate call handles this, but you must construct it correctly from your padded input_ids.
Get the Full Details

What the Documentation Does Not Tell You
Tokenizers add special tokens based on the model, not your preference. Llama 3 adds BOS and EOS tokens automatically through the chat template. Mistral does not. If you feed raw text into a Mistral model without handling tokens yourself, you get degraded quality. Always check what the tokenizer does with your input before assuming the model sees what you think it sees. Caching is another area where people waste hours. KV cache stores previous token states for faster autoregressive decoding. It doubles your GPU memory usage during generation and then releases it when the generation ends. If you are doing continuous inference in a loop without clearing the cache between calls, you will see memory creep until the process crashes. Call model.reset_cache() between independent requests or just reload the model if you are doing many separate conversations. Streaming output with iter() on the pipeline object works for real-time display but introduces a significant performance penalty. Each yielded token triggers a full forward pass including all the cached key-value computation. If you need speed more than interactivity, collect the full output and split it afterward. The difference is roughly 20 to 30 tokens per second slower with streaming enabled on a single A10G.
Context window limits are not just about input length. The total token count includes your system prompt, conversation history, tool definitions, and output. A model advertised with an 8K context might actually give you 4K of usable space after accounting for overhead. Count your tokens explicitly with len(tokenizer.encode(full_prompt)) before sending anything to the model. I learned this the hard way when a client complaint pipeline started truncating mid-response because the token count grew past the limit without any error being raised. Transformers silently truncates rather than throws an exception, so you get garbage output instead of a helpful error message.
When This Approach Breaks Down
Running LLMs locally in Python is fine for prototyping, small batches, and models up to about 13 billion parameters on decent hardware. It does not scale well for serving multiple users concurrently. The GPU sits idle between requests, and batch size is limited by memory. If you need concurrent inference, you should look at vLLM or TGI instead of raw HuggingFace transformers. These engines use PagedAttention and continuous batching to serve many requests from a single model copy with dramatically higher throughput. Quantized models lose capability on niche tasks even when they pass standard benchmarks. A 4-bit Qwen2.5 7B model will handle general chat, summarization, and basic coding fine. It will struggle with chain-of-thought reasoning, exact code execution, and multilingual tasks that require precise token-level control. Run your validation suite at both full precision and quantized before committing to a quantized deployment. The difference is usually not visible in a two-sentence test prompt but becomes obvious under sustained complex reasoning. Long context models exist but are not a free lunch. Processing 32K tokens on a local GPU takes 3 to 5 times longer than 4K tokens on the same hardware, and memory usage scales quadratically with sequence length due to the attention mechanism. If your application genuinely needs long context, consider chunking and retrieval instead of stuffing everything into one prompt. I switched a legal document analysis tool from 32K context to a retrieval-augmented pipeline and reduced average response time from 14 seconds to under 3 seconds while improving accuracy because the model could focus on relevant passages instead of scanning thousands of irrelevant tokens.

The ecosystem moves fast. A pattern that worked in early 2024 may not apply now. Check the model card for recommended generation parameters, read the pull request comments on the HuggingFace repo for known issues with your specific model and hardware combination, and validate your own use case before optimizing prematurely. Most performance problems are actually configuration problems in disguise.