Python Code Review Notes From a Bunch of Messy Projects
I've spent enough years reading Python code to know that most style guides are written by people who haven't maintained a codebase past six months. The Field Guide For Python Best Practices I'm about to describe comes from actual production experience, not from reading PEP 8 cover to cover and calling it a day. Let's talk about type hints first because this is where most teams fail, and it's a silent one. You've probably seen a function annotated like def fetch_data(url: str) -> dict: and been told that's good practice. It's not. A bare dict tells you exactly nothing about what's inside that return value. The kind of bug that costs a team three days of debugging usually looks like this: someone merges two datasets together and assumes the key names are identical because the type hint said "dict" for both. They're not. The fix is straightforward but requires discipline. Use typing.TypedDict for structured dictionary returns. For the fetch example above, define a small schema with the actual keys your API response contains. It adds maybe thirty seconds of typing per function and prevents an entire class of runtime KeyError disasters that show up in production at 2 AM.
Field Guide For Python Best Practices That Actually Work
Context managers deserve more attention than they get. Most people use them for file I/O and that's about it. But exception handling during resource cleanup is where they save your neck. I once had a background worker that would silently drop records because it was writing to a SQLite database without a proper context manager wrapping the commit. If the process crashed mid-write, those records were gone. No warning, no traceback, just missing data. Adding with blocks around transactional code cut our silent data loss incidents from roughly four per month to zero over the following quarter. Iterator protocol is another area where people casually write broken code. The __iter__ versus __next__ distinction matters more than most tutorials suggest. When you're building a generator-based pipeline that processes log files through multiple transformation stages, getting this wrong means you either materialize everything into memory at once or you accidentally consume the same source twice. I ran into this on a project processing about 2 GB of nginx logs daily. The initial implementation loaded everything into a list, which worked fine on my laptop and tanked the staging server. Moving to a proper generator chain with explicit yield statements reduced peak memory usage from roughly 1.8 GB to under 40 MB and cut processing time by about 40 percent since we could start writing output before reading everything in. Here's something nobody warns you about: mutable default arguments. You've definitely seen this one. It's in every beginner tutorial because it has to be there. But the edge case that catches experienced developers off guard is when the mutable default is nested inside a closure or a factory function. I was writing a caching decorator and accidentally used default_factory=list as a module-level default instead of inside the function signature. The cache accumulated entries across unrelated test runs because the default list was shared at module import time. The fix was moving it inside the function with an explicit None sentinel and constructing the default inside the body.
Dependency management in Python is still a mess and pretending otherwise helps nobody. Poetry, pipenv, uv, pdm — they all solve slightly different problems and each one has breaking changes that have cost me significant time. My working rule is simple: pin everything in a requirements.txt or pyproject.toml, never rely on ~= or ^ constraints in production code, and run pip freeze after every deployment to verify the lockfile matches what's actually running. This usually prevents the dependency hell that shows up when one service upgrades a library and breaks three others that depended on the old behavior. Testing is where the biggest waste of time happens. I've seen teams write integration tests that hit a real database, a real cache, and a real external API for features that should have been unit-tested in under a second. A realistic rule of thumb I've settled on: if a test takes longer than five seconds to run, it's probably doing too much. Mock the dependencies at the boundary. Use unittest.mock.patch or pytest.fixture with dependency injection. This doesn't make your tests less valuable. It makes them run fast enough that someone will actually run them before pushing code. Logging configuration is another area where the default Python behavior will quietly fail you. The standard library's logging module creates log files that grow without bound unless you configure a RotatingFileHandler or TimedRotatingFileHandler. I inherited a system where the root logger was writing to a single file with no rotation configured. After about three weeks the disk was full and the application stopped writing any logs at all, which meant we couldn't tell what was happening when things started failing. Setting rotation at 50 MB per file with five backup copies and a log level filter that routes ERROR and above to a separate stream took about twenty minutes to implement and has prevented every disk-full incident since.
Get the Full Details

There are genuine limitations to everything I've described here. Type hints don't catch dynamic attribute access. Context managers don't help with async resource leaks in event loop code. The generator pattern doesn't work well when you need random access to previously yielded values. None of these are reasons to skip them. They're reasons to know what each technique actually covers and what it doesn't, so you don't build false confidence into your architecture. If you're looking for a concrete starting point, the Python standard library documentation for the typing, contextlib, and logging modules is where most of the practical depth lives. The real-world patterns come from maintaining code long enough to see what breaks.