Working with Hamilton in production: what actually goes wrong
Hamilton is a Python library for building deterministic, typed execution graphs. You define a set of functions, annotate them with types, and Hamilton figures out the order they should run in based on dependencies. It was created at Stochastic AI and is widely used for feature engineering pipelines, data transformation workflows, and anything where you want reproducible graph execution without manually tracking which function depends on which output. The setup is straightforward. You install it with pip: pip install hamilton
Then you write functions. That's the core idea — you just write regular Python functions with type hints and Hamilton does the rest. Here's a minimal example: from hamilton import driver def base_feature(raw_data: pd.DataFrame) -> pd.Series:\n return raw_data['revenue'].mean()
def derived_feature(base: float, threshold: float) -> bool:\n return base > threshold dr = driver.Driver()\nresult = dr.execute(['derived_feature'], df_dict={'raw_data': my_dataframe, 'threshold': 1000}) Hamilton reads the type annotations, builds a directed acyclic graph, and executes functions in the correct topological order. You tell it what final outputs you want, and it works backward to figure out everything it needs to compute.
Get the Full Details

How Hamilton actually works under the hood
When you create a Driver, Hamilton inspects your module for functions. It builds an adjacency list from the type signatures — if function B's argument has the same name as function A's return annotation, Hamilton connects them. The execution engine then does a topological sort and runs each node exactly once. There's no caching by default, no parallel execution, and no error recovery built in. It's deliberately simple. The key strength is modularity. You can swap out one function without touching anything else, as long as the inputs and outputs match the types. This matters when you're iterating on feature engineering and have thirty functions that all feed into each other in non-obvious ways. Hamilton prevents you from accidentally running function X before function Y depends on it. I found this especially useful when I was rebuilding a feature pipeline that had grown organically over two years. Someone had added conditional branches inside individual functions, making the dependency graph invisible. Hamilton forced me to make every data flow explicit, which took about four hours of refactoring but caught three places where we were double-computing the same aggregation across different modules.
Common pitfalls beginners miss
The biggest issue is the naming convention. Hamilton matches arguments to return values by name, not by position or type alone. If your function returns df_clean but another function expects clean_df, Hamilton won't connect them regardless of whether they're the same shape. This caught me off guard when I was migrating code from a different framework where positional matching was the norm. Had to go through about twenty functions renaming variables to make everything wire up correctly. Took roughly an hour. Another thing that trips people up is that Hamilton doesn't handle side effects well. If a function modifies a DataFrame in place or writes to a database, the graph execution model gets confusing because the dependency tracking only covers function inputs and outputs. I learned this the hard way when a logging function that also wrote metrics to a file was being called at unexpected times during a rerun. The fix was to separate the side effect into its own function and explicitly include it in the execution list.
What Hamilton can't do
It doesn't support parallel execution out of the box. If you have fifty independent feature computations and want them to run concurrently, you're on your own — Hamilton will execute them sequentially. For small pipelines this isn't a problem. For anything larger you'd need to wrap Hamilton with something like Ray or Dask, which adds complexity that defeats some of the simplicity Hamilton provides. There's no built-in support for conditional execution either. Every function in the graph will be evaluated unless you can express the condition through data flow. I ran into this when building a pipeline where certain transformations only applied to specific date ranges. I ended up having every function accept a date range parameter and filter internally, which worked but made the function signatures noticeably more cluttered. Versioning and diffing between pipeline versions exists but feels underdeveloped compared to what you'd get from dedicated MLops platforms. Hamilton's git integration is functional for basic comparison but doesn't give you granular control over which nodes changed or what the impact was. If your org already uses something like Feast or Tecton for feature management, adding Hamilton on top creates more overhead than it saves.

When Hamilton is the right call
Use it when you have a feature engineering pipeline with many interdependent transformations and you want to eliminate manual dependency tracking. The sweet spot is mid-size teams — small enough that a full MLops platform is overkill, large enough that manually managing execution order is painful. If you're running a solo project with fewer than ten functions, you probably don't need it. If your pipeline involves distributed computing or complex branching logic, Hamilton will fight you rather than help. The library itself is lightweight, roughly a few hundred kilobytes, and has no heavy dependencies beyond pandas and numpy for most use cases. The documentation is decent but sparse on production patterns — you learn most of what matters by reading the source code and the GitHub issues. The maintainers are responsive but the project moves slowly, so if you hit a limitation there probably isn't a workaround coming soon.