The reality of building financial ML systems

Most people coming into this space think the hard part is the modeling. It isn't. The hard part is everything that happens before you even touch a dataset. I spent about three years building features for equity factors and realized I'd already wasted two of them because no one bothered to tell me that financial data has structural problems that don't exist in any textbook example. When you actually work with Advance In Financial Machine Learning, you quickly learn that the gap between academic papers and production is enormous. Papers show clean backtests with perfect leakage-free pipelines. Production has broken data feeds at 2am, survivorship bias in every universe you thought you had cleaned, and microstructure noise that makes feature importance scores look completely fabricated.

What actually matters in practice

The features that survive in live trading environments are almost never the ones you'd pick from a standard machine learning tutorial. Cross-sectional rank normalization alone fixed more broken pipelines than any hyperparameter search ever did. Most people don't realize that ranking returns within a universe eliminates a massive amount of daily regime variation that would otherwise swamp the signal. I once built a gradient boosting model on ~400 raw features for mid-cap equity selection. The validation backtest showed a Sharpe of 1.8, which sounded good until I realized the train-test split was done chronologically without blocking for sector exposure. When the model moved to production, the out-of-sample Sharpe dropped to 0.3. The actual problem wasn't overfitting in the traditional sense. It was that my training data had a temporary sector tilt that the model learned to exploit, and that tilt disappeared in live markets. Fixing it required implementing proper PurgedKFold cross-validation with embargo periods, which cut the effective training data by about 40 percent but brought live performance up to a realistic 0.9 Sharpe.

The leakage problem nobody talks about about honestly

Data leakage in financial ML doesn't usually look like the obvious kind where you accidentally include the target in a feature. It looks like subtle point-in-time mismatches. A common example is using market cap at the end of month when you should be using it at the beginning. Or aligning earnings dates with price data without accounting for the fact that prices move before the announcement becomes public in your dataset. These errors are nearly invisible in backtests and destroy PnL in production. The workaround I ended up using was building a strict point-in-time database layer where every feature is anchored to a specific timestamp, and then running a separate alignment check that verifies no future information leaks into any feature at any point. This took about two weeks of additional engineering and reduced my feature count by roughly 30 percent because a lot of what I thought was valid turned out to have some form of look-ahead bias. But it was the single highest-ROI investment I made in the entire project.

Get the Full Details

(PDF) Book Review: Marcos Lopez de Prado: Advances in Financial Machine Learning, Wiley, 2018
(PDF) Book Review: Marcos Lopez de Prado: Advances in Financial Machine Learning, Wiley, 2018

Feature construction that actually moves the needle

Volatility-adjusted momentum, cross-sectional dispersion measures, and mean-reversion signals based on short-term deviations from longer-term moving averages are the bread and butter. But the edge comes from how you combine and transform them. Simple things like winnowing your feature set down to the top 50 by a combination of SHAP importance and stability across multiple windows tend to work better than feeding 400 features into any model. I found that using elastic net regularization for feature selection before switching to tree-based models produced more stable portfolios than just ranking by importance. The L1 penalty pushes many coefficients exactly to zero, which gives you a cleaner feature set. Then the tree model captures the nonlinear interactions among the survivors. This two-step approach reduced portfolio turnover by about 25 percent compared to using the tree model alone for feature selection.

When trees fail and what to do about it

Gradient boosting models like XGBoost and LightGBM dominate this space because they handle heterogeneous feature types well and don't require heavy preprocessing. But they have a blind spot that catches a lot of people. They struggle with extrapolation, meaning they perform poorly on regimes that haven't appeared in training data. Market crashes, flash events, and sudden macro shifts all fall into this category. A model trained on five years of relatively calm markets will almost certainly overfit the calmness. The practical fix I use is ensemble stacking. Instead of relying on a single model, I run a lightGBM for the base predictions, a linear model on a smaller set of stable features, and then a simple meta-learner that combines them. The linear component acts as a fallback when the tree model encounters out-of-distribution inputs. It's not a perfect solution but it typically improves worst-case drawdowns by 15 to 20 percent during stress periods.

Execution and infrastructure

The infrastructure side of financial ML is where most teams either succeed quietly or fail loudly. You need proper walk-forward validation, not just a single train-test split. Rolling windows that expand as you move through time simulate real deployment far better than static splits. I usually run with 60-day expansion windows and 20-day forward steps, which gives me a reasonable balance between computational cost and realistic evaluation. Signal decay is another practical concern that is easy to underestimate. Most alpha signals in equities decay on a timescale of days to weeks at best. By the time your model trains and deploys, the signal you optimized for may have partially or fully evaporated. Shortening your rebalancing cycle, increasing transaction cost assumptions to at least 10 to 20 basis points per trade, and accepting lower raw returns in exchange for higher robustness is the tradeoff you need to make. A strategy with a backtested Sharpe of 1.5 after realistic costs is far more likely to survive than one showing 3.0 before costs.

Advances in Financial Machine Learning - Lopez de Prado, Marcos | 9781119482086 | Amazon.com.au ...
Advances in Financial Machine Learning - Lopez de Prado, Marcos | 9781119482086 | Amazon.com.au ...

Honest limitations

Financial ML will not replace the need for domain knowledge. A model trained purely on historical price and volume data will eventually converge to something close to a random walk because markets adapt. The edge comes from incorporating structural knowledge about how markets actually work, like understanding order flow, liquidity dynamics, and how different participant classes behave under stress. Pure data-driven approaches hit a ceiling that no amount of computing power can. Computational cost is also a real constraint. Full walk-forward backtests with purged cross-validation on multi-asset universes can take hours to days depending on your feature count and universe size. I typically use a combination of parallel feature generation, GPU-accelerated tree training, and caching intermediate results to keep iteration times under two hours. Anything slower and the feedback loop between hypothesis and validation becomes too long to be useful.