Building an Edgewater-Style Factor Model from Scratch
Chapter 35 covers multi-factor regression analysis applied to long-short equity portfolios. If you are coming from Chapter 34, this is where the math stops being textbook and starts looking like real work. The core concept is straightforward: you regress portfolio returns against a set of factor exposures and interpret what drives alpha versus what is just factor beta. Most people gloss over the mechanics because the theory seems simple. It is not. The chapter walks through orthogonalized factor construction, Fama-MacBeth procedures, and cross-sectional regression diagnostics. You learn how to build factors that do not overlap with each other, because overlapping factors produce multicollinearity that silently destroys your ability to tell which exposure is actually paying you. The textbook explanation spends about forty pages on least squares estimation. The practical part is figuring out why your model looks perfect in-sample and completely falls apart out-of-sample. Start with raw factor data. This means market cap, value, momentum, quality, volatility — however many factors your dataset supports. For each factor, construct a time series of returns. The simplest approach is to go long the top quintile and short the bottom quintile every month, rebalance at period end, and track the spread. That gives you one factor return observation per month. Do this for every factor over your full sample period.
Next, regress your portfolio's excess returns against these factor returns. The coefficients tell you your factor exposure. A coefficient of 1.5 on momentum means your portfolio behaves as if it is 150% levered into the momentum factor. The intercept is your alpha, or supposed alpha. But here is the part nobody emphasizes enough: you must also check the residual autocorrelation. If your residuals are autocorrelated, your standard errors are wrong. Your t-statistics are inflated. The significance you are seeing is probably a statistical artifact. I ran into this exact problem last year on a small quant team. We had built a mean-reversion strategy that showed a Sharpe of 1.8 in backtests. Chapter 35 diagnostics should have caught it immediately, but we skipped the Ljung-Box test on residuals because the R-squared was 0.82 and we were excited. When we finally ran the test, the residuals were clearly autocorrelated at lag one through lag six. The strategy's actual out-of-sample Sharpe was 0.6. We lost about four months of live trading capital before we fixed it. The workaround was simple: apply Newey-West standard errors to correct for the autocorrelation, then rebuild the strategy with a constraint that limits lag-one autocorrelation in the signal. The corrected Sharpe estimate was still positive, just not spectacular. It taught me to run diagnostic checks before celebrating any backtest result.
Advanced Nuances Beginners Miss
The first thing most people get wrong is factor orthogonality. They take standard factors from a provider like Axioma or Barra and assume they are independent. They are not. Value and quality are highly correlated because cheap stocks tend to be profitable ones. Momentum and volatility have a negative correlation because stocks that have dropped hard are volatile stocks. When you include correlated factors in a regression, the coefficients become unstable. Small changes in data produce wildly different loadings. The fix is to orthogonalize: run each factor against the previous ones and use the residuals. This guarantees independence but comes at a cost. Your orthogonalized factors become harder to interpret economically. That tradeoff is worth making, but you should know what you are giving up. The second mistake is timeframe mismatch. Some factors are monthly, some are quarterly. If you mix a monthly momentum factor with a quarterly profitability factor without aligning the observation dates, your regression will have missing values or stale data that biases the results. Always align your data to the same frequency and handle missing observations explicitly rather than letting the software drop them silently.
Get the Full Details

Practical Constraints and Where This Fails
Multi-factor regression assumes linear relationships. If your strategy has nonlinear payoffs — options-style exposure, for example — factor regression will misattribute risk and mask true alpha. It also assumes stationarity. Factor premia rotate. The value factor underperformed for nearly a decade after 2010. A model estimated on 2005–2015 data will give you very different exposure coefficients than one estimated on 2015–2024 data. You need to roll your estimation window periodically and track how factor loadings drift over time. If your portfolio has fewer than sixty months of history, factor regression becomes unreliable. The degrees of freedom are too low. Use principal component analysis as an alternative until you have enough observations. It does not give you interpretable economic factors, but it reduces dimensionality without the instability of OLS on thin data.
Step-by-Step Walkthrough
Step one is data collection. Gather daily or monthly portfolio returns and match them with factor returns over the same date range. Step two is descriptive statistics. Check means, standard deviations, and correlations between factors. If any factor pair has a correlation above 0.7, plan to orthogonalize. Step three is the regression. Run the OLS, save coefficients, standard errors, and residuals. Step four is diagnostics. Run Ljung-Box for autocorrelation, Breusch-Pagan for heteroskedasticity, and condition number analysis for multicollinearity. Step five is interpretation. Map significant coefficients to economic logic. An exposure to a factor you did not intend is a risk you need to decide whether to keep or hedge away. Step six is documentation. Write down every decision, especially any transformations you applied. You will forget why you orthogonalized a specific factor within six months if you do not record it. The downloadable material for this chapter includes a Python notebook with synthetic data, a ready-made regression pipeline, and a checklist for the diagnostic phase. I recommend using it not as a template to follow blindly but as a reference to compare against your own implementation. You will catch more errors by writing your own code and finding where it diverges from the example than by copying it verbatim.