Linear Models Beyond Return Prediction
How signal neutralization, risk modeling, and attribution fit alongside nonlinear return forecasts.
Contents
A stock-selection model can learn nonlinear interactions from thousands of inputs and still pass its forecasts to a portfolio built around linear factor exposures. The same investment process may use regression to examine a new signal, estimate common sources of risk, and explain a month’s returns. There is no contradiction here. These calculations solve different problems.
Discussions of linear regression in quantitative investing often begin with the difficulty of predicting returns: financial data are noisy, flexible models overfit, and simple models are easier to trust. That is a reasonable argument for a linear baseline. It is a much weaker explanation for why linear methods remain useful throughout an investment process. Better prediction does not eliminate the need to understand what a signal overlaps with, how a portfolio’s holdings move together, or which exposures contributed to a loss.
I find it more useful to follow a signal through these different tasks than to ask whether linear or nonlinear models are better in the abstract. What changes when we regress a signal on known exposures, replace that signal with realized returns, or aggregate the result into a portfolio? The algebra looks similar, but the interpretation changes at each step. Those changes reveal both the usefulness of linear structure and the limits of what it can tell us.
Prediction is only part of the problem#
Suppose we are researching a signal based on changes in analysts’ earnings estimates. An upward revision might be informative on its own, or its value might depend on valuation, recent price movements, and the breadth of revisions across analysts. Those interactions are a reason to consider a flexible predictor.
The statistical target is the conditional expected return,
where is the vector of stock returns over the next holding period and is the information available when the decision is made. A linear predictor approximates this target with a weighted sum of features. A tree or neural network can represent a richer relationship between the same information and future returns.
Neither choice settles the empirical question. A linear model can overfit a large feature set or repeated research on the same backtest. A nonlinear model can learn useful interactions when the data and validation support them. Gu, Kelly, and Xiu’s Empirical Asset Pricing via Machine Learning, for example, finds predictive gains from trees and neural networks in its study of equity returns, with interactions contributing to their performance. Such evidence makes it difficult to defend linearity as a universal answer to financial noise.
A linear baseline remains valuable because it makes the incremental contribution of complexity easier to see. The comparison should use the same information, holding horizon, evaluation periods, and trading assumptions. The question is whether the additional flexibility produces an improvement that survives subsequent observations and implementation costs.
Now imagine that the earnings-revision model passes that comparison. Its highest-ranked stocks are mostly small companies with strong recent price performance. We have a useful forecast, but we have also acquired a new question: how much of its value comes from those familiar exposures? Predictive accuracy alone cannot answer it. This is where regression takes on a different role.
A residual answers a conditional question#
At a particular date, collect the signal values for all stocks in a vector . Let the columns of contain the exposures we want to control for, such as industry membership, standardized log market capitalization, and momentum. Each row now represents one stock at the same date: this is a cross-sectional regression.
We can decompose the signal into a fitted component and a residual:
The diagonal matrix assigns positive weights to the stocks. Equal weights give ordinary least squares. The fitted component is the part of the signal that can be represented by a linear combination of the selected exposures. The residual is what remains after that projection.
For an unpenalized least-squares fit, the normal equations give
This is the useful guarantee: on this universe, under these weights, the residual is orthogonal to the included exposure columns. With an intercept, it also has zero weighted mean. No future returns enter the calculation. We are describing the signal we have today, before asking whether it predicts anything tomorrow.
For the earnings-revision signal, comparing the original score with its residual lets us ask a more focused question. Does the signal distinguish future winners after removing its linear overlap with size, industry, and momentum? If the residual remains useful, that is evidence of incremental predictive information relative to those controls. If its performance disappears, the signal may largely have repackaged an existing exposure.
The controls define the comparison. They do not define a universal boundary between skill and risk. Momentum may be an exposure the mandate deliberately seeks; removing it could discard useful return information. Conversely, a residual can still contain nonlinear dependence on size, an omitted common risk, or plain noise. Orthogonality is a geometric property of the fitted sample. It does not establish independence, causality, or “pure alpha.”1
There is a further distinction between a residual signal and a neutral portfolio. Ranking the residual, selecting its top names, or imposing position limits can change its exposures. A portfolio with weights has factor exposure , which must be evaluated on the actual holdings. Even weights proportional to need not be neutral when the regression used unequal weights: does not imply .
I would therefore read neutralization as a way of refining a research question. It tells us which component we have removed and lets us compare the investment value with and without it. It does not decide in advance which exposures the strategy ought to keep.
From stock returns to shared risk#
The same exposure matrix can serve another purpose once returns have been realized. Replace the signal on the left-hand side with stock returns:
In this fundamental factor model, exposures are constructed from stock characteristics available at the beginning of the period. A cross-sectional regression after the period ends estimates the realized factor returns and the stock-specific residuals . The coefficients now have a different meaning because the response has changed. A momentum coefficient describes how realized returns aligned with momentum exposure, conditional on the other model columns. It was not a forecast available at the beginning of the period.
This distinction also clarifies the overloaded word factor. A stock’s momentum characteristic is an exposure. The estimated payoff associated with that exposure over a completed period is a factor return. A characteristic can be useful for describing common variation even if it offers no reliable positive expected return. Risk modeling does not require every risk factor to be an alpha signal.
Repeating the regression gives a history of factor returns and specific returns. That history can support a structured covariance forecast:
where forecasts factor-return covariance and forecasts specific-return covariance. The usual simplified model makes diagonal and assumes factor and specific innovations are uncorrelated. These are assumptions about return variation; they do not follow from the orthogonality of residuals in a single cross section. The MOSEK factor-model chapter develops this covariance structure and the associated portfolio risk decomposition.
The reduction in dimension is substantial. For 4,000 stocks, an unrestricted symmetric covariance matrix contains 8,002,000 distinct entries. Given exposures to 50 factors, the factor covariance has 1,275 entries; adding 4,000 specific variances gives 5,275 covariance parameters. The exposure matrix still has to be constructed and maintained, but the covariance estimation problem is much smaller.
That economy comes from a substantive approximation: a limited set of common components accounts for the dependence we care about. It can fail when an omitted exposure becomes important or residual correlations rise during stress. Estimating factor returns also leaves the choice of covariance window, volatility adjustment, and regularization unresolved. A compact model makes those questions more tractable; it does not remove them.
For our earnings-revision strategy, the practical gain is the ability to see that apparently distinct stock picks may share the same risk. Fifty attractive forecasts are not fifty independent opportunities if they all load on the same industry or style. The return model ranks their appeal. The covariance model assesses how much diversification they actually provide.
The portfolio is where the models meet#
A forecast becomes economically meaningful through the positions it supports. The portfolio must reconcile expected returns with risk, existing holdings, trading costs, and the investment mandate. An illustrative objective is
subject to budget, position, and exposure constraints. The first term rewards forecast return, the second penalizes risk, and the third estimates the cost of moving from current holdings to the proposed portfolio. A benchmark-relative strategy can use active weights in its risk penalty and exposure constraints. Forecasts, risk estimates, and costs must refer to compatible horizons and units; a ranking score is not automatically an expected return.
Nothing in this formulation requires the forecast to be linear in its inputs. The optimizer consumes a vector of expected returns. A neural network and a ridge regression can both produce that vector. The covariance estimate and exposure constraints describe another part of the decision. Under suitable covariance, cost, and constraint assumptions, the problem can be solved as a convex quadratic program, as described in the MOSEK portfolio optimization guide.
This separation is especially useful for the earnings-revision example. Suppose the model finds a promising interaction among revisions, valuation, and momentum. A momentum limit may prevent the portfolio from expressing the entire forecast, but it need not erase the interaction from the prediction model. Instead, portfolio construction asks how much of that forecast can be used within the chosen risk budget. The resulting trade-off can be studied directly: what expected opportunity is being surrendered to reduce a particular exposure?
The same reasoning applies to turnover. A linear predictor based on rapidly decaying inputs can demand frequent trading. A nonlinear predictor can feed a slowly changing portfolio. Turnover and capacity depend on the signal’s horizon, portfolio constraints, liquidity, and trading costs. They cannot be inferred from the predictor’s model class.
A linear factor structure is therefore compatible with substantial predictive complexity. It supplies a tractable way to aggregate exposures and assess joint risk while leaving the forecast model free to represent useful nonlinear relationships. The quality of that arrangement depends on both components: a good forecast can be mishandled by a poor risk model, and a careful risk model cannot rescue a forecast with no investment value.
Attribution explains exposures, not decisions#
After the holding period, the factor model gives us a way to examine what happened. For beginning-of-period weights held over the period, gross portfolio return can be written as
The first contribution comes from the portfolio’s factor exposures and the realized factor returns. The second comes from the stock-specific residuals. This follows directly from substituting the fitted stock-return decomposition into portfolio return.
Suppose the earnings-revision portfolio loses money during a momentum reversal. Attribution can show how much of its loss is assigned to momentum under the chosen model and how much remains in the residual. That helps distinguish a loss associated with an accepted exposure from a loss the factor model does not explain. The distinction is useful even if a neural network generated every stock forecast.
It also has limits. The residual may contain omitted common risks, stock-specific news, or estimation error. It cannot automatically be labeled failed alpha. A contribution attributed to momentum is a model-based accounting statement, not proof that momentum was the underlying economic cause. Actual net performance also includes trading costs, financing, and changes in holdings; those require corresponding accounting adjustments.
There are consequently two different kinds of explanation. One concerns why the prediction model assigned a particular score to a stock. The other concerns how the resulting holdings were exposed when returns occurred. Factor attribution addresses the second. It can make a complex strategy’s realized exposures intelligible without explaining every internal decision of its predictor.
This is a more defensible reason to value interpretability than the promise that every loss can be explained away. A useful attribution system narrows the questions to investigate. It does not guarantee that the investment process is correct, or that the next loss will resemble the last one.
Choosing the approximation#
Following the earnings-revision signal leaves us with several distinct uses of linear structure. We used a projection to examine its overlap with familiar styles, a factor model to describe shared risk, and a return decomposition to interpret the portfolio’s outcome. None required the original forecast to be linear. Each made a different part of the investment process easier to reason about.
What matters to me is the choice of approximation at each stage. Neutralization chooses which relationships to set aside. A covariance model chooses which common movements to represent explicitly. Attribution chooses a vocabulary in which to describe gains and losses. These choices are useful precisely because they simplify the problem, but their omissions remain economically relevant. A portfolio can be neutral to the listed factors and still be concentrated in something the model has no name for.
That is why I would hesitate to describe linear regression either as an outdated predictor or as the inevitable solution to noisy markets. Both descriptions give too much weight to model class and too little to the question being asked. Nonlinear prediction and linear factor analysis can fit naturally into the same process, provided we keep their purposes separate and examine where their assumptions meet.
For later research, the question I want to carry forward is what an additional layer of complexity changes about the decision. It might improve a forecast, uncover dependence missing from the risk model, or show that a familiar attribution is misleading. Where a linear approximation remains adequate, its clarity is valuable. Where it hides something consequential, that same clarity should help us see why we need to go further.
The guarantee is specific to an unpenalized fit and the selected exposure span. With a ridge penalty , the normal equations instead give , so exact orthogonality generally disappears. Redundant columns, such as an intercept alongside a complete set of industry dummies, also require an identification convention if individual coefficients are to be interpreted. ↩︎