研究

Better by design: Why human choices matter for return predictions via machine learning

Machine learning (ML) models have become increasingly popular for predicting stock returns, both in academic research and industry practice. However, as a still developing field, we see a lot of variety when it comes to key design choices. Recent research systematically explores this and uncovers how these choices directly affect the performance of ML strategies.

作者

    Researcher

概要

  1. Humans still have many choices to make when designing machine learning strategies
  2. These choices have a substantial impact on the performance of machine learning strategies
  3. Machine learning models tend to outperform linear models only for certain design choices

Choices choices

The paper by Minghui Chen, Matthias Hanauer, and Tobias Kalsbach, titled ‘Design choices, machine learning, and the cross-section of stock returns’, identifies several key design choices researchers have to make when training ML models. For instance, when setting the prediction (target) variable, should the researcher employ the excess return over the risk-free rate or the abnormal return relative to the market? Is it better to use a continuous target variable or are categories, such as outperformers vs. underperformers, preferable? Is it better to train models based on a rolling window that leads to more adaptive models, or are models based on expanding windows superior, thanks to the availability of more training data?

To assess the importance of such choices, the authors identify seven such key design choices and examine all the ensuing possible combinations, resulting in a total of 1,056 ML models. In this way, the study trains each model on a common set of signals (features) for the US stock market and evaluates their out-of-sample performance using hypothetical top-minus-bottom decile portfolios.

Figure 1 reveals that portfolio returns vary substantially across different model designs, with monthly mean returns ranging from 0.13% to 1.98% and annualized Sharpe ratios ranging from 0.08 to 1.82.1 This variation highlights the substantial impact of human design choices on the performance of ML strategies.

Figure 1 | Cumulative performance of machine learning strategies

Source: Robeco, Chen et al. (2024). This figure shows the cumulative performance of a USD 1 initial investment in long-short ML portfolios for each possible combination of the research design choices. For each ML model and month, we first cross-sectionally sort all stocks based on their one-month-ahead return predictions. We then construct the value-weighted long-short portfolios by going long the top decile and short the bottom decile stocks. The solid black line represents the strategy with the median cumulative performance for each month, and the dashed black lines represent the 10th and 90th percentiles of each month, respectively. The sample period is from January 1987 to December 2021.

Machine learning models: Separating the wheat from the chaff

Having documented the substantial variation in the performance of ML models, the study also provides actionable guidance for ML model design:

  • Ensembles of ML models typically outperform individual algorithms.

  • The choice of target variable depends on the investment objective:
    o For identifying relative winners and losers among stocks, predicting stock returns over the market rather than the risk-free rate is better.
    o If the goal is to achieve high market-risk-adjusted returns, CAPM beta-adjusted returns are better.

  • Non-linear ML models are more likely to outperform their linear counterparts when:
    o using abnormal returns relative to the market as the target variable,
    o employing continuous target returns, or
    o adopting expanding training windows.

Conclusion

While computational infrastructure, ML algorithms, and data have become significantly more accessible over the past decade or two, model design remains a critical component of success. At first glance, it might seem that an ML investment strategy only requires a few basic elements: cloud computing space, generic factor data, some Python packages, and a couple of data scientists. However, this approach often lacks the crucial domain knowledge that Robeco has cultivated over 20 years in quant investing. That’s why in financial markets, where the signal-to-noise ratio is low and the risk of overfitting high, investment experience, and economic intuition still play a pivotal role. Robeco’s extensive expertise ensures that ML models focus on meaningful patterns and avoid common pitfalls, bridging the gap between technology and investment insight.

Active Quant: finding alpha with confidence

Blending data-driven insights, risk control and quant expertise to pursue reliable returns.

Find out more

Footnote

1Please note that these are hypothetical gross returns for long-minus-short strategies that do not consider any transaction costs. We investigated the impact of transaction costs on ML strategies in our study ‘The term structure of machine learning alpha’.

重要資料

本網站僅供《證券及期貨條例》(香港法例第571章)及其附屬法例所界定之專業投資者瀏覽及使用。 投資涉及風險。過往表現並不代表未來表現。本網站所載資料僅供參考之用,並不構成任何投資建議,亦非作出買賣任何證券或採納任何投資策略之要約或招攬。投資者不應僅憑本網站提供之資料作出投資決定,在作出任何投資決定前,應徵詢獨立意見(包括有關稅務影響之意見)。投資者應確保完全理解投資產品的相關風險,亦應考量自身投資目標及風險承受水平。投資乃閣下之個人決定。除非銷售投資產品的中介人已向閣下告知該投資產品適合閣下,並已解釋其符合閣下投資目標之原因,否則閣下不應投資。請參閱相關發售文件或其他法律文件,以獲取包括風險因素在內的進一步詳情。 本網站由荷寶投資管理香港有限公司發布,該公司受香港證券及期貨事務監察委員會(「證監會」)規管(中央編號:APU851)。本網站未經證監會審閱。 無法保證任何投資產品可實現其投資目標。概不就任何投資產品之表現或投資回報作任何聲明或承諾。投資的價值或會波動。本網站所載過往表現、推算或預測,均不應視作未來表現之保證或指標,且概不提供任何明示或暗示之保證。本網站內容建基於相信為可靠之來源,惟因應資料傳遞技術特性及須採用多項數據來源(包括第三方內容),故概不保證其準確性。所述觀點僅乃截至上述日期,或會隨市況變化而改變,可予更改而毋須另行通知。該等意見可能有別於其他荷寶投資專業人士之意見。因使用本材料或當中所載任何評論、意見或估算而引致之直接、間接或相應損失,荷寶概不承擔法律責任。荷寶並無責任更新本網站或任何網站內容。未經荷寶事先書面許可,不得複製、分發或刊發本網站任何材料。 除非另有說明,資料來源:荷寶。

警告 — 有不法分子在網站及社交媒體上冒用荷寳 了解更多