← Main site

Decision brief · Portfolio construction

Why “optimal” portfolios fail out of sample

In this frozen-window study, the maximum-Sharpe portfolio’s performance declined when fixed weights were evaluated on later data.

Depth 1 · Answer

Risk-based methods delivered a more stable risk shape—not a reliable return advantage.

Maximum-Sharpe mean–variance optimization fell from a Sharpe ratio of about 1.12 in the 2015–2021 fit period to 0.44 in the sealed 2022–2026 evaluation. Risk parity and hierarchical risk parity ran at roughly half the volatility and two-thirds the drawdown of the optimized books, but they did not reliably win on return. That narrower claim is what the evidence supports.

Published 8 August 2026 · Underlying investigation last updated 6 July 2026

MVO Sharpe · fit
1.12
MVO Sharpe · evaluation
0.44
Fit window
2015–2021
Sealed evaluation
2022–2026

Why it matters

Allocators, risk committees, and model reviewers

The result matters to anyone using an optimizer to convert noisy expected returns and covariances into precise weights—and to anyone evaluating whether an analytical system distinguishes numerical sophistication from decision reliability.

Decision affected

How much authority to give an optimized allocation

Do not interpret an in-sample efficient frontier or Sharpe ranking as a forecast. Decide first whether the mandate is return maximization, concentration control, or a stable risk budget, then evaluate the construction method against that stated objective.

Depth 2 · Evidence

The ranking changed when the estimates left the fit window

The experiment used 2,891 daily observations for 15 liquid ETFs from a frozen adjusted-close snapshot. Maximum-Sharpe MVO, box-uncertainty robust MVO, equal-risk-contribution risk parity, and hierarchical risk parity were fit on 2015–2021 data. Their weights were then frozen and applied to one untouched 2022–2026 window.

The robust portfolio posted the best out-of-sample Sharpe, about 0.6, but it remained concentrated in QQQ and benefited from the 2023–2025 mega-cap rally after drawing down more than 30% in 2022. The box uncertainty set reduced estimated return levels without changing their ranking, so it did not solve concentration. One 4.5-year window cannot distinguish a robust method from a fortunate regime.

Differences of ±0.2 in Sharpe are well inside an estimated sampling error near 0.5 for this evaluation window. The defensible conclusion is about realized risk shape, not which method will earn the highest future return.

Depth 3 · Application

Treat optimization as an assumption-sensitive proposal

  • Separate the return objective from the risk-budget objective before comparing methods.
  • Report concentration, risk contribution, drawdown, and out-of-sample behavior beside the in-sample optimum.
  • Make the uncertainty treatment explicit: shrinking return levels is not the same as changing uncertain rankings.
  • Use multiple rolling-origin evaluations before making comparative performance claims; this study did not perform them.
  • State each method's regime dependency. Risk-based allocations can embed a bond and stock–bond-correlation bet.

Depth 4 · Limits

What this result does not establish

  • There is one sealed evaluation window, not a distribution of rolling-origin outcomes.
  • The 15-ETF universe was selected in 2026 and is therefore survivorship-tilted.
  • Weights were frozen; rebalancing, transaction costs, and drift were outside this experiment.
  • All methods received the same raw covariance estimate. Covariance shrinkage was available but deliberately not used.
  • The study does not prove risk parity or HRP will outperform, nor that MVO will always underperform.

Depth 5 · Method and code

Move from the implication to the implementation

Related research