8  Minimum Track Record Length Under Fat Tails

How Long Must Performance History Be Before Skill Is Declared?

Author

Jean-Marc Choufani

Published

August 9, 2026

Status: Replication in progress — headline numerical estimate held


NoteChapter forthcoming

8.1 Motivation

Bailey and López de Prado (2012) derive the Minimum Track Record Length (MinTRL) — the number of observations needed to reject, at a given confidence level, the null hypothesis that a manager’s Sharpe ratio is zero. The derivation assumes finite fourth moment (\(\kappa > 4\)).

Under an AR(1)-GARCH(1,1) model, Kao et al. (2026) show that the estimation error around a nonzero population Sharpe has a model-dependent stable limit plus bias when \(\kappa \in (2, 4)\), with normalization \(n^{1-2/\kappa}\). That is narrower than a universal rule for testing a zero Sharpe.

Turning this limit result into a practical table requires more than its convergence exponent. The stable-limit scale, tail balance, dependence, volatility dynamics, target Sharpe, and finite-sample coverage must also be identified.

8.2 The Claim

At \(\hat{\alpha} = 2.69\) (the median in our exploratory equity sample), the tail index alone does not identify a minimum track record. The previously reported 1,100-year point estimate is withdrawn: it omitted the theorem’s bias, assumed a unit stable scale, and substituted a Gaussian quantile without establishing that the result was conservative.

A direct finite-sample diagnostic calibrated separate 5% one-sided tests under iid monthly Gaussian and Student-t(3) returns. For a population annual Sharpe of 0.5, a 3-year history had about 22–26% power, a 5-year history about 31–36%, a 20-year history about 72–75%, and a refined grid first exceeded 80% power near 25 years (300 months for Gaussian and 290 for Student-t(3)). Monte Carlo uncertainty and the discrete grid make this an approximate crossing, not a new point estimate. These figures are conditional on those full models and are not a replacement universal MinTRL.

The practical implication is more precise: short track records can have low power, but sufficiency must be reported as a size-and-power curve for an explicit return process, effect size, estimator, frequency, and dependence structure. A tail-exponent badge cannot make that decision.

8.3 Key prior work

  • Bailey and López de Prado (2012) — Probabilistic Sharpe Ratio and MinTRL
  • Bailey and López de Prado (2016) — Deflated Sharpe Ratio
  • Kao et al. (2026) — convergence rate under fat tails (critical new anchor)
  • Vlasiuk (2025) — Lévy-stable scaling of performance metrics
  • Gap: calibrated size, power, and coverage comparisons across explicit heavy-tailed and dependent return models. Novelty: promising, subject to replication and literature review.
Bailey, David H., and Marcos López de Prado. 2012. “The Sharpe Ratio Efficient Frontier.” Journal of Risk 15 (2).
Bailey, David H., and Marcos López de Prado. 2016. “The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting, and Non-Normality.” Journal of Investment Management 14 (3).
Kao, Lie-Jane, Cheng-Few Lee, and Han-Hsing Lee. 2026. “Estimated Sharpe Ratio of Asset Returns with Fat Tails: Theory and Empirical Evidence.” Review of Quantitative Finance and Accounting 67: 869–89. https://doi.org/10.1007/s11156-025-01474-6.
Vlasiuk, Dmytro. 2025. Lévy-Stable Scaling of Risk and Performance Functionals.