8 Minimum Track Record Length Under Fat Tails
How Long Must Performance History Be Before Skill Is Declared?
Status: Replication in progress — headline numerical estimate held
8.1 Motivation
Bailey and López de Prado (2012) derive the Minimum Track Record Length (MinTRL) — the number of observations needed to reject, at a given confidence level, the null hypothesis that a manager’s Sharpe ratio is zero. The derivation assumes finite fourth moment (\(\kappa > 4\)).
Under an AR(1)-GARCH(1,1) model, Kao et al. (2026) show that the estimation error around a nonzero population Sharpe has a model-dependent stable limit plus bias when \(\kappa \in (2, 4)\), with normalization \(n^{1-2/\kappa}\). That is narrower than a universal rule for testing a zero Sharpe.
Turning this limit result into a practical table requires more than its convergence exponent. The stable-limit scale, tail balance, dependence, volatility dynamics, target Sharpe, and finite-sample coverage must also be identified.
8.2 The Claim
At \(\hat{\alpha} = 2.69\) (the median in our exploratory equity sample), the tail index alone does not identify a minimum track record. The previously reported 1,100-year point estimate is withdrawn: it omitted the theorem’s bias, assumed a unit stable scale, and substituted a Gaussian quantile without establishing that the result was conservative.
A direct finite-sample diagnostic calibrated separate 5% one-sided tests under iid monthly Gaussian and Student-t(3) returns. For a population annual Sharpe of 0.5, a 3-year history had about 22–26% power, a 5-year history about 31–36%, a 20-year history about 72–75%, and a refined grid first exceeded 80% power near 25 years (300 months for Gaussian and 290 for Student-t(3)). Monte Carlo uncertainty and the discrete grid make this an approximate crossing, not a new point estimate. These figures are conditional on those full models and are not a replacement universal MinTRL.
The practical implication is more precise: short track records can have low power, but sufficiency must be reported as a size-and-power curve for an explicit return process, effect size, estimator, frequency, and dependence structure. A tail-exponent badge cannot make that decision.
8.3 Key prior work
- Bailey and López de Prado (2012) — Probabilistic Sharpe Ratio and MinTRL
- Bailey and López de Prado (2016) — Deflated Sharpe Ratio
- Kao et al. (2026) — convergence rate under fat tails (critical new anchor)
- Vlasiuk (2025) — Lévy-stable scaling of performance metrics
- Gap: calibrated size, power, and coverage comparisons across explicit heavy-tailed and dependent return models. Novelty: promising, subject to replication and literature review.