4 The Hill Estimator’s Domain of Validity
Asymptotic Validity, Finite Thresholds, and Source-of-Tail Contamination
SSRN: abstract_id=7022640 Status: Live working paper; corrected manuscript in preparation
The posted version incorrectly combined two bandwidth regimes. Persistent contamination at a fixed fraction such as \(k/n=0.10\) does not prove inconsistency in the Hill regime \(k\to\infty\), \(k/n\to0\). If a jump is bounded and the structural tail is unbounded and regularly varying, the Hill threshold eventually exceeds that fixed jump. The asymptotic inconsistency theorem is therefore withdrawn. This chapter states the corrected result and the finite-threshold question that remains.
4.1 Abstract
The Hill estimator characterizes regularly varying tails. Its usual consistency result requires an intermediate sequence of order statistics: \(k\to\infty\) while \(k/n\to0\). Operational studies often use a fixed fraction such as the top 10%, which is a different estimand. A bounded event jump can remain inside that fixed threshold region even though it eventually disappears from a proper intermediate tail. The valid concern is therefore pre-asymptotic and bandwidth-specific: how much do event-attributed jumps alter the estimate at the threshold actually used? A historical eight-name subgroup comparison motivates that question but is too small and too judgment-dependent to answer it alone. The registered revision varies bandwidth, jump frequency and magnitude, sample length, and source-attribution rules before making an economic claim.
4.2 1. The estimator and its asymptotic domain
For positive observations with descending order statistics, one common Hill parameterization is
\[ \hat{\alpha}_k = \left(\frac{1}{k}\sum_{i=1}^{k} \log\frac{X_{(i)}}{X_{(k+1)}}\right)^{-1}. \]
Its first-order domain is regular variation:
\[ \lim_{u\to\infty}\frac{1-F(ux)}{1-F(u)}=x^{-\alpha}, \qquad x>0. \]
Regular variation concerns the far tail, not absolute continuity of the entire distribution. An atom at a finite point can coexist with a regularly varying tail.
4.3 2. The mixture and the two bandwidth regimes
Consider
\[ X=(1-B)Z_S+BJ, \qquad B\sim\operatorname{Bernoulli}(p), \]
where \(Z_S\) has an unbounded regularly varying upper tail and \(J\) is a fixed finite positive jump.
4.3.1 2.1 Intermediate sequence
Proposition 4.1 Proposition (bounded jumps leave the far tail). If \(J<\infty\), \(Z_S\) is unbounded and regularly varying with exponent \(\alpha_S\), and \(k\to\infty\) with \(k/n\to0\), then the order-statistic threshold diverges in probability. Consequently the fraction of bounded \(J\) observations above the threshold converges to zero, and the bounded atom does not by itself invalidate consistency for \(\alpha_S\).
The proposition does not cover an unbounded jump distribution, jumps whose magnitude grows with sample size, a bounded structural distribution, or a threshold selected inside the event-jump region. Those are separate models and require separate results.
4.3.2 2.2 Fixed operational fraction
When \(k/n=q>0\) is fixed, the threshold approaches a finite mixture quantile rather than moving into the far tail. If that quantile lies below \(J\), event jumps remain among the selected observations. The resulting statistic is a finite-threshold mixture functional. It can be useful descriptively, but it should not be interpreted as the intermediate-sequence tail exponent without sensitivity analysis.
4.4 3. Reproducible diagnostic
The TrailMap correction diagnostic used a Pareto structural tail with exponent 3, a fixed jump of 4, and jump probability 2%. At a top-10% bandwidth, the median jump share stayed near 20% as \(n\) rose from 1,000 to 100,000, and the median Hill estimate stayed near 2.65 rather than 3. Under \(k=\lfloor n^{0.4}\rfloor\), the median jump share fell to zero after the threshold exceeded the fixed jump, and the estimate moved toward 3. This simulation illustrates the distinction; it is not a universal correction factor.
4.5 4. What the historical subgroup can support
An earlier historical split labeled 70 names FAT_STRUCTURAL and eight FAT_BINARY. Those counts and their reported economic difference are withdrawn because:
- the small subgroup does not identify a stable magnitude;
- classification depended on judgment rather than a versioned adjudication protocol;
- sector composition can explain some or all of the difference; and
- using the same sample to define and evaluate a category risks selection bias.
They are preserved only as the origin of a hypothesis, not as empirical evidence. A confirmatory study must persist event evidence, classification rules, adjudicator uncertainty, bandwidth sensitivity, and prospective outcomes.
4.6 5. Registered tests
The revision will vary:
- \(k/n\) and intermediate sequences;
- fixed, random, and heavy-tailed jump magnitudes;
- jump probabilities and clustering;
- sample lengths and structural exponents;
- event-attributed removal versus winsorization versus mixture modeling; and
- deterministic versus adjudicated source labels.
The claim is falsified operationally if source attribution and bandwidth changes leave estimates, uncertainty, and downstream decisions inside predeclared equivalence bounds.
4.7 6. Conclusion
The Hill estimator is not invalid merely because a return series contains bounded jumps. The defensible lesson is narrower and more useful: an operational fixed threshold can cross multiple data-generating regions, so the analyst must show threshold stability and source sensitivity before assigning a structural interpretation. That is a finite-sample measurement problem, not the asymptotic theorem previously claimed.
→ Chapter 4 considers a different separation: correcting dependence in inference does not correct bias in construction of the estimand.