5 HAC Repairs Precision, Not Contaminated Measurement
Construct the Right Estimand Before Correcting Its Standard Error
SSRN: abstract_id=7022678 Status: Corrected working-paper chapter
The measurement-versus-inference distinction survives. The previously reported approximately 1.6-times fat-tail amplification did not construct a Gaussian counterfactual and is withdrawn. The timing-clean aggregate PA relation also cannot be called a genuine independent signal because it later fails the Healthcare exclusion.
5.1 Abstract
Heteroskedasticity-and-autocorrelation-consistent standard errors can repair inference for a specified estimator under their regularity conditions. They cannot change a predictor or outcome constructed with future information, mismatched dates, or an invalid sample definition. This chapter separates measurement validity from inference validity. A rolling predictor that includes outcome-period observations estimates a contaminated quantity; HAC can estimate uncertainty around that quantity but cannot recover the point-in-time estimand. A second example comes from benchmark clocks: subtracting a benchmark return over nominal dates from a security return over shifted realized dates is not common-window excess return. Correct standard errors cannot repair either clock. Clean construction and robust inference are both necessary, and neither establishes economic independence, causality, or actionability.
5.2 1. Two layers of validity
- Measurement validity. Predictor, outcome, benchmark, eligibility, and sampling unit are constructed from information available at the declared time and over the intended common window.
- Inference validity. Uncertainty accounts for serial correlation, clustering, overlap, heteroskedasticity, and multiplicity appropriate to the actual evidence unit.
Newey-West HAC (Newey and West 1987) operates in the second layer. Passing it does not imply that the first layer, a sector test, a selection test, or a mechanism test has passed.
5.3 2. Overlapping construction
Let a trailing statistic be computed at formation date \(t\) from observations ending at \(t\) and let its outcome use sessions \(t+1\) through \(t+h\). A clean predictor is
\[ \theta^{clean}_{i,t}=g\!\left(r_{i,t-T+1},\ldots,r_{i,t}\right). \]
If the predictor is instead recomputed at \(t+h\), it includes observations from the outcome window:
\[ \theta^{cont}_{i,t}=g\!\left(r_{i,t-T+h+1},\ldots,r_{i,t+h}\right). \]
The second object may be useful descriptively, but it is not a point-in-time predictor. Shared returns mechanically couple the predictor to the outcome. HAC changes the estimated covariance of an estimator built from \(\theta^{cont}\); it does not transform \(\theta^{cont}\) into \(\theta^{clean}\).
The empirical sign reversal between clean and contaminated payoff-asymmetry constructions remains a useful demonstration of this problem. Its magnitude cannot be relabeled as a universal heavy-tail amplification factor without controlled model comparisons.
5.4 3. Clean timing is not signal certification
The timing-clean aggregate payoff-asymmetry relation also should not be called a genuine independent signal. It passes a timing and HAC check, then fails a sector-composition check:
| Sample | PA spread | Newey-West \(t\) |
|---|---|---|
| All sectors | +6.08% | +2.65 |
| Ex-Healthcare | -0.27% | -0.09 |
This is exactly why measurement and inference controls are necessary but not sufficient. Passing one or two gates cannot substitute for composition, multiplicity, implementation, and prospective tests.
5.5 4. The benchmark-clock problem
Suppose a security has no close on the intended entry date and its realized holding window begins on the next available session. Subtracting the benchmark’s nominal-window return compares different clocks. The result called excess return is then not the return difference over a common period. No HAC adjustment can repair the mismatch.
TrailMap now aligns benchmark entry and exit to the security’s realized dates or leaves the outcome unresolved. The registered prospective study stores intended and realized dates and compares nominal-window with exact-window excess returns at 20, 60, and 120 sessions. It reports continuous differences and direction-label flips, stratified by listing age, liquidity, halt/suspension, symbol event, and number of shifted sessions.
5.6 5. Diagnostic sequence
For any rolling-window claim:
- freeze eligibility and information timestamps;
- construct predictor, security outcome, and benchmark on declared common clocks;
- identify the actual evidence unit and dependence clusters;
- estimate uncertainty with the method appropriate to that construction;
- test sector, cohort, regime, multiplicity, and implementation sensitivity; and
- require new preregistered observations before promotion.
A large HAC \(t\)-statistic on a contaminated or composition-dependent estimate is evidence about that estimate, not proof of an independent economic signal.
5.7 6. Conclusion
Construct the right estimand first, then estimate its uncertainty. The 1.6-times amplification and genuine-signal interpretations are withdrawn. The surviving contribution is the logical separation between measurement validity and inference validity, now extended to exact benchmark clocks.
→ Chapter 5 generalizes this sequence into a data-gated falsification system.