The Alt/BTC Pairs Are Not Cointegrated - and the Half-Life Is 472 Days
The claim that needed settling
Cycle 2 of this series rejected a Kalman dynamic hedge ratio on 12 alt/BTC spreads, and traced the
failure to a specific cause: a log-level filter estimates a cointegration slope, and — in that
report's words — "these pairs are not cointegrated."
That was an aside, not a measurement. And it had quietly become load-bearing: it was the
reopening condition on the previous card, and it governed a whole class of future pairs-trading
work. So this cycle measured it.
Verdict: the pairs are not cointegrated, it is a property of the assets rather than of the
sample window, and the reversion that does exist is far too slow to pay 17.6 bps. All three
preregistered economic gates failed.
The trap, and how it was avoided
With 50,075 hourly bars per spread, an ADF or Johansen test has enormous power. It will reject
a unit root on a reversion so slow it could never pay a round trip. A p-value cannot answer a
tradability question at that sample size.
So the test's size and power were simulated at this exact sample length before a single real
p-value was read, and the statistical and economic axes were reported strictly apart, with the
economic one carrying the decision.
ADF size (500 simulated driftless random walks): 5.2% at n = 50,075 against a nominal 5%;
5.8% at the 2,160-bar rolling window. Correctly sized.
ADF power, 300 reps per cell:
| true half-life (bars) | 24 | 168 | 336 | 720 | 1,440 | 2,880 | 5,760 | 11,520 | 23,040 |
|---|---|---|---|---|---|---|---|---|---|
| full sample (n = 50,075) | 1.000 | 1.000 | 1.000 | 1.000 | 0.933 | 0.430 | 0.163 | 0.097 | 0.053 |
| rolling window (n = 2,160) | 1.000 | 0.237 | 0.083 | 0.073 | 0.057 | 0.067 | 0.050 | 0.060 | 0.040 |
The full-sample test sees everything down to a 60-day half-life. That matters both ways: a
rejection would have been consistent with a two-month half-life, and the absence of rejections
is informative rather than a power failure.
Statistical verdict
Stated before any count of what passed: family = 12 pairs, Holm–Bonferroni at FWER 0.05. At an
uncorrected α = 0.05 the expected number of false positives under the global null is
12 × 0.05 = 0.60 pairs, and P(at least one) = 45.96%.
| test | raw α = 0.05 | Holm FWER 0.05 | BH q = 0.10 |
|---|---|---|---|
| ADF on the traded (1,−1) spread | 1 / 12 | 0 / 12 | 0 / 12 |
| Engle–Granger, estimated β | 1 / 12 | 0 / 12 | 0 / 12 |
| Johansen trace | 5 / 12 (95%) | 3 / 12 (99%, Bonferroni proxy) | — |
| KPSS, null = stationary | 12 / 12 reject stationarity | — | — |
One raw rejection (DOGE, p = 0.0159, Holm-adjusted 0.191) against 0.60 expected is exactly what
the null predicts. The statsmodels ADF t-statistic was reproduced by an independent plain-numpy
OLS at the same lag order, max |difference| 1.7 × 10⁻¹².
Johansen disagrees — so its own vector was priced instead of argued with
Johansen rejects on 5 of 12 where ADF and Engle–Granger reject on 1. The obvious suspect was the
textbook k_ar_diff = 1 against ADF's AIC-selected 27–57 lags. It is not: the count is 5 at
k_ar_diff 1, 6, 12, 24, 48 and at the VAR-AIC order.
Rather than adjudicate, Johansen's own estimated cointegrating vector was taken at face value, the
implied spread formed, and measured:
- on the 5 pairs it accepts, the implied spread has a half-life of 2,064–5,260 bars (86–219
days), median 2,421 bars ≈ 101 days; - 0 of those 5 revert inside 336 bars;
- on 6 of 12 pairs the vector has a negative BTC coefficient — long the alt and long BTC.
That is not a hedge, not dollar-neutral, and cannot be traded as a spread at all.
The statistical axis can be granted in full and the economic answer does not move.
Economic verdict
AR(1) on the log spread gives b between 0.999718 and 1.000013:
| bars | days | |
|---|---|---|
| fastest pair (XRPBTC) | 2,455 | 102 |
| panel median | 11,326 | 472 |
| slowest (ATOMBTC) | 84,049 | 3,502 |
| DOTBTC | b > 1 — no half-life at all |
The fastest pair on the panel has a 102-day half-life, and the full-sample test has power
against everything down to 60 days. There was nothing to find in the range where the test can see.
Amplitude is the wrong question
For a near-unit-root series the "equilibrium" σ is dominated by diffusion, not by recoverable
reversion: σ_eq · (1 − 2^(−h/HL)) → σ_ε · √h as HL → ∞. A pure random walk passes any amplitude
test at any horizon, and the amplitudes here are indeed thousands of bps. You cannot trade a
random walk's variance. So the gate measured the predictable part directly: regress the h-bar
forward change of the spread on a causal z-score, HAC standard errors, and read −γ as the expected
capture of a round trip entered at |z| = 1.
Gate E1 — capture ≥ 52.85 bps (3 × the 17.616 bps round-trip commission), Holm-surviving over
the 84-test family, on ≥ 6 of 12 pairs → achieved on 1 of 12. And that one qualifier sits at a
220-day horizon with 8.3 independent observations.
At tradable horizons the sign is momentum
The horizons a 17.616 bps round trip can be paid at are h = 24, 168, 336 bars. Signed
(positive = reversion, negative = momentum):
28 of 36 cells are momentum. 8 are reversion. Zero reversion cells survive Holm. The largest
reversion capture anywhere at h = 24 is LTCBTC's +13.19 bps — below the round trip on its own,
before slippage.
The half-life-matched rule earns nothing
The arm's one free number was read off the measurement, never tuned: median half-life 11,326
bars ⇒ lookback 4,320, max hold 8,640.
| half-life-matched (preregistered) | tradable 168 (steelman, WATCH only) | |
|---|---|---|
| trades | 1,263 | 9,440 |
| GROSS | +0.043 %/yr | +0.399 %/yr |
| net | −0.504 %/yr | −3.917 %/yr |
| breakeven | 0.00 bps/side | 0.00 bps/side |
| years positive | 3/7 | 0/7 |
| net %/yr at maker 3.0 / 1.5 bps | −0.19 / −0.10 | −1.58 / −0.97 |
| gate | required | measured | |
|---|---|---|---|
| E1 capture ≥ 52.85 bps on ≥ 6 of 12 | 6 pairs | 1 pair | FAIL |
| E2 portfolio GROSS annualised | ≥ 9.36 %/yr | +0.043 %/yr | FAIL — 218× short |
| E3 portfolio breakeven per side | ≥ 13.2 bps | 0.00 bps | FAIL |
| validity floor | ≥ 300 trades | 1,263 | pass |
A breakeven of 0.00 bps is stronger than "too expensive." The rule does not clear any
commission — there is no gross edge for a cost to eat. At a 1.5 bps maker rate it is still
negative. Free execution does not manufacture an edge that is not there.
Is it regime-conditional? No.
2,160-bar (90-day) rolling windows, weekly step, 286 windows per pair:
| measured | simulated null | |
|---|---|---|
| rejection fraction, panel median | 6.29% | 5.80% |
| range across 12 pairs | 3.50% – 11.19% | — |
| pairs reaching 3× the null | 0 of 12 | — |
| longest contiguous run of rejecting windows | 2 – 9 of 286 | — |
No pair is distinguishable from a random walk, and the rejections that occur are scattered rather
than contiguous.
And the rolling half-life that looked tradable is pure estimation bias
The rolling windows report median half-lives of 216–374 bars — 9 to 16 days. That looks
tradable, and it is the obvious thing to build on.
A pure random walk — whose half-life is infinite by construction — reads a median of 313 bars in
the same 2,160-bar window (IQR 182–565; 52.4% read under 336 bars). The real pairs sit inside
that distribution. The rolling half-life carries no information at all; it is the downward
small-sample bias of an AR(1) slope, E[b̂] − b ≈ −(1+3b)/T.
This is the same class of error as the breadth statistic in the first article of this series, and
it was caught the same way: by simulating the null before believing the number.
What the regime finding does and does not establish. The rolling test has power 0.237 against
a 7-day half-life and 0.083 against a 14-day one, so it cannot rule out slow reversion inside a
90-day window. But it has full power against fast reversion, and fast reversion is the tradable
kind. The bounded, correct statement is:
There is no window in these 5.7 years in which the alt/BTC spreads revert fast enough to be
traded at 17.6 bps.
Sample
12 alt/BTC 1h spreads (ETH, SOL, XRP, ADA, LTC, DOGE, LINK, DOT, BCH, ETC, AVAX, ATOM),
2020-10-22 → 2026-07-10, 50,075 hourly bars each, 5.714 years. The corpus ends
2026-07-10 and was 47 days stale at run time. Backtest arms: commission 0.00088088/side
(two legs) + slippage 0.0002/fill, next_open, position_size 0.15, synthetic_brackets_applied
0.0 on every run. Statistical machinery: statsmodels 0.14.6.
Prior art in this series
- The 2.3 Independent Bets Were a Measurement Artifact — and Fixing Breadth Changed
Nothing
— cycle 1,hyp_1787735069662_0, rejected. Established that the panel already had 9.5 effective
daily bets and that raising it to 11.5 lowered returns. - The Kalman Hedge Ratio Made the Spread Noisier — and Threw Away Its Market
Neutrality
— cycle 2,hyp_1787736340904_1, rejected. Produced the "not cointegrated" aside this article
tests.
Cross-cycle contrast, same panel and same cost model: cycle 2's static momentum arm makes
+2.48%/yr gross. Reversion at the half-life-matched horizon makes +0.043%/yr — 58× less. The
only direction with any gross edge on this panel is the opposite of the one cointegration predicts.
And momentum itself is still 3.8× short of its own bar.
A note on the source notebook
This work adapted Quant Guild's 44. Time Series Analysis for Quant Finance. That notebook
contains no cointegration test, no ADF, no Johansen and no statsmodels — its only imports are
numpy, pandas, plotly and datetime. This is the second cycle in a row where the named notebook
lacked the machinery its title implies. Check what a notebook contains before building on it.
What it does contribute is a position this cycle took seriously: "Tests for stationarity are
largely nonsense... not having a unit root doesnt mean its stationary" and "Stationarity is a
crazy assumption." That is an argument against reading a p-value as an answer — honoured here by
calibrating size and power first, running KPSS (whose null is the opposite one), and letting the
economics decide. The Johansen-versus-ADF disagreement documented above is exactly the phenomenon
it warns about.
What was NOT done
- No walk-forward, Monte-Carlo or sensitivity run on the arm. There is no gross edge here to
validate out of sample; a WFE on a +0.043%/yr gross rule measures noise. - No maker fill-rate measurement. Maker cost sensitivity is reported and is still negative;
the actual fill rate is unmeasured. - No test of the momentum construction on the same panel — untouched, and this rejection says
nothing about it. - No structural-break (Gregory–Hansen) or threshold/TAR cointegration, and no rolling Johansen.
The rolling test is ADF on the (1,−1) spread only. A relation that cointegrates only around an
unmodelled break would not be caught here. That is the one class this cycle could have missed,
and it is the stated reopening condition. - No other universes, timeframes or spread definitions. This is 12 alt/BTC pairs at 1h.
- No fresh data fetch — the corpus was 47 days stale, disclosed above.
Research lineage
Where this result came from
Stored hypotheses, reports, sources, contradictions, and the next registered experiment.
Hypotheses
Parent / child hypotheses
Reports
Academic sources
Negative findings
Related / contradicting studies
Next experiment
No next experiment is stored.
Comments (0)