Observatory live · 1687 signals · 54 consilience · 110 confirmed · 291 killed Register interest · Newsletter
Long-cycle structural research
The Consiliences Institute

Unity of knowledge in an age of fragmentation

On Revising Our Own Findings: What the April Retest Battery Revealed

Confidence: confirmed

A Note on This Paper

The Observatory publishes findings. It also retracts them. This paper concerns the second of those activities. Specifically, it describes a battery of validation tests applied in April 2026 to a class of long-wave and cycle-period claims that had previously been reported as confirmed, and reports the outcomes — including the cases in which our prior published reading did not survive the stricter procedure.

The reason to publish a methodological retraction rather than silently update is that the public record otherwise overstates the strength of the surviving findings. A reader who encounters only the confirmations cannot calibrate against the prior rate at which the procedure produced confirmations that subsequently failed. The retract-in-public discipline is, in this respect, a precondition for the inductive value of the procedure itself.

The Retest Battery

The April retest battery applied four orthogonal tests to each long-wave or cycle-period claim previously reported as confirmed:

A direction test asks whether the empirical relationship operates in the direction the hypothesis claims, when applied to held-out data not used in the original fit. A claim that survives in-sample but reverses direction out-of-sample is not a robust empirical claim.

A specificity test asks whether the identified spectral peak or cycle period is distinguishable from competing candidate periods. A claim that a 3.4-year cycle is present is weak if a 3.0-year or 4.0-year cycle fits the data equivalently well; the specificity test requires the claimed period to substantially outperform nearby alternatives.

A scrambled-null test applies the same statistical procedure to time-shuffled or surrogate versions of the same data. A procedure that finds a “significant” cycle in scrambled data with high probability is not constraining the actual data; it is constraining the procedure.

A held-out test withholds a portion of the historical record from the original fit and asks whether the procedure correctly predicts the held-out portion. This is the simplest and, in long-wave work, the most demanding of the four.

Each surviving claim was required to pass at least three of four tests. Failing one is a downgrade to provisional status; failing two or more is grounds for retraction.

What Did Not Survive

Three classes of finding did not survive the April battery, and we describe them here in order of the public attention they had previously received.

The Kondratieff long-wave timing claim. Our prior reading had reported empirical confirmation of a fifty-five-year cycle in long-run commodity-price data, with spectral evidence consistent across three of four datasets at p=0.002. The retest preserves the spectral evidence at this strength but does not preserve the predictive claim. The held-out test — whether the cycle position correctly identifies the next phase transition — is underpowered, because the historical record contains too few full cycle realisations to constrain the prediction. We retain the descriptive finding: a long-period spectral feature is present in the data at the claimed period. We retract the predictive framing: the cycle does not constitute a forecasting tool at horizons relevant to investors or policy.

The Kitchin inventory-cycle predictive claim. A 3.4-year inventory cycle had previously been reported with strong in-sample statistical support (p=0.0003 pre-battery). The post-April battery returns mixed results: two of four tests pass, but the scrambled-null test returns z=1.7 and the cycle-period-specificity test returns 0.98 times the cleanest neighbouring period. The honest reading is that the timing-clock claim is weaker than the original p-value suggested. The directional classifier — whether the system is currently in expansion or contraction — remains coherent at twelve-month horizons. The phase-clock claim that the cycle predicts the timing of transitions, rather than merely the current phase, is not supported at the strength previously implied.

The geophysical-to-commodity-price trading claim. Our prior publication had connected geophysical forcings — ENSO state, volcanic aerosol loading, multidecadal ocean oscillations — to agricultural commodity prices through documented temperature-mediated yield channels. The scientific reading of this chain survives the retest. What does not survive is the trading-application layer: the specific instruments through which the relationship had been operationalised either failed their own batteries (specificity, in one case; mechanism-falsification on the underlying period claim, in another) or remained at provisional status awaiting regime confirmation that has not yet materialised. The geophysical-yield chain is real. There is no currently-deployed trading instrument that successfully expresses it.

These are not minor revisions. They are retractions of claims that had appeared in prior Observatory publications. The discipline of stating this directly is, we believe, the only appropriate response.

What Did Survive

The retest battery also produced confirmations, and the asymmetry between failure and survival is informative.

The debt-supercycle pattern reported by Reinhart and Rogoff, calibrated against two centuries of sovereign-crisis data, passes the surrogate-permutation test at p=0.028 over a peak period of approximately seventy-three years. The pattern survives the retest because the held-out period — the years since the original calibration — has unfolded in a way consistent with the original wave structure. The 2008-2015 sovereign-stress episode fits the expected position.

The Bank for International Settlements credit-gap indicator, which identifies banking crises via a threshold crossing on the deviation of private credit to GDP from its long-run trend, passes all four tests of the retest battery. The threshold is operational: a value above ten percent is associated with banking crisis onset within three years at an odds ratio of approximately thirty-five. The current value, approximately twelve percent below trend, places the system in the deleveraging phase. This indicator is, in our reading, the most operationally robust of the long-wave claims we have examined.

The asymmetry is informative because it distinguishes claims that derive from a clear physical or institutional mechanism — credit accumulation, debt forgiveness — from claims that derive primarily from spectral analysis of long time series. The mechanism-grounded claims survive at higher rates. The spectral-only claims do not. This is not a finding about cycles in general; it is a finding about the kinds of evidence that withstand stricter testing.

What the Retest Implies About the Original Procedure

A retest battery that fails several previously confirmed claims is informative about the original procedure. We owe the reader a candid reading of what that implies.

The original procedure relied on in-sample statistical tests applied to long historical datasets. The procedure is not wrong — the spectral peaks identified are real features of the data — but it is insufficient. A long historical dataset contains many candidate cycle periods, and the procedure of selecting the strongest one and reporting its p-value does not adequately correct for the search. A more demanding procedure, of the sort applied in the retest, requires the claim to survive tests the original procedure did not contemplate.

The implication for future Observatory work is that mechanism-grounded claims should be assigned higher prior weight than spectral-only claims, and that any cycle-period claim should pass a held-out predictive test before being treated as operationally usable. The retest battery is being incorporated into the Observatory’s standard procedure for this class of finding.

A Note on Public Retraction

The discipline of correcting one’s own published work in public is, in research more broadly, uneven. The institutional incentives run against it. A retraction is a marker of error rather than a marker of methodological care; a confirmation is publishable while a careful self-correction often is not. We believe this incentive structure is inverted relative to what serves the inductive value of the work.

The Observatory’s policy is to update published findings with explicit amendment notes when subsequent validation changes their status, and to publish methodology-of-revision papers — of which this is one — at intervals when the volume of revisions warrants a consolidated public statement. This procedure is not without cost; it makes the cumulative record longer and the synthesis of findings more difficult for casual readers. The alternative — silent updates and a polished public record — would produce a more readable archive at the cost of misrepresenting the actual epistemic state of the work.

We have chosen the longer and less polished archive. The choice is methodological, not aesthetic.

Limitations of This Paper

This paper is a methodology-of-revision paper, not a comprehensive retest of every prior Observatory finding. The April battery was applied specifically to long-wave economic claims and cycle-period claims; analogous batteries for other classes of finding — agricultural-yield correlations, frequency-biology pairs, traditional-ecological-knowledge validations — have not yet been comprehensively reported in equivalent form. A reader should not infer from this paper that other classes of finding have been comparably stress-tested. They have been tested, but the public reporting on those tests is distributed across the original signal documentation rather than consolidated as it is here.

A second limitation: the four-test battery is itself a methodological choice with degrees of freedom. A different battery — five tests instead of four, a different statistical-significance threshold, a different out-of-sample window — would produce a partially different set of confirmations and retractions. The boundary between confirmed and retracted is itself empirically calibrated, not philosophically fundamental. We have made our procedure explicit so that readers can apply their own judgement.


Evidence Strength and Disclosures

This paper documents methodology rather than presenting new empirical claims. The signals referenced — debt-supercycle, BIS credit-gap, Kondratieff long-wave, Kitchin inventory-cycle, and the geophysical-yield chain — are part of the Observatory’s published archive, and the retest results described here have been applied to each signal’s own status documentation as appropriate.

The retest battery is applicable beyond long-wave economics, and we expect to apply versions of it to additional classes of finding in subsequent work. Readers interested in the application of the battery to a specific signal can consult that signal’s own documentation, which carries the per-test results in addition to the canonical verdict.

The retraction-in-public discipline described here is offered as a methodological norm, not as a finished procedure. It will be revised as the Observatory’s experience with it accumulates.