ARCANE Research · The public record · Updated 25 Sep 2026

Research that says what would prove it wrong.

We test whether machines can read markets honestly. Every question is registered before the data can answer, every score comes with its margin of error, and a loss is printed in the same format as a win.

Publications
05
Lost or found nothing
02
Programmes
04
Registered, not yet answered
02

The record

Five publications. Two of them lost or found nothing.

  1. No. 05Evaluation

    Models without evidence lose to markets

    The walk-forward baseline for the ARCANE forecasting benchmark: 83 resolved questions, and the one gain that held.

    0.212 vs 0.284

    Brier score, market vs calibrated ensemble (lower is better)

    Negative result

    What would change itA prospective window, scored after publication, in which calibration no longer helps.

Programmes

Four questions we keep asking.

  1. 01

    Forecasting & calibration

    Can a system state a probability that still holds up once the world answers?

  2. 02

    Structure & learning

    Finding a pattern is cheap. The test is whether a learner can refuse one.

  3. 03

    Evidence systems

    Every claim should carry where it came from and what would break it.

  4. 04

    Risk & mandates

    Measure the loss a mandate actually governs, not a convenient proxy for it.

How we publish

Five rules every piece keeps.

  1. 01

    Registered first

    The question, the test and what would count as a failure are written down before the data can answer.

  2. 02

    Every number has a margin of error

    A score comes with its sample size and its margin of error, or it is not reported at all.

  3. 03

    Losses are published

    A result that found nothing, or went against us, is published in the same format as a win.

  4. 04

    What would change it

    Every piece names the observation that would overturn it, the same discipline as every ARCANE article.

  5. 05

    A stated boundary

    Every piece says what it publishes and what it holds back, so the gap is visible rather than hidden.

In preparation

Registered, running, not yet answered.

Fig. 2 · The record in order
AUG 2026SEPOCTNOVDECJAN 2027FEB23 AUG · CORPUS QUESTION REGISTERED17 SEP · BACKTEST AND LEARNER RUN; HEAD-TO-HEAD REGISTERED21 SEP · RESOLUTION RULES ADOPTED22 SEP · FIVE PUBLICATIONS RELEASED25 NOV – 25 DEC · PROSPECTIVE ENROLMENT25 JAN 2027 · REVIEW, NO EARLIER
Fig. 2Every dated event in the public record, in order. Solid marks happened; the hatched window and the open mark are planned, not run.
TypeWorkStatus
EvaluationAlpha against frontier models: a pre-registered head-to-headRegistered 17 September 2026. Remaining model arms and human judging in progress. Published whatever it shows.
EvaluationThe ARCANE forecasting benchmark: first prospective windowPlanned. Held-out enrollment 25 November to 25 December 2026; review no earlier than 25 January 2027.
ProgrammeOrder is not a forecastA reading programme on order, measure and complex systems that generates testable market questions without turning history into a signal.

Every public conclusion stays open to its evidence and its limits.

Research · ARCANE