Rolling benchmark · model vs. market · updated weekly

Which AI predicts
futures best?

A rolling, contamination-free benchmark that grades every frontier model against live prediction markets — on a future that keeps arriving, every day. The race to predict the future is on. We keep score.

See who's ahead. Then beat them.
FutureCapsule v1.2 Leaderboard

Skill score vs. the market.
100 = market parity.

Rank Model Skill score MAE (price) Last run

From this week's frozen set
Progress

Who's improving fastest at predicting futures.

Weekly skill scores, trending toward market parity. Early days of the most valuable capability in AI — and the gap is closing.

How it works · a rolling prediction-market benchmark v1.3
1 · Freeze

Frozen point-in-time information capsules — the model sees only what the market saw. Nightly immutable snapshots of live prediction markets: price history plus every news article published before each cutoff, checksummed, with a leakage gate guaranteeing nothing in a capsule post-dates its cutoff. Retrieval is a compiled program — BM25 + entity inverted index + dense RRF fusion, deliberately no GraphRAG.

2 · Predict

Rolling, not one-shot. Scoring only the final outcome extracts one bit per market; FutureCapsule scores the model's belief at every cutoff the data's time resolution allows — one evaluation point per cutoff × horizon (3-day / 7-day), temperature 0. The full benchmark set (v1.3) spans 778 markets and 17,290 evaluation points; the board below is graded on the live vault snapshot universe, sliced further per model by the contamination gate.

3 · Grade

Belief deltas graded against the market's future price trajectory — not binary outcomes. The score is skill vs. a no-change baseline ("the market is already right"): 100 = market parity, above 100 beats the market. 0 frontier models beat it so far. 95% CIs from a 1,000-resample block bootstrap.

Contamination discipline: a model is scored only on windows that resolve after its published training cutoff — enforced per model, not by trusting the prompt. That's why recent-cutoff models have fewer eligible points, and why cells with under 30 points are excluded. A human reference arm — a superforecaster panel scored on the same grid — is a required baseline, not an optional one. The trend chart replays the identical market universe at successive snapshot dates; scores are dated by snapshot, not by when the eval ran.

The human track

The same futures, priced by human strategists.

Behind this benchmark sits Principle's Strategic Intelligence Community — an invitation-only network of operational and strategy leaders who stress-test the same signal-driven futures: blind commits, probability spreads, scored as reality arrives. Machines on this page; humans in the room.

Submit your model

Think you can beat the market?

Your model, your scaffold, your fine-tune — run the same frozen questions. We'll reply with the frozen question set and submission details. Tournament track — tools and scaffolding welcome.

Weekly results in your inbox — subscribe in the footer.

Benchmark changelog
v1.3 2026-06-06 · current

Stratified rebuild of the v1.2 frozen set with structural market tags baked into every eval point. Noise filter: markets kept only if price range ≥ 0.12 and pinned fraction < 0.70 — 778 markets, 17,290 evaluation points. Price-relevance judged signal sidecar for all points.

v1.2 frozen base

Added the 24 h horizon (HTML-meta backfill certified 83.7% tight ≤ 6 h on n=200K), bringing the full set to 37,579 points across 1,681 markets. All 12 spec §10 acceptance gates pass. Remains the frozen base artifact under v1.3's stratification layer — the leaderboard below is scored on this methodology.

v1.1

Denser cutoff sampling — 7–9 cutoffs per market (up from the initial grid), raising the set to 24,861 evaluation points.

v1.0

Initial dense-price benchmark: frozen point-in-time capsules, rolling cutoff × horizon grid, no-change baseline, leakage gate, per-model contamination sidecars.