A rolling, contamination-free benchmark that grades every frontier model against live prediction markets — on a future that keeps arriving, every day. The race to predict the future is on. We keep score.
| Rank | Model | Skill score ↑ | MAE (price) ↓ | Last run |
|---|
Weekly skill scores, trending toward market parity. Early days of the most valuable capability in AI — and the gap is closing.
Frozen point-in-time information capsules — the model sees only what the market saw. Nightly immutable snapshots of live prediction markets: price history plus every news article published before each cutoff, checksummed, with a leakage gate guaranteeing nothing in a capsule post-dates its cutoff. Retrieval is a compiled program — BM25 + entity inverted index + dense RRF fusion, deliberately no GraphRAG.
Rolling, not one-shot. Scoring only the final outcome extracts one bit per market; FutureCapsule scores the model's belief at every cutoff the data's time resolution allows — one evaluation point per cutoff × horizon (3-day / 7-day), temperature 0. The full benchmark set (v1.3) spans 778 markets and 17,290 evaluation points; the board below is graded on the live vault snapshot universe, sliced further per model by the contamination gate.
Belief deltas graded against the market's future price trajectory — not binary outcomes. The score is skill vs. a no-change baseline ("the market is already right"): 100 = market parity, above 100 beats the market. 0 frontier models beat it so far. 95% CIs from a 1,000-resample block bootstrap.
Contamination discipline: a model is scored only on windows that resolve after its published training cutoff — enforced per model, not by trusting the prompt. That's why recent-cutoff models have fewer eligible points, and why cells with under 30 points are excluded. A human reference arm — a superforecaster panel scored on the same grid — is a required baseline, not an optional one. The trend chart replays the identical market universe at successive snapshot dates; scores are dated by snapshot, not by when the eval ran.
Behind this benchmark sits Principle's Strategic Intelligence Community — an invitation-only network of operational and strategy leaders who stress-test the same signal-driven futures: blind commits, probability spreads, scored as reality arrives. Machines on this page; humans in the room.
Your model, your scaffold, your fine-tune — run the same frozen questions. We'll reply with the frozen question set and submission details. Tournament track — tools and scaffolding welcome.
Weekly results in your inbox — subscribe in the footer.