NOAA's Mauna Loa monthly average CO2 for July 2026 is above 427.0 ppm.
- Our call
- 80% yes
- Outcome
- Yes
The 311 Prediction Benchmark
The best way to help our clients is to predict the future accurately. So we built a proprietary system that predicts global social, technology, economic, environmental and political events, and we publish its predictions before they happen. Anyone can check our record, hits and misses together.
Every prediction is dated and locked before the outcome is known, then marked against a named public source. How we score.
Part 1 · Our goal
Plenty of advisers offer a polished point of view on the future: frameworks, workshops and well-written scenarios. Very few put a date and a probability on what they expect, so no one can tell how often they are right.
We think clients deserve more than a good story. Our system produces specific, dated predictions, each tied to the public source that will settle it. We lock them, publish them and mark them in the open, so you can judge us on results.
A clear question, our answer and a probability, plus the public source and the date that will settle it.
The round is dated and sealed before the outcome can be known. A changed view is a new, dated prediction.
On the date, each prediction is checked against its named source. If the source cannot settle it, it is held, not guessed.
Hits and misses are published together, and what the misses teach us changes how the next round is written.
Predictions about public world events, such as the economy, technology, markets and geopolitics, so anyone can check our track record.
The predictions and foresight we create for clients. That work is our intellectual property and theirs, and no one can copy it or claim it as their own.
The benchmark proves the method behind our private work. If we can forecast public events well, and show our misses too, you can trust the forecasts we build for you.
Part 2 · Our results
Every round at a glance: how many questions, how many we got right, how good our probabilities were and how hard the set was. Open any round below for the full detail.
310 questions, 4 days ahead
350 questions, 24 days ahead
300 questions, 11 days ahead
Rounds 1 to 3
Totals to date cover Rounds 1 to 3. Round 1 is counted on its named-source basis (123 of 164). Round 3 is provisional. Hit-rate: share of scored calls we got right. Brier score: how good our probabilities were (lower is better; 0 is perfect; saying 50% on everything scores 0.25). Difficulty: the chance that a random guess is right (lower is harder).
Part 2 · Round by round
Each round shows what kind of questions we asked, how accurate we were, and how hard the set was. Round 3 is open below; tap any round to open it.
310 predictions across nine areas, by share of the round.
Not measured. Round 1 questions were not tagged by type, so we cannot say how hard the set was. Rounds 2 onwards are.
Correct calls by area, counting every resolved call (269 of 310). Our headline rate, 75%, counts only the 164 calls checked against a named source.
Markets and macro was the one weak area, and the largest block of the round.
Calls grouped by the confidence we stated before the outcome.
If anything we were slightly under-confident. The weakness was direction, not confidence.
Every miss, sorted by cause.
What changed: in later rounds, market and economic calls must state a calm or recovery scenario, with its own probability, before they are locked.
350 predictions: 150 foresight questions across society, technology, the economy, the environment and politics, plus 200 hard-number calls. Share of scored questions by type, from most to least predictable.
A random guess would be right 50% of the time. Most questions were predictable from history or were noisy measures; almost none were head-to-head contests.
Scored calls, split by how strongly we leaned one way.
What changed: from Round 3 we no longer make predictions at even odds. Every prediction must lean one way.
Six Round 2 predictions, locked on 7 August and due on 31 August 2026.
NOAA's Mauna Loa monthly average CO2 for July 2026 is above 427.0 ppm.
TSMC begins volume commercial shipping of 2nm (N2) chips.
The 3GPP freezes Release 21, the full 6G specifications.
A US commercial fusion facility delivers net electricity to the grid for 24 continuous hours.
China's GDP growth for the first half of 2026 meets or exceeds 4.8% year on year.
US utility-scale solar and wind exceed 22% of monthly generation in June or July 2026.
300 predictions: 75 economic, 75 political, 75 technology, 38 social and 37 environmental. Share of scored questions by type, from most to least predictable.
A random guess would be right 42% of the time, down from 50% in Round 2. A forecaster with our Round 2 skill would have scored 74.7% here: the set was 5.5 points harder, about 28% more expected errors.
Hit-rate on scored calls, by area. Provisional.
Economic was our weakest area, mainly short-term market calls during September's sell-off.
Provisional.
Our hit-rate fell 13 points, from 80.6% to 67.5%. Round 3 was a deliberately harder test, but we also got worse, and we are showing both.
Provisional: 15 calls are still awaiting mid-October data releases. 48 of 300 questions could not be fairly marked as written and were left out of the score (why). Method: each question is tagged with a predictability tier; Round 2 was tagged after the fact, without reference to its outcomes; the split of the fall uses a shift-share decomposition.
These rounds test how far ahead our accuracy holds, from about six weeks to three months.
Resolves 31 Oct 2026
About 6 weekspredicted aheadOpenResolves 30 Nov 2026
About 10 weekspredicted aheadOpenResolves 31 Dec 2026
About 3 monthspredicted aheadOpenEach round holds 300 predictions, by area.
Published with each round's results, on the same scale as Rounds 2 and 3.
Part 3 · Where we are going
So far we have tested how accurately we predict up to about three weeks ahead. Rounds 4 to 6 test how far out that accuracy holds. Our goal is a system that can predict accurately years ahead, and we will show every step of the way, whether the score holds or falls.
From the date a round was locked to the date it resolves.
Longer range: 106 further predictions were locked by 19 July 2026, most taken from forecasts in our Codex of the Future. The first are due in October 2026 and the furthest in 2051.
The rules
The scoring rules were fixed before any result was known, and they apply to every round.
Every prediction carries a probability. 50% or above means we said it would happen; below 50% means we said it would not. It is a hit when the named source shows that side happened.
A correct "it will not happen" counts the same as a correct "it will happen". But easy no's can flatter a score. Counting every resolved call, Round 1 scored 269 of 310. We lead with the harder figure: 123 of 164 checked against a named source.
If the named source cannot settle a prediction, it is held and left out of the score. In Round 2, 15 of 350 were held. In Round 3, 15 are waiting for official data.
Round 3 exposed a gap: some questions were not checked against their sources before they were locked. 48 of 300 could not be fairly marked as written, so we left them out of the score rather than mark them either way. Every round is now checked automatically before it is locked, and the affected questions in Rounds 4 to 6 have been replaced before any of them resolve.
Nothing is edited after a round is locked. A changed view is a new, dated prediction. In Round 1 our weakest area was markets and the economy, at 42 of 63, and we changed how we write those predictions.
We vary how hard each round is, and we show it. A higher score on easier questions proves less than a lower score on hard ones, so every round is published with its difficulty.
The benchmark tests our forecasting. It is not a recommendation to buy or sell anything.
Answers
Our hit-rate fell 13 points, from 80.6% to 67.5%. Round 3 was a deliberately harder test: predictable questions fell from 43% of the set to 16%, and the chance of guessing right fell from 50% to 42%. About 5 of the 13 points came from the harder questions and about 8 from weaker forecasting, mainly short-term market calls during September's sell-off. Round 3 is provisional.
Because the best way to help our clients is to predict the future accurately, and the only honest way to show that is in public. Our proprietary system predicts global social, technology, economic, environmental and political events. We publish those predictions before they happen and mark them against named public sources, so anyone can check our record.
No. We publish predictions about public world events so our accuracy can be tested. The predictions and foresight we create for clients stay private: they are our intellectual property and theirs, and are never published where anyone could copy them.
It is the public test of 311 Institute's forecasting. We write down dated predictions with a probability, lock them before the outcome can be known, then mark every one against a named public source and publish the hits and the misses together.
Each round is sealed before its resolution date and nothing in it is edited afterwards. If our view changes, that is a new, dated prediction. It never replaces the old one.
Every prediction carries a probability. 50% or above means we said it would happen; below 50% means we said it would not. It is a hit when the named source shows that side happened. Saying something will not happen counts the same as saying it will.
Round 3 is provisional because 15 of its calls are still waiting for official data releases due in mid-October. Its figures will be updated once those calls are marked.
It is held, not guessed. In Round 2, 15 of 350 predictions were held because the named source could not settle them. In Round 3, 48 of 300 questions could not be fairly marked as written and were left out. Held and voided questions are never counted either way.
A new round of 300 predictions comes due at the end of each month from September to December 2026. Results are added here once each round is marked and checked.
Tell us what you are facing and we will show you what our forecasting says about it, and how sure we are.