Forecast Calibration Playbook: How to Measure Honest Probabilities

Published on May 08, 202612 min read
Forecast Calibration Playbook: How to Measure Honest Probabilities

What calibration actually means

Calibration is the habit of checking whether your probabilities match reality. If events you call 70% likely happen only half the time, your model is overconfident. If they happen 90% of the time, you may be underpricing your own edge. Calibration is the foundation under every other prediction market discipline — without it, your position sizing is built on sand and your bankroll framework cannot save you from systematic mispricing.

Prediction markets make calibration visible because prices are continuous and outcomes are public. Traders can compare personal forecasts against market-implied probabilities and track performance bucket by bucket. The same data feeds the trader, the platform operator, and the academic researcher. This is one of the reasons the broader category exists: prediction markets are calibration machines.

This playbook walks through the practical steps to measure, improve, and operationalise your own calibration. It complements risk management for traders (which uses calibration to size positions) and the bankroll framework (which uses calibration to allocate capital across categories). The foundational platform context is in our introduction to prediction markets.

Step 1 — Build a prediction log

The first step is non-negotiable: log every forecast before resolution. Without the log, you cannot measure calibration; you can only narrate it.

A minimum-viable prediction log:

FieldWhy it matters
DateFor time-series tracking
QuestionThe forecast in plain words
My probabilityYour estimate, between 0 and 1
Market priceThe market-implied probability at the time
ReasonOne sentence — why your probability differs from the market
Resolution dateWhen you expect the question to resolve
OutcomeFilled in after resolution

The discipline is logging the probability before you trade, not after. A forecast written down after the trade is contaminated by the result. The integrity of your calibration depends on the integrity of the log.

You can use a spreadsheet, Manifold Markets tracking, or any of the public forecast tracking tools. The tool matters less than the consistency. A bad spreadsheet you fill in every day beats a perfect tool you abandon in two weeks.

Step 2 — Group forecasts into probability buckets

Once you have 50+ resolved forecasts, group them into buckets and check whether each bucket's hit rate matches the expected rate.

BucketExpected hit rateNumber of forecastsActual hit rateCalibration error
50–60%55%1250%-5%
60–70%65%1867%+2%
70–80%75%1560%-15%
80–90%85%1090%+5%
90–100%95%888%-7%

Calibration error in any bucket above ±5 percentage points is a signal worth investigating. The 70–80% bucket in this example is the worst — the trader thinks they have a 75% confidence on average but only gets 60% right. That is overconfidence at the moderate-conviction tier, which is exactly where most trading edges live.

The visual representation is the calibration plot — a scatter plot with expected probability on the x-axis and actual hit rate on the y-axis. A perfectly calibrated forecaster's points all lie on the y = x line. Overconfidence shows up as points below the line in the high-probability buckets; underconfidence shows up as points above the line.

Step 3 — Score yourself with Brier

Bucket analysis is intuitive but coarse. The Brier score gives a single number that captures both calibration and resolution (the ability to discriminate between events).

Brier score = mean((forecast_probability - actual_outcome)^2)

Where actual_outcome is 1 if the event happened and 0 if it did not. Lower is better. A perfect forecaster scores 0; a forecaster who predicts 50% on everything scores 0.25; a maximally bad forecaster (always predicting 0 when the answer is 1, or vice versa) scores 1.0.

A useful benchmark: the Iowa Electronic Markets have historically scored Brier values around 0.04–0.08 on election markets, which is better than most polls. Individual superforecasters in Tetlock's Good Judgment Project scored 0.10–0.15 on geopolitical questions. If your personal Brier score is above 0.20, you are systematically miscalibrated.

The Brier score can be decomposed into three terms: reliability (calibration error), resolution (ability to discriminate), and uncertainty (irreducible randomness). For working traders, reliability is the actionable component — that is what calibration improvement targets.

Step 4 — Find your calibration leaks

A Brier score of 0.18 is not actionable. "I am overconfident in tech earnings calls" is. The next step is finding the categories and contexts where your calibration breaks down.

Group your forecasts by:

  1. Category (politics, macro, crypto, sports, tech)
  2. Time to resolution (hours, days, weeks, months)
  3. Source of edge (news, model, expert, gut)
  4. Position size (small, medium, large)
  5. Direction of edge (long Yes vs long No)

Run the bucket analysis within each subgroup. Patterns will emerge. Common findings I have seen from real traders:

  • "I am calibrated on politics but overconfident on tech."
  • "My week-out forecasts are well calibrated; my day-of forecasts are wildly overconfident."
  • "I am calibrated when I have a model but overconfident when I trade on news."
  • "I am underconfident on No bets — when I think something is 30%, it is more like 22%."

Each pattern becomes a targeted calibration goal. The 90-day improvement loop is: identify a leak, design an experiment to address it, run for 90 days, re-measure.

Step 5 — Implement the superforecaster techniques

Philip Tetlock's Superforecasting research identified specific techniques that produce calibration improvements. The most important:

Triage your forecasts. Spend more time on questions where you can plausibly have an edge. A binary forecast on a coin flip is wasted effort. A binary forecast on a Fed meeting where you have read the dot plot, the meeting minutes, and the Fed funds futures curve is where calibration improves.

Decompose the question. "Will the Fed cut rates?" decomposes into "Will inflation be below 3%?" + "Will payrolls weaken?" + "What is the Fed reaction function?" Forecasting the sub-questions separately and combining is almost always more calibrated than forecasting the headline question directly.

Use base rates. Most events have a historical frequency. "What fraction of FOMC meetings have produced a rate cut in the last 30 years?" is a base rate. Anchor your forecast to the base rate and adjust based on current information.

Update incrementally. When new information arrives, update your probability by a moderate amount rather than swinging dramatically. The single most common calibration failure is overreacting to news. The fix is the Bayesian update — multiply the prior by the likelihood ratio, do not replace it.

Distinguish signal from noise. A 3-point poll shift two weeks before an election is mostly noise. A 3-point shift after a debate is more likely signal. Train yourself to recognize which inputs deserve a probability update and which do not.

Use the "outside view". When stuck, ask "what happens in similar situations?" rather than "what makes this case special?" Most cases are less special than the participants believe.

Step 6 — Calibration as a category review

The quarterly bankroll review (covered in the bankroll framework) should include a calibration audit by category. The combined output is a single dashboard:

  • Per-category Brier score and trend. Is calibration improving or degrading?
  • Per-category P&L. Profitable but miscalibrated is a red flag — you are getting lucky.
  • Per-category position count. Are you taking enough trades to make calibration measurable?
  • Per-category exposure. Is the category cap being hit, and is that limit justified by calibration?

The traders who compound over years are the ones who treat calibration as a measured, dashboarded discipline — not a vague self-assessment. Builders who run prediction platforms should publish category-level calibration data publicly; it is one of the strongest trust signals a platform can offer. Some platforms expose this in the resolution archive; the Iowa Electronic Markets and academic predictors like Metaculus are the public standards.

How calibration interacts with market price

Calibration is your edge over the market. If you and the market are both calibrated at 70% on a 70¢ contract, there is no trade. If you are calibrated at 70% and the market is at 60¢, you have a 10-point edge — assuming your calibration measurement is solid.

The trader's job is not to find market prices that "feel" wrong. It is to find market prices where your calibration is meaningfully better than the market's. That requires:

  1. Knowing your own calibration in the category (Step 4).
  2. Estimating the market's calibration in the category (compare market price to actual hit rate across similar markets).
  3. Trading only when your calibration advantage is large enough to overcome fees, spread, and variance.

The microstructure post goes deeper into how to translate a calibration advantage into an executable trade.

Common calibration mistakes

Three patterns I see repeatedly.

Confusing certainty with conviction. "I am 95% sure this will happen" is a probability. "I have done a lot of research" is not. Confidence in the process does not translate into confidence in the outcome. Many traders log 95% probabilities and then discover their 95% calls only hit 70% — they were measuring conviction, not probability.

Anchoring on the market. It is tempting to log your probability as "5 cents away from market". That is not your forecast; it is the market plus an arbitrary offset. Your honest forecast is the probability you would assign if the market did not exist. Anchoring on the market makes your "edge" mostly noise.

Skipping low-conviction trades. New traders skip the 50–60% bucket because the edge feels small. But the 50–60% bucket is where calibration is most valuable — small edges, repeated many times, compound. Your bankroll framework should include enough small-edge trades to make this bucket statistically meaningful.

Treating calibration as a one-time check. Calibration drifts. Categories change. Your edge in tech today is not your edge in tech a year from now. Re-measure at least quarterly.

Frequently Asked Questions

What is calibration in forecasting?

Calibration is the property that your stated probabilities match the actual frequencies of outcomes. A perfectly calibrated forecaster who calls events 70% likely will see those events happen exactly 70% of the time across many forecasts. Calibration is measurable and improvable, which makes it the foundation of every other trading discipline.

How do I measure my own calibration?

Log your forecasts before resolution, group them by probability bucket (50–60%, 60–70%, etc.), and check whether each bucket's hit rate matches the expected rate. Then compute a Brier score: mean((forecast - outcome)^2). Lower is better. A score above 0.20 indicates systematic miscalibration.

What is a Brier score and what is a good Brier score?

The Brier score is the mean squared error between your probability forecasts and actual binary outcomes. It ranges from 0 (perfect) to 1 (maximally bad). A score of 0.25 is what you get by predicting 50% on everything. Good forecasters score 0.10–0.15 on hard questions, 0.04–0.08 on easier ones. Individual superforecasters in academic studies have averaged around 0.10 across mixed domains.

How many forecasts do I need before calibration is meaningful?

You need at least 50 resolved forecasts to have rough calibration data and 200+ to draw category-level conclusions. Below 50 forecasts the noise dominates the signal. The good news is that 50–200 forecasts is achievable in a few months of active trading.

Can a prediction market be miscalibrated?

Yes. Markets are not perfectly efficient, especially in long-tail categories. The empirical literature — including Wolfers and Zitzewitz — shows that prediction markets are well calibrated in the aggregate but exhibit systematic biases at the extremes (longshots overpriced, near-certainties slightly underpriced). Your edge often lies in identifying these specific miscalibrations.

What is the difference between calibration and accuracy?

Calibration measures whether your probabilities match frequencies. Accuracy measures how often you are right. A coin flipper who predicts 50% on every coin flip is perfectly calibrated but useless because they have no resolution — they cannot discriminate between events. The Brier score captures both, which is why it is the standard composite metric.

How do superforecasters improve calibration?

Tetlock's research identified six techniques: triaging which questions to forecast, decomposing complex questions into simpler sub-questions, anchoring on base rates, updating incrementally on new information, distinguishing signal from noise, and using the "outside view". The full methodology is in the Good Judgment Project.

Should I publish my calibration data?

Publishing builds trust and accountability. Platforms benefit from publishing aggregate calibration; individual traders benefit from publishing their personal track records when they are confident in the discipline. The Metaculus leaderboard is a great example of public calibration data that builds the broader credibility of probabilistic forecasting.

Where to go next

You now have the calibration playbook. The natural next steps:

Calibration is uncertainty done honestly. The traders, builders, and analysts who internalize it are the ones who turn probability from a number into a discipline.

Written by Editorial Team
ET
Editorial TeamEditorial

We write about prediction markets, automated market makers and the math behind forecasting.

Related articles