XDRIPACADEMY
Sign in

Curriculum·O100 Placement and Priming·25 min

Calibration baseline

By the end of this lesson you can

  • Answer twenty questions with a confidence percentage attached to each
  • Read the resulting curve, which reports how much your confidence is worth
  • Explain why expertise and forecasting accuracy are separate quantities
  • Use the curve as an input to sizing rather than as a personality result
AutopsyExpert Political Judgment, published 2005expertise that did not convert into accuracy

Philip Tetlock spent years collecting explicit, scoreable forecasts from recognised experts and comparing them against the outcomes, against forecasts from well-informed non-specialists, and against simple extrapolation from current trends.

The experts performed poorly, and in many comparisons did not clearly beat the simpler benchmarks.

He also found a difference in style that predicted accuracy better than credentials did. Forecasters he called foxes, who draw on many small ideas and revise readily, outperformed those he called hedgehogs, who apply one large organising idea consistently.

And he noted that the hedgehog style, being confident, single-minded and combative, makes considerably better television.

Now the finding.

The people studied were genuinely expert. They knew more about their subjects than their audiences did, and more than the simple benchmarks that outperformed them.

What they lacked was calibration, meaning any reliable relationship between how sure they were and how often they were right.

That is a separate skill from knowledge, it is measurable in about twenty-five minutes, and per the course autopsy nobody can estimate their own.

Primary source

Per O100-01 you have a competency map and per O100-03 a ceiling. This lesson produces the third number, which is how much your own confidence is worth as evidence.

What the instrument does

Twenty questions, each answered with a confidence percentage.

The questions are not the point. Per part one below, what is scored is the relationship between the percentages and the outcomes, so a low knowledge score with honest percentages produces a good curve.

Which is why the exercise cannot be gamed by answering 50 to everything. Per part two a flat curve is itself a result, and per O100-L the instrument is re-run at every gate so a strategy of avoidance shows up as no information across all of them.

Worked example
Reading your own curve

Part one: the grouping.

Take the twenty answers and group them by the confidence you stated.

Suppose ten answers were marked 90 percent and you got seven right:

7 / 10 = 70 percent actual against 90 stated

A gap of 20 points, which per the course autopsy is the ordinary finding rather than an unusual one.

Part two: what a well calibrated curve looks like.

The questions you marked 60 percent should be right about 60 percent of the time. The ones you marked 90 should be right about 90.

A curve that sits below the diagonal is overconfident, which per the course autopsy is the common direction: accuracy stops rising with confidence above about three to one.

A flat curve carries no information. If every band comes out near the same accuracy, your percentages are not tracking anything, which is a different problem from being overconfident and is worth knowing.

Part three: what the number is for.

Per O100-03, a ceiling on exposure. Per part one, a discount on your own certainty.

If your 90 is really 70, then a decision you would only take at 90 percent confidence is a decision you are taking at 70, and per O100-03's recovery arithmetic the difference matters most on the largest commitments.

Part four: what it is not.

Not a personality result. Per the autopsy the experts studied were expert, and per part two the curve measures one specific relationship rather than a trait.

Not fixed. Calibration responds to feedback, which is why per O100-L the instrument is repeated at every gate and the curve is tracked rather than recorded once.

And not a substitute for the material. Per O100-01 the diagnostic measures what you know and this measures what your sense of knowing is worth. Both are needed and neither replaces the other.

Answer honestly rather than impressively, because the instrument scores the gap

The one instruction, and it is the opposite of how most tests are approached.

Put the percentage you actually believe. Per part two a low score with honest percentages produces a good curve and a high score with inflated ones does not.

Do not round everything to 50. Per part two a flat curve is a result and it is the least useful one, because per part three the discount it produces cannot be applied to anything.

Then keep the curve. Per part four it is re-measured at every gate, and per O100-05 the trajectory is the interesting part rather than today's value.

And apply it where it costs something. Per part three, if your stated 90 is a measured 70, the commitments you make at high confidence are the ones the discount changes.

Per the autopsy the trait that predicted accuracy was not credentials, and per the course autopsy the sensation of certainty stops carrying information above about 75 percent.

Common misconception

I am not overconfident, I just know what I know.

That is the sentence the instrument exists to test, and per the course autopsy it is the one that does not survive measurement.

What is true. Some people are well calibrated, the trait varies, and per part two the curve will say so if you are. Nobody is being told the answer in advance.

Why self-report cannot settle it. Per the course autopsy, subjects reporting near certainty at 98 or 99 percent were often right below 80 percent of the time, and the reporting felt the same in both cases. Per part one the gap is only visible when the percentages are compared with outcomes.

And expertise does not resolve it. Per the autopsy the experts studied knew their subjects and still did not clearly beat simple extrapolation, so per part four knowing more is not the same as knowing how much you know.

Nor does the claim cost anything to test. Per the instrument section it is twenty questions and twenty-five minutes, and per part three the output is a number you can use.

So per P6 the accurate framing: the claim is testable, cheap to test, and the test is the lesson. Per the callout, answer with the percentages you actually believe.

Key takeaway

The calibration baseline is twenty questions each answered with a confidence percentage, and what is scored is the relationship between the percentages and the outcomes rather than the answers. Group them afterwards: if ten answers marked 90 percent produced seven correct, your 90 is a measured 70, a gap of 20 points, which per the course autopsy is the ordinary finding rather than an unusual one. A curve below the diagonal is overconfident and a flat curve carries no information at all, which is a different problem worth knowing about. Then use it where it costs something, because a decision you would only take at 90 percent confidence is one you are actually taking at 70, and per O100-03 that matters most on the largest commitments. And expertise does not substitute for it: Tetlock's expert forecasters knew their subjects, did not clearly beat simple extrapolation, and the style that predicted accuracy was not the one that makes good television.

These come back later

What does the baseline measure?
The relationship between your stated confidence and your actual accuracy, across twenty questions each carrying a percentage.
How do you read the curve?
Group the answers by stated confidence and compare with the proportion correct. Saying 90 percent on ten questions and getting seven right means your 90 is really 70.
Why are expertise and calibration separate?
Because knowing more does not make the sense of certainty more informative. Per the autopsy expert forecasters did not clearly beat simple extrapolation.
What is the curve used for?
Sizing. A number that is right 70 percent of the time when you feel certain is an input to how much you commit, not a personality result.

Sources and review

Confidence medium·Volatility low·Reviewed 2026-08-07·Owner unassigned

Contested

Widely repeated figures for the number of experts and forecasts in Tetlock's study vary between accounts and are not asserted here. The load-bearing findings are that expert forecasters did not clearly beat well-informed non-specialists or simple extrapolation, and that a style difference predicted accuracy better than credentials did. Both are the book's central claims.

The fox and hedgehog distinction has been debated and later work, including the Good Judgment Project, refined it considerably. Per P6 this lesson uses it only to make the point that the trait predicting accuracy was not expertise, and does not present it as a settled taxonomy of thinkers.

O100-L produces the curve and the same instrument is re-run at every subsequent gate. This lesson owns the first measurement and the reading of it. Keep that split.

Track your progress

Create a free account to mark lessons complete and pick up where you left off.