Philip Tetlock spent years collecting explicit, scoreable forecasts from recognised experts and comparing them against the outcomes, against forecasts from well-informed non-specialists, and against simple extrapolation from current trends.
The experts performed poorly, and in many comparisons did not clearly beat the simpler benchmarks.
He also found a difference in style that predicted accuracy better than credentials did. Forecasters he called foxes, who draw on many small ideas and revise readily, outperformed those he called hedgehogs, who apply one large organising idea consistently.
And he noted that the hedgehog style, being confident, single-minded and combative, makes considerably better television.
Now the finding.
The people studied were genuinely expert. They knew more about their subjects than their audiences did, and more than the simple benchmarks that outperformed them.
What they lacked was calibration, meaning any reliable relationship between how sure they were and how often they were right.
That is a separate skill from knowledge, it is measurable in about twenty-five minutes, and per the course autopsy nobody can estimate their own.
Per O100-01 you have a competency map and per O100-03 a ceiling. This lesson produces the third number, which is how much your own confidence is worth as evidence.
What the instrument does
Twenty questions, each answered with a confidence percentage.
The questions are not the point. Per part one below, what is scored is the relationship between the percentages and the outcomes, so a low knowledge score with honest percentages produces a good curve.
Which is why the exercise cannot be gamed by answering 50 to everything. Per part two a flat curve is itself a result, and per O100-L the instrument is re-run at every gate so a strategy of avoidance shows up as no information across all of them.
Part one: the grouping.
Take the twenty answers and group them by the confidence you stated.
Suppose ten answers were marked 90 percent and you got seven right:
7 / 10 = 70 percent actual against 90 stated
A gap of 20 points, which per the course autopsy is the ordinary finding rather than an unusual one.
Part two: what a well calibrated curve looks like.
The questions you marked 60 percent should be right about 60 percent of the time. The ones you marked 90 should be right about 90.
A curve that sits below the diagonal is overconfident, which per the course autopsy is the common direction: accuracy stops rising with confidence above about three to one.
A flat curve carries no information. If every band comes out near the same accuracy, your percentages are not tracking anything, which is a different problem from being overconfident and is worth knowing.
Part three: what the number is for.
Per O100-03, a ceiling on exposure. Per part one, a discount on your own certainty.
If your 90 is really 70, then a decision you would only take at 90 percent confidence is a decision you are taking at 70, and per O100-03's recovery arithmetic the difference matters most on the largest commitments.
Part four: what it is not.
Not a personality result. Per the autopsy the experts studied were expert, and per part two the curve measures one specific relationship rather than a trait.
Not fixed. Calibration responds to feedback, which is why per O100-L the instrument is repeated at every gate and the curve is tracked rather than recorded once.
And not a substitute for the material. Per O100-01 the diagnostic measures what you know and this measures what your sense of knowing is worth. Both are needed and neither replaces the other.
The one instruction, and it is the opposite of how most tests are approached.
Put the percentage you actually believe. Per part two a low score with honest percentages produces a good curve and a high score with inflated ones does not.
Do not round everything to 50. Per part two a flat curve is a result and it is the least useful one, because per part three the discount it produces cannot be applied to anything.
Then keep the curve. Per part four it is re-measured at every gate, and per O100-05 the trajectory is the interesting part rather than today's value.
And apply it where it costs something. Per part three, if your stated 90 is a measured 70, the commitments you make at high confidence are the ones the discount changes.
Per the autopsy the trait that predicted accuracy was not credentials, and per the course autopsy the sensation of certainty stops carrying information above about 75 percent.
I am not overconfident, I just know what I know.
That is the sentence the instrument exists to test, and per the course autopsy it is the one that does not survive measurement.
What is true. Some people are well calibrated, the trait varies, and per part two the curve will say so if you are. Nobody is being told the answer in advance.
Why self-report cannot settle it. Per the course autopsy, subjects reporting near certainty at 98 or 99 percent were often right below 80 percent of the time, and the reporting felt the same in both cases. Per part one the gap is only visible when the percentages are compared with outcomes.
And expertise does not resolve it. Per the autopsy the experts studied knew their subjects and still did not clearly beat simple extrapolation, so per part four knowing more is not the same as knowing how much you know.
Nor does the claim cost anything to test. Per the instrument section it is twenty questions and twenty-five minutes, and per part three the output is a number you can use.
So per P6 the accurate framing: the claim is testable, cheap to test, and the test is the lesson. Per the callout, answer with the percentages you actually believe.
The calibration baseline is twenty questions each answered with a confidence percentage, and what is scored is the relationship between the percentages and the outcomes rather than the answers. Group them afterwards: if ten answers marked 90 percent produced seven correct, your 90 is a measured 70, a gap of 20 points, which per the course autopsy is the ordinary finding rather than an unusual one. A curve below the diagonal is overconfident and a flat curve carries no information at all, which is a different problem worth knowing about. Then use it where it costs something, because a decision you would only take at 90 percent confidence is one you are actually taking at 70, and per O100-03 that matters most on the largest commitments. And expertise does not substitute for it: Tetlock's expert forecasters knew their subjects, did not clearly beat simple extrapolation, and the style that predicted accuracy was not the one that makes good television.