Confidence is not evidence. It is a feeling your brain produces — often reliably, often not — and it does not track the difficulty of the decision in front of you. Ask someone how sure they are about a hiring call, a market entry, and whether to move their parent into care, and you will often get the same number: 80%, maybe 90%. The decisions have nothing in common except the person deciding.
That flat certainty is the calibration gap: the distance between how confident you feel and how often you are right at that level of confidence. Most people are overconfident. They claim high certainty on decisions where the base rate of being wrong is also high — and because they rarely keep score, they never discover the pattern.
The judgment literature has measured the gap for decades. When people give 90% confidence intervals, those intervals typically contain the true answer far less than 90% of the time — often under half, depending on domain and elicitation method.
Stated confidence vs. hit rate in interval estimates
Source: Soll & Klayman — Overconfidence in Interval Estimates (JEP:LMC, 2004), summarizing Klayman et al. (1999) — Classic finding across domains: subjective 90% intervals are systematically too narrow.If your "I'm 90% sure" intervals land closer to a coin flip, the feeling of certainty was not information. It was theater with a high production value.
Why confidence lies so consistently
Confidence is not designed to measure evidence. It is designed to reduce cognitive load and enable action. Evolutionarily, that is useful. A founder who is paralyzed at 60% confidence loses to one who acts at 90% — even when the 90% is unjustified.
Modern high-stakes decisions punish the same mechanism. The hire you were 90% sure about fails at 18 months. The investment thesis you never doubted drifts. The relationship decision you “just knew” was right becomes the conversation you avoided for three years.
None of these outcomes feel like calibration failures in the moment. They feel like bad luck, wrong timing, or information you could not have had. Sometimes that is true. Often it is a story that protects the feeling from examination.
What a Brier score actually measures
The Brier score is a proper scoring rule for probabilistic predictions. You make a forecast (“70% chance this hire works out”), the outcome resolves, and the score penalizes how far your stated probability was from reality.
- Predict 95% and you are wrong → large penalty
- Predict 55% and you are wrong → smaller penalty
- Predict 95% and you are right → still penalized if 95% was overkill
This matters because it rewards honest uncertainty. A person who says 60% when they mean 60% will outscore a person who says 95% whenever they feel motivated — even if their hit rate looks similar in a small sample.
Calibration is the habit Brier scoring reveals: are you right 70% of the time when you say 70%? 90% of the time when you say 90%? If not, your confidence is not information. It is theater.
Same decision. Two ways of knowing how sure you are.
Written probability + what would move it 10–20 points
Outcome resolved; Brier / hit-rate — narrative second
Twenty calls at 80% that resolve at 55% mean everything; one flop means little
The audit question confidence cannot answer
Before a decision, an audit asks a different question than “how confident am I?” It asks: what would have to be true for this to be the wrong call?
That inversion does not produce a feeling. It produces a list — often an uncomfortable one — of assumptions you were treating as facts. Confidence collapses the list. Examination expands it.
This is why calibration — tracking judgment against outcomes over time — matters. Not because feelings are worthless — they carry signal about avoidance, excitement, and dread — but because feelings do not self-correct. Records do.
What to do with this
- Before major decisions, state a probability out loud or in writing. Not “I feel good about this” — “I am 65% confident this hire succeeds in 12 months.”
- Resolve and score. When the outcome arrives, compare. No narrative first — score first.
- Look for patterns, not single outcomes. One wrong call at 80% means little. Twenty calls at 80% that resolve at 55% means everything.
DAUDIT tracks calibration across decisions when you use the instrument consistently — not as a grade, but as a mirror. The mirror is uncomfortable. That is the point.
For framing failures that happen before confidence even enters, read The most-examined decision wins. To examine a decision you are facing now, [open an audit](/#hp-decision-field).