← The JournalThe Science of Judgment

Confidence is not evidence

Your certainty stays flat across decisions that deserve very different levels of conviction. Calibration is the measure that survives contact with outcomes.

Confidence is not evidence. It is a feeling your brain produces — often reliably, often not — and it does not track the difficulty of the decision in front of you. Ask someone how sure they are about a hiring call, a market entry, and whether to move their parent into care, and you will often get the same number: 80%, maybe 90%. The decisions have nothing in common except the person deciding.

That flat certainty is the calibration gap: the distance between how confident you feel and how often you are right at that level of confidence. Most people are overconfident. They claim high certainty on decisions where the base rate of being wrong is also high — and because they rarely keep score, they never discover the pattern.

The judgment literature has measured the gap for decades. When people give 90% confidence intervals, those intervals typically contain the true answer far less than 90% of the time — often under half, depending on domain and elicitation method.

Stated confidence vs. hit rate in interval estimates

Source: Soll & Klayman — Overconfidence in Interval Estimates (JEP:LMC, 2004), summarizing Klayman et al. (1999) — Classic finding across domains: subjective 90% intervals are systematically too narrow.
Stated confidence target90%
Typical observed hit rate (~)45%

If your "I'm 90% sure" intervals land closer to a coin flip, the feeling of certainty was not information. It was theater with a high production value.

Why confidence lies so consistently

Confidence is not designed to measure evidence. It is designed to reduce cognitive load and enable action. Evolutionarily, that is useful. A founder who is paralyzed at 60% confidence loses to one who acts at 90% — even when the 90% is unjustified.

Modern high-stakes decisions punish the same mechanism. The hire you were 90% sure about fails at 18 months. The investment thesis you never doubted drifts. The relationship decision you just knew was right becomes the conversation you avoided for three years.

None of these outcomes feel like calibration failures in the moment. They feel like bad luck, wrong timing, or information you could not have had. Sometimes that is true. Often it is a story that protects the feeling from examination.

What a Brier score actually measures

The Brier score is a proper scoring rule for probabilistic predictions. You make a forecast (70% chance this hire works out), the outcome resolves, and the score penalizes how far your stated probability was from reality.

  • Predict 95% and you are wrong → large penalty
  • Predict 55% and you are wrong → smaller penalty
  • Predict 95% and you are right → still penalized if 95% was overkill

This matters because it rewards honest uncertainty. A person who says 60% when they mean 60% will outscore a person who says 95% whenever they feel motivated — even if their hit rate looks similar in a small sample.

Calibration is the habit Brier scoring reveals: are you right 70% of the time when you say 70%? 90% of the time when you say 90%? If not, your confidence is not information. It is theater.

Same decision. Two ways of knowing how sure you are.

Confidence as feeling
"I'm 90% sure." Flat across hire, market entry, and care. No record. When wrong, the story protects the feeling.
Confidence as score
Before

Written probability + what would move it 10–20 points

After

Outcome resolved; Brier / hit-rate — narrative second

Pattern

Twenty calls at 80% that resolve at 55% mean everything; one flop means little

The audit question confidence cannot answer

Before a decision, an audit asks a different question than how confident am I? It asks: what would have to be true for this to be the wrong call?

That inversion does not produce a feeling. It produces a list — often an uncomfortable one — of assumptions you were treating as facts. Confidence collapses the list. Examination expands it.

This is why calibration — tracking judgment against outcomes over time — matters. Not because feelings are worthless — they carry signal about avoidance, excitement, and dread — but because feelings do not self-correct. Records do.

What to do with this

  1. Before major decisions, state a probability out loud or in writing. Not I feel good about thisI am 65% confident this hire succeeds in 12 months.
  2. Resolve and score. When the outcome arrives, compare. No narrative first — score first.
  3. Look for patterns, not single outcomes. One wrong call at 80% means little. Twenty calls at 80% that resolve at 55% means everything.

DAUDIT tracks calibration across decisions when you use the instrument consistently — not as a grade, but as a mirror. The mirror is uncomfortable. That is the point.

For framing failures that happen before confidence even enters, read The most-examined decision wins. To examine a decision you are facing now, [open an audit](/#hp-decision-field).

Questions worth asking

What is the calibration gap?
The gap between how confident you feel and how often you are right at that confidence level. Most people are overconfident — they say they are 90% sure far more often than they are correct 90% of the time.
What does a Brier score measure?
The Brier score measures the accuracy of probabilistic predictions. Lower is better. It penalizes both wrong predictions and overconfidence — saying 95% when you should have said 60% hurts your score even if you were right.
Can I improve my calibration?
Yes — but only if you record predictions before outcomes resolve, score them honestly, and review the pattern. Feelings do not improve from reflection alone. Records do.

Thank you for reading. If this sharpened how you think about a decision you are facing, the instrument is one step away.

  • calibration
  • Brier score
  • confidence
  • judgment
The instrument

An essay sharpens how you think. An audit sharpens a decision you are actually facing.

One essay, one email. Not a newsletter unless you want one below.

No hype. No frequency promises. One quiet list for people who care about how they decide.

A response, a correction, a reframe from your own field — considered, not a comment thread.