AI Product Psychology

Trust Calibration: The Best Users Are Half-Skeptical

The first four lessons scored points for experience. This one hits the brakes: for trust as a metric, the target was never a perfect score. AI will err—that’s an unfixable physical fact—so a healthy user looks like this: use it freely where it’s safe to trust, keep a wary eye where you should check. Drift either way and you pay tuition. This lesson starts with two real tuition bills, then hands you three calibration tools you can tweak yourself.

Two ways to crash first
Full trust causes accidents; no trust wastes money. The foundational review of trust calibration (Lee & See, 2004) states the requirement clearly: trust must align with the system’s actual capability. Too much trust is misuse; too little is disuse—and each end of the spectrum has its own bill.
⚠ Overtrust: treating hallucination as truth

2023, New York—Mata v. Avianca: a practicing lawyer used ChatGPT to find case law and pasted 6 fabricated cases that don’t exist straight into court filings. After the judge checked each one, the lawyer was sanctioned and made global news. Similar accidents kept coming: AI output is fluent, confident, and well formatted—every surface cue nudges the judgment “this looks solid”, and those cues have nothing to do with whether the content is true.

⚠ Undertrust: AI becomes an expensive paperweight

The crash at the other end is quieter: a company buys AI tools, an employee hits an error once, and from then on every line of output gets sentence-by-sentence review. Review costs more than writing it yourself, so people stop using it. Procurement keeps paying; efficiency never rises. Undertrust doesn’t make the news—it only shows up in the internal postmortem titled “AI tool active usage: 8%.”

How reliable the system actually is → How reliable users think it is → Overtrust zone · misuse Undertrust zone · disuse Calibration line: trust = real capability
The x-axis is how reliable the system really is; the y-axis is how reliable users think it is. Product design’s job: pull users stuck in the two corners toward the diagonal. The three tools below all pull on that same line.
Six users, pulse-check one by one

Here’s the judgment mantra first: watch “risk × verifiability”, not “how powerful AI is.” The same full adoption is calibrated on a weekly report and overtrust in court. Six scenarios—you’re the calibrator.

Are these six users’ trust healthy? 0 / 6
Three options: overtrust, undertrust, well calibrated. Judge first, then read the explanation.
Make citations clickable

Calibration tool one: source citations—the kind you can open. Users who want to verify jump to the original in one click, which reins in overtrust; users too lazy to verify still see “there’s a source” and get reasonable confidence, which patches undertrust. The win-win requires citations that are real and clickable: a decorative fake citation exposed once is ten times worse than no citation. The AI answer below hangs three superscripts—open each one and compare with the source text.

Three citations—catch the one that doesn’t match 0 / 3 opened
Tap superscripts [1][2][3] to expand the source cards, and check whether the sentence in the answer matches what the source actually says.
Q: Fired verbally in the second month of probation—can I claim compensation?
Yes, you can claim. When an employer terminates a labor contract during probation, it must state the reasons to you[1]; if the termination is unlawful, the damages are twice the economic compensation standard[2]. Don’t wait: the limitation period for labor arbitration is one year, counted from the day you were dismissed[3]. Keep the dismissal notice, attendance records, and pay stubs as evidence.
One of the three citations doesn’t match the source—which?
Tune a confidence-signal set yourself

Tool two: confidence. Part One already covered it—the model doesn’t know what it doesn’t know, so self-reported certainty is unreliable. But the product layer has honest proxy confidence signals: how many docs retrieval hit, how relevant they are, whether sources agree, whether the knowledge cutoff covers the question. Below is the same medication answer; three switch groups map to three product decisions—flip them and watch how the answer on the right and the user’s trust calibration change.

Confidence-signal designer Flip the switches
Answer wording
Source citations
Warning bar
User trust calibration
22
Ibuprofen and acetaminophen—OK to alternate for a fever?
[1] Clinical guide to antipyretic analgesics [2] Pharmacopoeia · NSAIDs [3] Tertiary-hospital medication Q&A
Cry wolf five times in a row

Tool three does one job: that tiny line—“AI may err; please verify important information”—is anyone actually reading it? Psychology’s answer is banner blindness: Benway & Lane (1998) found with eye-tracking that elements constantly shown in a fixed spot get filtered out by the brain as background texture. Tap through five answers below and watch that line disappear with your own eyes.

Constant disclaimer vs. conditional warning Answer 0 / 5
First finish five answers in the “Constant disclaimer” group, then switch to “Conditional warning” to compare. Banner fade simulates how user attention dulls.
Attention to the warning
100
Sources and further reading: The foundational review of trust calibration is Lee & See (2004) Trust in Automation: Designing for Appropriate Reliance; misuse and disuse come from Parasuraman & Riley (1997); the lawyer-pasting-hallucinated-cases story is the real case Mata v. Avianca, Inc. (S.D.N.Y. 2023); banner blindness is from Benway & Lane’s (1998) eye-tracking work. How to repair trust after it collapses continues in Lesson 6 on algorithm aversion (Dietvorst et al., 2015); human gates for high-risk actions get a full design spectrum in Lesson 7 on defensiveness.

✅ What this lesson wants to share