AI-Rx - Your weekly dose of healthcare innovation

Estimated reading time: 3 minutes

TL;DR

  • A Mayo Clinic team traced the reasoning behind AI's reading of real oncology notes, not just the final answers.

  • GPT-4 produced a reasoning error in 23% of interpretations. Most weren't hallucinations, they were confirmation bias and anchoring.

  • Those errors were the ones most tied to guideline-discordant recommendations.

  • A correct answer can hide broken reasoning, and accuracy scores miss it entirely.

  • We built decision support to protect patients from human bias. We may be automating it.

Welcome to AI-Rx 👋

Most evaluations of clinical AI ask: was the answer right? A Mayo Clinic team asked a better question, was the reasoning right? The gap between those two is where this issue lives.

They scored the thinking, not the answer

The team gave models real oncology notes, breast, pancreatic, and prostate cancer, and instead of grading only the final interpretation, traced the logic behind it. That choice is the point: a model can reach a defensible conclusion through indefensible logic, and a standard accuracy score will still mark it correct.

The model made human errors, not machine ones

GPT-4 produced a reasoning error in 23% of its note interpretations. The most common weren't bizarre fabrications, they were confirmation bias and anchoring, the exact traps medicine trains clinicians to avoid.

The model latched onto an early detail, favored evidence that confirmed its first impression, and skipped findings that should have changed the plan. That's how humans fail under pressure, now reproduced fluently and at volume.

Those errors were the dangerous ones

This is the finding that matters. The reasoning failures weren't harmless, they were the errors most strongly associated with recommendations that broke with clinical guidelines. And because the final answer can still land in the right place, a benchmark won't catch it.

The uncomfortable irony

We built clinical decision support specifically to protect patients from human cognitive bias. If we now deploy AI that reproduces confirmation bias and anchoring, fluently, at scale, behind a reassuring accuracy number, we haven't removed the bias. We've automated it and hidden it.

Here's my final thought

The safest clinical AI isn't the one with the highest accuracy score. It's the one whose reasoning you can inspect. If your organization is piloting AI on clinical notes, the question isn't just whether the answers are right.

It's whether anyone is reading the reasoning behind them.

If you deploy clinical AI: is anyone auditing how it thinks, or only checking what it concludes?

Dr. Bhargav Patel, MD, MBA

Physician-Innovator | AI in Healthcare | Child, Adolescent, & Adult Psychiatrist | Medical & AI researcher

Source: Cognitive bias in LLM reasoning compromises interpretation of clinical oncology notes. npj Digital Medicine. 2026. (arXiv:2511.20680)

Enjoyed this? Forward it to a colleague, or subscribe: https://bhargavpatelmd.beehiiv.com/