AI-Rx - Your weekly dose of healthcare innovation
Estimated reading time: 4 minutes
TL;DR
Harvard, Oxford, and the Broad Institute built ATHENA-R1, an 8B-parameter treatment-reasoning agent.
Instead of answering from memory, it gathers evidence across 212 biomedical tools before committing.
It reportedly out-reasoned GPT-5 and a 671B model, despite being a fraction of the size.
The edge isn't scale. It's that the agent knows when it doesn't know, and checks.
GPT-5 had the same tools and reached for them ~1% of the time.
Welcome to AI-Rx 👋
After weeks on what big models can and can't do, here's a result that flips the story: a small model beat much larger ones at one of the hardest tasks in medicine, for a reason every clinician was trained on.
Why treatment reasoning is so hard
Choosing the right therapy isn't recall. You weigh a patient's other conditions, contraindications, interactions, and evidence that keeps changing, all at once.
A model answering from memory alone barely survives a real case, because its knowledge is frozen and it never stops to check.
What ATHENA-R1 does differently
Built at Harvard Medical School with Oxford and the Broad, ATHENA-R1 is an agent, not a one-shot answerer. Faced with a question, it gathers evidence step by step across 212 biomedical tools, drug databases, interaction checkers, population data, before committing.
It was trained with reinforcement learning over those live tools, across the full span of FDA-approved drugs.

The reported results: it beat GPT-5 on open-ended drug reasoning by a wide margin, and outperformed a 671-billion-parameter model, despite being roughly 80x smaller.

The part that matters isn't the size
Here's the detail that reframes everything. GPT-5 had the exact same tools available. It reached for them ~1% of the time, and kept answering from memory. ATHENA's whole edge is that it knows when it doesn't know, and goes to look it up. A smaller model that checks beat a larger one that assumes.

Why this is the right lesson
That habit, gathering evidence before committing, is what makes a good doctor good. We spend years training clinicians not to trust their first impression, to look things up, to verify before acting.
That discipline matters more in medical AI than raw scale does. Encouragingly, the agent's outputs also held up against millions of real patient records and won a blinded expert review, but the through-line is the design philosophy, not the leaderboard.
Here's my final thought
Stop asking only how big or how accurate a model is. Ask whether it knows the limits of what it knows, and whether it checks. The most important capability in medicine was never raw recall. It was the discipline to verify before you act, and you can build that into an agent.
If you practice medicine: how much of your skill is knowledge, and how much is the discipline to verify before you commit?
Dr. Bhargav Patel, MD, MBA
Physician-Innovator | AI in Healthcare | Child, Adolescent, & Adult Psychiatrist | Medical & AI researcher
Source: Gao S, Noori A, Zhu R, Clifton D, Zitnik M, et al. ATHENA-R1: an AI agent for treatment reasoning trained by reinforcement learning over a universe of 212 biomedical tools. 2026.
Enjoyed this? Forward it to a colleague, or subscribe: https://bhargavpatelmd.beehiiv.com/