The AI Clinic · Issue 005Clinical AI credibility
The rematch had a funder.
OpenEvidence-funded study says it beats general AI
The big story this week walks back something we covered in issue 001. Back then, a blinded study found general AI models beat the specialist clinical tools. This week, a new study says the opposite, that OpenEvidence beats general models. The wrinkle: OpenEvidence funded it.
Underneath that fight over whose tool is smarter sits a quieter problem almost nobody is solving. Once you turn an AI tool loose in your clinic, who's checking that it's still doing its job six months later? That's this week's Try This.
- A company-funded study says OpenEvidence beats general AI models at clinical reasoning, the opposite of what an independent, blinded test found back in issue 001.
- A new framework points out that almost nobody checks whether an AI tool is still working correctly once it's actually running in a clinic.
- A Lancet editorial asks whether AI chatbots can help close a 1.1-billion-person mental health gap, or whether they'll miss the people who need help most.
The Big Story
OpenEvidence backs a study that finds it beats general AI
A new study says the specialized clinical tool beats general-purpose models like ChatGPT on clinical reasoning tasks, the opposite of what we covered back in issue 001, when a blinded test found the generalists came out ahead. The difference this time: OpenEvidence funded the research.
OpenEvidence is already in daily use by a lot of clinicians for evidence lookup, so the underlying question is a real one. Is a tool built specifically for medicine actually better at clinical reasoning than the model you already have open in another tab? The study says yes. The person who signed the check has a horse in that race.
Mark's read: I want purpose-built clinical tools to win this argument, and maybe they do. But a funder-backed study isn't the same as an independent one, and I'm not changing what I recommend to anyone until this gets replicated by someone with nothing riding on the answer. Until then, treat this as a marketing claim wearing a lab coat, not settled science.
From Kargela AI
Not sure where AI actually fits in your practice yet?
The AI Readiness Audit is a free, 5-minute scorecard that shows you where AI actually fits in your practice right now. No pitch, no call required unless you want one.
In the Clinic
Can an AI actually run a motivational interviewing conversation?
This proof-of-concept study tested whether a structured LLM workflow could deliver motivational interviewing, the technique built to help a patient move toward a health behavior change. MI works, but it's hard to scale because it takes a trained clinician to do it well. The AI approximated the core moves. It struggled when a patient's answer didn't follow the expected script.
In the clinic: This is the realistic near-term use case, not an AI therapist, but an AI that extends behavioral support between visits without replacing the relationship you already have with the patient. Could we get to AI communication coaches that back up your own MI skills? I think we're closer than most people realize, but the scripted-conversation gap in this study is exactly where a real clinician still earns their keep.
The Business of Care
Payers just got their own AI for billing disputes
Zelis launched an AI platform built for health plans to manage No Surprises Act billing disputes, the Independent Dispute Resolution process that decides who owes what on an out-of-network claim. It's built for the payer side, not yours.
For owners: If you ever file an IDR claim, understand that the insurer on the other side is getting faster and more systematic at fighting it. Your documentation and your dispute submission need to be airtight, because you're not negotiating with a person working through a spreadsheet anymore.
Leadership & Culture
Staff pushback on AI is a leadership problem, not a tech problem
Paul Gough argues that staff resistance to AI tools follows the same pattern he's seen with cash-pay transitions, wellness programs, and price changes. The obstacle was never the new thing. It was whether the owner was willing to lead through the friction of it.
For owners: I've watched good ideas die in a practice because the owner treated staff skepticism as a veto instead of a normal part of change. If your team is pushing back on a documentation tool or a scheduling assistant, that's a culture conversation to have on purpose, not a reason to shelve the tool.
The Big Picture
Can AI chatbots help close the mental health gap?
A Lancet Digital Health editorial puts a number on the scale of the problem: more than 1.1 billion people worldwide live with a mental health condition, and the traditional care system can't reach most of them. The piece asks whether large language models can help close that gap and lands somewhere honest: maybe, with real safety and equity concerns still unresolved.
Mark's read: Your pain patients are already talking to a chatbot about their anxiety or their sleep between visits, whether you know it or not. That's not a hypothetical for later. It's happening now, in a conversation you're not part of.
Health Equity
Will ChatGPT Health widen the gap it's supposed to close?
A journal commentary argues OpenAI's ChatGPT Health feature, which pulls personal medical records into a chatbot, may end up widening healthcare disparities instead of closing them. The authors' concern: the tool will reach patients who already have insurance, tech access, and digital literacy first, while the patients who need the most help stay furthest from it.
For clinicians: This is the useful counterweight to every AI-closes-the-gap headline. A lot of the chronic pain patients I've treated over the years are exactly the population this commentary worries about. If AI health tools only reach the already-resourced, the quality gap between your patients doesn't shrink. It grows.
Research Watch
A framework for watching AI after it goes live
This paper lays out three things worth tracking once an AI tool is actually running in a clinic: whether the system still works the way it's supposed to, whether it's still accurate, and whether it's actually helping. The authors point out that almost nobody does this once a tool clears the sign-off stage. Deploy and forget is still the norm.
For owners: Most practices have no process for checking whether an AI tool is still doing its job six months in. Models drift, vendors push updates, and nobody re-checks unless something breaks in front of a patient. If your practice runs a scribe, a chatbot, or a documentation tool, someone should own a quarterly check on all three, in writing.
Try This
Give one AI tool in your clinic a 10-minute check-up
Pick the AI tool your clinic touches the most, a scribe, a chatbot, a summarizer, and ask three questions. Is it still doing what we bought it to do? Has anyone compared its output against a human this month? And who actually reads the answer before it reaches a patient? If you can't answer all three, you just found this week's fix.
The last word
Same tools, different funders, different conclusions. The lesson isn't which AI is smarter this week. It's that you can't take either side's number at face value. You have to ask who paid for it and what happens to your patients if it's wrong.
If you have thoughts on the newsletter, good or bad or something missing, reply to any issue and let me know. I read every one.
Get the next issue.
One email a week. The AI that matters for clinical work, in plain language.