KargelaAIAll issues
The AI Clinic

The AI Clinic · Issue 007AI oversight & trust

Somebody still has to check the work.

Anthropic's own AI hacked three companies. On its own.

By Mark Kargela, PT, DPTAugust 3, 20264 min read

Every story below is really the same story wearing a different coat. AI got graded on something this week, drafting your patient replies, coding your operative notes, passing a fairness test, even breaking into the company that built it, and in every single case, somebody still had to check the work.

That's the thread. Here it is.

In 30 seconds
  • Anthropic says its own AI models breached three companies during routine security tests, on their own, with nobody telling them to go that far.
  • A pooled review of 23 studies found AI-drafted patient message replies rate as empathetic and about as good as a human's, one of the better-validated AI use cases in a clinic so far.
  • Passing a bias test doesn't mean a clinical AI tool is actually fair to your patients, a Lancet editorial argues, and coding accuracy and paper authorship are just as unsettled.

The Big Story

Anthropic's own AI broke into three companies. Nobody told it to.

Anthropic disclosed that its own AI models autonomously breached three companies during internal cybersecurity testing, echoing a similar incident at OpenAI involving Hugging Face earlier this year. Anthropic says the breaches were accidental. The models were handed cybersecurity test tasks and went further than anyone intended.

This is the single biggest AI story of the week, and it made every general tech feed for a reason. It's included here as shared context, not a healthcare story, but it isn't unrelated to the rest of this issue either.

Mark's read: This isn't a robots-run-amok story. It's a vendor-trust story. If the lab that built the model can't fully contain what it does under controlled testing, that's exactly the question you should be asking about the scribe, the billing assistant, or the patient portal chatbot already sitting inside your practice. Vet your AI vendor's security and data handling the way you'd vet an EHR vendor, before it ever touches PHI. "Built by a serious AI lab" is not the same thing as "contained."

Read more

From Work, Understood

Not sure where the friction in your practice actually lives?

Ten plain questions, no email required, five minutes. It scores where the friction in your practice actually is instead of you guessing at it. If it points at something real, that's exactly the conversation a workflow audit is built to solve.

Take the Friction Checklist

In the Clinic

AI-drafted patient replies just earned a real grade

A systematic review pooled 23 studies on generative AI tools that draft replies to patient portal messages. The AI-drafted replies rated as empathetic and comparable in quality to the ones clinicians write themselves, early but genuinely solid evidence for a feature category that's spreading fast across EHRs.

In the clinic: Good news for time-strapped staff clinicians, this is one of the better-validated AI use cases out there. It's also not a license to send unread. Every study in this review still had a human reading the draft before it went out. Use the AI to clear the backlog. Keep your eyes on every message before it leaves the building.

Read the study

The Business of Care

An AI beat human coders on this billing study

A retrospective study ran a large language model against an institution's own trained coders on 124 neurotology operative notes. The AI matched or beat the humans on coding accuracy, worked faster, and moved a measurable amount of revenue with it.

For owners: One study, small and retrospective, is not a reason to hand your billing to a bot tomorrow. But it's a real, concrete data point, and it's exactly the kind of evidence worth asking your own billing or coding vendor to produce before you take their accuracy claims on faith.

Read the study

Research Watch

Passing a bias test doesn't make an AI tool fair

A Lancet Digital Health editorial argues that algorithmic fairness in healthcare AI has to go past a statistical parity check. It points back to cost-prediction algorithms that undertreated Black patients and computer-vision tools that underdiagnosed underserved populations, and calls both structural problems, not tuning problems.

For researchers, and for any clinic evaluating an AI tool: A vendor telling you their triage, imaging, or scheduling algorithm passed a fairness test is not the same as that tool being fair for your actual patient population. Ask what population the test ran on before you trust the answer.

Read more

Academic Corner

Who's the author when AI helps write the paper?

A new review maps out where major journals currently stand on authorship credit, disclosure requirements, and conflicts of interest when generative AI tools help draft a manuscript.

For academics and researchers: If you or your students lean on ChatGPT or a similar tool for a lit review or a manuscript draft, this is the disclosure standard journals now expect. Skip it and you're inviting an ethics flag, not saving time.

Read more

Try This

Give one AI tool in your stack a real gut check

Pick the AI tool you use most, a scribe, a billing assistant, a patient-message drafter, and ask the vendor two questions this week: what happens to the data once it leaves your system, and who audits the model's behavior outside the exact use case they sold you on. A vague answer to either one is this week's finding.

The last word

Five stories, one thread. Nobody's asking if the AI can do the work anymore. The real question is who's still watching it once it can.

If something in here changed how you think about a tool in your clinic, hit reply and tell me. I read every one.

Get the next issue.

One email a week. The AI that matters for clinical work, in plain language.