It does a really good job with recognizing medications though, which most-patients butcher the name on.
Hallucinations are present, but usually they’re pretty minor (screwing up gender, years).
It doesn’t really seem to understand what the most important part of the conversation is, it treats all the information equally as important when that’s not really the case. So you end up with long text of useless information that the patient thought was useful but not at all relevant to their current presentation. That’s where having an actual physician is useful to parse through what is important or not.
At baseline it doesn’t take me long to write a note so it really wasn’t saving me that much more time.
What I do use it for is recording the conversation and then referencing back to it when I’m writing the note. Useful to “jog my memory” in a structured format.
I have to put a disclaimer in my note saying that I was using it. I also have to let the patient know upfront that the conversation is getting recorded and I’m testing something for Microsoft, etc. etc. You can tell who the programmer patients are because they immediately ask if it’s “copilot“ lol
LLM's have so much potential in medicine, and I think one of the most important applications they will have is the ability to ingest a patient's medical chart within their context and present key information to clinicians that would've otherwise been overlooked in the bloated mess that most EMR's are nowadays (including Epic).
There's been so many times where I've found critically important details hidden away as a sidenote in some lab/path note overlooked for years that very likely could've been picked up by an LLM. Just a recent example - a patient with repeated admissions over the years due to severe anemia, would usually be scoped and/or given a transfusion without much further workup and discharged once Hgb >7. Blood bank path note from 10 years ago mentions presence of warm autoantibodies as a sidenote; for some reason the diagnosis of AIHA is never mentioned nor carried forward in their chart. A few missed words which would've saved millions of dollars in prolonged admissions and diagnostic costs over the years.
I don't mean to come off antagonistic here. But surely the more important benefit is the patient who would've avoided years of sickness and repeated hospital visits?
I've noticed that seems to be a common trend for any AI-generated text in general.
And the prompt landscape in the field is vast. And fascinating. Every specialist has their own preference for what is important to include in a note vs what should be excluded; and this preference changes by disease - what a neurologist want in an epilepsy note is very different from what they need in a dementia note for eg.
Note preferences also change widely between physicians, even in the same practice and same specialty! I'm the founder of Marvix AI (www.marvixapp.ai), an AI assistant for specialty care, we work with several small specialty care practices where every physician has their own preferences on which details they want to retain in their note.
But if you can get the prompts to really align with a physician's preferences, this tech is magical - physicians regularly confess to us that this tech saves them ~2 hours every day. We have now had half a dozen physicians tell us in their feedback calls that their wives asked them to communicate their 'thanks' to us for getting their husbands back home for dinner on an important occasion!
[Edit: typo and phrasing]
And if all hospitals were doing was having doctors treat patients, this would be ok. But healthcare is fueled by these "minor" details and this will result in delays in payment and reimbursent, trouble with patient identification, corruption of clinical coding, etc.
One would image those to be the biggest dangers.
Overall, I was quite impressed. It definitely made writing notes much faster, which all doctors hate to do. While it had some problems with where to put key pieces of information (like putting details from the physical exam back in the history), it only took 5 mins of rearrangement after the visit to complete the note.
For simple diagnoses, it does a decent job coming up with the assessment and plan, probably because all the simple diagnoses were in the training set. For more complex ones though, it needs to be exactly dictated by the doctor. I can see this being used very well in primary care.
Edit: When I said “coming up with an assessment and plan” I mean documenting the assessment and plan based on the ai’s recorded conversation with the patient. The conversation with the patient is meant to be understandable. The “assessment and plan” documentation on the other hand is jargony and meant to be read by other physicians.
And let me make this clear. I, as your patient, I never NEVER want the AI's treatment plan. If you aren't capable of thinking with your own brain, I have no desire to trust you with my health, just like I would never "trust" an AI to do any technical job I was personally responsible for due to the fact that it doesn't care at all if it causes a disaster. It's just stochastic word picker. YOU are a doctor.
I wrote a little bit more of my thoughts here, in case it’s of interest to anyone: [0]
On that same vein, I recently made a tool I wrote for myself public [1] - it’s a “copilot” for writing medical notes that’s heavily focused on letting the clinician do the clinical reasoning, with the tool exclusively augmenting the flow rather than attempting to replace even a little bit of it.
[0] https://samrawal.substack.com/p/the-human-ai-reasoning-shunt
"Having trouble processing a medical claim with 50+ pages of notes? Not to worry, Dragon Copilot Claim Review(tm) trims the fluff and tells you what really happened!"
"Having trouble understanding a large convoluted PR? Not to worry, Copilot(tm) Automated Review has your back!"
"Having trouble decided which cordless vacuum to buy? Not to worry, Amazon's Customers Say(tm) shows you what people think!"
There is definitely _some_ world utility to this arms race, but is it enough?
Medical claims won't be growing in pages just because a doctor can parse them a bit faster. They may grow initially, because it's likely that people's mental capacity is what keeps other factors from ballooning the claims further - but it'll level out when some other practical limit is reached. Same with coding and PRs, same with research and all kinds of activities - except advertising.
There, AI will (already is) causing an arms race, because the "high volume producer"'s goal is to overwhelm their victims, so if the victims start protecting themselves with AI tools, the producer will keep increasing production to compensate. But that's not the fault of AI - it's the fault of allowing the advertising industry to exist.
This seems like a tool that insurance companies would love to get a copy of the data stream, and that could get very sticky quite quickly.
To my understanding, notes would otherwise largely be written from memory after the visit - which adds a fairly significant opportunity for omissions and errors to sneak in.
It seems plausible to me that by fixing that low-hanging fruit, this tool could potentially reach current human levels of accuracy overall even if it has shortcomings in other areas, like not being as good at non-shallow reasoning. Not to necessarily say it's currently at human-level.
> Would it really save time writing paperwork if you have to go through it anyways and check if there's anything wrong?
Five minutes saved per encounter, allegedly[0]. The decrease in clinician burnout and patient satisfaction also seem pretty significant. But, not sure how much Microsoft have massaged those figures.
Some other companies in this space are Epic, Freed, Nuance, DeepScribe, Nabla, Ambience, Tali, Augmedix
The outcome of this is essentially that AI generated healthcare decisions will be superficially laundered through a human doctor, rather than a human doctor simply using the AI as a tool
This may be compounded by insurance companies using the AI as their guideline for their plans and payouts, and government healthcare agencies using the AI as their guideline for acceptable treatment practices
Healthcare is already not an ideal industry from a patient perspective in many places. It is difficult to imagine AI making this situation better for patients
I tested 5 other Ambient AI tools in the past 6 months. All of them can extract the Chief Complaint and do a decent job with the HPI, but as soon as you get a more complicated case, they all fell apart. Sections like Physical Exam became a mess and I have to rewrite the whole thing.
So far the best, in my opinion, is still LucasAI by Lucas Health. Super simple to use, very basic interface. It just works and produces the best notes. I barely touch them. With the ICD10 codes, sometimes it doesn't pick the very best. This thing has been improved over the past month, but it's always much better than Copilot.
The AVS are both good, but with LucasAI I can translate immediately in Spanish and send it over to the patient via email/sms.
I haven't explored the integration with Copilot. Bad news, I know for a fact LucasAI is not fully integrated with Meditech yet, but I just copy and paste the whole note in 10 seconds. Not a problem. My buddy is using ECW and he said it's integrated. Haven't seen with my eyes.
I hear the sound of 100 defamation lawyers cold calling Nabla right now
In most of the rest of the world we have functioning healthcare systems, and I imagine this will be useful in reducing the time doctors spend on paperwork and increasing the time they're providing healthcare.
Marginalized groups will receive worse care because they weren’t in the training set.
Lazy doctors will get lazier.
Errors will be deeply buried in a bureaucracy with no available remedy because nobody at the ground level understands the technology.
Scheduling will become less transparent. Escalations will be automatically denied.