Transcript
PLEASE NOTE: While an effort has been made to correct errors in this AI-generated transcript, some mistakes may have been missed. Users should take responsibility for checking anything drawn from here. Further, the guests are speaking in their own capacity and presenting opinions in an open discussion. This podcast does not represent the RACP’s position and is not the authoritative last word on the subject matter.
MIC CAVAZZINI: Welcome to Pomegranate Health, and part two part series on the application of Large Language Models to medical records. I’m Mic Cavazzini for the Royal Australasian College of Physicians. In the last episode we talked about ambient scribes in the consulting room which are making their way from general and private practice, to clinics and wards near you. We heard about some pilot studies to determine the feasibility and reliability of this tool, what kind of time savings can be expected, but also what some of the caveats are. My guests were three early adopters who have asked some rigorous questions of the technology. They were consultant generalist Professor Ian Scott, who piloted ambient scribes in Brisbane’s Metro South network.
IAN SCOTT: I'm currently a clinical consultant in AI at Metro South Hospital and Health Service and a professorial research fellow at the Queensland Digital Health Centre.
MIC CAVAZZINI: Adelaide haematologist and general physician Dr Andrew Vanlint.
ANDREW VANLINT: And I've unintentionally fallen into being somewhat of an expert doctor in documentation and clinical coding, possibly the most unsexy thing to be a specialist of, but probably getting more spotlight in the last five to ten years than ever before.
MIC CAVAZZINI: And Dr Adam Brand, who was behind another scribing trial on the Gold Coast.
ADAM BRAND: I'm an emergency physician at Gold Coast Hospital and I'm also the medical director for digital and information. Yeah, a very keen user of new and current technology.
Reliable discharge summaries
MIC CAVAZZINI: In the scribing examples we heard about in the last episode there’s one AI that transcribes every word of a consult, and then a Large Language Model that makes sense of that transcript. It could generates a summary note of the episode, or fill out the main entries into the electronic medical record. The information only flows in one direction, from the immediate consult up into the EMR.
But what if the LLM were let out of the box, and given access to all recorded notes and test results? The usefulness of this comes clearly into perspective when we think about discharge summaries. Discharge from hospital is a notorious example of discontinuity in care, which is made all the more risky when there are gaps in these guiding documents.
In a 2003 study conducted at a large hospital in New York it was found that almost half of patients discharged over a year subsequently experienced at least one medical error. In a significant portion of cases, important test results acquired in hospital, or new medications initiated, were not communicated to the community practitioner receiving the discharge summary. Or the specialists had intended for further work up to take place that was not actioned.
In some cases the miscommunication occurs because of delay in sending the discharge letter or a failure to do so entirely. In a 2013 study conducted through the Flinders University Medical Centre in South Australia, it was found that the absence of a discharge summary was associated with an almost 80% increase in the rate of readmission within a week. And a discharge letter sent with a week’s delay may as well have not been sent at all, so little did it improve outcomes.
Discharge summaries are time consuming to write, so clearly there is a place for using Large Language Models to help get them out quickly and to make them more comprehensive. An LLM with full access to the patient record is less likely to forget some charted medication or test result and it’s going to be more consistent in the format in which it spits these letters out.
A few years ago, Dr Vanlint conducted a survey of GP preferences with regards to discharge letters. General Practitioners told the researchers they wanted information about, new or changed diagnoses with key investigations mentioned, medication changes with explanatory rationale and items to follow up on. Among the complaints were inconsistent structure, not enough information and heavy use of jargon and abbreviations that were specialty specific.
According to an audit conducted at Royal Melbourne Hospital one in five words in a discharge summary is an abbreviation or an initialism, 7% of which could be ambiguous. Does MS stand for Multiple Sclerosis or Mitral Stenosis? Is PE a Pulmonary Embolism or a Pleural Effusion? Was that meant to be QD, QOD or QID, potentially an eightfold difference in dosing! Another study published in the MJA in 2015 found that common initialisms such as TTE, EST and NKDA were misinterpreted by around a third of the 155 GPs surveyed. If you too are scratching your head, they’d be transthoracic “echocardiogram”, “exercise stress testing” and “no known drug allergies”.
Okay, so we want standardisation and clarity. This is core business for Large Language Models trained for the task and bolted onto the electronic medical record. At a pinch, you could even try using Chat GPT, in a simulated setting at least, where the data privacy and sovereignty isn’t an issue. In a webinar for the RACP EVOLVE series, Professor Scott shared a few examples of published trials demonstrating that ChatGPT-4 was found to help could, incrementally, improve the clarity of discharge letters and reduce the time taken to write them. I asked how confident he was in this use of LLMs from the evidence to date.
IAN SCOTT: Well, I think they definitely have place. I think the literature is now there's a number of studies that have suggested that the discharge summaries can be as accurate and complete and in more structured format and also tend to include, then, the information that the recipient really wants to see. And in a way that can be put into lay summary form so the patient actually gets a copy of it as well, I think that's a real plus.
We just did a recent study involving a number of our interns where we asked them to look at a narrative of three or four patient cases. Okay, here's the notes from day one through to discharge. So, we got them to look at that and then we asked and then we asked them to look at a generated discharge summary using Chat GPT 4.0. And interestingly enough, they found that the factual accuracy was really good. In other words, there was no factual errors. Things were as they as the as the notes described them.
But what they really spent some time on is that they wanted to make sure that the information was in a way that was most informative to the GP. So, they kind of rephrased certain things, they restructured it as well because the note didn't quite come out in the way that they thought kind of had a logical flow. The LLM tended to chunk certain things whereas they thought that needed to be separated out so that different parts of the script needed to go in different places.
So, to me that was I think quite insightful. It was a small sample, but it suggested that people may not necessarily again save a lot of time, and we didn't save a lot of time with the scribes either. You don't save time necessarily, but you are able then to reflect and make sure that the note and the or the discharge summary that comes out at the at the end is really of higher quality than perhaps you would have done yourself trying to write it. And I think that for again for interns to write discharge summaries, sometimes of some fairly complicated long admissions, that's a big cognitive ask. Perhaps that using the scribe to generate a discharge summary gives them time then to sit back and look at this and then perhaps better understand what actually did happen. And also at the end of the day they've had some cognitive input to structure it in a way that they know then meets the interests of the recipient.
MIC CAVAZZINI: I think we keep echoing the same messages in different ways. And I think this might this might be another example, but an interesting paper I came across in PLOS Digital Health, written in Tokyo, took a ground-up approach to identifying the various sources of information that contribute to a physician written discharge summary. They found that 61% of information was drawn directly from inpatient records. A significant chunk was described as external information, mostly from prior referral documents, which may or may not get uploaded to the EMR properly. But 11% of the content of a human written discharge summary was not documented elsewhere. And as they write, it “contained speculation and post-discharge plans that were generated by the author”. So, the researchers were concerned that asking an LLM for a discharge summary when there were gaps in available information would make hallucinations more likely. But Andrew, is this risk managed by training the humans as well as the models? I mean i is it the same problem as we've already described.
ANDREW VANLINT: I think it's similar to that concept I mentioned before of going, we're no longer drafting, we're editing and it's not necessarily that editing means we just check it and go, “Yep, done”. It's an invested process of editing where we go, “Okay, this is this is an 80% complete document. Let's finish off that 20% and we need to make sure we're allocating time for people to do that in a thoughtful way.
ANDREW VANLINT: We're doing a similar study at the moment with where we have a large language model in-house on our servers in public, which reads through admission, ward round, consult notes from the document and then drafts a discharge summary, the intern RMO/BPT on that particular team, will then read that, edit it, and then put that into the record. So, there's still the accountability, there's still the agency to change things and improve things but we've programmed that large language model to ascribe to the handover to GP format that we developed and is our network-wide approach to discharge summaries.
Which if you wind back like five, six years ago, we didn't have an approved format. And now all our metropolitan and most of our rural networks in South Australia have ascribed to this handover to GP format in theory. So, we're all agreed that that's what it wants and that's what most GPs are wanting to have the format in. LLM will do that and then we can still edit it and add those little anecdotal insights or things that you've been verbally discussed but have never made it to the digital record.
MIC CAVAZZINI: So, in your health network, you've already made the scribe all seeing it can dip into the EMR?
ANDREW VANLINT: Yeah, so this one's not an ambient listening, it's just a large language model that is doing the discharge summary.
MIC CAVAZZINI: Okay. It's another model again, yeah.
ANDREW VANLINT: But you know, in the future we will combine both and you would have even patients now ask now, sometimes they'll say, “Does the scribe that's listening to our consultation have access to my records to incorporate that stuff?” And I say, “No, not yet, but that's going be the next generation of this software, which will then require much bigger computing power as well”. So, there'll be an environmental and power cost and infrastructure required for that next step forward.
MIC CAVAZZINI: Yeah, I’m sort of eliding, now, those different models. Adam, from my understanding of the work up in Queensland or the rollout now in New South Wales, there's that doesn't sound like that two way access for the scribe or other language models is anywhere on the near horizon. It's more the governance, isn't it?
ADAM BRAND: Yeah.
MIC CAVAZZINI: We talk a lot about consenting patients to the use of a scribe. Do we need to be consenting patients to permitting access of the EMR to another model? And then you get into that whole minefield of where is the data being stored and for how long et cetera. Do we have any clarity on that?
ADAM BRAND: Yeah, look, we'd been trying to do some form of AI scribe, I think, prior to 2022, before they even existed. And so we were in talks with Nuance and then Microsoft and trying to you know be first on the scene with that. I think when LLMs made it into the into the mindset of the world it was then a race to see who could get data residency in Australia.
I think we have very great privacy rules in Australia, which I'm really grateful for, which ensures that everything stays onshore, whether that's in a cloud onshore that sits in Sydney or Brisbane or Melbourne, or in our own hospital infrastructures. But there's also the idea of ensuring that any of these companies do not learn from any of the content that goes in because that is patients' medical records and we want to benefit from it, but without giving up that privacy. So, those things were incredibly important for us and still are. And thankfully the industry is listening.
Improved clinical coding
MIC CAVAZZINI: So far, we’ve talked about how automation in the compiling of clinical notes can help improve clinical practice. But there is also a very low-hanging fruit in improving accuracy and rigor in medical billing. Most listeners will be aware that Australia’s Medical Benefits Schedule pays for outpatient services on according to unique items numbers. The MBS subsidises the cost of the service moments after the patient walks out the door, and it’s only much later that the Professional Services Review might identify where erroneous billing has taken place. In episodes 56 I gave examples of the many mistakes in billing made by well-intentioned providers trying to navigate a Byzantine system, or the outright fraud perpetrated by some less well-intentioned ones. And in episode 58, billing expert Loryn Einstein explained just how the PSR identifies and prosecutes these issues.
After hearing a draft of today’s podcast she told me that nearly every practitioner that gets dragged before the PSR has inadequate clinical records to explain the MBS items that were billed. It doesn’t matter if you’ve got thorough work-flow that you don’t bother documenting because it’s so routine. The PSR wants to know what happened in every unique consultation. Loryn told me that better documentation, perhaps with assistance from ambient scribes, would likely reduce "clawback" from the PSR for private practitioners who have done the work but have failed to document it to sufficient standards.
Billing for inpatient care occurs in a completely different way, but here too there is scope for Large Language Models to improve accuracy and enhance reimbursement of hospital services. Secondary and tertiary services are funded in a job lot based on the care provided over the entire previous year. Activity-Based Funding, as it’s called, depends on the summation of all episodes of patient care, whose complexity is coded from the clinical notes available in the EMR. Poor quality documentation can lead to services being shortchanged thousands of dollars per treatment episode. Be grateful to the clinical documentation specialists in the back room at your hospital that perform this nitpicky coding work once your patient is discharged.
It kind of works like this. First, based on the primary diagnosis, the encounter is assigned one of the 17000 or more codes in the ICD-10, the International Classification of Diseases. The primary diagnosis related-group for a pneumonia is distinct to that for a myocardial infarction. But what kind of pneumonia? Is it community-acquired or hospital-acquired? What antibiotics were prescribed? Were there other comorbidities or complications? So, even within a single DRG there are subcodes which indicate complexity of care, and this can be only be extracted from a detailed patient record.
The final DRG code you end up with translates to a weighting number, such that a simple pneumonia gets paid out at $6300 while the complex one is more like $12,100. A simple bowel resection will get earn close to $38,500 while a complex one is around $72,500. Accurate coding doesn’t just make sure that public services get funded appropriately, it also helps with tracking hospital-acquired complications, hospital benchmarking, planning clinical trials and tracking of disease outbreaks. Dr Andrew Vanlint is a clinical documentation evangelist, and has his own YouTube series called Coding Matters. I asked him to explain the most common mistakes made in documenting patient encounters.
ANDREW VANLINT: So, the two biggest errors that most doctors make in public or private inpatients is firstly that we are not using specific diagnostic terminology. We're using a lot of inferences and other language. So, a really common one would be we talk about falls, but we don't talk as much about the cause of and the consequence of the falls, because the falls is an event, it's not a diagnosis. But the thing that causes the fall often is a diagnosis, the Parkinson's, the previous stroke, the hy postural hypotension.
And the consequence of the fall, drastically can be, you know, intracerebral haemorrhages, the more specific the better, subarachnoid/intracranial. It can be fractures, displaced or undisplaced that may or may not require surgical intervention. And it can be simpler things like bruises, hematomas, sprains, soft tissue injuries often not documented properly. So, we need to use more diagnostic terminology and we use a lot of things commonly like, they're hypoxic or they're oxygen-dependent. The coders are legally not allowed to use that language. What they need us to say is at least once in the record, “acute type one respiratory failure” or “acute T1RF” is acceptable too.
So, that's the first one. They're easily correctable and AI can fix that because it we can we've trained our models to do that for our discharge summaries is to autocorrect those. The second one is about causal relationships, which again, the coders by law, are only allowed to use very direct causal language. So, the most common thing is that doctors, we love to sound intellectual and nuanced. I love it. We all love it. Maybe a little bit less so in in ED. They're just more like, “Let's get it done”.
And so we use a lot of these terms like “on the background” of or “in the setting of”. That's indirect language. So, a classic is, “this person is presented with a delirium on the background of dementia”. And we're kind of going, “they're kind of linked, but they're kind of not”. We can't actually use that language for coding, we need to use direct language. So, I'll give you my common example of this is the who broke the window incident. So, if I say, “the window broke in the setting of Ian throwing a rock”, Ian could be guilty. It's a bit suspicious, but it's not a hundred percent certain. But if I say, “the window broke due to/secondary to/caused by Adam throwing a rock”, Adam is guilty as hell. He definitely broke that window.
So, we really need specific diagnostic terminology and we need, direct causal language to link these things together. And unfortunately the culture that we have, we need to culturally shift that and teach our juniors and students the right way to document from the start. Or we can use AI and program our models to correct those things, to use the right terminology and to use direct causal language to replace what we're already doing. So, there's those two methods and at the moment I'm pursuing both paths so that the hospital can run financially sustainably as well, because that's important to be able to continue funding high quality care. We're seeing more patients who are older and more complex, but we're not getting any more money. And the biggest issue is that we are not documenting that complexity in a way that can be captured through coding and remunerated through activity-based funding.
MIC CAVAZZINI: Yeah, I'm going to send listeners to your YouTube videos which unpack common mistakes in each specialty. And the take home message is be explicit; be specific, you know, join the dots, show you're working. Andrew a and yeah.
ANDREW VANLINT: It doesn't have to take a long time. It's not working harder, it's working smarter. Good documentation does not have to be necessarily be lots more paragraphs or sentences. It's just being more purposeful with the words you choose.
MIC CAVAZZINI: Again, I'll go to Adam for a practical example. I don't know if you're familiar with a tool that Andrew developed called eCoder, which codes ED presentations. ED presentations go by a different process to the one we've described earlier. But it's a web-based tool that anyone can use and as I understand it, you have to enter the patient summary by hand and it will suggest to you the most appropriate ICD codes. As an emergency physician have you have you got any evaluation of how well this is affected remuneration?
ADAM BRAND: We have something similar internally that we've developed on our own—Lakehouse. That sort of analyses the medical record and suggested alternative, better coding. So, we have a data quality role that exists that goes through the records every day to manage things that don't map properly from SnowMed to ICD ten for the technical minded. But, in that process there's also an improvement in the coding, which is slightly different from the inpatient, which is again different from the outpatient.
But yeah, our clinical coding team quite heavily uses LLMs over our medical record to find opportunities that are not adequately coded. And they will reach out to those teams and work with them to improve the quality of the documentation but also enhance their coding at the same time.
MIC CAVAZZINI: Has anyone put a dollar figure on how much better that your ED is getting remunerated?
ADAM BRAND: Yeah, it it's definitely got seven figures in it. You know, it's been a process for us over the years to improve, essentially like proving the work that we done we have done. So, in very much the same way when you get marked for a maths problem at school you get you get marks for your workings. And I think we often do a lot of hard work in any of the care spaces and it is a real shame if you're underselling yourself as a health service by not conveying the complexity with which the patients are there.
But you know, my hope in this is that this is all about activity-based funding and coding, and we're all proving that we're very active. The panacea really is to go to value-based funding where we're actually talking about the outcomes, and that comes with another level of data quality and universality of the data before we can get there. But I think with these tools and moving in a digital space, we could hopefully get to a stage where we're actually talking about outcomes, complications, and you know, better care so that we're not just getting interested in, you know, activity.
Safety regulation
MIC CAVAZZINI: I’ve just got a couple more questions. I don't want to end with doom and gloom, but a couple of the caveats that we haven't talked about yet. We mentioned a moment a moment ago about how human coders are forbidden from making inferences, even if two pieces of information seem to sit together. Well, AI scribes and LLMs love to make inferences. And Adam, in your study with Gold Coast Health, 16% of staff surveyed had observed bias in the way that the scribe assigned relevance and meaning to parts of the clinical discussion. You know, which aspects of the history to elevate. And this is just a scribe, it's not meant to be an assistive diagnostic tool. So, even in the act of choosing what features to from the transcript to go into the record, could the scribe be nudging clinicians a clinician's diagnostic thinking in one direction or another?
ADAM BRAND: Yeah, look, I think that that is a totally valid question. I suppose my response to it is that I have noticed in the, you know, two to three years since we did that study, there's been a significant improvement in just the general quality overall; the prompt engineering, the output that is there, where actually it doesn't try and fill in gaps. It's more honest about whether there where there is a hole in what has been talked about in there and that's a key part of all of this.
We we're using LLMs at the moment over our intranet around clinical guidelines and we've had to be incredibly cautious to start off with that we are adamant that the AI assistant should not be trying to guess or fill in gaps. And so, I think we're learning as we go along. And I think it takes research and publishing and high quality review to put some rigour over what comes through in anecdote or in conferences or in conversations between people or what we're talking about at the moment. But it always comes down to that person that is reading the note and ensuring that it feels right to them and is factually accurate.
MIC CAVAZZINI: And AIs that are explicitly involved in some stage of diagnostic assistance have to be approved by Australia’s Therapeutic Goods Administration as a medical device, to demonstrate that patient outcomes are improved not compromised. But the TGA is not too fussed about ambient scribes. By contrast, the UK regulator has recently made moves to bring them under the formal standards and registration of medical devices. In the GP magazine, the Medical Republic this intel was shared from a panel about AI scribing in the UK.
So, in a trial of Heidi AI across 55 British GP practices, more than 30,000 consultations were transcribed over 15 months. In these, a 10% error rate was observed. And the presenters expressed concern that apart from being corrected on the fly by users, as we’ve described, there was no “flag this error” button, to alert not just the developers but collect real world data for the regulators. So, how much more or less cautious do we need to be about this? I don't know if Ian or Andrew is a better place to answer regulation question.
IAN SCOTT: Well, look, the regulators are really going to have a tough time, because these models continue to develop, they evolve, they're taking on different functions and the and the lines between say a purely assistive scribe versus a scribe that gives you sort of diagnostic suggestions is also becoming more blurred. I think at the end of the day we're going have to rely on health services and people who are actually using these tools to come up with their governance structures and have external creditors making sure that as an AI ecosystem we're using these tools properly, rather than just rely on the regulator to say, well, “you can use this and you can't use that”.
So, my feeling is that over time the regulators are going to say, “Well, we'll have some definite decisions about certain tools, but there'll be other areas in which it's going to be more grey and we're going to have to rely on the users”- that is health service organizations, general practices, to be accredited to actually use AI systems. And who actually does that accreditation? Well, it might be the Commission of Quality and Safety and Healthcare, for example, it might be some other some other body.
But I think it comes back to Andrew's point as well that we're expecting AI to be perfect. Why are we doing this? I mean we know that there's a lot of error and a lot of bias and a lot of systemic problems that we're dealing with right now. If AI can help us remove some of that and actually improve our quality of care and our and our quality of documentation, then let's start using it. Why do we have to wait it to for to be perfect? And these models and the companies themselves do have skin in the game and it's in their vested interest to make sure that they actually are improving the model. They'll be wanting to improve the model and I think we should assist them in doing that.
I think in terms of we have a role to audit what we're doing. So, for example, in rolling out our scribe, we're going to have certain markers that we'll be auditing on a regular basis to make sure that one, we're asking the users themselves to notify us and we've set up a call line to say if you have errors or you see hallucinations or you see a constant pattern of erroneous material coming out of the AI scribe, then please let us know and we'll feed that back immediately to the to the vendor.
Number two, we'll be watching and seeing how many people are actually getting in and editing their notes, for example, rather than just posting it straight into the EMR and we'll be giving them a little touch on the head and say, “Look, you know, we've noticed that in the last ten entries you haven't done anything in terms of editing”. So, there are ways that we can I think monitor, and that to me is part of the governance process that we're involved as a health service in making sure that our users are using the tool properly. So, I think over time things will get better. That's how I see it.
And I think, all of us probably confront the issue how do we spend time and money on actually generating the evidence base that indicates how efficacious and how safe these tools are. I struggle to find funding within our health service to actually do this work. And yet that's the work we need to do. We've made a rule that at our service we're going to pilot test everything that comes in. So, nothing's going to be off the shelf and just deployed. We'll be doing clinical trials to make sure that we do have an evidence base, if only to convince our executive that this is a worthwhile investment, alright?
MIC CAVAZZINI: And I think that local picture is telling given some of the investigative journalism from the Ninefax papers recently about actually how poorly the TGA does keep an eye on safety of medical devices that you know, and that that's that includes intrathecal pumps and portable defibrillators and insulin pumps, which, apparently they wave through quite a lot of them without testing just based on European certifications. And, a quote here from that SMH article, “The TGA conducted detailed conformity assessments on just 4.3 per cent of the 55,223 devices it had approved for use since 2016. In that decade, it rejected just 830.” So, no pressure on you guys, but takes alert minds like your own. Andrew, do you want to have a final word here?
ANDREW VANLINT: I think for me what it comes down to is trust. I want to work in an environment where I trust the leaders, I trust my peer consultants, I trust my junior staff, I trust the medications and equipment that I use. And I want to do the same for AI scribes and other AI tools. We need to trust them. And in order to have that trust, similar to what Ian said, I need credibility. So, I need evidence that these tools are safe and effective. I need transparency, which means I need to know where are the weak points, just like I need a trainee who's strong in some areas but still developing others. I need to know where are those areas and how are we going to improve them, how do I report that stuff as well? And then I need accountability. I need to know who ultimately is accountable when something goes wrong, because something will go wrong. Even in the best clinicians, even the best equipment, there will be failures and errors from time to time. And I need all three of those.
So, I suppose with that in mind and that trust, and at the risk of sounding a friend of bureaucracy, I would welcome a lot of these tools to come under the scrutiny of the TGA, because I think that is the fastest way for them to build credibility and trust for the public and the local communities we serve and our fellow clinicians and health leaders to then be able to bring them in. And we're seeing very much there's always been a risk aversity in our public systems, particularly. Because of good clinical governance, we are often very slow to adopt technology in healthcare way behind other industries because it needs to be really good and proven to be that way.
In the case of AI scribes it's really – it's enhanced the way that I practice and many patients have commented that they feel it's a higher quality consultation because it's not a consultation between myself, the computer and them. It's me and them and the computer is somewhat separate.
I need to consent the patient. So, I do so verbally and a lot of practices have that in writing when the patient joins the practice. But I actually find it allows me to focus more on the patient. Because in previous times I may have spent some time focusing on the patient, some time typing things or looking things up. And I find the ambient scribe means I just sit and really give good eye contact and good focus with my body language and attention almost purely to the patient until such time I need to order some tests. And then I conclude my time with the patient, they leave the room, I click stop, and within five seconds I've got a note that I can then use as is or edit accordingly.
The patient experience
MIC CAVAZZIN: Most of the clinicians participating in the Gold Coast and Brisbane pilots described the same experience that Dr Vanlint just has—that the tool helped improve the quality and intimacy of their consultations and around 60% of patients agreed with this too. It should be noted that this feeling isn’t universal and there can be unexpected outcomes. In MJA Insight, consumer health advocate Elizabeth Deveny described how some patients didn’t share as much about their personal lives when sensed that the machine was listening. Or they felt that the doctor remembered more about them personally when they took the notes themselves.
Maybe it will take some getting used to, and consumers are yet to see some of the other potential benefits of harnessing generative language models for documentation. Take patient-facing letters. Did you know that 44% of Australian adults are functionally handicapped by low literacy and that the federal government’s standard for accessible writing is year 7 level. Well, you could translate your discharge letter into instructions suitable for such a patient with just one more click. In a future-gazing video on Youtube, paediatric surgeon Associate Professor Bhavesh Patel demonstrates several other exciting examples of generative AIs in various consumer-facing roles that could elicit greater engagement in self-management and follow up.
And remember that infamous study published in the JAMA press about an early version of ChatGPT that was rated as more informative and more empathetic that real human clinicians? That was based on responses provided to an “Ask the doctor” web forum. Rather than worrying about AI taking your job, imagine instead an ambient tool that could “read the room” and give you post-hoc feedback on how clear your language and your explanations were, and how your patients responded. Maybe that sounds a bit dystopian—machines teaching clinicians how to be more human—but think of it this way. These machines haven’t just been trained on every digitally available medical text-book, but also every novel you never had time to read, and every chat forum that you never knew existed. Their entire reason for being is to parse language.
If you’re still not convinced about the real linguistic and conceptual intelligence demonstrated by large language models I suggest you go back to episode 100, titled “Conversations with ChatGPT”. From about the 25 minute mark, I summarise some theory of mind experiments conducted by research psychologists working with Microsoft that have some really spooky outcomes. That was the final episode in a five part series that also discussed the ergonomics of diagnostic aides, questions around governance of assistive AI and the legal mine field when medical errors will occur.
I better wrap up now, but not before thanking Andrew Vanlint, Adam Brand and Ian Scott, for their very engaging contribution to Pomegranate Health. Many thanks also to the members of the podcast editorial who provided feedback on early drafts of this story. They are RACP physicians Aidan Tan, Simeon Wong and Joseph Lee and also Paul Cooper PhD and Loryn Einstein. At our website I’ve provided links to Professor Scott’s lectures for the RACP EVOLVE Series and Dr Vanlint’s Youtube Channel on Coding Matters. It’s also worth checking out the resources at Clinical Documentation Improvement Australia which I found very helpful, starting with those from director Mike Kertes. Just go to racp.edu.au/podcast, then click on episode 158 for all the links. There you’ll find a complete transcript of this episode embedded with links to all the academic literature. And there’s more on AI at the Medflix video library which you can find at elearning.racp.edu.au, or just search for RACP online learning.
If you found these podcasts useful, please tell a colleague or leave a review at your favourite pod browser. It’s easy to subscribe with any podcasting app on your smartphone, or via the email alerts request at our website. This podcast was produced on the lands of the Gadigal clan of the Yura nation. I pay respect to their elders past and present. I’m Mic Cavazzini, thanks for listening.