pmcJAMA Netw OpenJAMA Netw Open3211jamasd101729235JAMA Network Open2574-3805pmc-is-collection-domainyespmc-collection-titleJAMA NetworkPMC11581472PMC11581472.111581472115814723940104110.1001/jamanetworkopen.2024.38573zld2401841ResearchResearch LetterOnline OnlyHealth InformaticsUtility of Artificial Intelligence–Generative Draft Replies to Patient MessagesUtility of Artificial Intelligence–Generative Draft Replies to Patient MessagesUtility of Artificial Intelligence–Generative Draft Replies to Patient MessagesEnglishEdenMD 1 LaughlinJanelleMD 2 SippelJeffreyMDMPH 3 DeCampMatthewMDPhD 1 4 LinChen-TanMD 1 Division of General Internal Medicine, University of Colorado School of Medicine, DenverUniversity of Colorado Health Medical Group, LongmontDivision of Pulmonary Medicine, University of Colorado School of Medicine, DenverCenter for Bioethics and Humanities, University of Colorado, DenverArticle Information

Accepted for Publication: August 11, 2024.

Published: October 14, 2024. doi:10.1001/jamanetworkopen.2024.38573

Correction: This article was corrected on November 26, 2024, to fix an error in an institution name in the Methods.

Open Access: This is an open access article distributed under the terms of the CC-BY License. © 2024 English E et al. JAMA Network Open.

Corresponding Author: Eden English, MD, University of Colorado School of Medicine, 360 S Garfield St, Denver, CO 80206 (eden.english@cuanschutz.edu).

Author Contributions: Dr Sippel had full access to all of the data in the study and takes responsibility for the integrity of the data and the accuracy of the data analysis.

Concept and design: English, Laughlin, Sippel, Lin.

Acquisition, analysis, or interpretation of data: English, DeCamp, Sippel, Lin.

Drafting of the manuscript: All authors.

Critical review of the manuscript for important intellectual content: All authors.

Statistical analysis: English, Sippel, Lin.

Administrative, technical, or material support: English, DeCamp, Sippel, Lin.

Supervision: English, Laughlin, Lin.

Conflict of Interest Disclosures: Dr DeCamp reported serving as a consultant on ethics policy issues for the American College of Physicians outside the submitted work. No other disclosures were reported.

Data Sharing Statement: See Supplement 2.

Additional Contributions: We would like to thank the Epic analysts at UCHealth who worked tirelessly to enable this functionality for our users, with special thanks to Cortney Arellano, Jordan Phillips, Aubrey Clark, and Sandra Segura. We would also like to thank Alex Constantinides, DO (UC Health Medical Group), for his clinical leadership and help with deployment in the early pilot clinics These individuals did not receive compensation for their contributions.

1410202410202426112024710472784e243857311620241182024141020242111202428112024Copyright 2024 English E et al. JAMA Network Open.https://creativecommons.org/licenses/by/4.0/This is an open access article distributed under the terms of the CC-BY License.jamanetwopen-e2438573.pdfError in Methods71126112024e2451866JAMA Network Open10.1001/jamanetworkopen.2024.51866PMC1160022539589750Are Artificial Intelligence-Generated Replies the Answer to the Electronic Health Record Inbox Problem?7101102024e2438528e2438528JAMA Netw Open10.1001/jamanetworkopen.2024.3852839401042

This quality improvement study analyzes the usefulness of patient message replies drafted by artificial intelligence for various health care practitioners.

pmc-status-qastatus0pmc-status-liveyespmc-status-embargonopmc-status-releasedyespmc-prop-open-accessyespmc-prop-olfnopmc-prop-manuscriptnopmc-prop-legally-suppressednopmc-prop-has-pdfnopmc-prop-has-supplementyespmc-prop-pdf-onlynopmc-prop-suppress-copyrightnopmc-prop-is-real-versionnopmc-prop-is-scanned-articlenopmc-prop-preprintnopmc-prop-in-epmcyespmc-license-refCC BY
Introduction

Messages from patients to care teams via patient portals are increasing1 and may be associated with burnout.2 Health systems are deploying large language models (LLMs) to address this issue by drafting replies to patient messages.3,4,5

In contemporary team-based care, how useful are LLM’s to different team members? What iterative changes in an LLM prompt might improve the usefulness of the draft reply?

Methods

This prospective quality improvement (QI) study was determined to be exempt from review and the requirement of informed consent by the University of Colorado institutional review board. The QI study was conducted from September 2023 to March 2024 within a large academic health system (UCHealth), following the Standards for Quality Improvement Reporting Excellence (SQUIRE) reporting guideline. We deployed an LLM, GPT-4 (Generative Pretrained Transformer 4 [OpenAI]), via the electronic health record (Epic Systems). Internally dubbed PAM Chat (Patient Advice Message Chatbot reply), the LLM drafted a reply to incoming messages from the patient portal in 9 clinics (6 primary care and 3 specialty), for nurses, medical assistants (MAs), and clinicians (physicians and advance practice clinicians [APCs]).

Based on feedback during the project, we iteratively improved prompts (see the eMethods in Supplement 1 for prompt versions and rationale). For example, we included the assessment and plan (A/P) of the last progress note written by the recipient clinician, emphasizing conciseness. We disclosed that content was computer generated before editing by the clinician to patients.

We surveyed all users 2 weeks after all clinics were live with PAM Chat, including the net promoter score (NPS) and unvalidated de novo items assessing perceptions of the tool. Analysis of variance was used to test the null hypothesis that attitudes toward PAM Chat would not vary by professional role. Data analysis was conducted using JMP Pro software version 17.0.0 (SAS Institute) and took place from April to June 2024. The threshold for significance was a 1-sided P < .05.

Results

PAM Chat generated 21 323 drafts, and 2596 drafts were used (12%) across all 166 users (12 nurses, 14 MAs, and 93 clinicians). Nurses had more favorable views of PAM Chat (Table 1). Nurses were more likely to agree that PAM Chat reduced the need to forward messages to physicians and APCs and helped them to stay within their scope of practice. Of all nurses, 11 (92%) believed that PAM Chat helped improve efficiency, empathy, and tone. Nurses were most likely to recommend PAM Chat to others (NPS, 58). MAs, physicians, and APCs were less favorable.

Comparison of Survey Respondents (Registered Nurses, Medical Assistants, and Clinicians) Who Replied Agree or Strongly Agree
QuestionRespondents, No. (%) (N = 69)P value (χ2)
Registered nurses (n = 12)Medical assistants (n = 14)Clinicians (n = 43)a
“I can reply more quickly.”11 (92)7 (50)20 (46).03 (10.9)b
“It’s easier to express written empathy.”11 (92)6 (43)18 (42).01 (13.1)b
“The message tone is appropriate.”11 (92)11 (79)28 (65).21 (5.9)b
“PAM Chat is fun to use.”11 (92)8 (57)22 (51).08 (8.2)b
“Minimal risk of giving incorrect information.”8 (67)2 (14)12 (28).04 (10.1)b
“Significant edits are rarely needed.”NANA5 (12)<.001 (NA)c
“I can address more messages without forwarding.”9 (75)3 (21)NA.002 (12.8)b
“This helps me to stay within my scope of practice.”6 (50)4 (29)NA.01 (9.0)b
Net promoter score (likely to recommend to others)58 −29−43.004 (NA)d

Abbreviations: NA, not applicable; PAM Chat, patient advice message chatbot.

Clinicians included physicians and advance practice clinicians.

χ2 Likelihood ratio.

Two 1-sided t test for single-group measurement.

Analysis of variance.

Comments from colleagues about PAM Chat were mixed: “Sometimes the answers are astoundingly helpful, other times they are not helpful at all.” Including the A/P from the last note made some replies useful. Illustrative comments are in Table 2.

Example PAM Chat Replieson
Incoming patient messagePAM Chat draft replyaComments
“Subject: Knuckle swallow in the morningMessage: What should i do about this...hope to hear from you soon...thank you.”“Hello [Name],Swelling in the knuckles in the morning is a common symptom of rheumatoid arthritis. It’s important to continue taking your prescribed medications, as they can help manage this symptom. You can also try some gentle hand exercises and applying a warm compress to your hands in the morning to help reduce the swelling. If the swelling continues or worsens, please make an appointment so we can evaluate your symptoms and adjust your treatment plan if necessary.Take care.”The incoming patient message was incredibly brief and had a misspelled word (swallow vs swollen). AI was able to elucidate the question and used the diagnosis from the last progress note in the response.
“Subject: Check inMessage: Hi Dr, I have a note on my calendar today to update you about menopause symptoms. The hot flashes have really subsided and I’m doing well overall. Hope things are good with you! Thank you.”“Hello [Name],I’m glad to hear that your menopause symptoms, particularly the hot flashes, have subsided and that you’re doing well overall. It’s important to continue with the plant-based diet and regular exercise as these can help manage any remaining or future symptoms.Take care.”The last progress note suggested a plant-based diet to manage symptoms.
“Subject: Unusual discharge from noseMessage: Hi. Sunday morning I had a spontaneous discharge from my left nostril unlike anything I’ve experienced before. It was a thin yellowish liquid that steadily trickled out like a pretty good bloody nose…I’m scheduled to see you on the 20th for my annual. If there is anything you think I should do before then, let me know. It just seems so unusual. Thanks.”“Hello [Name],I’m sorry to hear about your unusual nasal discharge. It could be due to a variety of reasons, including a sinus infection or allergies. However, it’s also important to rule out a cerebrospinal fluid leak, which can sometimes present with clear, watery nasal discharge. I recommend you come in for an appointment as soon as possible so we can evaluate your symptoms and determine the best course of action.Take care.”Although potentially appropriate for a neurosurgery clinic reply, this response was not clinically relevant here. PAM Chat does not have access to clinic specialty.
“Subject: WegovyMessage: That makes sense. I will avoid the compounded medication. I wanted to send you a Harry and David basket for the holidays. What is a good address to send it to? Thanks!!”“Hello [Name], Thank you for your message.I appreciate your kind gesture, but as a health care provider, I must maintain professional boundaries and cannot accept personal gifts. Your thoughtfulness is more than enough.Take care.”AI was empathetic and appropriate and gently declined the offered gift.

Abbreviations: AI, artificial intelligence; PAM Chat, patient advice message chatbot.”

All outgoing PAM Chat drafts include the language: “Automatically generated and then edited by [Clinician or staff name].”

Discussion

In this QI study, nurses liked PAM Chat more than MAs and physicians or APCs, with more than 90% believing that it improved efficiency, empathy, and tone. Our results diverge from other reports3,4,5 that describe primary care nurses’ negative views of LLMs. This finding may be because the specific messages each group sees are different. Physicians and APCs may preferentially receive complex messages, which are more difficult for the LLM; MAs may view messages containing clinical information as out of scope. This role-based difference and our overall use rate of 12% suggests that, moving forward, LLMs may need to be tuned to recognize who will receive the message (MA, nurse, or physician or APC) and create a reply accordingly.

Our findings also suggest that implementing PAM Chat with ethics principles (eg, transparency) and including the last A/P note can improve utility. Interestingly, in response to a message about a gift, the LLM generated a response that suggested that gift-receiving is not permitted—an AI oversimplification given that the American Medical Association Code of Ethics permits gifts in limited circumstances.6

This evaluation was conducted in 9 clinics in 1 health system and may not be generalizable to all health systems. We hope sharing our prompts and results will support ongoing improvement of LLMs in the pursuit of high-quality, safe, and effective patient care while reducing care team burden.

ReferencesHuang M, Khurana A, Mastorakos G, . Patient portal messaging for asynchronous virtual care during the COVID-19 pandemic: retrospective analysis. JMIR Hum Factors. 2022;9(2):e35187. doi:10.2196/3518735171108 PMC9084445Akbar F, Mark G, Warton EM, . Physicians’ electronic inbox work patterns and factors associated with high inbox work duration. J Am Med Inform Assoc. 2021;28(5):923-930. doi:10.1093/jamia/ocaa22933063087 PMC8068414Garcia P, Ma SP, Shah S, . Artificial intelligence-generated draft replies to patient inbox messages. JAMA Netw Open. 2024;7(3):e243201. doi:10.1001/jamanetworkopen.2024.320138506805 PMC10955355Tai-Seale M, Baxter SL, Vaida F, . AI-generated draft replies integrated into health records and physicians’ electronic communication. JAMA Netw Open. 2024;7(4):e246565. doi:10.1001/jamanetworkopen.2024.656538619840 PMC11019394Baxter SL, Longhurst CA, Millen M, Sitapati AM, Tai-Seale M. Generative artificial intelligence responses to patient messages in the electronic health record: early lessons learned. JAMIA Open. 2024;7(2):ooae028. doi:10.1093/jamiaopen/ooae02838601475 PMC11006101American Medical Association. Gifts from patients: code of ethics 1.2.8. Accessed September 9, 2024. https://code-medical-ethics.ama-assn.org/sites/amacoedb/files/2022-08/1.2.8.pdf

eMethods. PAM Chat Prompt Changes and Current General Prompt

Data Sharing Statement