pmc JAMA Netw Open JAMA Netw Open 3211 jamasd 101729235 JAMA Network Open 2574-3805 pmc-is-collection-domain yes pmc-collection-title JAMA Network PMC12059973 PMC12059973.1 12059973 12059973 40332938 10.1001/jamanetworkopen.2025.8874 zoi250324 1 Research Original Investigation Online Only Geriatrics Machine Learning Multimodal Model for Delirium Risk Stratification Machine Learning Multimodal Model for Delirium Risk Stratification Machine Learning Multimodal Model for Delirium Risk Stratification Friedman Joseph I. MD 1 2 Parchure Prathamesh MSC 3 Cheng Fu-Yuan MS 3 Fu Weijia MS 3 Cheertirala Satyanarayana MS 3 Timsina Prem ScD 3 Raut Ganesh MS 3 Reina Katherine DNP 4 Joseph-Jimerson Josiane DNP 5 Mazumdar Madhu PhD 3 6 Freeman Robert DNP 3 Reich David L. MD 7 Kia Arash MD 3 7 1 Department of Psychiatry, Icahn School of Medicine at Mount Sinai, New York, New York 2 Department of Neuroscience, Icahn School of Medicine at Mount Sinai, New York, New York 3 Institute for Healthcare Delivery Science, Icahn School of Medicine at Mount Sinai, New York, New York 4 Nursing Administration, Mount Sinai Morningside Hospital, New York, New York 5 Department of Nursing, The Mount Sinai Hospital, New York, New York 6 Department of Population Health Science and Policy, Icahn School of Medicine at Mount Sinai, New York, New York 7 Department of Anesthesiology, Perioperative, and Pain Medicine, Icahn School of Medicine at Mount Sinai, New York, New York Article Information Accepted for Publication: March 5, 2025. Published: May 7, 2025. doi: 10.1001/jamanetworkopen.2025.8874 Open Access: This is an open access article distributed under the terms of the CC-BY License . © 2025 Friedman JI et al. JAMA Network Open . Corresponding Author: Joseph I. Friedman, MD, Department of Psychiatry, Mount Sinai Hospital, One Gustave L. Levy Place, Box 1230, New York, NY 10029 ( joseph.friedman@mountsiani.org ). Author Contributions: Drs Timsina and Kia had full access to all of the data in the study and take responsibility for the integrity of the data and the accuracy of the data analysis. Concept and design: Friedman, Timsina, Reina, Mazumdar, Freeman, Reich, Kia. Acquisition, analysis, or interpretation of data: Friedman, Parchure, Cheng, Fu, Cheertirala, Raut, Reina, Joseph-Jimerson, Mazumdar, Kia. Drafting of the manuscript: Friedman, Fu, Reina, Mazumdar, Kia. Critical review of the manuscript for important intellectual content: Parchure, Cheng, Fu, Cheertirala, Timsina, Raut, Reina, Joseph-Jimerson, Mazumdar, Freeman, Reich, Kia. Statistical analysis: Friedman, Parchure, Cheng, Fu, Cheertirala, Timsina, Raut, Mazumdar, Kia. Administrative, technical, or material support: Friedman, Cheng, Raut, Reina, Joseph-Jimerson, Freeman, Reich, Kia. Supervision: Friedman, Mazumdar, Reich, Kia. Conflict of Interest Disclosures: None reported. Data Sharing Statement: See Supplement 2 . 7 5 2025 5 2025 8 5 488197 e258874 25 11 2024 5 3 2025 07 05 2025 09 05 2025 09 05 2025 Copyright 2025 Friedman JI et al. JAMA Network Open . https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the CC-BY License. jamanetwopen-e258874.pdf Key Points Question Can a machine learning model be used to accurately stratify risk of hospital delirium in live clinical practice? Findings This quality improvement study including 32 284 inpatient admissions developed an automated multimodal machine learning delirium risk stratification model that demonstrated acceptable discriminative performance in live clinical practice. Additional analyses using 7023 admissions assessed for delirium with the Confusion Assessment Method showed that model deployment was associated with a significant 4-fold increase in delirium detection rates and significant reductions in daily doses of benzodiazepine and antipsychotic medications. Meaning These findings suggest that a machine learning model may be used to automate delirium risk stratification in live clinical practice and may enhance delirium identification and care. This quality improvement study describes the development, operationalization, and validation of a multimodal machine learning–based model for delirium risk stratification in clinical practice and its associations with workflow and clinical outcomes. Importance Automating the identification of risk for developing hospital delirium with models that use machine learning (ML) could facilitate more rapid prevention, identification, and treatment of delirium. However, there are very few reports on the performance of ML models for delirium risk stratification in live clinical practice. Objective To report on development, operationalization, and validation of a multimodal ML model for delirium risk stratification in live clinical practice and its associations with workflow and clinical outcomes. Design, Setting, and Participants This quality improvement study developed an ML model supported by automated electronic medical records to stratify the risk of non–intensive care unit delirium in live clinical practice using the Confusion Assessment Method as the diagnostic reference standard, with an iterative model update method. Data from patients aged at least 60 years admitted to non–intensive care units at Mount Sinai Hospital between January 2016 and January 2020 were used to train and test the ML model presented. The model was validated in live clinical practice from March 2023 to March 2024. Analysis of the model’s associations with workflow and clinical outcomes was conducted retrospectively in 2024, comparing hospitalized patients prior to deployment of any model version (pre-ML cohort) and during model clinical deployment (post-ML cohort). Main Outcomes and Measures Outcomes of interest were area under the receiver operating characteristic curve, monthly delirium detection rates, median length of hospital stay, and daily doses of opiate, benzodiazepine, and antipsychotic medications administered. Results The overall sample included 32 284 inpatient admissions (mean [SD] age, 73.56 (9.67) years, 15 157 [46.9%] women). A total of 25 261 inpatient admissions of older patients with both medical and surgical primary diagnoses represented the combined model testing and training cohort (median age, 73.37 [66.42-81.36] years) and live clinical deployment validation cohort (median [IQR] age, 72.11 [62.26-78.97] years), while 7023 inpatient admissions of older patients with both medical and surgical primary diagnoses represented the combined pre-ML (median [IQR] age, 74.00 [68.00-81.00] years) and post-ML (median [IQR] age, 75.33 [68.34-82.91] years) cohorts. The model presented is a fusion of electronic medical record patient data features and clinical note features processed by natural language processing. The results of model validation in live clinical practice included an area under the curve of 0.94 (95% CI, 0.93-0.95). Median (IQR) monthly delirium detection rates of inpatients assessed for delirium with the Confusion Assessment Method increased from 4.42% (95% CI, 3.70%-5.14%) in the pre-ML cohort to 17.17% (95% CI, 15.54%-18.80%) in the post-ML cohort ( P < .001). Post-ML vs pre-ML cohorts received lower daily doses of benzodiazepines (median [IQR] 0.93 [0.42-2.28] diazepam dose equivalents vs 1.60 [0.66-4.27] diazepam dose equivalents; P < .001) and olanzapine (median [IQR], 1.09 [0.38-2.46] mg vs 2.50 [1.17-6.65] mg; P < .001). Conclusions and Relevance This quality improvement study demonstrates the feasibility of a novel multimodal ML model to automate delirium risk stratification in live clinical practice. The model demonstrated acceptable performance in live clinical practice and may facilitate resource allocation to enhance delirium identification and care. pmc-status-qastatus 0 pmc-status-live yes pmc-status-embargo no pmc-status-released yes pmc-prop-open-access yes pmc-prop-olf no pmc-prop-manuscript no pmc-prop-legally-suppressed no pmc-prop-has-pdf no pmc-prop-has-supplement yes pmc-prop-pdf-only no pmc-prop-suppress-copyright no pmc-prop-is-real-version no pmc-prop-is-scanned-article no pmc-prop-preprint no pmc-prop-in-epmc yes pmc-license-ref CC BY Introduction Delirium is a frequently occurring neuropsychiatric syndrome in hospitalized individuals, precipitated by various medical, surgical, pharmacological, and environmental factors. It manifests as acute mental status changes and is associated with short-term and long-term negative consequences, including increased morbidity, mortality, hospital readmission rates, functional decline, and extended hospital stays. 1 , 2 Despite its frequent occurrence and negative consequences, the diagnosis of delirium is often missed and delayed in the hospital setting. Therefore, more timely diagnosis of delirium and accurate assessment of delirium risk followed by rapid treatment of existing delirium and preventative measures in patients with high risk are essential to mitigating these adverse outcomes. To facilitate the more timely diagnosis of delirium in the hospital setting, there has been an increasing interest in applying artificial intelligence (AI) to risk stratification for the development of delirium in hospitalized patients. Consequently, there is a rapidly evolving body of research exploring the development of prediction models using machine learning (ML) and supported by electronic medical record (EMR) data for delirium that develops in hospital settings outside the intensive care unit (ICU). 3 , 4 , 5 , 6 , 7 , 8 , 9 , 10 , 11 , 12 , 13 , 14 , 15 , 16 , 17 , 18 , 19 , 20 , 21 , 22 , 23 , 24 , 25 , 26 , 27 , 28 , 29 , 30 , 31 , 32 , 33 , 34 , 35 , 36 , 37 , 38 , 39 , 40 , 41 , 42 , 43 , 44 , 45 , 46 , 47 , 48 However, most of these models have not been tested in live clinical settings, resulting in a paucity of evidence for AI’s added value to workflow and clinical outcomes associated with delirium in the hospitalized patient. More recently, a few publications of ML-based delirium risk models have investigated the benefit of using natural language processing (NLP) features in the model being evaluated. 11 , 25 , 32 , 33 , 46 Furthermore, 2 publications have reported on the fusion of the traditional EMR-based features with NLP features in their ML-based delirium risk stratification model, with both supporting the model-boosting efficacy of the NLP component. 11 , 32 Despite these advances, to our knowledge, only 4 publications have reported on performance of ML-based delirium prediction models in live clinical practice, with 3 reports testing the same model (without NLP). 9 , 22 , 48 Two of these reports compared nursing- and physician-derived delirium risk assignment vs ML-based delirium risk assignment, 9 , 22 and 1 study tested model predictive performance against a structured delirium assessment tool. 48 Clinical outcomes were reported in only 1 of these studies: user acceptance, as assessed by an author-developed questionnaire. 22 The fourth report described clinical deployment of ML-based prediction models (with an NLP component) for acute kidney injury, sepsis, and delirium, with delirium diagnosis based on EMR-derived diagnostic codes assigned at discharge and no report of clinical impact. 25 To address these shortcomings, in this study, we aimed to report on the development, training, validation, and operationalization of a multimodal ML-based model for delirium risk stratification in live clinical practice, with sequential model optimization, using a vertical integration approach 49 and present the results of the optimized model’s performance in live clinical practice and its associations with clinical workflow and clinical outcomes. Methods Study Design and Setting This quality improvement study was approved by the Mount Sinai Hospital (MSH) institutional review board, which also granted a waiver of informed consent due to minimal risk based on the protected health information that was accessed. We followed the Standards for Quality Improvement Reporting Excellence ( SQUIRE ) reporting guideline. This study was conducted at MSH from July 2019 through March 2024 using an interrupted time-series analysis 50 with a focus on evaluating the association of the EMR-supported ML-based delirium risk stratification model we developed with workflow and clinical outcomes. Model Development Since 2016, a novel delirium intervention program has been operational within the MSH system, targeting patients who have already developed delirium. This program uses a modified smaller team compared with traditional multicomponent delirium prevention programs. 51 The team is supported by extensive workflow automation, facilitated by custom tools using data connectivity, real-time monitoring, and automated multistep processes integrated into our EMR. 51 Trained team members certified as reliable assessors 51 use a digitized version of the Confusion Assessment Method (CAM), 52 integrated into our EMR, to enter their patient assessments. These assessments are scored by an EMR-based smart tool that uses automated multistep processes to generate a diagnosis of delirium based on the preprogrammed criteria outlined in the CAM Manual and Training Guide. 53 On a diagnosis of delirium, additional delirium smart tools are activated, and team members initiate treatment of identified patients. 51 All patients aged 60 years and older admitted to non-ICU medical and surgical units who met the inclusion criteria 51 received a CAM assessment by the delirium service prior to deployment of any version of our ML-based delirium risk stratification model. Due to the high volume of assessments and limited resources, patients received only a single CAM assessment during their hospitalization. To address these inefficiencies, the delirium service collaborated with the clinical data science team to develop and deploy an EMR-supported ML model to stratify risk of developing delirium in our hospital. We used a vertical integration approach for the development, training, testing, and deployment training of our ML-based delirium risk stratification model as opposed to a model-centric approach, outlined by Zhang and colleagues. 49 Briefly, work on model design and deployment at all stages was carried out by a cross-disciplinary team that considered whether the existing infrastructure worked synergistically and supported the model during development, testing, and deployment, with ongoing evaluation of clinical impact, user and other stakeholder feedback, and model upgrades responsive to all these elements. 49 Consequently, 3 successive versions of the delirium risk stratification model were developed, tested, and deployed (eFigure 1 in Supplement 1 ). The currently deployed model at MSH presented here represents a fusion of optimized model versions leveraging both EMR patient data features and NLP-processed clinical note features, henceforth referred to as the fusion multimodal with NLP model. Cohorts Data for fusion multimodal with NLP model training, testing, postdeployment validation, and analyses of model deployment associations with workflow and clinical outcomes were derived from visit-level EMR data from patients aged at least 60 years admitted or transferred to MSH non-ICU units from January 2016 to January 2020 and March 2023 to March 2024. We used 4 cohorts for these analyses. The fusion model training/testing cohort included inpatient admissions from January 2016 to January 2020 with at least 1 CAM assessment performed. The fusion model live clinical deployment validation cohort included inpatient admissions assessed by the fusion model from March 1, 2023, to March 31, 2024. The pre-ML cohort included inpatient admissions with at least 1 CAM assessment performed during a 13-month period before model deployment (March 1, 2018, to March 31, 2019) to assess workflow and clinical outcomes before any model version was deployed. Finally, the post-ML cohort included inpatient admissions with at least 1 CAM assessment performed during a 13-month period after model deployment (March 1, 2023, to March 31, 2024) to assess workflow and clinical outcomes during the model’s live clinical deployment. The pre-ML and post-ML cohorts were created to ensure that comparison of workflow and clinical outcomes was conducted on groups derived from equal periods of time. Model Development Data Sources Visit-level data from patients aged 60 years and older admitted or transferred to MSH non-ICU units between January 2016 and January 2020 were sourced from multiple data feeds: the Admission-Discharge-Transfer platform provided demographic and admission data, structured clinical assessments (including laboratory results and vital signs) were extracted from our EMR software (Epic; Epic Systems), electrocardiogram measurements were obtained from MUSE version 9 (GE HealthCare Technologies), and unstructured clinical data (eg, progress notes and care notes) were extracted from the EMR. Training Label Patients were classified as either delirium-positive or delirium-negative based on standardized scoring of the CAM. 53 The time of delirium onset ( t 0 ) was defined as the timestamp of the first positive CAM assessment indicating delirium, or the timestamp of the final negative CAM assessment. Structured and Semistructured EMR Data Processing The model’s input relies on clinical observations. To standardize input data, a sampling module was developed to apply adaptive logic, ensuring a fixed number of observations within predefined intervals. This approach creates consistent, reproducible observation arrays for each type, accommodating diverse clinical scenarios. 54 A time series was constructed by specifying a sampling window and frequency relative to the risk stratification time ( t p ). The sampling window was determined based on variable availability, optimizing data completeness and minimizing missing values (eFigure 2 in Supplement 1 ). The risk stratification time ( t p ) was set to 24 hours prior to the CAM assessment, while the sampling frequency defined the standard intervals between clinical measurements, ensuring consistency across observations. To handle the missing values in the numerical variables, the across-cohort median value was imputed for each variable. For each categorical variable, a specific encoding map was used to handle missing values. For example, the encoding map for gender was: men: (0, 0, 1), women: (0, 1, 0), unknown: (1, 0, 0), and missing value: (0, 0, 0). Patient race and ethnicity were identified in accordance with our EMR’s documentation workflow and recorded by nursing staff at admission. Race and ethnicity were categorized as Asian, Black or African American, Hispanic, White, and other (eg, American Indian or Alaska Native, Native Hawaiian or Pacific Islander, multiple races, and other). For each variable, an array of sampled observations was constructed and subsequently assembled into a feature vector. Clinical Note Preprocessing For each inpatient admission included in the fusion model training and testing cohort, clinical notes, including care notes (submitted by registered nurses) and progress notes (submitted by residents, fellows, attending physicians, nurse practitioners, physician assistants), were aggregated into a text corpus. This corpus encompassed care and progress notes sampled from the 12-hour window preceding the risk stratification time (eFigure 3 in Supplement 1 ). The text corpus was processed using a sentence detection module to segment the text into individual sentences, which were then input into an NLP pipeline for tokenization, stemming, lemmatization, and the creation of 1-g and 2-g bag-of-words models. Term frequency rate was calculated at the encounter level across the cohort. Words with a term frequency rate of at least 0.3 in notes of patients classified as delirium-positive were selected as candidate features. The resulting feature list was categorized into 3 primary categories: diagnoses, signs and symptoms, and medications. Expert clinical feedback was used to refine the feature selection, focusing on those with relevance to the presentation of delirium. These selected features were then assembled into a feature vector. Clinical Note Classifier Development The historical cohort dataset was randomly split into training (70%) and testing (30%) subsets. Due to the significant class imbalance (95% delirium-negative vs 5% delirium-positive), random undersampling was applied to the training set to achieve a balanced distribution of 50% negative and 50% positive. Visits without clinical notes were excluded from the training set. A 10-fold cross-validation was used to train the model using the random forest algorithm from the open-source Apache Spark project ML library, 55 and recursive feature elimination was used for feature selection. After hyperparameter tuning, recursive feature elimination was implemented to reduce the number of features. For feature elimination, we applied an area under the receiver operating characteristic curve (AUROC) score threshold of 2.5% or less to permanently remove features with minimal contribution. Fusion Multimodal With NLP Model Development To develop the fusion multimodal with NLP model, the undersampled training set was used, incorporating feature vectors derived from both structured and semistructured observational data. The NLP risk score was appended to these feature vectors. A 10-fold cross-validation procedure was used to train the model using the random forest algorithm. Following hyperparameter optimization, the recursive feature elimination method was applied to reduce the number of features (eFigure 4 in Supplement 1 ). Fusion Multimodal With NLP Model Deployment in Live Clinical Practice Since the live clinical deployment of the fusion model on February 27, 2023, delirium risk stratification at MSH has been conducted daily for every patient aged at least 60 years admitted or transferred to non-ICU medical and surgical units. The resulting delirium risk value is visualized in our EMR’s patient lists with color coding, with peach indicating high risk (risk ≥0.55) and green, low risk (risk <0.55). Hovering over these risk values opens a popup providing specific features of the model-based risk prediction. MSH delirium service assessors use this risk visualization to prioritize CAM screening of patients at high risk for delirium. Once the CAM assessments are completed and documented in the digitized version of the CAM in our EMR, 51 these patients are suppressed from the prediction model for 5 days. Subsequently, they are once again risk stratified by the model. Model use and performance are tracked prospectively using a real-time dashboard. Statistical Analyses Differences in demographic and clinical characteristics between pre-ML and post-ML cohorts were analyzed using median values (due to significant skewness), with the Kruskal-Wallis test for continuous variables and the χ 2 test for categorical variables. Elixhauser Comorbidity Index was calculated for all patients based on all secondary International Statistical Classification of Diseases, Tenth Revision, Clinical Modification ( ICD-10-CM ) diagnoses. 56 Model Assessment Sensitivity (recall), specificity, F1 score (how good the model is at identifying patients with and without the diagnosis), and the AUROC were calculated with 95% CIs using the scikit-learn library and custom Python scripts 57 to evaluate and compare the discriminatory performance of the different models developed during fine-tuning and during clinical deployment of the fusion model. Workflow and Clinical Outcomes Workflow changes were analyzed by comparing monthly delirium detection rates, calculated as the ratio of CAM-positive delirium screening results to total CAM assessments performed each month. Delirium detection rates were compared using logistic regression adjusted for age, sex, race and ethnicity, surgical or medical primary diagnosis, dementia present at assessment, and Elixhauser Comorbidity Index. Length of stay (LOS) in the hospital (in days) and quantities of medication administration were chosen as the clinical outcomes, as previously reported in association with our delirium service deployment. 51 Opiate and benzodiazepine medications were specifically chosen due to their known delirium-inducing effects 58 and antipsychotic medications were chosen because of the Food and Drug Administration’s black box warning about increased death risk in patients aged 65 years and older with dementia. 59 Opiate medication administration was analyzed by comparing the proportion of patients receiving these medications and then comparing the intravenous morphine dose equivalents administered in milligrams per hospital day. Benzodiazepine medications administration were similarly compared using the diazepam dose equivalents administered in milligrams per hospital day. Similar analyses were conducted for the administration of the antipsychotic medications haloperidol, risperidone, olanzapine, and quetiapine. Given the highly skewed distribution of the selected clinical outcomes, we applied the probabilistic index model to compare the LOS and medication administrations, adjusting for the same covariates. P values were 2-sided, and statistical significance was set at P < .05. Data were analyzed using R software version 4.3.3 (R Project for Statistical Computing). Results The overall sample included 32 284 inpatient admissions (mean [SD] age, 73.56 (9.67) years, 15 157 [46.9%] women). A total of 25 261 inpatient admissions of older patients with both medical and surgical primary diagnoses represented the combined model testing and training cohort (median age, 73.37 [66.42-81.36] years) and live clinical deployment validation cohort (median [IQR] age, 72.11 [62.26-78.97] years), while 7023 inpatient admissions of older patients with both medical and surgical primary diagnoses represented the combined pre-ML (median [IQR] age, 74.00 [68.00-81.00] years) and post-ML (median [IQR] age, 75.33 [68.34-82.91] years) cohorts. A total of 3992 inpatient admissions with at least 1 CAM assessment in the pre-ML cohort and 3031 admissions in the post-ML cohort were used to analyze the workflow and clinical outcomes. Differences of note in the pre-ML vs post-ML cohorts included higher proportions of Asian (260 patients [6.5%] vs 86 patients [2.8%]) and White (1758 patients [44.0%] vs 974 patients [32.1%]) patients; lower proportions of Black or African American patients (755 patients [18.9%] vs 789 patients [26.X%]), Hispanic patients (818 patients [20.6%] vs 827 patients [27.3%]), and patients with other race and ethnicity (342 patients [8.5%] vs 264 patients [8.7%]); a higher proportion of surgical patients (2045 patients [51.2%] vs 1179 patients [38.9%]); a lower Elixhauser Comorbidity Index (median [IQR], 15 [3-27] vs 24 [12-38]); and a lower LOS (median [IQR], 6.86 [4.07-13.13] days vs 13.11 [7.65-22.49] days) ( Table 1 ). The demographic and clinical data for the fusion model training and testing cohort (5646 patients) and fusion model live clinical deployment validation cohort (19 615 patients), used to analyze model performance, did not show such differences and are shown the eTable in Supplement 1 . Table 1. Clinical and Demographic Characteristics of the Cohorts Before and After Deployment of the Multimodal With Natural Language Processing Delirium Prediction Model Characteristic Admissions, No. (%) P value Predeployment a Postdeployment b Admissions, No. 3992 3031 NA Unique patients, No. 3587 2605 NA Admissions with ≥1 CAM assessment, No. 3992 3031 NA Delirium prevalence 184 (4.6) 519 (17.1) <.001 Age, median (IQR), y 74.00 (68.00-81.00) 75.33 (68.34-82.91) <.001 Gender Women 1933 (48.4) 1516 (50.0) <.001 Men 2059 (51.6) 1492 (49.2) Missing 0 23 (0.8) Race and ethnicity Asian 260 (6.5) 86 (2.8) <.001 Black or African American 755 (18.9) 789 (26.0) Hispanic 828 (20.6) 827 (27.3) White 1758 (44.0) 974 (32.1) Other c 342 (8.5) 264 (8.7) Unknown 69 (1.5) 68 (2.2) Missing 0 (0.0) 23 (0.8) Surgical or medical primary diagnosis Medical 1944 (48.4) 1816 (59.9) <.001 Surgical 2045 (51.2) 1179 (38.9) Missing 3 (0.1) 36 (1.2) Dementia present at assessment Yes 449 (11.3) 325 (10.7) <.001 No 3540 (88.6) 2670 (88.1) Missing 3 (0.1) 36 (1.2) Elixhauser Comorbidity Index, median (IQR) 15.00 (3.00-27.00) 24.00 (12.00-38.00) <.001 Length of stay, median (IQR), d 6.78 (3.85-11.66) 13.11 (7.65-22.49) <.001 Abbreviations: CAM, Confusion Assessment Method; NA, not applicable. a March 1, 2018, to March 31, 2019. b March 1, 2023, to March 31, 2024. c Other contains races including American Indian or Alaska Native, Native Hawaiian or Pacific Islander, multiracial, and other. Model Performance Table 2 presents the discriminative performance statistics of the fusion multimodal with NLP ML-based delirium risk stratification model during the training, testing, and live clinical deployment periods. Figure 1 shows the corresponding AUROC curves. During live clinical deployment, the model achieved an AUROC of 0.94 (95% CI, 0.93-0.95), and at an operational risk probability threshold of 0.55, the model achieved a sensitivity of 83% (95% CI, 79%-87%), specificity of 90% (95% CI, 89%-91%), and an F1 score of 0.31 (95% CI, 0.28-0.34). Table 2. Performance of the Fusion Multimodal With Natural Language Processing Delirium Risk Prediction Model During Training, Testing, and Live Clinical Deployment Periods Period Time period Threshold Total admissions, No. Estimate (95% CI Sensitivity Specificity PPV NPV Accuracy F1 Score AUROC Training January 2016 to January 2020 0.55 1008 0.82 (0.78-0.87) 0.83 (0.79-0.88) 0.84 (0.79-0.88) 0.82 (0.77-0.87) 0.83 (0.80-0.86) 0.83 (0.79-0.86) 0.92 (0.89-0.94) Testing January 2016 to January 2021 0.55 4638 0.75 (0.67-0.83) 0.76 (0.74-0.78) 0.14 (0.11-0.16) 0.98 (0.98-0.99) 0.76 (0.74-0.78) 0.23 (0.19-0.27) 0.82 (0.78-0.86) Validation during fusion model live clinical deployment March 1, 2023, to March 31, 2024 0.55 19 615 0.83 (0.79-0.87) 0.90 (0.89-0.91) 0.19 (0.17-0.21) 0.99 (0.99-1.00) 0.90 (0.89-0.90) 0.31 (0.28-0.34) 0.94 (0.93-0.95) Abbreviations: AUROC, area under the curve receiver operating characteristic curve; NPV, negative predictive value; PPV, positive predictive value. Figure 1. Receiver Operating Characteristic Curves for the Fusion Multimodal With Natural Language Processing Model to Predict Delirium AUROC indicates area under the curve receiver operating characteristic curve; post-ML indicates the period after model deployment. Model Features and Their Importance The top 20 variables ranked by Gini importance are summarized for the NLP model and fusion multimodal with NLP model in eFigure 5 in Supplement 1 . The accuracy of the patients’ orientation to person, time, and place was identified as the strongest variable in the fusion model, closely followed by the composite NLP prediction score. Workflow Change Monthly delirium detection rates significantly increased during the deployment period. Median (IQR) rates increased from 4.42% (3.70%-5.14%) to 17.17% (15.54%-18.80%) ( P < .001) ( Figure 2 ). Figure 2. Comparison of Delirium Detection Rates Box-and-whisker plot comparing 5 summary statistics for the monthly delirium detection rates before any machine learning (ML) model deployment and following deployment of the multimodal with natural language processing ML-based delirium risk prediction model in live clinical practice. Delirium detection rates were calculated by dividing the number of positive Confusion Assessment Method (CAM) delirium screening results by the number of total CAM assessments each month. Whiskers indicate range; boxes, IQR; bold line, median. Clinical Outcomes LOS was significantly higher in the post-ML cohort vs pre-ML cohort (median [IQR] LOS, 13.11 [7.65-22.49] days vs 6.78 [3.85-11.66] days). Compared with the pre-ML cohort, higher proportions of patients in the post-ML cohort received opiates (1352 patients [44.6%] vs 852 patients [21.3%]; P < .001) and olanzapine (31 patients [0.8%] vs 162 patients [5.3%]; P < .001). However, patients receiving these medications in the post-ML vs pre-ML cohorts received significantly lower daily doses of benzodiazepines (median [IQR] dosage, 0.93 [0.42-2.28] diazepam dose equivalents vs 1.60 [0.66-4.27] diazepam dose equivalents; P < .001) and olanzapine (median [IQR] dosage, 1.09 [0.38-2.46] mg vs 2.50 [1.16-6.65] mg; P < .001) ( Table 3 ). Table 3. Opiate, Benzodiazepine, and Antipsychotic Medication Administration Before Any Version of the ML-Delirium Risk Stratification Model and Following Clinical Deployment of the Fusion ML-Delirium Risk Stratification Model Medication Proportion receiving medication, No. (%) a Dose administered per hospital d, median (IQR) P value Opiates, IV morphine dose equivalents Pre-ML model deployment 2462 (61.7) 4.80 (1.49-15.54) .12 Post-ML model deployment 2003 (66.1) 3.87 (1.00-12.87) Benzodiazepine, diazepam dose equivalents Pre-ML model deployment 852 (21.3) 1.60 (0.66-4.27) <.001 Post-ML model deployment 1352 (44.6) 0.93 (0.42-2.28) Haloperidol, mg Pre-ML model deployment 351 (8.8) 0.27 (0.10-0.64) .12 Post-ML model deployment 498 (16.4) 0.23 (0.09-0.60) Quetiapine, mg Pre-ML model deployment 214 (5.4) 9.64 (3.70-25.49) .13 Post-ML model deployment 370 (12.2) 8.44 (2.65-25.45) Olanzapine, mg Pre-ML model deployment 31 (0.8) 2.50 (1.17-6.65) <.001 Post-ML model deployment 162 (5.3) 1.09 (0.38-2.46) Risperidone, mg Pre-ML model deployment 30 (0.8) 0.53 (0.23-1.83) .75 Post-ML model deployment 43 (1.4) 0.48 (0.27-2.00) Abbreviations: IV, intravenous; ML, machine learning. a Includes 4068 admissions in the pre-ML period and 3031 admissions in the post-ML period. Discussion In this quality improvement study, we present a fusion multimodal with NLP ML-based delirium risk stratification model developed using a vertical integration approach and demonstrated acceptable discriminative performance statistics over a 13-month period in live clinical practice. Model deployment was associated with a significant 4-fold increase in delirium detection rates. This shift in workflow allowed for focused recurrent screening of a patient population with higher risk of delirium, optimizing resource allocation. While at least 1 other ML-delirium risk stratification model with NLP tested in live clinical practice has been published, 25 it lacked reporting on associated workflow or clinical outcomes. Similarly, 3 reports of an ML-based delirium risk stratification model without NLP tested in live clinical practice also did not report on such outcomes. 9 , 22 , 48 We believe the observed increased hospital LOS in the post-ML cohort is attributable to a higher prevalence of delirium and a higher Elixhauser Comorbidity Index, both known factors associated with increased hospital LOS. 1 , 60 Additionally, while a higher proportion of patients in the post-ML cohort received benzodiazepines, haloperidol, and olanzapine, they received significantly lower daily doses of these medications. Our service has focused on disseminating delirium prevention and treatment best practices, including psychotropic and opiate medication reduction. 51 Given a higher Elixhauser Comorbidity Index and delirium prevalence in the post-ML cohort, it is possible that an increased prevalence of associated pain, anxiety, and agitation in this refined cohort necessitated medication use in more patients, while concomitantly, efforts were made to reduce dosing. However, this remains speculative and cannot be definitively tested with our current data. A major strength of our model development process is the use of the highly sensitive and specific CAM delirium assessment tool 52 by raters with proven reliability 51 as the reference standard for the delirium diagnosis. This type of criterion standard delirium diagnostic method, which is so infrequently used in published ML-based delirium prediction models, 5 , 8 , 10 , 17 , 19 , 20 , 24 , 27 , 28 , 37 , 38 , 39 , 40 , 43 , 44 , 45 , 47 , 48 distinguishes our approach from most published reports that rely solely on EMR review and ICD-10-CM coding. Another strength of the model presented here is its development and validation in live clinical practice, leveraging a vertical integration pipeline. 49 The iterative model upgrade process facilitated the resolution of practical challenges associated with model deployment. Furthermore, our model was developed using a diverse cohort of medical and surgical patients, in contrast to most published ML-based delirium predictive models that were developed and validated on more narrowly defined cohorts of medical 6 , 18 , 20 , 23 , 24 , 32 , 36 , 46 , 47 or surgical 3 , 8 , 10 , 14 , 15 , 16 , 17 , 22 , 26 , 27 , 28 , 29 , 30 , 34 , 38 , 39 , 40 , 41 , 42 , 43 , 44 , 45 , 48 patients. This enhances the potential generalizability of our model. We will continue to advance our model through rigorous evaluation and clinical application. Our immediate plans involve a formal analysis of model performance at Mount Sinai Morningside Hospital in New York, New York. To ensure optimal utilization and impact, we will implement a real-time monitoring system at Mount Sinai Morningside Hospital. Additionally, we are actively exploring opportunities to expand deployment to affiliated hospitals, particularly those without dedicated delirium services. By broadening the reach of our model, we aim to improve patient care and streamline clinical workflows in diverse health care settings. Limitations This study has some limitations. Any enthusiasm for this model’s generalizability should be tempered by the absence of external validation at other hospitals. Another limitation of this model’s generalizability is imposed by the fact that our model was developed to work in synergy with our unique form of multicomponent-based treatment of extant delirium, which diverges from the more common application of multicomponent-based delirium programs that focus on delirium prevention. 61 Given that many hospital systems lack dedicated delirium programs, the model’s adoption and effective implementation would depend on the willingness of health care practitioners to adopt this delirium risk stratification tool and conduct proper patient assessments. While this does not preclude the model’s utility in such settings, further study is necessary to confirm its applicability. Conclusions This quality improvement study developed and demonstrated the feasibility and clinical utility of a novel multimodal ML model to automate delirium risk stratification in live clinical practice. This model demonstrated acceptable performance in live clinical practice and may facilitate resource allocation to enhance delirium identification and care. References 1 Siddiqi N , House AO , Holmes JD . Occurrence and outcome of delirium in medical in-patients: a systematic literature review . Age Ageing . 2006 ; 35 ( 4 ): 350 - 364 . doi: 10.1093/ageing/afl005 16648149 2 Leslie DL , Zhang Y , Holford TR , Bogardus ST , Leo-Summers LS , Inouye SK . Premature death associated with delirium at 1-year follow-up . Arch Intern Med . 2005 ; 165 ( 14 ): 1657 - 1662 . doi: 10.1001/archinte.165.14.1657 16043686 3 Davoudi A , Ebadi A , Rashidi P , Ozrazgat-Baslanti T , Bihorac A , Bursian AC . Delirium prediction using machine learning models on preoperative electronic health records data . Proc IEEE Int Symp Bioinformatics Bioeng . 2017 ; 2017 : 568 - 573 . 30393788 10.1109/BIBE.2017.00014 PMC6211171 4 Veeranki SPK , Hayn D , Kramer D , Jauk S , Schreier G . Effect of nursing assessment on predictive delirium models in hospitalised patients . Stud Health Technol Inform . 2018 ; 248 : 124 - 131 . doi: 10.3233/978-1-61499-858-7-124 29726428 5 Corradi JP , Thompson S , Mather JF , Waszynski CM , Dicks RS . Prediction of incident delirium using a random forest classifier . J Med Syst . 2018 ; 42 ( 12 ): 261 . doi: 10.1007/s10916-018-1109-0 30430256 6 Halladay CW , Sillner AY , Rudolph JL . Performance of electronic prediction rules for prevalent delirium at hospital admission . JAMA Netw Open . 2018 ; 1 ( 4 ): e181405 . doi: 10.1001/jamanetworkopen.2018.1405 30646122 PMC6324279 7 Veeranki SPK , Hayn D , Jauk S , . An improvised classification model for predicting delirium . Stud Health Technol Inform . 2019 ; 264 : 1566 - 1567 . 31438234 10.3233/SHTI190537 8 Mufti HN , Hirsch GM , Abidi SR , Abidi SSR . Exploiting machine learning algorithms and methods for the prediction of agitated delirium after cardiac surgery: models development and validation study . JMIR Med Inform . 2019 ; 7 ( 4 ): e14993 . doi: 10.2196/14993 31558433 PMC6913743 9 Jauk S , Kramer D , Großauer B , . Risk prediction of delirium in hospitalized patients using machine learning: An implementation and prospective evaluation study . J Am Med Inform Assoc . 2020 ; 27 ( 9 ): 1383 - 1392 . doi: 10.1093/jamia/ocaa113 32968811 PMC7647341 10 Wang Y , Lei L , Ji M , Tong J , Zhou CM , Yang JJ . Predicting postoperative delirium after microvascular decompression surgery with machine learning . J Clin Anesth . 2020 ; 66 : 109896 . doi: 10.1016/j.jclinane.2020.109896 32504969 11 Sun H , Depraetere K , Meesseman L , . A scalable approach for developing clinical risk prediction applications in different hospitals . J Biomed Inform . 2021 ; 118 : 103783 . doi: 10.1016/j.jbi.2021.103783 33887456 12 Wong A , Otles E , Donnelly JP , . External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients . JAMA Intern Med . 2021 ; 181 ( 8 ): 1065 - 1070 . doi: 10.1001/jamainternmed.2021.2626 34152373 PMC8218233 13 Zhao Y , Luo Y . Unsupervised learning to Subphenotype delirium patients from electronic health records. IEEE International Conference on Bioinformatics and Biomedicine. arXiv . Preprint posted online October 31, 2021. doi: 10.48550/arXiv.2111.00592 14 Oosterhoff JHF , Karhade AV , Oberai T , Franco-Garcia E , Doornberg JN , Schwab JH . Prediction of postoperative delirium in geriatric hip fracture patients: a clinical prediction model using machine learning algorithms . Geriatr Orthop Surg Rehabil . Published online December 13, 2021 . doi: 10.1177/21514593211062277 34925951 PMC8671660 15 Racine AM , Tommet D , D’Aquila ML , ; RISE Study Group . Machine learning to develop and internally validate a predictive model for post-operative delirium in a prospective, observational clinical cohort study of older surgical patients . J Gen Intern Med . 2021 ; 36 ( 2 ): 265 - 273 . doi: 10.1007/s11606-020-06238-7 33078300 PMC7878663 16 Xue B , Li D , Lu C , . Use of machine learning to develop and evaluate models using preoperative and intraoperative data to identify risks of postoperative complications . JAMA Netw Open . 2021 ; 4 ( 3 ): e212240 . doi: 10.1001/jamanetworkopen.2021.2240 33783520 PMC8010590 17 Zhao H , You J , Peng Y , Feng Y . Machine learning algorithm using electronic chart-derived data to predict delirium after elderly hip fracture surgeries: a retrospective case-control study . Front Surg . 2021 ; 8 : 634629 . doi: 10.3389/fsurg.2021.634629 34327210 PMC8313764 18 Castro VM , Sacks CA , Perlis RH , McCoy TH . Development and external validation of a delirium prediction model for hospitalized patients with coronavirus disease 2019 . J Acad Consult Liaison Psychiatry . 2021 ; 62 ( 3 ): 298 - 308 . doi: 10.1016/j.jaclp.2020.12.005 33688635 PMC7933786 19 Son CS , Kang WS , Lee JH , Moon KJ . Machine learning to identify psychomotor behaviors of delirium for patients in long-term care facility . IEEE J Biomed Health Inform . 2022 ; 26 ( 4 ): 1802 - 1814 . doi: 10.1109/JBHI.2021.3116967 34596563 20 Cano-Escalera G , Graña M , Irazusta J , Labayen I , Besga A . Risk factors for prediction of delirium at hospital admittance . Expert Syst . 2022 ; 39 ( 4 ): 1 - 10 . doi: 10.1111/exsy.12698 21 Gutheil J , Donsa K . SAINTENS: self-attention and intersample attention transformer for digital biomarker development using tabular healthcare real world data . Stud Health Technol Inform . 2022 ; 293 ( 293 ): 212 - 220 . doi: 10.3233/SHTI220371 35592984 22 Jauk S , Veeranki SPK , Kramer D , . External validation of a machine learning based delirium prediction software in clinical routine . Stud Health Technol Inform . 2022 ; 293 : 93 - 100 . doi: 10.3233/SHTI220353 35592966 23 Kurisu K , Inada S , Maeda I , ; Phase-R Delirium Study Group . A decision tree prediction model for a short-term outcome of delirium in patients with advanced cancer receiving pharmacological interventions: a secondary analysis of a multicenter and prospective observational study (Phase-R) . Palliat Support Care . 2022 ; 20 ( 2 ): 153 - 158 . doi: 10.1017/S1478951521001565 35574912 24 Li Q , Zhao Y , Chen Y , Yue J , Xiong Y . Developing a machine learning model to identify delirium risk in geriatric internal medicine inpatients . Eur Geriatr Med . 2022 ; 13 ( 1 ): 173 - 183 . doi: 10.1007/s41999-021-00562-9 34553310 25 Sun H , Depraetere K , Meesseman L , . Machine learning-based prediction models for different clinical risks in different hospitals: evaluation of live performance . J Med Internet Res . 2022 ; 24 ( 6 ): e34295 . doi: 10.2196/34295 35502887 PMC9214618 26 Bishara A , Chiu C , Whitlock EL , . Postoperative delirium prediction using machine learning models and preoperative electronic health record data . BMC Anesthesiol . 2022 ; 22 ( 1 ): 8 . doi: 10.1186/s12871-021-01543-y 34979919 PMC8722098 27 Hu XY , Liu H , Zhao X , . Automated machine learning-based model predicts postoperative delirium using readily extractable perioperative collected electronic data . CNS Neurosci Ther . 2022 ; 28 ( 4 ): 608 - 618 . doi: 10.1111/cns.13758 34792857 PMC8928919 28 Menzenbach J , Kirfel A , Guttenthaler V , ; PROPDESC Collaboration Group . Pre-Operative Prediction of Postoperative Delirium by Appropriate Screening (PROPDESC) development and validation of a pragmatic POD risk screening score based on routine preoperative data . J Clin Anesth . 2022 ; 78 : 110684 . doi: 10.1016/j.jclinane.2022.110684 35190344 29 Oosterhoff JHF , Oberai T , Karhade AV , . Does the SORG Orthopaedic Research Group Hip Fracture Delirium Algorithm perform well on an independent intercontinental cohort of patients with hip fractures who are 60 years or older? Clin Orthop Relat Res . 2022 ; 480 ( 11 ): 2205 - 2213 . doi: 10.1097/CORR.0000000000002246 35561268 PMC10476833 30 Xue X , Chen W , Chen X . A novel radiomics-based machine learning framework for prediction of acute kidney injury-related delirium in patients who underwent cardiovascular surgery . Comput Math Methods Med . 2022 ; 2022 : 4242069 . doi: 10.1155/2022/4242069 35341014 PMC8956431 31 Castro VM , Hart KL , Sacks CA , Murphy SN , Perlis RH , McCoy TH Jr . Longitudinal validation of an electronic health record delirium prediction model applied at admission in COVID-19 patients . Gen Hosp Psychiatry . 2022 ; 74 : 9 - 17 . doi: 10.1016/j.genhosppsych.2021.10.005 34798580 PMC8562039 32 Wang L , Zhang Y , Chignell M , . Boosting delirium identification accuracy with sentiment-based natural language processing: mixed methods study . JMIR Med Inform . 2022 ; 10 ( 12 ): e38161 . doi: 10.2196/38161 36538363 PMC9812273 33 Fu S , Lopes GS , Pagali SR , . Ascertainment of delirium status using natural language processing from electronic health records . J Gerontol A Biol Sci Med Sci . 2022 ; 77 ( 3 ): 524 - 530 . doi: 10.1093/gerona/glaa275 35239951 PMC8893184 34 Choi JY , Yoo S , Song W , . Development and validation of a prognostic classification model predicting postoperative adverse outcomes in older surgical patients using a machine learning algorithm: retrospective observational network study . J Med Internet Res . 2023 ; 25 : e42259 . doi: 10.2196/42259 37955965 PMC10682929 35 St Sauver J , Fu S , Sohn S , . Identification of delirium from real-world electronic health record clinical notes . J Clin Transl Sci . 2023 ; 7 ( 1 ): e187 . doi: 10.1017/cts.2023.610 37745932 PMC10514685 36 Pagali SR , Kumar R , Fu S , Sohn S , Yousufuddin M . Natural language processing CAM algorithm improves delirium detection compared with conventional methods . Am J Med Qual . 2023 ; 38 ( 1 ): 17 - 22 . doi: 10.1097/JMQ.0000000000000090 36283056 37 Matsumoto K , Nohara Y , Sakaguchi M , . Development of machine learning prediction models for self-extubation after delirium using emergency department data . Stud Health Technol Inform . 2024 ; 310 : 1001 - 1005 . doi: 10.3233/SHTI231115 38269965 38 Zhao X , Li J , Xie X , . Online interpretable dynamic prediction models for postoperative delirium after cardiac surgery under cardiopulmonary bypass developed based on machine learning algorithms: a retrospective cohort study . J Psychosom Res . 2024 ; 176 : 111553 . doi: 10.1016/j.jpsychores.2023.111553 37995429 39 Li Q , Li J , Chen J , . A machine learning-based prediction model for postoperative delirium in cardiac valve surgery using electronic health records . BMC Cardiovasc Disord . 2024 ; 24 ( 1 ): 56 . doi: 10.1186/s12872-024-03723-3 38238677 PMC10795338 40 Rössler J , Shah K , Medellin S , . Development and validation of delirium prediction models for noncardiac surgery patients . J Clin Anesth . 2024 ; 93 : 111319 . doi: 10.1016/j.jclinane.2023.111319 37984177 41 Yang T , Yang H , Liu Y , . Postoperative delirium prediction after cardiac surgery using machine learning models . Comput Biol Med . 2024 ; 169 : 107818 . doi: 10.1016/j.compbiomed.2023.107818 38134752 42 Song Y , Zhang D , Wang Q , . Prediction models for postoperative delirium in elderly patients with machine-learning algorithms and Shapley Additive Explanations . Transl Psychiatry . 2024 ; 14 ( 1 ): 57 . doi: 10.1038/s41398-024-02762-w 38267405 PMC10808214 43 Sadlonova M , Hansen N , Esselmann H , ; FINDERI investigators . Preoperative delirium risk screening in patients undergoing a cardiac surgery: results from the prospective observational FINDERI study . Am J Geriatr Psychiatry . 2024 ; 32 ( 7 ): 835 - 851 . doi: 10.1016/j.jagp.2023.12.017 38228452 44 Sheng W , Tang X , Hu X , . Random forest algorithm for predicting postoperative delirium in older patients . Front Neurol . 2024 ; 14 : 1325941 . doi: 10.3389/fneur.2023.1325941 38274882 PMC10808713 45 Matsumoto K , Nohara Y , Sakaguchi M , . Temporal generalizability of machine learning models for predicting postoperative delirium using electronic health record data: model development and validation study . JMIR Perioper Med . 2023 ; 6 : e50895 . doi: 10.2196/50895 37883164 PMC10636625 46 Miyazawa Y , Katsuta N , Nara T , . Identification of risk factors for the onset of delirium associated with COVID-19 by mining nursing records . PLoS One . 2024 ; 19 ( 1 ): e0296760 . doi: 10.1371/journal.pone.0296760 38241284 PMC10798448 47 Lee SH , Hur HJ , Kim SN , . Predicting delirium and the effects of medications in hospitalized COVID-19 patients using machine learning: a retrospective study within the Korean Multidisciplinary Cohort for Delirium Prevention (KoMCoDe) . Digit Health . Published online January 5, 2024 . doi: 10.1177/20552076231223811 38188862 PMC10771056 48 Jauk S , Kramer D , Sumerauer S , Veeranki SPK , Schrempf M , Puchwein P . Machine learning-based delirium prediction in surgical in-patients: a prospective validation study . JAMIA Open . 2024 ; 7 ( 3 ): ooae091 . doi: 10.1093/jamiaopen/ooae091 39297150 PMC11408728 49 Zhang J , Budhdeo S , William W , . Moving towards vertically integrated artificial intelligence development . NPJ Digit Med . 2022 ; 5 ( 1 ): 143 . doi: 10.1038/s41746-022-00690-x 36104535 PMC9474277 50 Kim KN . Improving causal inference in observational studies: interrupted time series design . Cardiovasc Prev Pharmacother. 2020 ; 2 ( 1 ): 18 - 23 . doi: 10.36011/cpp.2020.2.e2 51 Friedman JI , Li L , Kirpalani S , . A multi-phase quality improvement initiative for the treatment of active delirium in older persons . J Am Geriatr Soc . 2021 ; 69 ( 1 ): 216 - 224 . doi: 10.1111/jgs.16897 33150615 52 Inouye SK , van Dyck CH , Alessi CA , Balkin S , Siegal AP , Horwitz RI . Clarifying confusion: the confusion assessment method: a new method for detection of delirium . Ann Intern Med . 1990 ; 113 ( 12 ): 941 - 948 . doi: 10.7326/0003-4819-113-12-941 2240918 53 Hospice Elder Life. Hospice Elder Life Program. Accessed April 1, 2025. http://www.hospitalelderlifeprogram.org 54 Kia A , Timsina P , Joshi HN , . MEWS++: enhancing the prediction of clinical deterioration in admitted patients through a machine learning model . J Clin Med . 2020 ; 9 ( 2 ): 343 . doi: 10.3390/jcm9020343 32012659 PMC7073544 55 Apache Spark 2.3.0. Machine Learning Library (MLlib) Guide. Accessed May 3, 2024. https://spark.apache.org/docs/2.3.0/ml-guide.html 56 Agency for Healthcare Research and Quality . Elixhauser Comorbidity software refined for ICD-10-CM . Accessed April 1, 2025. https://hcup-us.ahrq.gov/toolssoftware/comorbidityicd10/comorbidity_icd10.jsp 57 Pedregosa F , Varoquaux G , Gramfort A , . Scikit-learn: machine learning in Python . JMLR . 2011 ; 12 ( 85 ): 2825 - 2830 . 58 Gaudreau JD , Gagnon P , Roy MA , Harel F , Tremblay A . Association between psychoactive medications and delirium in hospitalized patients: a critical review . Psychosomatics . 2005 ; 46 ( 4 ): 302 - 316 . doi: 10.1176/appi.psy.46.4.302 16000673 59 Yan J . FDA extends black-box warning to all antipsychotics . Psychiatr News . 2008 ; 43 ( 14 ): 1 - 27 . doi: 10.1176/pn.43.14.0001 60 Elixhauser A , Steiner C , Harris DR , Coffey RM . Comorbidity measures for use with administrative data . Med Care . 1998 ; 36 ( 1 ): 8 - 27 . doi: 10.1097/00005650-199801000-00004 9431328 61 Hshieh TT , Yang T , Gartaganis SL , Yue J , Inouye SK . Hospital Elder Life Program: systematic review and meta-analysis of effectiveness . Am J Geriatr Psychiatry . 2018 ; 26 ( 10 ): 1015 - 1033 . doi: 10.1016/j.jagp.2018.06.007 30076080 PMC6362826 Supplement 1. eFigure 1. Timeline for the MSH Quality Improvement Initiative eFigure 2. Sampling Strategy for EMR Features eFigure 3. Sampling Strategy for Clinical Notes: eFigure 4. Model Fusion Architecture of the Multimodal (+NLP) Application eFigure 5. Fusion Multimodal and Natural Language Processing Model Variables Ranked by Gini Importance eTable. Clinical And Demographic Characteristics of Fusion Model Test/Train and Fusion Model Live Clinical Deployment Validation Cohorts Supplement 2. Data Sharing Statement