Corresponding author.
Diabetic retinopathy (DR) is a leading cause of preventable blindness, and the growing global burden of diabetes is placing increasing pressure on ophthalmic services. Artificial intelligence (AI)–based retinal image analysis offers a promising strategy to scale up DR screening while reducing reliance on specialist graders. We assessed the performance of an AI-based DR screening system implemented in a real-world endocrinology clinic at the Erasmus Hospital, Belgium. Adult patients with diabetes underwent non-mydriatic fundus photography, and images were analyzed by the AI system for referable DR and diabetic macular edema. All images were independently graded by a retinal specialist using the Early Treatment Diabetic Retinopathy Study (ETDRS) classification as the reference standard. Of 405 patients screened, 353 (86.7%) were included in the primary analysis. The AI system achieved an area under the curve of 96.5%, sensitivity of 88.9%, specificity of 98.7%, and high predictive values for referable DR detection. Subgroup analyses showed consistently high accuracy across demographic and clinical strata. Multivariate analysis identified higher HbA1c at diagnosis and longer diabetes duration as significant predictors of referable DR for both AI and human grading. These findings support the robustness, generalizability and operational feasibility of this AI system for DR screening in routine clinical care.
Received 2025 Nov 16; Accepted 2026 Jan 21; Collection date 2026.
Diabetes mellitus is characterized by chronically elevated blood glucose levels that progressively damage the inner lining of blood vessels and lead to long-term microvascular complications. Diabetic retinopathy (DR) is the most common microvascular complication of diabetes and a leading cause of vision loss and preventable blindness in working-age adults worldwide. The global prevalence of diabetes is rising at an alarming rate: 589 million adults were living with diabetes in 2024, with projections reaching 853 million by 2050
Early detection through regular screening is essential to prevent vision loss, but the growing demand for screening challenges an already under-resourced ophthalmic workforce
Traditional DR screening is based on dilated fundus examination by ophthalmologists, which is labor-intensive, resource-demanding and costly for health systems
Several AI systems have since been developed for automated, real-time DR detection and have generally shown accuracy comparable to or exceeding that of human graders
The primary objective of this study was to evaluate the accuracy of a recently implemented AI-based screening system for referable DR in a Belgian population with diverse ethnic backgrounds at Erasmus Hospital, Brussels University Hospital. Secondary objectives were to assess the generalizability of the system across subgroups and to investigate its ability to identify systemic risk factors associated with referable DR, hypothesizing that AI performance would be at least equivalent to that of human graders.
The study was approved by the Erasmus Hospital HUB Ethics Committee (reference P2023/464). All participants provided informed consent, and all procedures were conducted in accordance with the Declaration of Helsinki and relevant institutional guidelines and regulations.
All adult patients (≥ 18 years) with a diagnosis of diabetes attending consultations at the Erasmus Medical Center were invited to participate between 10 January and 20 April 2024. Recruitment took place immediately following their usual diabetes consultation.
The AI-based screening model was developed by MONA.health, a Belgian healthcare AI company supported by KU Leuven and the Health Unit of VITO. The system autonomously screens for referable DR and diabetic macular edema and its development and validation details have been reported previously
For each participant, a single 45-degree, non-dilated, color fundus photograph of each eye was acquired using a Crystalvue camera, centered between the optic disc and macula. The MONA.health software independently performs an automated image quality assessment, and requires at least one image of sufficient quality per eye to generate a report on the patient’s referral status for DR and diabetic macular edema. If the maximum predicted value for at least one eye exceeded the fixed thresholds (1.371 for DR and 0.38 for diabetic macular edema, corresponding to an operating point targeting 90% sensitivity), the patient was flagged for referral for one or both conditions. Referable DR was defined as moderate non-proliferative DR or worse, diabetic macular edema, or an ungradable image, based on the Early Treatment Diabetic Retinopathy Study (ETDRS) classification
Prior to patient recruitment, nurses and a medical student (LB) received standardized one-to-one training by the same trainer, covering retinal image acquisition, troubleshooting for poor image quality, automated grading using the AI system and communication of results to patients. Training and supervision were provided until staff were comfortable performing data collection and imaging independently.
Following image acquisition, retinal images were uploaded to the AI system, which generated an automated referral status report. Results were orally communicated to patients and incorporated into their medical records at the time of screening. Patients who tested positive for referable DR and/or diabetic macular edema and were previously undiagnosed, or who had insufficient image quality, were invited to consult an ophthalmologist. The Ophthalmology Department at Erasmus Hospital contacted these patients for follow-up if needed.
To determine the accuracy of the AI screening model, all retinal images were manually graded for referable DR by a retinal specialist (EM) at Erasmus Hospital using the ETDRS classification. As the reference standard, referable DR was defined as moderate non-proliferative DR or vision-threatening stages (severe non-proliferative and proliferative DR) in at least one eye. When grading differed between eyes, the presence of referable DR in at least one eye was considered sufficient to classify the patient as having referable DR, in accordance with ETDRS criteria. The AI system applies the same clinical definition of referable DR but provides a binary referable/non-referable output without further severity stratification. Human grading was performed exclusively for reference standard evaluation and was not part of the AI screening workflow. Patients identified as false negatives (no DR reported by AI in either eye but DR detected on manual grading) were contacted by telephone and advised to attend an ophthalmologic consultation.
All analyses were performed using RStudio (version 4.3.3). Patients with missing reference standard grading due to missing images in at least one eye (small pupil, poor fixation, technical error, upload error) or ungradable images according to either the AI system or manual grading were excluded from the primary analysis.
Sensitivity, specificity, positive predictive value and negative predictive value for referable DR detection were estimated with 95% exact binomial confidence intervals. AUC and corresponding confidence intervals were estimated using a bootstrap technique with 1,000 replicates. All diagnostic accuracy analyses were performed at the patient level. The highest predicted probability across both eyes was used for classification and AUC estimation, as referral is triggered when referable DR is detected in at least one eye.
The same diagnostic accuracy analysis was performed in predefined subgroups based on age, sex, ethnicity, diabetes type, HbA1c at diagnosis, diabetes duration, BMI and image quality. Image quality was categorized as excellent, good, adequate or “dirty lens”; for subgroup analysis, “excellent” corresponded to excellent quality in both eyes, “good” to good quality in at least one eye, and “adequate/dirty lens” to adequate or dirty lens quality in both eyes (Fig.
Examples of images of dirty lens and insufficient, adequate, good, excellent quality. Images collected during the present study using the MONA color fundus machine.
A sensitivity analysis including all patients was undertaken to explore the potential impact of excluding individuals with missing reference standard diagnosis. In the worst-case scenario, all missing/unreadable images were treated as false negatives for referable DR; in the best-case scenario, all missing/unreadable images with referable features were treated as true positives and the remaining as true negatives.
Two multivariate logistic regression models were constructed to assess the association between systemic risk factors and referable DR, using AI-based and human grader-based classifications, respectively. Candidate variables were selected using a bootstrap procedure (n = 500), retaining only those that were significantly associated with referable DR in at least 50% of resamples. Model adequacy followed the rule of at least 10 events per explanatory variable
Parts of the text (grammar and style) were revised with the assistance of a large language model (ChatGPT, OpenAI). All scientific content, data interpretation and conclusions were determined and verified by the authors, who retain full responsibility for the work.
A total of 405 patients with diabetes were invited to participate in the study. Of these, 353 (87.2%) were included in the primary analysis, while 52 (12.8%) were excluded due to missing or ungradable images. Reasons for exclusion included insufficient image quality (n = 8; 15.4%), missing images related to small pupils (n = 14; 26.9%), poor fixation (n = 1; 1.9%), technical issues such as incorrect positioning or duplicate images (n = 8; 15.4%), and images considered ungradable by the human grader (n = 21; 40.4%).
Among the 353 included patients, 192 (54.4%) were male. One hundred sixty-eight (47.6%) had type 2 diabetes, 113 (32.0%) had type 1 diabetes, and 42 (11.9%) had diabetes of unknown etiology. In the excluded group, 29 (55.8%) were male; 34 (65.4%) had type 2 diabetes, 7 (13.5%) had type 1 diabetes, and 9 (17.3%) had diabetes of unknown etiology. The median age of included patients was 56 years (range 19–89), and the median duration of diabetes was 13 years (0–56). Excluded patients were significantly older, with a median age of 68 years (25–88), and had a longer duration of diabetes (median 15 years; 1–45).
A majority of included and excluded patients, 213 (60.4%) and 28 (53.8%) respectively, reported that more than 12 months had elapsed since their last eye examination. Patient characteristics are summarized in Table
Characteristics of patients included and excluded from the primary analysis.
| Characteristic | Included | Excluded | Total | p value |
|---|---|---|---|---|
| n = 353 | n = 52 | n = 405 | ||
| 1.1 × 10⁻1⁰ | ||||
| Median (years, range) | 56 (19–89) | 68 (25–88) | - | |
| Mean (years) | 54.80 | 69.15 | - | |
| 0.97 | ||||
| Male | 192 (54.4%) | 29 (55.8%) | 221 (54.6%) | |
| Female | 161 (45.6%) | 23 (44.2%) | 184 (45.4%) | |
| 0.97 | ||||
| Caucasian North African | 139 (39.4%) | 20 (38.5%) | 159 (39.2%) | |
| Caucasian European | 188 (53.2%) | 29 (55.8%) | 217 (53.6%) | |
| Other | 26 (7.4%) | 3 (5.8%) | 29 (7.2%) | |
| 0.01 | ||||
| Type 1 | 113 (32.0%) | 7 (13.5%) | 120 (29.6%) | |
| Type 2 | 168 (47.6%) | 34 (65.4%) | 202 (49.9%) | |
| Other | 30 (8.5%) | 2 (3.8%) | 32 (7.9%) | |
| Unknown | 42 (11.9%) | 9 (17.3%) | 51 (12.6%) | |
| 0.83 | ||||
| ≤ 12 months | 140 (39.7%) | 24 (46.1%) | 164 (40.5%) | |
| 12–24 months | 73 (20.7%) | 10 (19.2%) | 83 (20.5%) | |
| ≥ 24 months | 140 (39.7%) | 18 (34.6%) | 158 (39.0%) | |
| 0.02 | ||||
| Median (years, range) | 13 (0–56) | 15 (1–45) | - | |
| Mean (years) | 15 | 19 | - |
Other type of diabetes includes: Monogenic diabetes, diabetes secondary to chronic pancreatitis, corticotherapy-related diabetes, post-transplantation diabetes. Other ethnicity includes: Sub-Saharan African and Asian.
Referable DR was identified by the human grader in 54 patients (15.3%) and by the AI system in 59 patients (15.8%). A total of 31 discrepancies between AI and human grader diagnoses were observed: 6 false negatives, 4 false positives, and 21 cases that were ungradable according to human grader because of image quality issues. Among the ungradable cases, 13 were classified as non-referable and 8 as referable by the AI system (Table
Cross tabulation results between AI system and human grading.
| Primary analysis (n = 353) | Sensitivity analysis scenario (n = 405) | ||
|---|---|---|---|
| Complete-case set (95% CI) | Best-case (95% CI) | Worst-case (95% CI) | |
| Prevalence | 15.3% | – | - |
| Sensitivity (%) | 88.9 [77.4–95.8] | 90.2 [79.8–96.3] | 48.5 [38.3–58.7] |
| Specificity (%) | 98.7 [96.7–99.7] | 98.8 [97.0–99.7] | 96.4 [93.6–98.2] |
| AUC (%) | 96.5 [91.8–99.7] | – | - |
| Predictive value | |||
| Positive (%) | 92.3 [81.4–97.9] | 93.2 [83.5–98.1] | 81.3 [69.1–90.3] |
| Negative (%) | 98.0 [95.7–99.3] | 98.3 [96.3–99.4] | 85.3 [81.1–88.8] |
A flow chart of the screening procedure and patient follow-up is shown in Fig.
Flow chart of screening procedures and follow-up of patients. DR = diabetic retinopathy. NR = non referable. RDR = referable diabetic retinopathy.
Examples of grading reports generated by AI-based screening system, including DR and diabetic maculopathy referral status. Details of the development and validation of this AI system have been previously described
In the primary analysis, the prevalence of referable DR was 15.3% based on the human grader. The AI system demonstrated excellent diagnostic performance for the detection of referable DR. The area under the curve (AUC) was 96.5% (95% CI: 91.8–99.7), sensitivity was 88.9% (77.4–95.8), and specificity was 98.7% (96.7–99.7). The positive predictive value was 92.3% (81.4–97.9), and the negative predictive value was 98.0% (95.7–99.3). These values indicate strong agreement between the AI system and human grader for classifying patients as referable or non-referable (Table
Diagnostic accuracy for referable diabetic retinopathy.
| AI grading | Total | ||||
|---|---|---|---|---|---|
| NR | RDR | Ungradable (excluded) | |||
| Human grading | NR | 295 | 4 | 0 | 299 |
| RDR | 6 | 48 | 0 | 54 | |
| Ungradable (excluded) | 14 | 7 | 31 | 52 | |
| Total | 315 | 59 | 31 | 405 |
AUC = Area under the curve. CI = confidence interval. NR = Non Referable. RDR = Referable Diabetic Retinopathy. AI = Artificial intelligence.
To assess the impact of excluding patients with missing or ungradable reference standard grading, a sensitivity analysis was performed on all 405 patients. In a worst-case scenario, all ungradable and non referable cases were assumed to be false negatives, leading to a substantial reduction in estimated sensitivity to 48.5%. In a best-case scenario, all ungradable and referable cases were treated as true positives and all ungradable and non-referable cases as true negatives, resulting in an estimated sensitivity of 90.2%. These analyses suggest that missing data may influence sensitivity estimates but do not fundamentally alter the high overall performance of the AI system.
Subgroup analyses were conducted according to age, sex, ethnicity, diabetes type, HbA1c at diagnosis, duration of diabetes, body mass index (BMI) and image quality (Table
Sub-group analysis (n = 353). AUC = area under the curve. DR = diabetic retinopathy.
| Subgroup | Prevalence | Sensitivity % (95% CI) | Specificity% (95% CI) | AUC % | Comparison AUC p value |
|---|---|---|---|---|---|
| 19–52 years | 5.4% | 84.2 [60.4–96.7] | 99.1 [95.1–100] | 99.8 [99.1–100] | Reference |
| 53–64 years | 5.4% | 89.5 [66.9–98.7] | 99.0 [94.7–100] | 91.0 [77.2–100] | 0.16 |
| 65–89 years | 4.5% | 93.7 [69.8–99.9] | 97.6 [91.7–99.7] | 98.7 [96.3–100] | 0.23 |
| Female | 6.2% | 81.9 [59.7–94.9] | 98.5 [94.9–99.9] | 99.3 [98.2–100] | 0.19 |
| Male | 9.1% | 93.7 [79.2–99.2] | 98.7 [95.5–99.9] | 94.6 [86.5–100] | |
| European Caucasian | 8.2% | 86.2 [68.3–96.1] | 99.3 [96.5–100] | 96.5 [90–100] | 0.82 |
| North African Caucasian | 5.1% | 88.9 [65.3–98.6] | 98.3 [94.1–99.8] | 95.3 [85.2–100] | |
| Type 1 | 6.0% | 81.0 [58.1–94.5] | 99.0 [94.1–100] | 99.2 [97.8–100] | Reference |
| Type 2 | 6.5% | 95.6 [78.0–99.9] | 98.0 [94.1–99.6] | 96.2 [88.6–100] | 0.39 |
| Other* | 2.8% | 90 [55.5–99.7] | 100 [94.22–100] | 90.5 [70.48–100] | 0.36 |
| < 8% | 4.8% | 88.2 [63.5–98.5] | 99.5 [97.2–100] | 94.4 [83.2–100] | 0.63 |
| ≥ 8% | 10.2% | 88.9 [74–96.9] | 97.0 [91.6–99.4] | 97.2 [92.5–99.9] | |
| < 10 years | 4.2% | 86.7 [59.5–98.3] | 99.1 [95.4–100] | 89.4 [72.5–100] | 0.17 |
| ≥ 10 years | 11.0% | 89.7 [75.8–97.1] | 98.3 [95.2–99.6] | 99.4 [98.6–100] | |
| < 28 kg/m2 | 8.2% | 86.2 [68.3–96.1] | 98.7 [95.2–99.9] | 96.1 [89.4–99.8] | 0.90 |
| ≥ 28 kg/m2 | 7.4% | 92.3 [74.0–99.0] | 98.1 [94.6–99.6] | 96.7 [89.7–100] | |
| Excellent | 5.7% | 85.0 [62.1–96.8] | 98.3 [94.1–99.8] | 99.6 [98.7–100] | Reference |
| Good | 8.0% | 92.8 [76.5–99.1] | 99.2 [95.8–100] | 93.6 [83.8–100] | 0.17 |
| Adequate or dirty lens | 1.7% | 83.3 [35.8–99.5] | 97.9 [88.7–100] | 98.9 [96.4–100] | 0.40 |
BMI = body mass index. HbA1c = glycated hemoglobin. CI = confidence interval. *Other type of diabetes includes: Monogenic diabetes, diabetes secondary to chronic pancreatitis, corticotherapy-related diabetes, post-transplantation diabetes.
Multivariate logistic regression models were constructed to evaluate systemic predictors of referable DR for both AI-based and human grader-based classifications. Elevated HbA1c at diabetes diagnosis and longer duration of diabetes emerged as significant predictors of referable DR in both models. These two factors are well-established determinants of DR progression and were confirmed as strong risk markers in our cohort. Other variables, such as BMI and smoking status, were not significantly associated with referable DR, possibly due to the relatively low prevalence of referable DR and limited statistical power for these factors. Interestingly, other ethnicities (sub-Saharan African and Asian) were associated with an increased risk of referable DR in the AI model but not in the human grader model (Table
Multivariate analysis of systemic risk factors with referable diabetic retinopathy diagnosed by the ensemble AI model, as compared with human graders.
| Artificial intelligence model | Human grader model | |||
|---|---|---|---|---|
| OR (95% CI) | P value | OR (95% CI) | P value | |
| Age (per 1-year increase) | 1.02 [1.00–1.04] | 0.12 | 1.00 [0.98–1.02] | 0.99 |
| European Caucasian | Reference | |||
| North African Caucasian (vs reference) | 0.97 [0.50–1.84] | 0.91 | 0.77 [0.38–1.52] | 0.46 |
| Other (vs reference) | 2.96 [1.07–7.79] | 0.03 | 2.34 [0.81–6.25] | 0.1 |
| Diabetes duration (per 1-year increase) | 1.03 [1.00–1.06] | 0.04 | 1.03 [1.00–1.06] | 0.04 |
| HbA1c (per 1% increase) | 1.29 [1.13–1.47] | 1.4 × 10⁻4 | 1.30 [1.15–1.49] | 6.3 × 10⁻4 |
| BMI (per 1 kg/m2 increase) | 0.99 [0.94–1.04] | 0.65 | 1.00 [0.95–1.06] | 0.92 |
Patients were the units of analysis (n = 353). BMI = body mass index. HbA1c = glycated hemoglobin. OR = odds ratio. CI = confidence interval.
Overall, 83 patients (23.5%) received a referral recommendation from the human graders for newly identified ocular findings. Among these, 36 (43.3%) were referred for retinal vascular tortuosity, 21 (25.3%) for asymmetric or atrophic optic discs, 14 (16.9%) for suspected age-related macular degeneration, 5 (6.0%) for optic disc hemorrhage or retinal lesions requiring further investigation, 5 (6.0%) for choroidal naevi and 4 (4.8%) for epiretinal membranes. Of all referred cases, 13 (15.7%) had referable DR. These results highlight the added value of fundus photography for detecting non-DR ocular pathologies that also warrant specialist assessment.
In this real-world evaluation of an AI-based DR screening system, we observed high diagnostic accuracy, indicating robust performance in a routine clinical setting. Despite a smaller sample size compared with pivotal trials, the AI model exceeded FDA performance thresholds for DR screening systems (sensitivity > 85% and specificity > 82.5%)
The MONA.health model was trained on a large US-based screening dataset including a majority of Latin American and unspecified ethnicities, as well as publicly available benchmarking data
Our findings are consistent with previous work showing that AI algorithms for DR detection can perform at least as well in real-world conditions
In routine clinical practice, however, ungradable images are systematically referred for ophthalmologic evaluation, as they may reflect ocular media opacities or advanced retinal disease. This real-world management pathway is therefore better reflected by the best-case scenario considered in our sensitivity analysis. Enhanced imaging protocols and targeted staff training may help to mitigate these limitations.
The majority of people with diabetes have either no DR or only mild DR and are therefore at low risk of imminent vision loss. In such low-risk populations, variation in human grader performance has limited impact on clinical outcomes, as management typically involves routine follow-up at one- to two-year intervals
In multivariate analysis, high HbA1c at diagnosis and longer diabetes duration were strongly associated with referable DR for both AI and human grader, confirming their role as major risk factors for DR progression
An important observation in our study was that 23.5% of patients were referred for newly discovered ocular findings, yet only 15.7% of these referrals were due to DR. This underscores the potential of fundus photography to uncover a broader spectrum of eye diseases, including retinal vascular anomalies, optic disc pathology, age-related macular degeneration and epiretinal membranes. In this context, while fundus photography enabled the identification of various non-diabetic ocular findings requiring specialist referral, the AI system was specifically designed for DR and diabetic macular edema detection and does not aim to identify other ocular pathologies, underscoring its role as a triage rather than a comprehensive diagnostic tool. Future research should explore AI models capable of multi-disease detection (e.g. DR, glaucoma, Age-related macular degeneration, hypertensive retinopathy, high myopia and cataract) to avoid overlooking clinically significant non-DR pathologies
This study has several strengths. To our knowledge, it is the first to evaluate the implementation of an AI-assisted DR screening software in real-world settings in Belgium. The system demonstrated clear time benefits in clinical workflow and maintained high diagnostic performance across diverse ethnic groups. Although time metrics were not formally recorded, real-world observations suggest that the AI-assisted screening workflow was efficient. The complete process, including patient installation, image acquisition, automated analysis and communication of results, required around 10 min per patient, while specialist review of retinal images required around 1 min per case. In comparison, ophthalmology consultations generally require substantially more time, often 30–45 min, due to waiting times, pupil dilation and comprehensive examination. The concordance between AI and human graders in identifying patients at highest risk of DR, based on HbA1c and duration of diabetes, supports its clinical relevance for risk stratification.
Several limitations should also be acknowledged. First, the sample size is modest compared with larger multicenter studies
In conclusion, our evaluation of an AI-based DR screening model in a real-world hospital setting demonstrates high accuracy, robustness and operational feasibility, with performance exceeding regulatory benchmarks for sensitivity and specificity. The system offers clear time savings and appears generalizable across ethnically diverse populations. Further optimization of imaging protocols, long-term follow-up of referral pathways and health-economic analyses will be crucial to fully realize the benefits of AI-assisted DR screening and to guide its integration into national screening strategies.
The authors thank the nursing and medical staff of the Department of Endocrinology at Erasmus Hospital for their support in patient recruitment and imaging.
The M.C. laboratory is supported by the European Union Horizon Health project NEMESIS, the Belgian Fonds National de la Recherche Scientifique (FNRS), and the Walloon Region strategic axis Fonds de la Recherche Scientifique (FRFS)–Walloon Excellence in Life Sciences and Biotechnology (WELBIO).
L.B., E.M. and M.C. designed the study. L.B. collected the data, performed the analyses and drafted the manuscript. E.M. performed the reference grading and contributed to the methodological design. M.L., M.C., L.C. and A.B. referred patients for the screening. E.M. and M.C. supervised the study.
The datasets generated and analyzed during the current study are available from the corresponding author upon reasonable request. Access is restricted due to GDPR and institutional data-protection policies.
The authors declare no competing interests.
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
The datasets generated and analyzed during the current study are available from the corresponding author upon reasonable request. Access is restricted due to GDPR and institutional data-protection policies.