Corresponding author.
Diabetic retinopathy (DR) is one of the most common complications of diabetes mellitus (DM) with a prevalence that varies from 10% to 61% in different countries. From healthcare to the precise prevention, diagnosis, and management of diseases, Artificial Intelligence (AI) is progressing rapidly in interdisciplinary fields, including ophthalmology. We aimed to explore the efficacy of artificial intelligence (AI)-based screening for diabetic retinopathy in type 2 DM patients. High-resolution colour fundus photographs were obtained for all patients using a non-mydriatic camera (Nidek AFC-330). DAIRET software uses a machine learning algorithm to identify the most important signs of DR. The software provides an output with a binary result: negative, in case of no referable disease or positive, when the presence of signs of DR is detected. The photographs were also manually graded negative or positive for DR by an expert ophthalmologist, masked to AI responses. Expert clinical opinion was used as the reference (gold standard), and the results were compared with those of the AI responses. A significant association was observed between DAIRET and expert diagnosis (χ2 = 18.6, p < 0.001). Disease prevalence, according to expert assessment, was 20.2% (33/163). Observed agreement between DAIRET and Expert diagnosis was 69.3%. Agreement metric was: Gwet’s AC1 = 0.47 (95% CI: 0.33–0.61), indicating moderate agreement. Diagnostic performance of the DAIRET test indicates good sensitivity (73%), a specificity (68%) and high negative predictive value (NPV=90.8%), suggesting that DAIRET is more effective at excluding diabetic retinopathy than at confirming it. ROC analysis stratified by gender showed no difference in performance (p>0.05). Overall, these results suggest that DAIRET may be effectively integrated into screening workflows to rule out diabetic retinopathy, reduce specialist workload, and prioritize referrals. In particular, for daily practical use, DAIRET is managed within the MètaClinic electronic medical record. In-house data processing without going to an external cloud is a significant strength. Larger, prospective studies are needed to validate these findings, optimise thresholds, and assess performance across diverse populations and clinical settings.
The online version contains supplementary material available at
Received 2026 Mar 12; Accepted 2026 Jun 11; Collection date 2026.
According to data from the International Diabetes Federation, 589 million adults (20-79 years) are living with diabetes – 1 in 9. This number is predicted to rise to 853 million by 2050 (
Diabetic retinopathy (DR) is one of the most common complications of diabetes mellitus (DM) and it is the primary eye disease that causes blindness in the population
From healthcare to the precise prevention, diagnosis, and management of diseases, Artificial Intelligence (AI) is progressing rapidly in various interdisciplinary fields
In Germany, AI-based screening for diabetic retinopathy is viewed positively and is considered ‘fundamentally suitable for future use’ according to the National Care Guideline
In the United States of America, the first automated screening system for DR was approved back in 2018
Indeed, AI algorithms could improve the efficacy of screening and might be implemented for clinical use after thorough validation in a real-life setting
We aimed to explore the efficacy of artificial intelligence (AI)-based screening for diabetic retinopathy (DR) in type 2 diabetes mellitus (T2DM) patients.
This observational pilot study aimed to evaluate the diagnostic performance of DAIRET (acronym for
Inclusion and exclusion criteria are shown in Fig.
Participant Flow Diagram. Selection process of the study population. A total of 170 patients with Type 2 Diabetes Mellitus were initially enrolled between September 2024 and June 2025. Out of these, 7 subjects were excluded: 3 due to ungradable fundus images (not performant) and 4 due to dubious clinical records, resulting in a final sample of 163 participants for AI-based statistical analysis.
High-resolution colour fundus photographs were obtained for all patients using a non-mydriatic camera (Nidek AFC-330). Usually, two 45° fields per eye were acquired: one centred on the optic disc and one on the macula (Fig.
Color fundus photographs centered on the optic disc (Left) and on the macula (Right). The images illustrate the standard fields captured during the screening process to allow for both automated AI analysis and clinical validation of the optic nerve head and macular region.
Fundus colour photograph of a patient without any signs of DR (Left) and of a patient with Mild DR (Right). The right image demonstrates the presence of microaneurysms and small red macular dots, representing characteristic bulges in blood vessel walls identified by the DAIRET software during the screening phase.
The photographs were also manually graded as negative or positive for DR by an expert ophthalmologist, masked to AI responses. Expert clinical opinion was used as the reference (gold standard), and the results were compared with those of the AI responses.
Sample size was calculated using a power analysis for a single proportion based on Cohen’s arcsine transformation. Due to the absence of preliminary pilot data or established literature regarding the expected performance of the AI software, a medium effect size (Cohen’s h = 0.3) was assumed. With a significance level of 0.01 and a statistical power of 90% in a two-sided test, the minimum required sample size was determined to be 166 patients (n = 165.33).
A total of 170 consecutive diabetic patients enrolled in the DR screening in the Eye Unit of San Salvatore Hospital, L'Aquila (Italy), between September 2024 and June 2025, were recruited. Patients not performant or whose tests were classified as ambiguous or inconclusive were excluded from the analysis. Specifically, 7 DAIRET tests (4.1%) were excluded, yielding a final analytic sample of 163 patients.
Descriptive statistics (mean and standard deviation) were computed for demographic and clinical variables, including age, glycaemic control, lipid profile, renal function, liver enzymes, blood pressure, and diabetes duration (Table
Distribution of clinical characteristics according to gender.
| Clinical features | Males | Females | ||||
|---|---|---|---|---|---|---|
| Observed | Mean | SD | Observed | Mean | SD | |
| Age | 111 | 64.9 | 12.9 | 52 | 67.2 | 11.8 |
| HbA1c (%) | 110 | 7.2 | 1.2 | 52 | 7.0 | 1.0 |
| Body mass index (kg/m2) | 108 | 28.0 | 4.7 | 52 | 28.3 | 5.6 |
| Total cholesterol (mg/dL) | 104 | 152.9 | 43.0 | 49 | 162.9 | 38.5 |
| HDL cholesterol (mg/dL) | 105 | 47.3 | 14.7 | 51 | 52.2 | 13.4 |
| LDL cholesterol (mg/dL) | 104 | 83.3 | 34.3 | 49 | 85.8 | 32.9 |
| Triglycerides (mg/dL) | 105 | 116.9 | 76.9 | 49 | 123.7 | 55.6 |
| Serum creatinine (mg/dL)* | 110 | 1.1 | 0.7 | 52 | 0.8 | 0.2 |
| Glomerular filtration rate | 110 | 86.1 | 25.3 | 52 | 83.1 | 20.2 |
| AST (U/L) | 97 | 22.5 | 8.2 | 45 | 21.6 | 6.1 |
| ALT (U/L) | 97 | 23.6 | 14.0 | 45 | 20.2 | 7.3 |
| Systolic blood pressure (mmHg) | 104 | 128.2 | 13.1 | 50 | 129.0 | 15.5 |
| Diastolic blood pressure (mmHg) | 104 | 73.1 | 8.0 | 50 | 74.6 | 9.1 |
Using expert opinion as the reference standard, a 2×2 confusion matrix was constructed to compare DAIRET test results with expert diagnosis. From this matrix, diagnostic performance measures were calculated and summarised in Table
DAIRET performance.
| DAIRET test parameters | Estimate | 95% CI |
|---|---|---|
| Sensitivity (Se) | 0.73 | 0.56 – 0.85 |
| Specificity (Sp) | 0.68 | 0.60 – 0.76 |
| Positive Predictive Value (PPV) | 0.37 | 0.26 – 0.49 |
| Negative Predictive Value (NPV) | 0.91 | 0.83 – 0.95 |
| Accuracy | 0.69 | 0.62 – 0.76 |
| Positive Likelihood Ratio (LR+) | 2.31 | 1.66 – 3.20 |
| Negative Likelihood Ratio (LR−) | 0.4 | 0.23 – 0.70 |
A χ2 test was used to assess the association between DAIRET and expert diagnosis using Cramér’s V to quantify the strength of the statistical association.
The Inter-rater agreement was assessed using percent agreement, Cohen’s kappa, and Gwet’s AC1
Given the imbalanced disease prevalence, Gwet’s AC1 was considered the primary agreement metric.
In our data, disease prevalence according to the reference standard is 20.2%, resulting in unbalanced marginal distributions. The Inter-rater agreement was assessed using percent agreement, Cohen’s kappa and Gwet’s AC1. Cohen’s kappa is well known to be sensitive to prevalence and marginal imbalance, often producing artificially low values despite substantial observed agreement (
Gwet’s AC1 is robust to prevalence effects and provides a more stable estimate of agreement. The discrepancy between observed agreement (k=0.69, 95% CI=(0.16, 0.44)) and kappa (0.30), typically when kappa is affected by skewed prevalence, may underestimate the true level of agreement, describing a fair agreement. The Gwet’AC1 (AC1=0.47, 95% CI=(0.33, 0.61)).
Diagnostic accuracy measures (sensitivity, specificity, predictive values, accuracy, and likelihood ratios) were estimated with 95% confidence intervals (CIs). Proportion CIs were calculated using the exact binomial method. Likelihood ratio CIs were calculated on the log scale and back-transformed. Using expert opinion as the reference standard, a 2×2 confusion matrix was constructed to compare DAIRET test results with expert diagnosis. From this matrix, diagnostic performance measures were calculated and summarised in Table
Subgroup analysis ROC analyses were additionally performed stratified by sex to assess potential gender differences in diagnostic performance. All statistical analyses were performed using Stata software, version 17 (StataCorp, College Station, TX, USA). Statistical significance was set at
A total of 170 consecutive diabetic patients were initially enrolled in the study. Following the screening protocol, 7 subjects (4.1%) were excluded: 3 due to ungradable fundus images (media opacities) and 4 due to incomplete or dubious clinical records. This resulted in a final analytic sample of 163 participants. Mean age was 65.7 ± 12.6 years. Descriptive statistics for clinical variables are reported in Table
A significant association was observed between DAIRET and expert diagnosis (χ2 = 18.6, p < 0.001), with a Cramér’s V of 0.34, indicating a moderate association.
Disease prevalence, according to expert assessment, was 20.2% (33/163).
Observed agreement between DAIRET and Expert diagnosis was 69.3%.
Agreement metrics were: Cohen’s kappa = 0.30 (95% CI: 0.16–0.44), indicating
Given the unbalanced prevalence, Gwet’s AC1 was considered a more reliable measure of agreement than Cohen’s kappa.
Using expert opinion as the reference standard, the DAIRET test showed the following performance (Table
These results indicate good sensitivity and high negative predictive value, suggesting that DAIRET is more effective at excluding diabetic retinopathy than at confirming it.
ROC analysis yielded an AUC of 0.71, indicating acceptable overall diagnostic accuracy. The cut-off was 0.5, corresponding to a Sensitivity of 73% and a Specificity of 68%. Bootstrap validation confirmed the robustness of these estimates. ROC analysis stratified by gender (Fig.
Gender-Stratified ROC Curve Analysis. Comparison of AI diagnostic performance between Female participants and Male participants. The analysis highlights the software’s consistency across genders, with an Area Under the Curve (AUC) of 71.5% for females and 69.3% for males (
The integration of Artificial Intelligence (AI) into medical screening represents a fundamental turning point in modern healthcare, acting as a powerful ally in improving the accuracy of early diagnosis. Its importance lies primarily in its ability to analyse huge amounts of data in real time. For example, in mammography screening, it improves breast cancer detection
Artificial intelligence (AI) is revolutionizing diabetic retinopathy screening by rapidly analysing fundus images with high accuracy and safety. While this pilot study was conducted in a high-resource setting, the workflow efficiency observed supports the broader global application of such tools in resource-limited environments where specialist access is constrained. This is further supported by a recent scoping review highlighting the growing role of AI in diabetic eye care, particularly for screening populations at risk of sight loss in low-income and middle-income countries (LMICs)
In our pilot study, the AI-based DAIRET test demonstrated moderate diagnostic accuracy for detecting diabetic retinopathy. The observed sensitivity (73%) and high negative predictive value (91%) indicate that the software performs acceptably in excluding disease, supporting its potential role as a preliminary first-line triage aid within organised diabetic retinopathy screening programmes. When placing these findings in the context of existing AI literature, large-scale clinical trials and validation studies—such as those by Gulshan et al.
The lower baseline sensitivity and specificity observed in our pilot study can be attributed to the preliminary sample size, real-world screening conditions, and the single-grader reference standard. However, the high NPV (90.8%) achieved by DAIRET matches the clinical requirements for entry-level public health screening workflows, where a primary objective is to safely identify low-risk individuals to reduce specialist burden and optimize referral pathways
The agreement analysis showed moderate concordance between DAIRET and expert diagnoses. Although Cohen’s kappa suggests only fair agreement, Gwet’s AC1 indicates a more robust level of agreement after accounting for the imbalanced disease prevalence (20.2%), highlighting the importance of selecting appropriate agreement metrics in diagnostic accuracy studies.
Subgroup analysis showed consistent diagnostic performance across genders, with an AUC of 71.5% for females and 69.3% for males. The lack of a statistically significant difference (p=0.84) suggests that the DAIRET algorithm maintains a stable diagnostic accuracy regardless of the patient’s sex. This consistency is a relevant finding for the implementation of the tool in unselected populations within organized screening programs.
Overall, these results suggest that DAIRET may be effectively integrated into screening workflows to rule out diabetic retinopathy, reduce specialist workload, and prioritise referrals. In particular, for daily practical use, DAIRET is managed within the MètaClinic electronic medical record: DR classification data are available within a few minutes. Regarding data protection regulations, in-house data processing without use of an external cloud is a significant strength.
Undoubtedly, some limitations of this study should be noted. First, DAIRET can diagnose and grade DR
Through an AI algorithm, it is not suitable for some patients. For example, it is not possible to obtain fundus photographs from some DM patients due to small pupils or poor image quality caused by cataract opacity. Second, larger, prospective studies are needed to validate these findings, optimise thresholds, and assess performance across diverse populations and clinical settings. This is especially true in population-based organised screening.
Third, the reference gold standard in this pilot study relied on the manual clinical evaluation of a single expert ophthalmologist. Although highly experienced, this approach introduces a potential limitation regarding subjective inter-observer variability, whereas a multi-reader consensus panel or central reading center would provide an even more robust reference standard for large-scale trials.
On the other hand, it should be emphasized that there are points highlighting the importance of AI alongside screening are: rapid and effective screening, especially distinguishing between healthy patients and those requiring specialist intervention, diagnostic support for ophthalmologists in managing large patient volumes, and improving access to care.
In conclusion, this pilot study demonstrates that the DAIRET software provides a reliable and consistent diagnostic output for diabetic retinopathy screening. The high Negative Predictive Value (90.8%) supports its use as an effective 'rule-out’ tool, capable of identifying healthy individuals and reducing the workload for ophthalmologists. Furthermore, our findings indicate that the software’s performance is stable across genders, with no statistically significant differences in diagnostic accuracy between male and female participants (p=0.84).
While further large-scale studies are needed to refine these results, this automated workflow represents a promising step toward more efficient and accessible screening programs, both in high-resource settings and in resource-constrained environments.
Supplementary Information.
EA designed the study, coordinated the research, and drafted the manuscript. MB and MC were responsible for clinical data collection and ophthalmological assessments. PC contributed to the clinical supervision. FM performed the statistical analysis, ROC curve modelling, and contributed to drafting the methodology. All authors read and approved the final manuscript.
The datasets generated and analysed during the current study are stored within the MètaClinic electronic medical record system at the University of L’Aquila. Due to privacy and data protection regulations regarding in-house processing, the data are not publicly available but can be made available from the corresponding author on reasonable request.
The authors declare no competing interests.
This study adhered to the tenets of the Declaration of Helsinki. The authorization protocol to process the data was obtained on 3 April 2024 (protocol number FFORIC24.01) by Institutional Review Board of Department of Life, Health and Environmental Sciences of the University of L’Aquila, Italy. In addition it was approved by official protocol for the implementation of tele-retinography for the early diagnosis of diabetic retinopathy in the Abruzzo Region, Italy. (
The patients enrolled in the study provided consent for the use of their clinical data and fundus images for research and publication purposes. All data were anonymized prior to analysis to ensure patient privacy.
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
The datasets generated and analysed during the current study are stored within the MètaClinic electronic medical record system at the University of L’Aquila. Due to privacy and data protection regulations regarding in-house processing, the data are not publicly available but can be made available from the corresponding author on reasonable request.