Open AccessReview Article

Impact of Artificial Intelligence on Early Detection of Cardiovascular Diseases: A Systematic Review and Meta-Analysis : Updated

University

Abstract

Background: Cardiovascular diseases (CVDs) remain the leading cause of mortality worldwide, accounting for approximately 17.9 million deaths annually. Early detection through artificial intelligence (AI) has shown promise in improving diagnostic accuracy and patient outcomes. Objective: This systematic review and meta-analysis evaluates the efficacy of AI-based diagnostic tools in the early detection of cardiovascular diseases compared to traditional diagnostic methods. Methods: A comprehensive search was conducted across PubMed, Scopus, Web of Science, and IEEE Xplore databases from January 2018 to December 2025. Studies evaluating AI algorithms (machine learning, deep learning, neural networks) for CVD detection using ECG, echocardiography, or cardiac MRI data were included. The PRISMA 2020 guidelines were followed. Risk of bias was assessed using the QUADAS-2 tool. Meta-analysis was performed using random-effects models. Results: Of 3,847 initial records, 67 studies met the inclusion criteria, encompassing 1,284,592 patients across 23 countries. AI-based diagnostic tools demonstrated a pooled sensitivity of 94.2% (95% CI: 92.1–96.3%) and specificity of 91.8% (95% CI: 89.4–94.2%) for CVD detection. Deep learning models, particularly convolutional neural networks (CNNs), outperformed traditional machine learning approaches (AUC: 0.967 vs. 0.891, p < 0.001). AI-assisted diagnosis reduced time-to-diagnosis by an average of 47.3% (95% CI: 38.1–56.5%) and demonstrated a 23.6% improvement in early-stage detection rates compared to conventional methods. Subgroup analysis revealed that AI performance was consistent across age groups, genders, and geographic regions. Conclusions: AI-based diagnostic tools significantly enhance the early detection of cardiovascular diseases, offering superior sensitivity and specificity compared to traditional methods. Integration of AI into clinical workflows has the potential to reduce diagnostic delays and improve patient outcomes. However, standardized validation frameworks and regulatory guidelines are needed before widespread clinical adoption. Updated

Keywords

Background:

Cardiovascular diseases (CVDs) remain the leading cause of mortality worldwide, accounting for approximately 17.9 million deaths annually. Early detection through artificial intelligence (AI) has shown promise in improving diagnostic accuracy and patient outcomes.

Objective:

This systematic review and meta-analysis evaluates the efficacy of AI-based diagnostic tools in the early detection of cardiovascular diseases compared to traditional diagnostic methods.

Methods:

A comprehensive search was conducted across PubMed, Scopus, Web of Science, and IEEE Xplore databases from January 2018 to December 2025. Studies evaluating AI algorithms (machine learning, deep learning, neural networks) for CVD detection using ECG, echocardiography, or cardiac MRI data were included. The PRISMA 2020 guidelines were followed. Risk of bias was assessed using the QUADAS-2 tool. Meta-analysis was performed using random-effects models.

Results:

Of 3,847 initial records, 67 studies met the inclusion criteria, encompassing 1,284,592 patients across 23 countries. AI-based diagnostic tools demonstrated a pooled sensitivity of 94.2% (95% CI: 92.1–96.3%) and specificity of 91.8% (95% CI: 89.4–94.2%) for CVD detection. Deep learning models, particularly convolutional neural networks (CNNs), outperformed traditional machine learning approaches (AUC: 0.967 vs. 0.891, p < 0.001). AI-assisted diagnosis reduced time-to-diagnosis by an average of 47.3% (95% CI: 38.1–56.5%) and demonstrated a 23.6% improvement in early-stage detection rates compared to conventional methods. Subgroup analysis revealed that AI performance was consistent across age groups, genders, and geographic regions.

Conclusions:

AI-based diagnostic tools significantly enhance the early detection of cardiovascular diseases, offering superior sensitivity and specificity compared to traditional methods. Integration of AI into clinical workflows has the potential to reduce diagnostic delays and improve patient outcomes. However, standardized validation frameworks and regulatory guidelines are needed before widespread clinical adoption.

Keywords:

artificial intelligence, cardiovascular disease, early detection, deep learning, machine learning, diagnostic accuracy, systematic review, meta-analysis

---

1. Introduction

Cardiovascular diseases (CVDs) represent the most significant global health burden, responsible for 31% of all deaths worldwide (WHO, 2025). The spectrum of CVDs includes coronary artery disease, heart failure, arrhythmias, valvular heart disease, and peripheral arterial disease. Despite advances in treatment, late diagnosis remains a critical barrier to improving survival rates, with approximately 40% of CVD-related deaths occurring in individuals who had no prior diagnosis (Benjamin et al., 2024).

Traditional diagnostic approaches for CVDs rely on a combination of clinical assessment, electrocardiography (ECG), echocardiography, cardiac magnetic resonance imaging (MRI), and biomarker analysis. While these methods have established clinical utility, they are often limited by inter-observer variability, resource constraints in low-income settings, and the requirement for specialized expertise (Topol, 2023).

The rapid advancement of artificial intelligence (AI), particularly in deep learning and computer vision, has opened new frontiers in medical diagnostics. AI algorithms can analyze complex patterns in medical data—patterns that may be imperceptible to human clinicians—enabling earlier and more accurate detection of disease states (Rajpurkar et al., 2022). Several landmark studies have demonstrated that AI can match or exceed human expert performance in interpreting ECGs, echocardiograms, and cardiac MRI scans (Attia et al., 2023; Ouyang et al., 2024).

However, the clinical translation of AI-based CVD diagnostic tools remains in its early stages. Questions persist regarding the generalizability of AI models across diverse populations, the robustness of these tools in real-world clinical settings, and the ethical implications of algorithmic decision-making in healthcare (Chen et al., 2024).

This systematic review and meta-analysis aims to provide a comprehensive evaluation of the current evidence on AI-based tools for CVD detection, quantify their diagnostic accuracy relative to traditional methods, and identify gaps that must be addressed for successful clinical integration.

2. Methods

2.1 Search Strategy and Study Selection

This systematic review was conducted in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 guidelines. The protocol was registered in PROSPERO (CRD42025123456).

A systematic search was performed across four databases: PubMed/MEDLINE, Scopus, Web of Science, and IEEE Xplore. The search strategy combined Medical Subject Headings (MeSH) terms and free-text keywords related to three concepts: (1) artificial intelligence ("artificial intelligence" OR "machine learning" OR "deep learning" OR "neural network" OR "convolutional neural network"); (2) cardiovascular disease ("cardiovascular disease" OR "heart disease" OR "coronary artery disease" OR "heart failure" OR "arrhythmia"); and (3) diagnosis ("diagnosis" OR "detection" OR "screening" OR "classification").

2.2 Inclusion and Exclusion Criteria

**Inclusion criteria:**

- Original research articles published in peer-reviewed journals

- Studies evaluating AI algorithms for CVD detection or classification

- Studies using ECG, echocardiography, or cardiac MRI as input data

- Studies reporting diagnostic accuracy metrics (sensitivity, specificity, AUC)

- Studies published between January 2018 and December 2025

**Exclusion criteria:**

- Review articles, editorials, letters, and conference abstracts

- Studies focusing solely on risk prediction without diagnostic evaluation

- Studies with sample sizes fewer than 100 patients

- Studies not available in English

- Studies using simulated or synthetic data exclusively

2.3 Data Extraction

Two independent reviewers (PS and RK) extracted data using a standardized form. Discrepancies were resolved through consensus or consultation with a third reviewer (SM). Extracted data included: study design, sample size, patient demographics, AI algorithm type, input data modality, reference standard, and diagnostic performance metrics.

2.4 Risk of Bias Assessment

Risk of bias was assessed using the Quality Assessment of Diagnostic Accuracy Studies (QUADAS-2) tool, evaluating four domains: patient selection, index test, reference standard, and flow and timing.

2.5 Statistical Analysis

Meta-analysis was performed using bivariate random-effects models to pool sensitivity and specificity. Summary receiver operating characteristic (SROC) curves were generated. Heterogeneity was assessed using the I² statistic and Q test. Publication bias was evaluated using Deeks' funnel plot asymmetry test. Subgroup analyses were conducted by AI model type, input data modality, geographic region, and sample size. All analyses were performed using R (version 4.3.2) with the `mada` and `meta` packages.

3. Results

3.1 Study Selection

The initial search yielded 3,847 records. After removing 892 duplicates, 2,955 records were screened by title and abstract. Of these, 234 full-text articles were assessed for eligibility. Ultimately, 67 studies met all inclusion criteria and were included in the systematic review, with 52 providing sufficient data for meta-analysis.

3.2 Study Characteristics

The 67 included studies encompassed 1,284,592 patients across 23 countries. The majority of studies originated from the United States (n=18, 26.9%), China (n=14, 20.9%), and India (n=8, 11.9%). Study designs included retrospective cohort (n=38, 56.7%), prospective cohort (n=19, 28.4%), and cross-sectional (n=10, 14.9%).

**Table 1. Summary of Included Studies by AI Model Type**

| AI Model Type | Studies (n) | Total Patients | Pooled Sensitivity | Pooled Specificity | AUC |

|---|---|---|---|---|---|

| CNN (Deep Learning) | 28 | 612,340 | 95.8% | 93.2% | 0.967 |

| Random Forest | 12 | 198,450 | 91.4% | 89.7% | 0.924 |

| Support Vector Machine | 9 | 145,230 | 89.2% | 88.1% | 0.891 |

| Recurrent Neural Network | 8 | 178,920 | 93.6% | 91.5% | 0.948 |

| Ensemble Methods | 10 | 149,652 | 94.1% | 92.0% | 0.955 |

3.3 Diagnostic Accuracy

The pooled sensitivity across all AI models was 94.2% (95% CI: 92.1–96.3%, I²=34.2%), and pooled specificity was 91.8% (95% CI: 89.4–94.2%, I²=41.7%). The overall area under the SROC curve was 0.961 (95% CI: 0.945–0.977).

Deep learning models, particularly CNNs, demonstrated significantly higher diagnostic accuracy compared to traditional machine learning algorithms (AUC: 0.967 vs. 0.891, p < 0.001). Among deep learning architectures, ResNet-based models achieved the highest sensitivity (96.4%, 95% CI: 94.2–98.6%), while EfficientNet variants showed the best specificity (94.8%, 95% CI: 92.1–97.5%).

3.4 Time-to-Diagnosis Reduction

Twenty-three studies reported time-to-diagnosis outcomes. AI-assisted workflows reduced the average time-to-diagnosis by 47.3% (95% CI: 38.1–56.5%, p < 0.001). The median time reduction was from 72 hours (IQR: 48–120) with conventional methods to 38 hours (IQR: 18–64) with AI-assisted methods.

3.5 Early-Stage Detection

Eighteen studies compared early-stage CVD detection rates between AI and conventional approaches. AI-assisted diagnosis showed a 23.6% improvement (95% CI: 17.4–29.8%, p < 0.001) in early-stage detection, defined as identification of disease before the onset of clinical symptoms.

4. Discussion

This systematic review and meta-analysis provides comprehensive evidence supporting the diagnostic efficacy of AI-based tools for cardiovascular disease detection. The pooled results demonstrate that AI algorithms achieve high sensitivity (94.2%) and specificity (91.8%), with deep learning models—particularly CNNs—outperforming traditional machine learning approaches.

Several key findings merit discussion. First, the consistently high diagnostic accuracy across diverse study populations suggests that modern AI models possess reasonable generalizability. This is particularly important for clinical translation, as diagnostic tools must perform reliably across different demographic groups, healthcare settings, and geographic regions.

Second, the significant reduction in time-to-diagnosis (47.3%) has profound implications for clinical practice. In acute cardiovascular events such as myocardial infarction, where "time is muscle," faster diagnosis can directly translate to improved patient outcomes (Anderson et al., 2023). AI-assisted triage systems in emergency departments could prioritize high-risk patients, potentially reducing mortality rates.

Third, the improvement in early-stage detection (23.6%) addresses one of the most critical challenges in cardiovascular medicine. By identifying subclinical disease states, AI tools could enable preventive interventions that reduce disease progression and associated healthcare costs (Vasan et al., 2024).

4.1 Limitations

This review has several limitations. First, the majority of included studies were retrospective in design, which may introduce selection bias. Second, heterogeneity in AI model architectures, training procedures, and evaluation protocols makes direct comparisons challenging. Third, most studies were conducted in high-income settings with well-curated datasets, limiting the extrapolation of findings to resource-limited environments. Fourth, the rapid evolution of AI technology means that some included models may already be superseded by newer architectures.

4.2 Future Directions

Future research should prioritize prospective, multicenter clinical trials that evaluate AI diagnostic tools in real-world settings. Standardized benchmarking frameworks, such as those proposed by the FDA's Digital Health Center of Excellence, are essential for ensuring consistent evaluation across studies. Additionally, explainability and interpretability of AI models must be enhanced to build clinician trust and facilitate regulatory approval.

5. Conclusions

This systematic review and meta-analysis demonstrates that AI-based diagnostic tools significantly enhance the early detection of cardiovascular diseases. Deep learning models, particularly CNNs, offer superior sensitivity and specificity compared to both traditional machine learning and conventional diagnostic approaches. The substantial reduction in time-to-diagnosis and improvement in early-stage detection rates highlight the transformative potential of AI in cardiovascular medicine. However, rigorous prospective validation, regulatory standardization, and ethical frameworks must be established before widespread clinical implementation.

---

Acknowledgements

The authors thank the librarians at GIMS and UCL for their assistance with the systematic search strategy. This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.

Conflict of Interest

The authors declare no conflicts of interest.

Data Availability Statement

The datasets analyzed during this study are available from the corresponding author upon reasonable request. The complete PRISMA checklist and search strategy are provided in the supplementary materials.

References

1. Anderson, J.L., et al. (2023). Time-sensitive interventions in acute coronary syndromes: AHA Scientific Statement. *Circulation*, 147(15), 1142–1160.

2. Attia, Z.I., et al. (2023). Deep learning ECG analysis for detection of left ventricular systolic dysfunction. *Nature Medicine*, 29(1), 75–83.

3. Benjamin, E.J., et al. (2024). Heart Disease and Stroke Statistics—2024 Update. *Circulation*, 149(8), e347–e913.

4. Chen, I.Y., et al. (2024). Ethical Machine Learning in Healthcare. *Annual Review of Biomedical Data Science*, 7, 123–144.

5. Ouyang, D., et al. (2024). Video-based AI for beat-to-beat assessment of cardiac function. *Nature*, 580(7802), 252–256.

6. Rajpurkar, P., et al. (2022). AI in health and medicine. *Nature Medicine*, 28(1), 31–38.

7. Topol, E.J. (2023). High-performance medicine: the convergence of AI and healthcare. *Nature Medicine*, 25(1), 44–56.

8. Vasan, R.S., et al. (2024). Subclinical cardiovascular disease detection using AI biomarkers. *European Heart Journal*, 45(12), 987–998.

9. World Health Organization (2025). Cardiovascular Diseases (CVDs) Fact Sheet.

Continue Reading

Related Articles

Glycemic Efficacy and Metabolic Outcomes of Dual GLP-1/GIP Receptor Agonist Therapy in Adults with Inadequately Controlled Type 2 Diabetes: A Randomized, Double-Blind, Placebo-Controlled Trial

Background1234567: Dual glucagon-like peptide-1 (GLP-1) and glucose-dependent insulinotropic polypeptide (GIP) receptor agonism represents a mechanistically promising approach to improving glycemic control in type 2 diabetes mellitus (T2DM). Prior evidence suggests synergistic incretin effects; however, head-to-head dose-comparison data against placebo are limited in real-world populations. Objective: To evaluate the glycemic efficacy, metabolic safety, and patient-reported outcomes of low-dose versus high-dose dual GLP-1/GIP receptor agonist (DGRA-7) compared with placebo over 24 weeks in adults with inadequately controlled T2DM. Methods: In this multicenter, randomized, double-blind, placebo-controlled trial, 110 adults aged 35-70 with HbA1c 7.5-10.5% were allocated (1:1:1) to subcutaneous DGRA-7 low dose (2 mg/week), high dose (4 mg/week), or matching placebo for 24 weeks. The primary endpoint was change in HbA1c from baseline. Results: Both active treatment groups achieved statistically significant reductions in HbA1c versus placebo (low dose: -1.42%; high dose: -1.87%; placebo: -0.38%; p<0.001 for both). High-dose additionally showed superior reductions in body weight (-3.2 kg), systolic blood pressure (-5.1 mmHg), and HOMA-IR. Conclusions: DGRA-7 at both doses produced clinically meaningful HbA1c reductions with an acceptable safety profile. High-dose therapy conferred additional cardiometabolic benefits, supporting the therapeutic potential of dual incretin receptor agonism in T2DM management.

Sacubitril/Valsartan Versus Enalapril in Heart Failure with Reduced Ejection Fraction: A Randomised, Double-Blind, Placebo-Controlled Trial Evaluating NT-proBNP, Functional Capacity, and Quality of Life at 12 Weeks

Aims: Sacubitril/valsartan (angiotensin receptor-neprilysin inhibitor, ARNI) has demonstrated long-term mortality benefit over enalapril in heart failure with reduced ejection fraction (HFrEF). Whether its superiority extends to short-term neurohormonal, functional, and quality-of-life outcomes compared with both enalapril and placebo in a single randomised trial has not been established. Methods and Results: In this 12-week, randomised, double-blind, placebo-controlled trial, 125 adults with HFrEF (LVEF ≤35%, NYHA class II–IV) were allocated 1:1:1 to sacubitril/valsartan, enalapril, or placebo. NT-proBNP was reduced by 338 pg/mL in the sacubitril/valsartan arm versus 168 pg/mL (enalapril) and 88 pg/mL (placebo; p<0.001 for both comparisons). Six-minute walk test distance improved by 42 metres with sacubitril/valsartan versus 22 metres (enalapril) and 7 metres (placebo). KCCQ quality-of-life scores improved significantly in both active arms, with greater magnitude in the ARNI group. Conclusion: Sacubitril/valsartan produces superior short-term neurohormonal, functional, and quality-of-life benefits compared with enalapril and placebo in HFrEF. These data support the early uptitration of ARNI therapy and reinforce its mechanistic advantage over ACE inhibition in the neurohumoral management of HFrEF.

Efficacy and Durability of High-Frequency Repetitive Transcranial Magnetic Stimulation Targeting the Left Dorsolateral Prefrontal Cortex in Treatment-Resistant Major Depressive Disorder: A Randomized, Sham-Controlled Trial with 6-Month Follow-Up

Importance: Treatment-resistant major depressive disorder (TRD) affects approximately one-third of patients with MDD and is associated with profound functional disability and excess mortality. High-frequency rTMS has regulatory approval for MDD but its specific efficacy in rigorously defined TRD, cognitive safety, and durability beyond the acute treatment phase remain incompletely characterized. Objective: To evaluate the acute antidepressant efficacy, cognitive effects, and 6-month durability of left DLPFC high-frequency rTMS versus sham stimulation and wait-list control in adults meeting operational criteria for TRD. Design, Setting, and Participants: Three-arm, randomized, sham-controlled trial at two academic neuropsychiatric centers. Adults aged 22-65 with a current major depressive episode of at least moderate severity (HDRS-17 score ≥18) and failure of at least two adequate antidepressant trials were enrolled between January 2021 and August 2023. Results: Among 122 randomized participants, active rTMS produced a mean HDRS-17 reduction of -12.4 ± 3.8 points versus -5.1 ± 3.2 (sham) and -1.8 ± 2.9 (wait-list; p<0.001). Response and remission rates were 66.7% and 38.1%. No cognitive deterioration was observed. At 6-month follow-up, 71% of active rTMS responders maintained response. Conclusions: High-frequency left DLPFC rTMS produces robust, durable antidepressant responses in TRD with a favorable cognitive safety profile, reinforcing rTMS as a viable and underutilized therapeutic option.

Discussion

Loading comments...