Trained on 135,475 coronary angiography examinations comprising 812,850 multi-view videos from two Shanghai hospitals, CAG-MIND — a domain-specific vision-language foundation model — achieved AUROC scores of 0.940 for coronary stenosis detection, 0.907 for balloon/stent prediction, and 0.875 for CABG recommendation across both internal and external validation cohorts. Notably, the model attained mean AUROCs of 0.686–0.745 in zero-shot settings without any task-specific supervision, rising to 0.827–0.846 after fine-tuning, and outperformed competing architectures even when trained on just 10% of labeled data.

Coronary artery disease remains the leading cause of global mortality, and angiography interpretation has long required years of subspecialty training — a bottleneck that limits timely revascularization decisions worldwide. CAG-MIND's contrastive vision-language pretraining approach, aligning multi-view imaging with structured report semantics, represents a meaningful architectural advance over prior task-specific AI tools that could not generalize across diagnostic domains. The 0.94 stenosis detection AUROC is clinically significant, approaching the inter-observer agreement range among experienced interventional cardiologists. However, critical limitations warrant caution: the training data originates entirely from two Chinese tertiary centers, raising questions about generalizability to diverse imaging equipment, patient populations, and reporting conventions globally. The model has not been tested in prospective clinical workflows or against patient outcomes. As a preprint posted on medRxiv and not yet peer-reviewed, these results require independent validation before clinical adoption. Still, the data-efficient fine-tuning performance suggests genuine foundation-model potential — an incremental but substantive step toward AI-assisted catheterization laboratory support.