Ethics approval
For the retrospective study, ethical approval was obtained individually from the institutional review board of SYSUCC, SCCH, SIPD, First Affiliated Hospital of Zhejiang University, Hunan Cancer Hospital (HCH), Shantou Central Hospital, Xinjiang Medical University Affiliated Cancer Hospital (XMUACH), Fujian Cancer Hospital, General University Hospital in Prague, GenesisCare, SNCH and LCH. The study was registered at https://www.chictr.org.cn/ (Chictr.org.cn identifier: ChiCTR2300074806).
Dataset
This study aims to classify patients into positive and negative categories, where the ‘positive’ category includes patients with esophageal malignancy, including HGIN and EC. Data were used for model training and for validation in three clinical scenarios—opportunistic screening in hospitals, opportunistic screening in established LDCT programs and the population-based endoscopic screening program. For the retrospective test cohorts, patient labels were defined as follows: cases of esophageal malignancy were confirmed through surgical or biopsy histopathology, positive patients in all cohorts were staged according to the eighth edition of the American Joint Committee on Cancer pathological staging system. Negative controls were defined as individuals free of malignant esophageal lesions by at least 2 years of clinical follow-up or negative endoscopic results within 1 year. It remains possible that a minor fraction of indolent lesions remained undetected during the follow-up period, thereby introducing potential label noise. Detailed information on the dataset is provided in Supplementary Methods—Training, internal and external validation. The standard of truth for each cohort is provided in Supplementary Table 3.
Training data
Internal training cohort
The internal training cohort contains 6,813 NC CT scans, including 3,744 esophageal malignant patients and 3,069 negative controls. Among these, 2,716 esophageal malignant cases and 1,539 negative controls were recruited from SYSUCC (center A), while 1,028 esophageal malignancy cases and 1,530 negative controls were sourced from SCCH (center B).
Annotation protocol
Early-stage esophageal malignant lesions are often extremely small, and the esophagus itself is a long, narrow tubular organ that frequently undergoes physiological collapse and is subject to motion from the heart and great vessels. These factors make subtle lesions difficult to distinguish from normal tissues on CT, even on contrast-enhanced scans, posing substantial challenges for accurate lesion identification and annotation. To address these challenges, we established a rigorous annotation protocol in which strong reference information was provided to two annotators and one expert reviewer to ensure high-quality data for EAGLE model development (Extended Data Fig. 1). For each patient, paired contrast-enhanced CT and NC CT scans with corresponding surgical pathology reports were available. We selected the most recent CT scan before the first surgery or endoscopy examination. This examination is typically a contrast-enhanced CT, and in Chinese cancer hospitals, the same session also includes an accompanying NC series. Pathology reports were obtained from that surgical or endoscopy episode. Lesion annotations were performed on contrast-enhanced CT volumes due to their superior visualization of lesion boundaries. Two radiologists with ≥3 years’ experience in esophageal imaging independently annotated lesions using standardized platforms. Interannotator agreement was quantified through DSC—masks with DSC >0.6 were combined as gold-standard annotations, while lower agreement cases underwent senior radiologist review (>8 years’ expertise). Reviewers integrated imaging-pathology data to refine annotations, and excluded cases from the training cohort if they could not confidently delineate the lesion. Final contrast-enhanced CT annotations were then precisely transferred to NC CT through image registration45, applying transformation matrices to achieve voxel-wise alignment. This systematic workflow ensured accurate cross-modal ground truth data generation for model training and validation.
Validation cohorts for opportunistic screening in hospital
Internal test cohort
The internal test cohort was used to evaluate the model’s performance within the same institutional framework and comprised 2,500 patients, including 1,281 patients with esophageal malignant lesions, and 1,219 negative controls. Among these, 767 esophageal malignant patients and 715 negative controls were sourced from the SYSUCC; the remaining 223 esophageal malignant patients and 504 negative controls originated from the SCCH. Notably, the internal test cohort also included some challenging cases of early-stage EC or HGIN from the concurrent training cohort for which radiologists could not provide lesion annotations (Extended Data Fig. 1). The inclusion of these cases substantially increased the diagnostic complexity of the test dataset, allowing for a more rigorous validation of the model’s classification performance.
External multicenter test cohort
The external test cohort was used to assess the cross-institutional generalizability and was established from six centers in China, one center in the Czech Republic (General University Hospital in Prague, center I), and one center in Australia (GenesisCare, center J). Among centers in China, four centers were located in the east (SIPD, center C; First Affiliated Hospital of Zhejiang University, center D; Shantou Central Hospital, center F; Fujian Cancer Hospital, center H), one center in the central (HCH, center E) and one center in the northwest (XMUACH, center G). The inclusion criteria were as follows: NC CT scans must fully cover the chest region, encompassing the esophagus and EGJ regions. Positive cases were confirmed through surgical or biopsy histopathology, while negative controls were confirmed based on at least 2 years of clinical follow-up. In total, the external test cohort comprised 2,812 patients with esophageal malignancy and 8,654 negative controls.
Model calibration and real-world retrospective validation cohort
The real-world study collected NC CT scans with continuous time series data from SYSUCC, SCCH and SIPD, consisting of 35,402 patients. For the real-world cohorts, cases of esophageal malignancy were confirmed through surgical or biopsy histopathology, while negative controls were determined by clinical follow-up. The dataset from each center was divided into two distinct continuous series—one for model calibration (RW1) and the other for model validation (RW2). The overall RW1 comprised 20,758 patients for model calibration, and the refined model was subsequently evaluated on RW2, comprising 14,644 patients. The standard of truth was defined by two time points for each patient—the initial SOC, which was the diagnosis at the first visit when the NC CT was acquired, and the follow-up SOC, which was the diagnosis during subsequent follow-up before model evaluation.
Prospective validation cohort
A total of 17,446 patients were enrolled from January 2025 to April 2025 at SIPD. This study aimed to validate the clinical reliability of the calibrated EAGLE-Plus model for opportunistic screening in real-world clinical settings, particularly uncontrolled, high-throughput clinical environments. Esophageal malignancy was confirmed by biopsy or surgical pathology. Follow-up of positive predictions was completed on 31 July 2026.
Validation cohorts for opportunistic screening in established LDCT programs
Two LDCT validation cohorts were established to evaluate the model’s accuracy and clinical utility in LDCT-based physical examination and lung-cancer screening programs.
LDCT test cohort
This cohort enrolled 147 patients with esophageal malignancy and 1,460 negative controls recruited from SCCH and SIPD, which were used to assess the generalizability of the EAGLE on LDCT. Specifically, from SCCH, we retrospectively collected 1,464 individuals who had undergone both chest LDCT and endoscopy as part of routine physical examinations, of whom 4 were confirmed to have esophageal malignancy by biopsy histopathology. To enhance the estimation of LDCT sensitivity, we additionally retrieved 85 patients with esophageal malignancy confirmed at SIPD, all of whom had undergone chest LDCT within 1 month before confirmation.
Real-world retrospective LDCT validation cohort
This cohort enrolled 10,959 consecutive individuals between July and December 2024 at SIPD aged 45–75 years. The original purpose of those CT scans was for a physical examination. This validation was designed to derive specificity thresholds under real-world healthcare conditions and estimate the specificity of esophageal malignancy in the LDCT-based program. The standard of truth for negative controls was established by first identifying cases with radiology reports showing no esophageal abnormalities, then including only those with follow-up records and no clinical diagnosis of EC.
Validation cohorts for population-based endoscopic screening
External test cohort for endoscopic screening scenario
This cohort included 552 patients with esophageal malignancy and 150 negative controls, recruited from SNCH (center K) and LCH (center L). All patients underwent both NC CT scans and endoscopic examinations. Positive cases were confirmed through surgical or biopsy histopathology; negative controls were confirmed through endoscopic results within 1 year. The cohort was designed to evaluate EAGLE’s cross-institutional generalizability in high-risk populations and ability to differentiate esophageal malignancy from benign esophageal conditions.
Prospective enrolled cohort from endoscopic screening program
This cohort was completed and registered at http://www.chictr.org.cn (Chictr.org.cn identifier: ChiCTR2300074806), and prospectively enrolled 530 patients between 7 and 26 July 2023 at SNCH. All patients were initially referred for endoscopic examination and provided written informed consent to undergo an additional NC CT scan (Supplementary Fig. 7). The cohort was specifically established to evaluate EAGLE’s potential clinical utility in real-world endoscopic screening workflows for high-risk individuals, focusing on its role as a risk-stratification tool for endoscopy.
Model: EAGLE
EAGLE is a two-stage framework (Extended Data Fig. 2a). The first stage focuses on esophageal localization by using a segmentation UNet. The second stage combines classification and segmentation using a joint UNet-based model, which takes the cropped region of interest as input to simultaneously predict esophageal and malignant lesion segmentation masks, diagnostic outcomes (positive or negative) and class activation mapping heatmaps46 that highlight regions driving diagnostic decisions.
Esophageal localization
Stage 1 focuses on identifying the esophageal region. Given that esophageal lesions typically occupy a small volume in CT scans, isolating the esophagus expedites lesion detection and eliminates nontarget anatomical structures to enhance region-specific analysis. Here an nnU-Net V2 framework47 was trained to segment the entire esophagus, including both healthy tissue and pathological anomalies—from NC CT inputs. Supervised learning uses voxel-level annotations of both esophageal boundaries and lesions. A 3D low-resolution architecture was adopted for computational efficiency, leveraging a downsampled UNet variant and a ResEnc backbone within the nnU-Net V2 framework48.
Malignant lesion detection
Stage 2 aims to identify the malignant lesion. Given a 3D region-of-interest CT, we feed it into a 3D UNet-like backbone and obtain multiscale feature maps from the image decoder. Then, we standardize the spatial dimensions of all feature maps from the decoders to a uniform spatial size using trilinear interpolation. Subsequently, we concatenate these maps by aligning them along their depth dimension, forming an aggregated feature representation that serves as the input for the subsequent cancer screening task. The segmentation head takes the last feature map as input and outputs the malignant lesion segmentation result. The cancer screening head takes the aggregated feature map as input and outputs the probability of the patient having cancer. It consists of one 3D convolution layer with kernel size 3 × 3 × 3, one global average pooling layer and one fully connected layer, as illustrated in Extended Data Fig. 2b. The overall loss of stage 2 is:
$${{\mathscr{ \mathcal L }}}_{\mathrm{EAGLE}}={{\mathscr{ \mathcal L }}}_{\mathrm{seg}}+\alpha {{\mathscr{ \mathcal L }}}_{\mathrm{cls}}$$
(1)
where \({{\mathscr{ \mathcal L }}}_{\mathrm{seg}}\) is the loss for the 3D UNet segmentation network; \({{\mathscr{ \mathcal L }}}_{\mathrm{cls}}\) is cross-entropy loss for the cancer screening task and α is its weight.
Interpretability
Our model’s interpretability is achieved through dual mechanisms (Extended Data Fig. 2c). First, stage 2 of EAGLE generates precise segmentation masks delineating detected malignant lesions. Second, we implemented class activation mapping46 to visualize diagnostic heatmaps derived from the classification head’s convolutional feature maps. These heatmaps highlight anatomical regions that influenced the model’s diagnostic decisions, providing spatial justification for classification outcomes.
Operation point and threshold selection
Because the model outputs continuous risk scores, the operating point can be adjusted by varying the decision threshold. For both EAGLE and EAGLE-Plus, thresholds were determined during fivefold cross-validation and were fixed after training. In each fold, the threshold was first selected on the corresponding validation split according to the predefined operating criterion, and the final threshold was fixed as the mean value across the five folds after training. In the opportunistic screening setting, the threshold for each fold was selected to achieve a specificity of 99% on its respective validation split. The final fixed thresholds were 0.93 for EAGLE and 0.88 for EAGLE-Plus. Conversely, in the population-based screening setting, EAGLE was evaluated as a risk-stratification tool for endoscopy EC screening procedures like narrow band imaging endoscopy or Lugol’s chromoendoscopy49,50, and thresholds were calibrated to achieve a sensitivity of 98% for each fold to prioritize sensitivity and minimize the risk of missed cancers. Notably, EAGLE’s operating point can be flexibly adjusted according to application requirements; specifically, a lower threshold increases sensitivity while decreasing specificity.
Performance comparison
As illustrated in Extended Data Table 3, we first compared five state-of-the-art backbone architectures integrated into EAGLE’s implementation leveraging the nnU-Net framework47, including convolutional neural network-based backbone ResNet encoder (ResEnc48), two Transformer-based backbones Swin-UNETR51 and Mask2Former52, and two Mamba-based architectures53. In the internal validation cohort, ResEnc demonstrated superior performance. In the eight-center external validation cohort, although ResEnc’s AUC was marginally lower (0.976 versus 0.979), it achieved the highest sensitivity (89.5%) and specificity (98.5%), outperforming other models. Given its simplicity (pure convolutional neural network design) and robust generalization across multicenter data, ResEnc48 was selected as the default backbone for EAGLE. We further compare EAGLE (ResEnc) with another widely adopted framework, nnDetection54, which demonstrates markedly inferior performance with an AUC of 0.954 in the eight-center external validation cohort. Detailed descriptions of the compared methods are provided in Supplementary Methods—Comparison to other methods.
LDCT simulation strategy
To enhance the model’s performance on LDCT scans, we developed EAGLE with LDCT simulation. A low-dose simulation tool31 was used to generate synthetic LDCT data from the NC CT scans in the training set, which were then used for model training. Validation on real-world LDCT scans demonstrates that this strategy consistently improves sensitivity while maintaining the same level of specificity.
Reader study
This study establishes a multicenter (n = 4) reader study cohort to conduct blinded comparisons between EAGLE predictions and interpretations by experienced radiologists, thereby validating the potential clinical integration of AI-driven workflows. This reader study comprised 300 patients collected from four centers, including two internal centers (SYSUCC and SCCH) and two external centers (HCH and XMUACH), including 200 patients diagnosed with esophageal malignancy and 100 negative controls. NC CT scans from additional external centers were not included in this analysis because stage information or imaging data were unavailable from those institutions until the initiation of the first-round reader study.
A total of 17 readers from three institutions participated in this study, comprising 3 subspecialized esophageal radiologists, 6 general radiologists and 8 radiology residents. The cohort demonstrated a mean clinical experience of 7.6 years (range = 3–20) in radiology practice, with radiologists having interpreted an average of 1,141 esophageal CT scans (range = 100–1,000) before the year preceding the study (Supplementary Table 1).
This reader study comprised two sequential sessions designed to evaluate AI-assisted diagnostic performance. The initial phase established baseline interpretation accuracy without AI support, followed by an AI-assisted evaluation phase after at least a 3-month washout period to mitigate learning effects. During both sessions, readers independently interpreted cases presented in randomized order through our customized online platform (Supplementary Fig. 2), blinded to clinical data and required to classify each case as esophageal malignant or not.
Opportunistic screening: real-world retrospective and prospective validations
Clinical applications of opportunistic screening impose rigorous requirements on the model’s FP rate. To address this, we calibrated the model using large-scale real-world retrospective cohorts and subsequently assessed its performance changes in independent real-world cohorts, yielding the optimized EAGLE-Plus framework. The overall real-world model calibration and validation collected 35,402 NC CT scans from SYSUCC, SCCH and SIPD. Data were collected from four scenarios, that is, physical examination, emergency, outpatient and inpatient. Each center contributed two distinct continuous data series.
Iterative training of EAGLE-Plus
Initially, the EAGLE model was evaluated using the first data series of each center (RW1, n = 20,758). Subsequently, the model was calibrated using hard cases from RW1 and four internal and external test cohorts to fine-tune the EAGLE model. Finally, these additional training samples were annotated by the same expert radiologist to ensure consistency, and were merged with the original training data for further fine-tuning of EAGLE, with an additional 200 epochs using the same training pipeline. The upgraded EAGLE-Plus model was evaluated on the second real-world continuous data series from each center (RW2, n = 14,644). Performance on RW2 reflects its effectiveness in uncontrolled, real-world opportunistic screening settings, particularly its stable sensitivity and low FP rate.
Hard cases were selected as follows. First, all FP cases identified by the MDT were categorized as hard negatives. Second, hard positives were defined by the following two components: (1) all false-negative cases; and (2) borderline positives (that is, true-positive cases with scores just above the threshold). The latter were included to compensate for the lower frequency of false negatives and to maintain an overall ratio of hard negatives to hard positives at approximately 1:1. Detailed results are provided in Supplementary Methods—Details of real-world model calibration and validation.
Prospective validation in hospital
We conducted a prospective validation to evaluate the real-world clinical implementation of the EAGLE-Plus model in clinical settings. The study enrolled patients from 1 January 2025 to 30 April 2025, with a final data analysis conducted after a follow-up period ending on 31 July 2026. EAGLE-Plus was seamlessly integrated into the DAMO MED system, where it prospectively analyzed routine NC CT scans and flagged high-risk individuals for further review.
The primary endpoint was PPV among AI-positive cases. Secondary endpoints were sensitivity for HGIN and EC and the AI-triggered referral rate. PPV was defined as the proportion of confirmed true-positive cases among all AI-positive predictions. Sensitivity was defined as the proportion of confirmed HGIN and EC cases identified during the study and follow-up period that had been classified as AI-positive. The AI-triggered referral rate was defined as the proportion of screened individuals referred for further clinical evaluation on the basis of AI findings after MDT review. Study oversight was provided by the principal investigator and the MDT, which included radiologists, endoscopists and thoracic surgeons. The SOC pathway consisted of routine double reading by radiologists blinded to AI outputs and follow-up outcomes. AI-positive cases were independently reviewed by the MDT, which intervened only when suspicious findings had not been captured by the initial SOC pathway. No feedback from the AI system or MDT review was provided to SOC radiologists during image interpretation. When further clinical action was considered necessary, the MDT communicated directly with the treating clinicians or patients.
Esophageal malignancy was confirmed by histopathological examination of endoscopic biopsy or surgical specimens. For cases without pathological confirmation, outcomes were ascertained by endoscopy, electronic medical record review or telephone follow-up. Among individuals with negative findings by both AI and SOC, the MDT randomly reviewed approximately 1% of cases using available electronic medical records through 31 July 2026. Detailed methodology, results, workflow and follow-up procedures are provided in Extended Data Table 4 and Supplementary Methods—Details of prospective validation in hospital.
Real-world validation in LDCT program
This study aims to evaluate the performance of EAGLE-Plus in the large-scale LDCT real-world cohort. Individuals aged 45–75 who underwent LDCT for routine health examination program at SIPD between 1 July 2024 and 31 December 2024 were eligible. A total of 10,959 participants were included. Diagnostic labels were confirmed for all cases flagged by either SOC or EAGLE-Plus-positive, based on initial radiology reports, clinical records and telephone follow-up through 31 July 2026. For participants with concordant negative findings by both SOC and AI, an additional team randomly selected approximately 1% for review to confirm the absence of malignancy, provided their clinical records were available through 31 July 2026. Further details are provided in Supplementary Methods—Details of real-world retrospective LDCT validation.
Population-based endoscopic screening: EAGLE for pre-endoscopy risk stratification
Suining is a high-incidence area for EC. As part of the National Screening Program for upper gastrointestinal cancers, SNCH provides free endoscopic screening to residents aged 40–69 with high-risk factors. Only individuals who provided informed consent forms for NC CT were included in this study (Supplementary Fig. 7). Between 7 July 2023 and 26 July 2023, 530 participants were prospectively enrolled. Each participant first underwent NC CT, followed by standard upper gastrointestinal endoscopy with iodine staining (1.2% Lugol’s solution). All focal lesions detected during endoscopy were biopsied, and two board-certified pathologists independently evaluated the specimens using established diagnostic criteria55.
In the prospectively enrolled cohort, six cases of HGIN were enrolled, but no invasive cancers were found. Because the low number of malignant events limited direct evaluation of EAGLE as a risk-stratification tool, we performed an exploratory simulation study using a bootstrapping approach. In each iteration, and based on the reported 2:1 incidence ratio between HGIN and EC reported in high-risk populations undergoing endoscopic screening2, three cancer cases were randomly sampled with replacement from the 420 EC cases of the SNCH retrospective test cohort. These sampled EC cases were combined with the prospectively enrolled cohort to construct a hybrid dataset for analysis. For each simulated hybrid cohort, we estimated the detection rate achievable when using EAGLE to triage only high-risk patients for endoscopic examination. This process was repeated 1,000 times to reflect the expected disease prevalence and variability in real-world screening settings. Detailed methodology and results were provided in Supplementary Methods—Details of validation on prospectively enrolled endoscopic screening cohort.
Statistical analysis
All statistical analyses were performed in Python (v3.9.19) using scikit-learn (v1.4.2), NumPy (v1.26.4) and SciPy (v1.13.0). Binary classification performance was evaluated using the AUC, sensitivity and specificity. PPV was used to assess the precision of the model’s positive predictions in clinical practice. Balanced accuracy was additionally used in the reader study to account for class imbalance. Unless otherwise stated, each metric is reported as a point estimate with a 95% CI. Bootstrap resampling with 1,000 iterations was used to estimate 95% CIs. In each iteration, patients were resampled with replacement from the cohort under analysis and the metric was recalculated; the 95% CI was defined by the 2.5th and 97.5th percentiles of the bootstrap distribution. AUCs of two models evaluated on the same patients were compared using the DeLong test for correlated receiver operating characteristic curves. Permutation tests were used for statistical comparisons of sensitivity, specificity, stage-specific sensitivity and balanced accuracy, with at least 1,000 random permutations. The McNemar test was used to compare paired binary outcomes, specifically the esophageal malignancy detection rate between the EAGLE risk-stratified pathway and direct endoscopy in the population-based screening cohort. Comparisons between models or between readers were performed on the same set of patients and were therefore treated as paired, repeated measurements, whereas all other analyses were based on measurements from distinct, independent patients. The bootstrap and permutation procedures are nonparametric and do not assume normality of the underlying data distributions. All tests were two sided, no adjustment was made for multiple comparisons and a P value of <0.05 was considered statistically significant. Exact P values are reported in the Figs. 1c, 2f,g, 3c,d, 4a and 5e, and P values below 0.001 are reported as P < 0.001. A fixed random seed of 65,535 was used throughout.
Reporting summary
Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.

