A deep learning-clinical combined model with SHapley Additive exPlanations (SHAP) method for assessing the tumor spread through air spaces in lung adenocarcinoma: a multicohort retrospective study
Original Article

A deep learning-clinical combined model with SHapley Additive exPlanations (SHAP) method for assessing the tumor spread through air spaces in lung adenocarcinoma: a multicohort retrospective study

Rong Zhang1#, Jing-Song Sun2#, Zi-Wei Liu1#, Yun-Ying Lin1, Xiao-Yan Cai3, Li-Jun Kang1, Bao-Liang Guo1, Ya-Xuan Song4, Chen-Zhengren Bao1, Kui Chen5, Yi Sun6, Yi-Fan Chen1, Zhi-Ping Cai1, Han Xie2, Hai-Xiong Chen1*, Shao-Min Yang7*, Qiu-Gen Hu1*

1Department of Radiology, The Eighth Affiliated Hospital of Southern Medical University (The First People’s Hospital of Shunde), Foshan, China; 2Department of Radiology, Lecong Hospital of Shunde, Foshan, China; 3Department of Science and Education, The Eighth Affiliated Hospital of Southern Medical University (The First People’s Hospital of Shunde), Foshan, China; 4School of Medicine, South China University of Technology, Guangzhou, China; 5Department of Radiology, The Third Xiangya Hospital of Central South University, Changsha, China; 6Department of Radiology, Chenzhou First People’s Hospital, Chenzhou, China; 7Department of Radiology, Xingtan Hospital Affiliated to Shunde Hospital of Southern Medical University, Foshan, China

Contributions: (I) Conception and design: R Zhang, JS Sun, YY Lin, ZW Liu, XY Cai, HX Chen, SM Yang, QG Hu; (II) Administrative support: R Zhang, JS Sun, YY Lin, ZW Liu, XY Cai, HX Chen, SM Yang, QG Hu; (III) Provision of study materials or patients: R Zhang, ZW Liu, K Chen, Y Sun; (IV) Collection and assembly of data: R Zhang, ZW Liu, CZ Bao, K Chen, Y Sun, YF Chen, ZP Cai; (V) Data analysis and interpretation: H Xie, CZ Bao, YF Chen, ZP Cai; (VI) Manuscript writing: All authors; (VII) Final approval of manuscript: All authors.

#These authors contributed equally to this work as co-first authors.

*These authors contributed equally to this work.

Correspondence to: Hai-Xiong Chen, MD, PhD. Department of Radiology, The Eighth Affiliated Hospital of Southern Medical University (The First People’s Hospital of Shunde), No. 1 Jiazi Road, Lunjiao, Shunde District, Foshan 528308, China. Email: 13825553451@139.com; Shao-Min Yang, MD, PhD. Department of Radiology, Xingtan Hospital Affiliated to Shunde Hospital of Southern Medical University, 338 Nanguo West Road, Xingtan Town, Shunde District, Foshan 528325, China. Email: ysmsin@aliyun.com; Qiu-Gen Hu, MD, PhD. Department of Radiology, The Eighth Affiliated Hospital of Southern Medical University (The First People’s Hospital of Shunde), No. 1 Jiazi Road, Lunjiao, Shunde District, Foshan 528308, China. Email: hu6009@163.com.

Background: Predicting the tumor spread through air spaces (STAS) pattern in patients with lung adenocarcinoma (LUAD) is important for timely preventive intervention, selection of appropriate treatment, and improvement in their quality of life. Therefore, it is critical to develop a machine learning (ML) model that can evaluate STAS in LUAD based on longitudinal study data. This study aims to develop and validate a computed tomography (CT) based deep learning model for preoperative prediction of STAS in lung adenocarcinoma.

Methods: In total, 689 patients diagnosed with LUAD after computed tomography (CT) and surgery at four hospitals from January 2019 to December 2023 were included and divided into the training, internal validation, external validation I, and external validation II cohorts. A deep learning (DL) radiomics score (radscore) was developed based on the ResNet-101 framework using CT images of LUAD. Seven ML algorithms [logistic regression (LR), extreme gradient boosting (XGBoost), support vector machine (SVM), random forest (RF), decision tree (DT), k-nearest neighbor (KNN), and artificial neural networks (ANN)] were used to distinguish between STAS and non-STAS. SHapley Additive exPlanations (SHAP) was used for individualized and visual interpretations. SHAP addressed the cognitive opacity of the ML models.

Results: When comparing seven different ML models in the training cohort, the optimal combined model showed the best performance on various evaluation measures, with area under the receiver operating characteristic (ROC) curve (AUC) values of 0.906 [95% confidence interval (CI): 0.873–0.936] and 0.903 (95% CI: 0.855–0.944) in the training cohort and internal validation cohort, respectively. Based on the importance of the characteristics identified by the model interpretation method (SHAP), the important characteristics were the DL-score, lobulation, vacuole sign, and microvascular sign. In addition, SHAP summary diagrams were used to illustrate the positive and negative effects of the features influenced by the combined model. The SHAP dependency diagrams explain how individual features affect the output of a predictive model. Satisfactory generalization performance was shown with AUCs of 0.841 (95% CI: 0.757–0.923) and 0.882 (95% CI: 0.806–0.953) in the two external validation cohorts, respectively.

Conclusions: A combined model based on the DL-score and clinically independent risk factors can accurately evaluate disease STAS in patients with LUAD. Additionally, an interpretable framework can increase the transparency of the model, provide clear explanations for personalized risk prediction, and offer a more intuitive understanding of the effects of the key features in the model.

Keywords: Lung adenocarcinoma (LUAD); spread through air spaces (STAS); artificial neural networks (ANN); deep learning (DL); SHapley Additive exPlanations (SHAP)


Submitted Apr 14, 2025. Accepted for publication Jul 15, 2025. Published online Sep 11, 2025.

doi: 10.21037/qims-2025-887


Introduction

Non-small cell lung cancer (NSCLC) is the leading cause of cancer-related death and seriously threatens the health and safety of people worldwide (1). Lung adenocarcinoma (LUAD) is the most common histological subtype of lung cancer, and surgical treatment is the most important treatment for patients with LUAD. However, the long-term recurrence rate remains high in patients with LUAD, even after surgery (2,3). One reason for this may be tumor metastasis through the spread through air spaces (STAS) pathway (4), which is a newly discovered mode in recent years in addition to direct invasion, lymphatic metastasis, and blood metastasis (3). A large number of clinical studies have confirmed that the presence of STAS indicates that the tumor biology is more aggressive, the risk of tumor recurrence and lymph node metastasis is higher, and the survival rate of patients is lower (3-6). Therefore, the assessment of STAS in patients with LUAD can influence clinical decisions, such as the choice of surgical modalities, the degree of lymph node dissection, and the need for postoperative chemotherapy. Currently, the evaluation of STAS is based on histopathological analyses. However, owing to the diversity of STAS pathology, inconsistent standards, and intraoperative impact on specimens, the sensitivity of intraoperative frozen sections for STAS detection is too low to provide reliable guidance for clinical, especially surgical, decision-making (7,8). Therefore, the accurate evaluation of STAS in patients with LUAD before surgery is of importance for the guidance of auxiliary clinical surgery, the degree of lymph node dissection, and postoperative management.

Previous studies have suggested that STAS can be predicted by computed tomography (CT) imaging, which is associated with solid nodules, central low attenuation, ill-defined opacity, air bronchogram, and high consolidation/tumor ratio (CTR) (9,10). However, there are many limitations in traditional imaging diagnostic indicators and signs, such as differences in scanning technical parameters (CT thickness), inconsistencies in STAS imaging standards, and subjective errors in viewer judgment, which lead to insufficient accuracy in the evaluation of image morphological models. Several studies highlight that the development of deep learning (DL) technology provides a tool for quantitative deep mining of image information (11-13). Lin et al. (14) developed a DL model based on CT images for predicting STAS of ground glass-predominant LUAD, with an area under the receiver operating characteristic (ROC) curve (AUC) of 0.82 and an accuracy of 74%. Wang et al. (15) built a novel, fully automated artificial intelligence system (FAIS) based on CT images to predict epidermal growth factor receptor (EGFR) genotype and targeted therapy, and a convenient method that was validated in a large cohort to assist targeted therapy planning. These studies confirmed the feasibility of predicting STAS expression in LUAD by CT DL. However, the “black box” model of image-omics machine learning (ML) affects doctors’ trust in its results (16), and the internal decision-making mechanism and deduction process of the prediction model were not clear, limiting the popularization and application of the model (17-19). The SHapley Additive exPlanation (SHAP) concept was introduced to solve the inexplicability bug. SHAP is a local interpretation method derived from game theory that uses a predictive model based on a combination of all possible feature subsets containing a particular feature, quantifying the contribution of each feature to the predictive model, thereby demonstrating the decision-making process for each case (20,21). SHAP was successfully used to assess the therapeutic effect of whole brain radiotherapy (22), prognosis of coronavirus disease of 2019 (COVID-19) (23), and cardiac surgery-associated acute kidney injury (24).

However, no study has developed explainable DL models targeting the prediction of STAS in LUAD. Based on a multicenter cohort, this study aimed to construct and validate a DL combined model for the accurate prediction of STAS in LUAD using CT images and SHAP to complete individualized visual interpretation. The objective was for the model to improve the acceptability and practicability of clinical decision support tools obtained by clinicians, provide a theoretical basis and direction for personalized and precise treatment of patients with LUAD, and reduce overtreatment. We present this article in accordance with the TRIPOD+AI reporting checklist (available at https://qims.amegroups.com/article/view/10.21037/qims-2025-887/rc).


Methods

Patients and data collection

This retrospective study was conducted according to the Declaration of Helsinki and its subsequent amendments. Ethical approval was obtained from the Ethics Committee of The Eighth Affiliated Hospital of Southern Medical University (The First People’s Hospital of Shunde) (No. KYLS20231203). The institutional review board waived the requirement for informed consent due to the retrospective design of the study. All participating hospitals/institutions were informed and agreed with the study. Patients who underwent surgical resection for LUAD at The Eighth Affiliated Hospital of Southern Medical University (The First People’s Hospital of Shunde) (hospital I), Xingtan Hospital Affiliated to Shunde Hospital of Southern Medical University (hospital II), The Third Xiangya Hospital of Central South University (hospital III), and Chenzhou First People’s Hospital (hospital IV) from 1 January 2019 to 31 December 2023 were included. All participating hospitals agreed to participate in the study. Figure 1 shows the inclusion and exclusion criteria and process in detail.

Figure 1 Flowchart of inclusion and exclusion criteria. Hospital I: The Eighth Affiliated Hospital of Southern Medical University (The First People’s Hospital of Shunde); hospital II: Xingtan Hospital Affiliated to Shunde Hospital of Southern Medical University; hospital III: The Third Xiangya Hospital of Central South University; hospital IV: Chenzhou First People’s Hospital. CT, computed tomography; STAS, tumor spread through air spaces.

To conduct the study, patients from hospitals I and II were randomly stratified into a training cohort and an internal validation cohort at a ratio of 7:3, whereas those from hospitals III and IV were considered the two external validation cohorts.

Histological evaluation of LUAD

All patients with LUAD enrolled in this study were divided into different groups based on the lung cancer pathologic classification defined by the World Health Organization (WHO) and the International Association for the Study of Lung Adenocarcinoma. Staging was performed according to the eighth edition of the International Association for the Study of Lung Cancer (IASLC) international tumor, node, metastasis (TNM) staging standard for lung cancer (25) and the pathological excision specimens were obtained. The excised specimen was fixed with 10% formalin, placed on a paraffin block, sliced into 5-µm sections, and stained with hematoxylin and eosin. STAS was considered present when the tumor cells were detached from the main mass and identified within the alveolar spaces beyond the circumferential edge of the main tumor. All pathological tissue slides were initially evaluated by mid-level pathologists and subsequently reviewed by senior pathologists.

CT image acquisition and tumor annotation

CT scans were performed on each patient using a multi-slice spiral CT device. The parameters of CT scanners vary across different centers. To mitigate the center-related variability of CT images obtained from diverse scanners and hospitals, all original CT images underwent appropriate pre-processing. Initially, the images were resampled to a voxel size of 1×1×1 mm3 (x, y, z) through the application of a linear interpolation algorithm. This step was taken to standardize the voxel spacing, ensuring consistency in the spatial representation of the images. Subsequently, a bin width of 25 Hounsfield units (HU) was established to discretize the voxel intensity. This not only helped in reducing the inherent noise within the images but also enhanced the overall quality and interpretability of the CT scans. Table S1 provides the details of the imaging protocols used in each hospital.

Preoperative CT images were retrieved from the Picture Archiving and Communication Systems (PACS) of the four hospitals. CT images were then imported into ITK-SNAP software (https://keyan.ITK-SNAP.com/login) for annotation. The region of interest (ROI) was manually delineated using a bounding box that included the entire tumor volume. Three radiologists with 8–10 years of experience independently performed tumor annotations in the lung window setting (mean, −500 HU; width, 1,500 HU), and existing divergence were resolved by consulting a senior radiologist with more than 15 years of experience. Based on the information for each patient, we assessed the baseline clinical-radiological factors of LUAD, such as age, sex, tumor diameter, primary tumor site, tumor-lung boundary, lobulation, tumor type, spiculation, pleural retraction, air bronchogram, vacuole sign, crescent sign, microvascular sign, CT value, and CTR; both radiologists were blinded to the presence or absence of STAS.

Development of a DL model

The transfer learning model used in this study was ResNet-101, and the initial weight values were pretrained on ImageNet. For model training, the slices were resized to 224×224, including the largest tumor area and adjacent layers of each CT image, which were selected and assembled. Subsequently, z-score normalization was performed on all the images. Furthermore, data augmentation strategies were applied, namely random horizontal and vertical flipping and random cropping. With an initial learning rate of 0.0005, we utilized AdamW to update the model parameters. The epochs were set to 100, batch size was set to 32, loss function was CrossEntropyLoss, and learning rate scheduler was Cosine Annealing. The output of the ResNet-101 model was used as the DL-score for further analysis. The DL analysis process is illustrated in Figure 2. Gradient-weighted class activation mapping (Grad-CAM) was applied to make the model’s decision-making process more transparent and investigate its interpretability. We used the gradient information of the last convolutional layer of ResNet-101 for weighted fusion to obtain a class activation map that highlighted the important regions of the classification target image.

Figure 2 Overall workflow in this study. (I) Data acquisition and image preprocessing. (II) Construction of STAS prediction model for pulmonary tumors. (III) Model performance analysis. (IV) Multicenter-model validation. (V) Interpretability analysis of deep learning. (VI) SHAP interpretive analysis. DL, deep learning; SHAP, SHapley Additive exPlanations; STAS, spread through air spaces.

Independent risk factors selection and construction of a clinical model

Underlying clinical information, including age and sex, was obtained from the medical records. To construct a clinical model, we used univariate logistic regression (LR) analysis to identify independent risk factors that were significantly correlated with STAS (P<0.05). Subsequently, multivariate LR was used to select factors with a high correlation, followed by stepwise LR to construct a clinical model based on the independent risk factors.

Construction of multi-ML combined models

Using seven ML classifiers, namely, LR, extreme gradient boosting (XGBoost), support vector machine (SVM), random forest (RF), decision tree (DT), k-nearest neighbor (KNN), and artificial neural networks (ANN), we integrated prominent clinical-radiology factors and the DL-score. In the model development and validation stages, we first determined the optimal hyperparameters of the models: the activation function was rectified linear unit (ReLU), the number of epochs was 300, the learning rate was 0.0015, the number of hidden neurons was 24, and the loss function was CrossEntropyLoss. A brief description of these ML classifiers is provided in Appendix 1.

To verify the generalization performance of the optimal model and study the overfitting effect, all ML models were tested in two independent external validation cohorts. These validation cohorts were drawn from different hospitals (III and IV). Subsequently, we evaluated the diagnostic performance of all ML classifiers and selected the classifier with the best performance as our final combined model to predict STAS in patients with adenocarcinoma. The predictive performance of the optimal model was evaluated by calculating the AUC, accuracy, sensitivity, and specificity.

The interpretability of the optimal model

To improve the interpretability of the combined model, we calculated the SHAP value, which explains the positive and negative contributions of each signature during the model construction. The contribution of each feature to the optimal model was allocated based on their contribution, and the SHAP values were generated based on the axioms. To interpret the combined model with the best performance, we used the SHAP analysis to quantitatively explain the contribution composition of the combined model and visually determine the effect of each feature for each patient.

The SHAP value is a decomposition algorithm that objectively allocates the final result (prediction result) to a feature. When interpreting the model, the SHAP value can be understood as the importance of the contributions of individual input characteristics to the model prediction value. The higher the SHAP value, the greater is the contribution of the input features to the prediction results. The SHAP values can be sampled to estimate the contribution of each feature to the prediction, which shows how they are fairly distributed among the features. The visualization was such that attribution of features such as SHAP values can be visualized as “forces”, where each feature value was a force that increased or decreased the prediction. The predictions started from a baseline, and the baseline SHAP value was the average of all predictions. The size of the arrow indicates the contribution of the feature to the SHAP value. The red and blue arrows indicate positive and negative values, respectively.

Statistical analysis

The parametric Student’s t-test and nonparametric Mann-Whitney U test were used to compare the differences between the positive and negative STAS groups, and the chi-squared test was performed for categorical variables. Continuous and categorical variables are presented as medians (interquartile ranges) and frequencies (percentages), respectively. The performance of the models was compared using the AUC, accuracy, sensitivity, and specificity, and the best cutoff value was determined using the Youden index. The “SHAP” algorithm was implemented using SHAP Python packages. All the statistical analyses were performed using Python (version 3.7.3; Python Software Foundation, Wilmington, DE, USA) and R (version 4.1.3; R Foundation for Statistical Computing, Vienna, Austria). Statistical significance was defined as a two-sided P value <0.05.


Results

Clinical characteristics

The baseline clinical and radiological information of the patients is summarized in Table 1. In total, 689 patients with LUAD were included in this study. Patients of hospital I and hospital II were integrated and randomly divided into a training cohort (n=351, mean age, 62.62±11.38 years) and an internal validation cohort (n=151, mean age, 61.58±11.65 years) at a ratio of 7:3. The two external validation cohorts consisted of 91 patients (mean age, 61.87±10.71 years) from hospital III and 96 patients (mean age, 63.70±11.18 years) from hospital IV. Among the 689 patients, 44.6% (307/689) patients were positive for STAS and 55.4% (382/689) patients were negative for STAS.

Table 1

Baseline characteristics of patients in the total cohort

Items Training cohort (N=351) Internal validation cohort (N=151) External validation cohort I (N=91) External validation cohort II (N=96)
STAS (−) (N=190) STAS (+) (N=161) P value STAS (−) (N=82) STAS (+) (N=69) P value STAS (−) (N=50) STAS (+) (N=41) P value STAS (−) (N=60) STAS (+) (N=36) P value
Gender <0.001 0.678 0.053 0.126
   Male 69 (36.3) 93 (57.8) 33 (40.2) 31 (44.9) 18 (36.0) 24 (58.5) 24 (40.0) 21 (58.3)
   Female 121 (63.7) 68 (42.2) 49 (59.8) 38 (55.1) 32 (64.0) 17 (41.5) 36 (60.0) 15 (41.7)
Age (years) 63.0 (53.0, 71.0) 66.0 (58.0, 71.0) 0.020 61.0 (51.8, 70.3) 63.0 (56.0, 72.0) 0.086 62.0 (56.0, 70.0) 65.0 (57.0, 69.0) 0.696 64.5 (57.0, 72.0) 67.5 (57.5, 72.0) 0.578
Primary site of tumor 0.778 0.957 0.486 0.049
   LLL 22 (11.6) 25 (15.5) 8 (9.8) 9 (13.0) 8 (16.0) 8 (19.5) 9 (15.0) 2 (5.6)
   LUL 50 (26.3) 37 (23.0) 23 (28.0) 21 (30.4) 10 (20.0) 8 (19.5) 10 (16.7) 13 (36.1)
   RLL 42 (22.1) 32 (19.9) 15 (18.3) 11 (15.9) 7 (14.0) 11 (26.8) 11 (18.3) 8 (22.2)
   RML 11 (5.8) 11 (6.8) 8 (9.8) 6 (8.7) 6 (12.0) 4 (9.8) 4 (6.7) 5 (13.9)
   RUL 65 (34.2) 56 (34.8) 28 (34.1) 22 (31.9) 19 (38.0) 10 (24.4) 26 (43.3) 8 (22.2)
Tumor type <0.001 <0.001 <0.001 <0.001
   Solid 39 (20.5) 104 (64.6) 18 (22.0) 33 (47.8) 15 (30.0) 29 (70.7) 8 (13.3) 27 (75.0)
   Nonsolid 68 (35.8) 5 (3.1) 25 (30.5) 4 (5.8) 18 (36.0) 0 (0.0) 23 (38.3) 1 (2.8)
   Part solid 83 (43.7) 52 (32.3) 39 (47.6) 32 (46.4) 17 (34.0) 12 (29.3) 29 (48.3) 8 (22.2)
Lobulation <0.001 <0.001 <0.001 <0.001
   No 145 (76.3) 21 (13.0) 64 (78.0) 16 (23.2) 32 (64.0) 6 (14.6) 35 (58.3) 6 (16.7)
   Yes 45 (23.7) 140 (87.0) 18 (22.0) 53 (76.8) 18 (36.0) 35 (85.4) 25 (41.7) 30 (83.3)
Spiculation 1.000 1.000 1.000 0.011
   No 114 (60.0) 27 (16.8) 45 (55.6) 18 (26.1) 25 (50.0) 11 (26.8) 27 (45.0) 6 (16.7)
   Yes 76 (40.0) 134 (83.2) 36 (44.4) 51 (73.9) 25 (50.0) 30 (73.2) 33 (55.0) 30 (83.3)
Pleural retraction <0.001 0.543 0.027 0.028
   No 86 (45.3) 37 (23.0) 36 (43.9) 26 (37.7) 26 (52.0) 11 (26.8) 20 (33.3) 4 (11.1)
   Yes 104 (54.7) 124 (77.0) 46 (56.1) 43 (62.3) 24 (48.0) 30 (73.2) 40 (66.7) 32 (88.9)
Air bronchogram 1.000 1.000 0.006 <0.001
   No 86 (45.3) 58 (36.0) 36 (43.9) 31 (44.9) 38 (76.0) 20 (48.8) 14 (23.3) 18 (51.4)
   Yes 104 (54.7) 103 (64.0) 46 (56.1) 38 (55.1) 12 (24.0) 21 (51.2) 46 (76.7) 17 (48.6)
Crescent sign 0.031 0.270 0.320 0.052
   No 143 (75.3) 137 (85.1) 66 (80.5) 61 (88.4) 43 (86.0) 31 (75.6) 39 (65.0) 30 (85.7)
   Yes 47 (24.7) 24 (14.9) 16 (19.5) 8 (11.6) 7 (14.0) 10 (24.4) 21 (35.0) 5 (14.3)
Vacuole sign <0.001 <0.001 0.201 <0.001
   No 111 (58.4) 135 (83.9) 40 (48.8) 61 (88.4) 36 (72.0) 35 (85.4) 23 (38.3) 32 (91.4)
   Yes 79 (41.6) 26 (16.1) 42 (51.2) 8 (11.6) 14 (28.0) 6 (14.6) 37 (61.7) 3 (8.6)
Microvascular sign 0.011 0.103 0.012 0.250
   No 12 (6.3) 1 (0.6) 5 (6.1) 0 (0.0) 9 (18.0) 0 (0.0) 7 (11.7) 8 (22.9)
   Yes 178 (93.7) 160 (99.4) 77 (93.9) 69 (100.0) 41 (82.0) 41 (100.0) 53 (88.3) 27 (77.1)
Tumor-lung boundary 0.409 0.959 0.011 0.014
   No 37 (19.5) 25 (15.5) 20 (24.4) 18 (26.1) 25 (50.0) 9 (22.0) 20 (33.3) 3 (8.6)
   Yes 153 (80.5) 136 (84.5) 62 (75.6) 51 (73.9) 25 (50.0) 32 (78.0) 40 (66.7) 32 (91.4)
Tumor diameter (mm) 15.5 (11.2, 21.9) 18.0 (14.0, 26.0) 0.001 14.2 (11.3, 20.7) 15.0 (12.0, 23.0) 0.312 15.0 (10.0, 19.2) 18.0 (11.0, 31.0) 0.017 18.4 (13.0, 24.9) 20.0 (15.0, 27.0) 0.639
CT value −382.3 (−551.6, −187.3) 5.4 (−90.4, 26.7) <0.001 −392.9 (−500.8, −222.8) −29.0 (−240.2, 23.0) <0.001 −330.1 (−570.7, −84.9) 15.4 (−60.3, 40.4) <0.001 −280.4 (−502.7, −17.5) 16.6 (−68.7, 35.2) <0.001
CTR 0.3 (0.0, 0.6) 1.0 (0.7, 1.0) <0.001 0.3 (0.0, 0.8) 0.9 (0.4, 1.0) <0.001 0.4 (0.0, 1.0) 1.0 (0.8, 1.0) <0.001 0.4 (0.0, 0.8) 1.0 (0.9, 1.0) <0.001

Data are presented as median (interquartile range) or n (%). −, negative; +, positive; CT, computed tomography; CTR, consolidation/tumor ratio; LLL, left lower lobe; LUL, left upper lobe; RLL, right lower lobe; RML, right middle lobe; RUL, right upper lobe; STAS, spread through air spaces.

Performance of the DL model

The output of the DL model was the DL-score, which was considered significantly associated with the risk of LUAD and used to distinguish between people at high and low risk. In the training cohort and the internal validation cohort, respectively, the DL-model achieved performance AUCs of 0.861 [95% confidence interval (CI): 0.822–0.897] and 0.735 (95% CI: 0.652–0.815), accuracies of 0.792 and 0.662, sensitivities of 0.851 and 0.681, and specificities of 0.742 and 0.646, respectively. In the external validation cohorts, the DL-model showed excellent diagnostic performance with AUCs of 0.753 [95% confidence interval (CI): 0.652–0.861] and 0.855 (95% CI: 0.770–0.926), accuracies of 0.692 and 0.719, and sensitivities of 0.878 and 0.889, and specificities of 0.540 and 0.617, respectively.

Performance of the clinical model

Clinical and radiological factors, including sex, age, tumor diameter, tumor density, lobulation, pleural retraction, crescent sign, vacuole sign, microvascular sign, CT value, and CTR, were significantly associated with STAS in the univariate logistic analysis (P<0.05). Subsequently, three independent risk factors associated with STAS, namely lobulation sign, vacuole sign, and microvascular sign, were identified using multivariate LR (Table 2). The AUCs of the clinical model in the training cohort and the internal validation cohort, respectively, were 0.859 (95% CI: 0.821–0.893) and 0.837 (95% CI: 0.779–0.887), accuracy 0.818 and 0.781, sensitivity 0.863 and 0.768, and specificity 0.779 and 0.793. In the external validation cohorts I and II, the AUCs of the clinical model were 0.800 (95% CI: 0.713–0.878) and 0.781 (95% CI: 0.682–0.878), accuracy 0.769 and 0.656, sensitivity 0.854 and 0.667, and specificity 0.700 and 0.650, respectively (Table 3).

Table 2

Univariate and multivariate logistic regression analyses for clinical radiological factors

Factors Univariate logistic analysis Multivariate logistic analysis
OR 95% CI P value OR 95% CI P value
Age 1.02 1.01–1.04 0.013 1.00 0.98–1.03 0.845
Tumor diameter 1.04 1.01–1.06 0.004 0.99 0.96–1.03 0.709
Lobulation (yes) 21.48 12.18–37.90 <0.001 13.41 6.64–27.08 <0.001
Spiculation (yes) 7.44 4.49–12.34 <0.001 1.98 0.95–4.16 0.070
Pleural retraction (yes) 2.77 1.74–4.41 <0.001 1.19 0.62–2.29 0.604
Air bronchogram (yes) 1.47 0.96–2.26 0.080
Vacuole sign (yes) 0.27 0.16–0.45 <0.001 0.29 0.15–0.55 <0.001
Microvascular sign (yes) 10.79 1.39–83.76 0.023 15.03 1.61–139.94 0.017

CI, confidence interval; OR, odds ratio.

Table 3

Evaluation of the performance of the seven algorithms

Cohort Models AUC (95% CI) Accuracy Sensitivity Specificity
Training LR 0.926 (0.898–0.951) 0.860 0.783 0.926
XGBoost 0.914 (0.885–0.943) 0.846 0.807 0.879
SVM 0.923 (0.896–0.949) 0.840 0.907 0.784
RF 0.942 (0.918–0.964) 0.877 0.845 0.905
DT 0.949 (0.927–0.968) 0.872 0.907 0.842
KNN 0.925 (0.900–0.950) 0.843 0.789 0.889
ANN 0.906 (0.873–0.936) 0.849 0.839 0.858
Internal validation LR 0.847 (0.787–0.905) 0.775 0.638 0.890
XGBoost 0.867 (0.813–0.916) 0.801 0.710 0.878
SVM 0.851 (0.790–0.906) 0.775 0.783 0.768
RF 0.833 (0.765–0.895) 0.762 0.652 0.854
DT 0.833 (0.770–0.895) 0.768 0.739 0.793
KNN 0.841 (0.778–0.903) 0.788 0.696 0.866
ANN 0.903 (0.855–0.944) 0.821 0.739 0.890
External validation I LR 0.828 (0.740–0.916) 0.791 0.780 0.800
XGBoost 0.827 (0.747–0.907) 0.758 0.780 0.740
SVM 0.828 (0.741–0.915) 0.769 0.878 0.680
RF 0.834 (0.751–0.919) 0.791 0.805 0.780
DT 0.819 (0.730–0.905) 0.769 0.829 0.720
KNN 0.810 (0.720–0.903) 0.758 0.805 0.720
ANN 0.841 (0.757–0.923) 0.791 0.854 0.740
External validation II LR 0.855 (0.770–0.932) 0.812 0.639 0.917
XGBoost 0.849 (0.759–0.936) 0.802 0.667 0.883
SVM 0.829 (0.744–0.912) 0.667 0.722 0.633
RF 0.834 (0.749–0.914) 0.771 0.694 0.817
DT 0.822 (0.725–0.915) 0.771 0.778 0.767
KNN 0.875 (0.799–0.940) 0.823 0.806 0.833
ANN 0.882 (0.806–0.953) 0.833 0.778 0.867

ANN, artificial neural network; AUC, area under the curve; CI, confidence interval; DT, decision tree; KNN, k-nearest neighbor; LR, logistic regression; RF, random forest; SVM, support vector machine; XGBoost, extreme gradient boosting.

Performance of the multi-ML combined models

Based on the DL model and the clinical model, we used seven ML methods to construct combined prediction models from the training cohort. In the model development and validation stages, we first determined the optimal hyperparameters of the ANN model: the activation function was ReLU, epoch =300, learning rate =0.0015, number of hidden neurons =24, and the loss function was the cross-entropy function. The results showed that among the seven models, the ANN model exhibited the best prediction performance and was used as the final combined model. In the training cohort and the internal validation cohort, the AUCs of the ANN model were 0.906 (95% CI: 0.873–0.936) and 0.903 (95% CI: 0.855–0.944), respectively (Table 3, Figure 3A,3B). In the external validation cohorts I and II, a satisfactory generalization performance showed that the AUCs of the ANN model were 0.841 (95% CI: 0.757–0.923) and 0.882 (95% CI: 0.806–0.953), respectively (Table 3, Figure 3C,3D). Therefore, compared with the DL and clinical models, the optimal combined model (ANN model) based on the DL-score and clinically independent risk factors showed the best predictive performance for STAS of LUAD (Table 4, Figure 4).

Figure 3 Comparison of ROC curves for the seven models. (A) ROC curves for the training cohort of the seven models. (B) ROC curves for the internal validation cohort of the seven models. (C) ROC curves for the external validation cohort I of the seven models. (D) ROC curves for external validation cohort II of the seven models. ANN, artificial neural network; AUC, area under the curve; CI, confidence interval; DT, decision tree; KNN, k-nearest neighbor; LR, logistic regression; RF, random forest; ROC, receiver operating characteristic; SVM, support vector machine; XGBoost, extreme gradient boosting.

Table 4

Evaluation of the performance of the clinical, DL, and combined models

Cohort Models AUC (95% CI) Accuracy Sensitivity Specificity
Training Clinical 0.859 (0.821–0.893) 0.818 0.863 0.779
DL 0.861 (0.822–0.897) 0.792 0.851 0.742
Combined 0.906 (0.873–0.936) 0.849 0.839 0.858
Internal validation Clinical 0.837 (0.779–0.887) 0.781 0.768 0.793
DL 0.735 (0.652–0.815) 0.662 0.681 0.646
Combined 0.903 (0.855–0.944) 0.821 0.739 0.890
External validation I Clinical 0.800 (0.713–0.878) 0.769 0.854 0.700
DL 0.753 (0.652–0.861) 0.692 0.878 0.540
Combined 0.841 (0.757–0.923) 0.791 0.854 0.740
External validation II Clinical 0.781 (0.682–0.878) 0.656 0.667 0.650
DL 0.855 (0.770–0.926) 0.719 0.889 0.617
Combined 0.882 (0.806–0.953) 0.833 0.778 0.867

AUC, area under the curve; CI, confidence interval; DL, deep learning.

Figure 4 Comparison of ROC curves for the clinical, DL, and combined models. (A) ROC curves in the training cohort. (B) ROC curves in the internal validation cohort. (C) ROC curves in the external validation cohort I. (D) ROC curves in the external validation cohort II. AUC, area under the curve; CI, confidence interval; Clinic. Model, clinical model; DL. model, deep learning model; ROC, receiver operating characteristic.

The visualization of feature importance and case application analysis

To examine the relationship between the model and predictors, we used the SHAP algorithm to provide a more intuitive explanation of the optimal combined model, to illustrate how these predictors affect the STAS of LUAD in the optimal combined model. The SHAP summary diagram (Figure 5A,5B) shows the contribution of the factors (DL-score, lobulation, vacuole sign, and microvascular sign) in predicting the STAS of each patient with LUAD. Each bar corresponds to a specific radiological feature used by the model. Bar length represents the magnitude of the feature’s impact (SHAP value). Color gradient (blue to red) indicates relative feature values across samples. SHAP values (X-axis): value range (−0.3 to 0.4) represents each feature’s contribution to model output. Positive values (right side): features that increase probability of STAS-positive prediction. Negative values (left side): features that decrease probability of STAS-positive prediction. The color gradient provides immediate visual cues. Red features: present/pronounced in STAS-positive cases. Blue features: present/pronounced in STAS-negative cases. In addition, the larger the absolute distribution range of the SHAP value, the greater the importance of the features in the evaluation of STAS. Figure 5C,5D show two representative patients of the correct prediction of STAS-negative and -positive, with the Grad-CAM visualized to illustrate our model decision-making process, highlighting differences in longitudinal image changes captured by the model. As shown in the Grad-CAM heatmap produced using the Grad-CAM method, the red and yellow regions represent areas activated by the ResNet-101 model and have the greatest predictive significance, whereas the green and blue backgrounds reflect areas with weaker predictive values. The deeper the feature color, the higher the degree of overlap between the attention area identified by our model and the actual lesion location, and the greater the possibility of STAS-positive prediction. Conversely, STAS-negative recognition of the attention area relative to the actual lesion area is diffuse distribution. Figure 5C shows images of a 63-year-old man with STAS-negative LUAD. CT showed a 15-mm mass with microvascular sign, no vacuole sign and lobulation sign, indicating probable STAS-negative with a low DL-score (−0.938), which strongly suggested STAS-negative and was consistent with the final pathological results. The Grad-CAM heatmap is an example of a STAS-negative case which shows that the attention regions we identified are diffusely distributed in relation to the actual lesion areas. Figure 5D shows images of a 46-year-old woman with STAS-positive LUAD. CT images showed a 26-mm mass, positive lobulation, vacuole, microvascular sign, and a high DL-score (2.100), indicating underlying STAS-positivity, which was consistent with the final pathological results. The Grad-CAM heatmap is an example of a STAS-positive case which shows that the attention regions identified by our method exhibited a higher degree of overlap with the actual lesion locations.

Figure 5 Combined model visualization and case application analysis. (A) SHAP plots of the ANN model beeswarm plot; its scatter plot shows the relationship between the characteristic value and the predicted probability through colors, including positive and negative predictive effects. (B) Force plot shows the SHAP values for all features and samples. The left Y-axis is the ranking of important features, and the features are sorted according to their influence from greatest to smallest. The right Y-axis is the visualization. The color depth in the image represents the size of SHAP value, that is, the influence of the value under the feature on the model. (C,D) Application analysis of two patients with STAS negative (C) and positive (D) of lung adenocarcinoma. −, negative; +, positive; ANN, artificial neural network; DL, deep learning; SHAP, SHapley Additive exPlanations; STAS, spread through air spaces.

Discussion

In this multicenter retrospective study, we developed an optimal combined model based on DL to predict STAS in LUAD, the performance of which was validated in an internal validation cohort and two external validation cohorts. A reasonable visual interpretation of the prediction was provided to improve the diagnostic confidence of clinicians and achieve an accurate diagnosis of LUAD.

CT is a routine modality used in clinical practice to evaluate LUAD. Previous studies have reported that CT imaging features have potential value in the evaluation and prediction of STAS in LUAD (9,26,27). In our study, lobulation, vacuole sign, and microvascular sign were critical features for predicting STAS, which is consistent with some previous studies. We speculate that this may be because of the following reasons. The lobulation sign is a radiological feature of tumor biological heterogeneity. It is caused by the gradient differences in the degree of tumor cell differentiation, the heterogeneity of proliferation rate driven by the local microenvironment, and the dynamic remodeling of the tumor-stroma interface resulting in uneven mechanical stress. Patients with positive STAS have a high malignant potential, and their tumors exhibit more aggressive proliferation, reflecting the multi-dimensional pattern of tumor cell invasive growth. The vacuole is an important sign in the differentiation of LUAD from other benign nodules, especially in the differential diagnosis of early lung cancer (28,29). Microvascular sign can indicate the presence of tumor-associated angiogenesis, which is the formation of new blood vessels to supply the growing tumor with nutrients and oxygen (30,31). The tumor vascular microenvironment shows characteristic alterations, mainly manifested in two aspects: first, the microvascular density within and around the tumor significantly increases; second, pathological vascular remodeling occurs. In the diagnosis and treatment of LUAD, these microvascular signs have dual important clinical significance. From a diagnostic perspective, they help improve the accuracy of assessing the malignancy of the tumor; in terms of prognosis, these signs are significantly correlated with the progression-free survival of patients (31). Specifically, the degree of vascular proliferation is positively correlated with the invasiveness of the tumor, that is, the more obvious the vascular proliferation, the stronger the invasiveness of the tumor. Moreover, the characteristic vascular invasion pattern can predict the response rate to targeted therapy, thereby providing key imaging evidence for formulating individualized treatment plans. Therefore, STAS-positive patients have a poor prognosis, which is associated with microvascular angiogenesis in tumors.

ResNet-101 is a deep residual network architecture, and the key innovation is the introduction of skip connections or residual connections (32-36). These connections allow the gradient to flow more easily through the network during training, mitigating the vanishing gradient problem and enabling the training of very deep networks, leading to state-of-the-art performance on many benchmark datasets when applied to tasks, such as medical image analysis in radiomics. ResNet-101 can be used as a feature extractor to automatically learn and extract meaningful features from medical images, and is introduced to address the degradation problem encountered when the deep network cannot obtain better performance because of the disappearance of gradients during training (37,38). Therefore, it has advantages as a feature extraction tool that other convolutional neural networks (CNNs) do not have. In this study, we used the ResNet101 model pre-trained using ImageNet, a large computer vision dataset, to extract DL features. We evaluated the ANN model against six widely utilized ML models to provide comprehensive results, facilitating clinicians in selecting the optimal model for their needs. Our findings demonstrated that the ANN model outperformed the ML models, as evidenced by the higher AUC, with a sensitivity of 90% and specificity of 96%. Additionally, in the external validation cohorts I and II, the ANN model exhibited superior predictive performance and generalizability, yielding satisfactory results. Therefore, the ANN model has clear advantages over traditional ML models. This study confirmed the reliability of a combined model based on the DL score and clinically independent risk factors for predicting STAS in LUAD with a high degree of accuracy.

The main drawback of the DL model is its inability to be interpreted, which posed a stubborn conundrum on the deployment of this black-box technique in clinical practice (16), namely, an apparent conflict between the performance of the complex model and the clinical interpretability. In this study, we used the SHAP method to comprehensively analyze the complex relationship between features and STAS, including the positive and negative effects. The overall SHAP summary plot helps us to understand which features positively and negatively affect the prediction results, whereas the importance feature plot provides an average assessment of the feature importance for the entire dataset. The SHAP force plot provides examples of how different features contribute to the STAS prediction. In this study, the DL-score, lobulation, vacuole sign, and microvascular sign were the four most important factors for predicting STAS in LUAD, with the maximum width of the SHAP distribution interval. The detailed contribution of each feature was visualized for each patient with LUAD. Overall, the SHAP values used in this study provide a method to unravel the black box of ML models, enhancing their interpretability and transparency. This allowed us to better understand the predictive value of the combined model for STAS in LUAD. By analyzing the SHAP values, we can quantitatively evaluate the extent to which these factors influence the prediction results and provide a basis for intervention measures and personalized prevention.

Our study had several limitations. First, the number of samples was relatively low; potential future research directions may be to collect more samples for hyperparameter optimization and iterative training to verify the accuracy of the predictions and generalization of the models. Meanwhile, a prospective analysis is required to further identify the performance of the ANN model, even for this study. Second, the collection of clinical data was not sufficiently comprehensive and may have ignored potential predictors. Third, no prognostic analysis was performed for any of the included patients. Therefore, the lack of follow-up data prevented us from evaluating the effect of STAS on patient prognosis. It is unclear whether DL features correlate with survival outcomes. Future studies should include survival endpoints to fully evaluate the predictive efficiency of the proposed model. Finally, the foundation cancer image biomarker model specifically designed for medical images was not fully utilized for research. In future studies, we will introduce this basic radiomics model into the study.


Conclusions

In this study, we used ResNet-101 as the transfer learning model, generated a DL-score, and constructed a combined model (ANN) that outperformed traditional ML models in predicting STAS in LUAD. By integrating the model with SHAP, we visually demonstrated the effect of different variables on STAS, offering potential value for the early identification and intervention of high-risk patients. Thus, our combined model may assist clinicians to improve the assessment and management of patients with LUAD.


Acknowledgments

None.


Footnote

Reporting Checklist: The authors have completed the TRIPOD+AI reporting checklist. Available at https://qims.amegroups.com/article/view/10.21037/qims-2025-887/rc

Data Sharing Statement: Available at https://qims.amegroups.com/article/view/10.21037/qims-2025-887/dss

Funding: The study was supported by the grants of the Guangdong Medical Science and Technology Research Fund (Nos. A2024112 and A2024022); Research Launch Project of Shunde Hospital of Southern Medical University (Nos. SRSP2023003, SRSP2024013, and SRSP2024032); The Science and Technology Planning Project of Foshan (Nos. 2320001006640 and 2220001005383); Southern Medical University image Alliance research fund project (No. FS202402016).

Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://qims.amegroups.com/article/view/10.21037/qims-2025-887/coif). The authors have no conflicts of interest to declare.

Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. Ethical approval was obtained from the Ethics Committee of The Eighth Affiliated Hospital of Southern Medical University (The First People’s Hospital of Shunde) (No. KYLS20231203), and the institutional review board waived the requirement for informed consent due to the retrospective nature of the study. All participating hospitals/institutions were informed and agreed with the study.

Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.


References

  1. Saad MB, Hong L, Aminu M, Vokes NI, Chen P, Salehjahromi M, et al. Predicting benefit from immune checkpoint inhibitors in patients with non-small-cell lung cancer by CT-based ensemble deep learning: a retrospective study. Lancet Digit Health 2023;5:e404-20. [Crossref] [PubMed]
  2. Wang J, Yao Y, Tang D, Gao W. An individualized nomogram for predicting and validating spread through air space (STAS) in surgically resected lung adenocarcinoma: a single center retrospective analysis. J Cardiothorac Surg 2023;18:337. [Crossref] [PubMed]
  3. Kadota K, Nitadori JI, Sima CS, Ujiie H, Rizk NP, Jones DR, Adusumilli PS, Travis WD. Tumor Spread through Air Spaces is an Important Pattern of Invasion and Impacts the Frequency and Location of Recurrences after Limited Resection for Small Stage I Lung Adenocarcinomas. J Thorac Oncol 2015;10:806-14. [Crossref] [PubMed]
  4. Villalba JA, Shih AR, Sayo TMS, Kunitoki K, Hung YP, Ly A, Kem M, Hariri LP, Muniappan A, Gaissert HA, Colson YL, Lanuti MD, Mino-Kenudson M. Accuracy and Reproducibility of Intraoperative Assessment on Tumor Spread Through Air Spaces in Stage 1 Lung Adenocarcinomas. J Thorac Oncol 2021;16:619-29. [Crossref] [PubMed]
  5. Wang Y, Ding Y, Liu X, Li X, Jia X, Li J, Zhang H, Song Z, Xu M, Ren J, Sun D. Preoperative CT-based radiomics combined with tumour spread through air spaces can accurately predict early recurrence of stage I lung adenocarcinoma: a multicentre retrospective cohort study. Cancer Imaging 2023;23:83. [Crossref] [PubMed]
  6. Jiang C, Luo Y, Yuan J, You S, Chen Z, Wu M, Wang G, Gong J. CT-based radiomics and machine learning to predict spread through air space in lung adenocarcinoma. Eur Radiol 2020;30:4050-7. [Crossref] [PubMed]
  7. Shih AR, Mino-Kenudson M. Updates on spread through air spaces (STAS) in lung cancer. Histopathology 2020;77:173-80. [Crossref] [PubMed]
  8. Jin W, Shen L, Tian Y, Zhu H, Zou N, Zhang M, Chen Q, Dong C, Yang Q, Jiang L, Huang J, Yuan Z, Ye X, Luo Q. Improving the prediction of Spreading Through Air Spaces (STAS) in primary lung cancer with a dynamic dual-delta hybrid machine learning model: a multicenter cohort study. Biomark Res 2023;11:102. [Crossref] [PubMed]
  9. Kim SK, Kim TJ, Chung MJ, Kim TS, Lee KS, Zo JI, Shim YM. Lung Adenocarcinoma: CT Features Associated with Spread through Air Spaces. Radiology 2018;289:831-40. [Crossref] [PubMed]
  10. de Margerie-Mellon C, Onken A, Heidinger BH, VanderLaan PA, Bankier AA. CT Manifestations of Tumor Spread Through Airspaces in Pulmonary Adenocarcinomas Presenting as Subsolid Nodules. J Thorac Imaging 2018;33:402-8. [Crossref] [PubMed]
  11. Zheng X, He B, Hu Y, Ren M, Chen Z, Zhang Z, Ma J, Ouyang L, Chu H, Gao H, He W, Liu T, Li G. Diagnostic Accuracy of Deep Learning and Radiomics in Lung Cancer Staging: A Systematic Review and Meta-Analysis. Front Public Health 2022;10:938113. [Crossref] [PubMed]
  12. Zhan Y, Dai R, Li F, Cheng Z, Zhuo Y, Shan F, Zhou L. Repeatability and reproducibility of deep learning features for lung adenocarcinoma subtypes with nodules less than 10 mm in size: a multicenter thin-slice computed tomography phantom and clinical validation study. Quant Imaging Med Surg 2024;14:5396-407. [Crossref] [PubMed]
  13. Huang Y, Liu Z, He L, Chen X, Pan D, Ma Z, Liang C, Tian J, Liang C. Radiomics Signature: A Potential Biomarker for the Prediction of Disease-Free Survival in Early-Stage (I or II) Non-Small Cell Lung Cancer. Radiology 2016;281:947-57. [Crossref] [PubMed]
  14. Lin MW, Chen LW, Yang SM, Hsieh MS, Ou DX, Lee YH, Chen JS, Chang YC, Chen CM. CT-Based Deep-Learning Model for Spread-Through-Air-Spaces Prediction in Ground Glass-Predominant Lung Adenocarcinoma. Ann Surg Oncol 2024;31:1536-45. [Crossref] [PubMed]
  15. Wang S, Yu H, Gan Y, Wu Z, Li E, Li X, et al. Mining whole-lung information by artificial intelligence for predicting EGFR genotype and targeted therapy response in lung cancer: a multicohort study. Lancet Digit Health 2022;4:e309-19. [Crossref] [PubMed]
  16. Ma M, Liu R, Wen C, Xu W, Xu Z, Wang S, Wu J, Pan D, Zheng B, Qin G, Chen W. Predicting the molecular subtype of breast cancer and identifying interpretable imaging features using machine learning algorithms. Eur Radiol 2022;32:1652-62. [Crossref] [PubMed]
  17. Castelvecchi D. Can we open the black box of AI? Nature 2016;538:20-3. [Crossref] [PubMed]
  18. Rudin C. Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Instead. Nat Mach Intell 2019;1:206-15. [Crossref] [PubMed]
  19. Nam JG, Park S, Park CM, Jeon YK, Chung DH, Goo JM, Kim YT, Kim H. Histopathologic Basis for a Chest CT Deep Learning Survival Prediction Model in Patients with Lung Adenocarcinoma. Radiology 2022;305:441-51. [Crossref] [PubMed]
  20. Liu Y, Fu Y, Peng Y, Ming J. Clinical decision support tool for breast cancer recurrence prediction using SHAP value in cooperative game theory. Heliyon 2024;10:e24876. [Crossref] [PubMed]
  21. Fransvea P, Fransvea G, Liuzzi P, Sganga G, Mannini A, Costa G. Study and validation of an explainable machine learning-based mortality prediction following emergency surgery in the elderly: A prospective observational study. Int J Surg 2022;107:106954. [Crossref] [PubMed]
  22. Wang Y, Lang J, Zuo JZ, Dong Y, Hu Z, Xu X, Zhang Y, Wang Q, Yang L, Wong STC, Wang H, Li H. The radiomic-clinical model using the SHAP method for assessing the treatment response of whole-brain radiotherapy: a multicentric study. Eur Radiol 2022;32:8737-47. [Crossref] [PubMed]
  23. Baek S, Jeong YJ, Kim YH, Kim JY, Kim JH, Kim EY, Lim JK, Kim J, Kim Z, Kim K, Chung MJ. Development and Validation of a Robust and Interpretable Early Triaging Support System for Patients Hospitalized With COVID-19: Predictive Algorithm Modeling and Interpretation Study. J Med Internet Res 2024;26:e52134. [Crossref] [PubMed]
  24. Tseng PY, Chen YT, Wang CH, Chiu KM, Peng YS, Hsu SP, Chen KL, Yang CY, Lee OK. Prediction of the development of acute kidney injury following cardiac surgery by machine learning. Crit Care 2020;24:478. [Crossref] [PubMed]
  25. Detterbeck FC, Nicholson AG, Franklin WA, Marom EM, Travis WD, Girard N, Arenberg DA, Bolejack V, Donington JS, Mazzone PJ, Tanoue LT, Rusch VW, Crowley J, Asamura H, Rami-Porta R; IASLC Staging and Prognostic Factors Committee; Advisory Boards; Multiple Pulmonary Sites Workgroup; Participating Institutions. The IASLC Lung Cancer Staging Project: Summary of Proposals for Revisions of the Classification of Lung Cancers with Multiple Pulmonary Sites of Involvement in the Forthcoming Eighth Edition of the TNM Classification. J Thorac Oncol 2016;11:639-50.
  26. Sun F, Huang Y, Yang X, Zhan C, Xi J, Lin Z, Shi Y, Jiang W, Wang Q. Solid component ratio influences prognosis of GGO-featured IA stage invasive lung adenocarcinoma. Cancer Imaging 2020;20:87. [Crossref] [PubMed]
  27. Song H, Cui S, Zhang L, Lou H, Yang K, Yu H, Lin J. Preliminary exploration of the correlation between spectral computed tomography quantitative parameters and spread through air spaces in lung adenocarcinoma. Quant Imaging Med Surg 2024;14:386-96. [Crossref] [PubMed]
  28. Ding Y, Chen Y, Wen H, Li J, Chen J, Xu M, Geng H, You L, Pan X, Sun D. Pretreatment prediction of tumour spread through air spaces in clinical stage I non-small-cell lung cancer. Eur J Cardiothorac Surg 2022;62:ezac248. [Crossref] [PubMed]
  29. Liu J, Yang X, Li Y, Xu H, He C, Qing H, Ren J, Zhou P. Development and validation of qualitative and quantitative models to predict invasiveness of lung adenocarcinomas manifesting as pure ground-glass nodules based on low-dose computed tomography during lung cancer screening. Quant Imaging Med Surg 2022;12:2917-31. [Crossref] [PubMed]
  30. Deng L, Tang HZ, Luo YW, Feng F, Wu JY, Li Q, Qiang JW, Preoperative CT. Radiomics Nomogram for Predicting Microvascular Invasion in Stage I Non-Small Cell Lung Cancer. Acad Radiol 2024;31:46-57. [Crossref] [PubMed]
  31. Monteiro AS, Araújo SRC, Araujo LH, Souza MC. Impact of microvascular invasion on 5-year overall survival of resected non-small cell lung cancer. J Bras Pneumol 2022;48:e20210283. [Crossref] [PubMed]
  32. Rajkomar A, Dean J, Kohane I. Machine Learning in Medicine. N Engl J Med 2019;380:1347-58. [Crossref] [PubMed]
  33. Shi Y, Zou Y, Liu J, Wang Y, Chen Y, Sun F, Yang Z, Cui G, Zhu X, Cui X, Liu F. Ultrasound-based radiomics XGBoost model to assess the risk of central cervical lymph node metastasis in patients with papillary thyroid carcinoma: Individual application of SHAP. Front Oncol 2022;12:897596. [Crossref] [PubMed]
  34. Kang D, Gweon HM, Eun NL, Youk JH, Kim JA, Son EJ. A convolutional deep learning model for improving mammographic breast-microcalcification diagnosis. Sci Rep 2021;11:23925. [Crossref] [PubMed]
  35. Kumar V, Prabha C, Sharma P, Mittal N, Askar SS, Abouhawwash M. Unified deep learning models for enhanced lung cancer prediction with ResNet-50-101 and EfficientNet-B3 using DICOM images. BMC Med Imaging 2024;24:63. [Crossref] [PubMed]
  36. Haennah JHJ, Christopher CS, King GRG. Prediction of the COVID disease using lung CT images by Deep Learning algorithm: DETS-optimized Resnet 101 classifier. Front Med (Lausanne) 2023;10:1157000. [Crossref] [PubMed]
  37. Jiang Y, Chen L, Zhang H, Xiao X. Breast cancer histopathological image classification using convolutional neural networks with small SE-ResNet module. PLoS One 2019;14:e0214587. [Crossref] [PubMed]
  38. Rahaman MM, Millar EKA, Meijering E. Breast cancer histopathology image-based gene expression prediction using spatial transcriptomics data and deep learning. Sci Rep 2023;13:13604. [Crossref] [PubMed]

(English Language Editor: J. Jones)

Cite this article as: Zhang R, Sun JS, Liu ZW, Lin YY, Cai XY, Kang LJ, Guo BL, Song YX, Bao CZ, Chen K, Sun Y, Chen YF, Cai ZP, Xie H, Chen HX, Yang SM, Hu QG. A deep learning-clinical combined model with SHapley Additive exPlanations (SHAP) method for assessing the tumor spread through air spaces in lung adenocarcinoma: a multicohort retrospective study. Quant Imaging Med Surg 2025;15(10):8833-8849. doi: 10.21037/qims-2025-887

Download Citation