Original Article


Ultrasound-based deep learning radiomics with explainable machine learning for predicting benign and malignant thyroid nodules in lymphocytic thyroiditis: a multicenter study

Ji Zhou, Wenwu Lu, Wei Wei, Ziyi Liu, Wang Zhou, Wei Peng, Wenbo Ding, Xin Wu, Chaoxue Zhang

Abstract

Background: Thyroid nodules (TNs) with lymphocytic thyroiditis (LT) complicate ultrasound (US) assessment, and few prediction models have been validated in this population. The aim of this study is to develop and validate a US-based deep learning radiomics nomogram (DLRN) model for predicting malignancy in TNs coexisting with LT.

Methods: This multicenter hybrid study enrolled 640 patients with pathologically confirmed TNs and concomitant LT. The cohort comprised a training set from Centers 1–3 (retrospective, n=411), an external test set from Center 4 (retrospective, n=162), and a validation set (prospective, n=67). Deep learning (DL) and handcrafted radiomics features were extracted separately from US images. After feature fusion, the features were Z-score normalized and selected using the Mann-Whitney U test or Chi-squared test, with Fisher’s exact test used when appropriate, followed by the least absolute shrinkage and selection operator (LASSO) regression. A DL radiomics (DLR) model was developed based on 10 machine learning algorithms and subsequently incorporated with independent clinical factors through logistic regression to build a DLRN. Model performance was evaluated by the area under the receiver operating characteristic curve (AUC), calibration and decision curve, along with SHapley Additive Explanations (SHAP) analysis for interpretability.

Results: Logistic regression was selected as the optimal DLR algorithm, achieving AUCs of 0.90 [95% confidence interval (CI): 0.87–0.93] in the training set, 0.90 (95% CI: 0.84–0.94) in the external test set, and 0.82 (95% CI: 0.71–0.91) in the prospective validation set. The DLRN, refined by integrating independent clinical factors via logistic regression, demonstrated superior discrimination with AUCs of 0.93, 0.94, and 0.92 in the training, external test, and prospective validation sets, respectively. The DLRN model demonstrated robust clinical applicability, as evidenced by calibration and decision curve analyses. SHAP analysis enhanced DLRN interpretability by quantifying feature contributions to individual predictions and overall model output.

Conclusions: The DLRN model was designed for the preoperative identification of malignant TNs in LT, thereby facilitating personalized treatment.

Download Citation