Artificial intelligence-enhanced ultrasound imaging for thyroid nodule detection and malignancy classification: a study on YOLOv11
Original Article

Artificial intelligence-enhanced ultrasound imaging for thyroid nodule detection and malignancy classification: a study on YOLOv11

Jiaqi Yang1, Zhigang Luo2, Yanting Wen3, Jing Zhang4

1Operating Room, West China Hospital, Sichuan University/West China School of Nursing, Sichuan University, Chengdu, China; 2Glory Wireless Co. Ltd., Chengdu, China; 3Department of Ultrasonography, The Fifth People’s Hospital of Chengdu, Chengdu, China; 4School of Computer Science and Technology, Chongqing University of Posts and Telecommunications, Chongqing, China

Contributions: (I) Conception and design: J Yang, Z Luo; (II) Administrative support: Y Wen; (III) Provision of study materials or patients: All authors; (IV) Collection and assembly of data: J Yang, J Zhang; (V) Data analysis and interpretation: All authors; (VI) Manuscript writing: All authors; (VII) Final approval of manuscript: All authors.

Correspondence to: Zhigang Luo, ME. Glory Wireless Co. Ltd., Room 307 & 309, Electronics and Information Industry Building, No. 159 Section 1, East 1st Ring Road, Chengdu 610000, China. Email: lzg0179@163.com.

Background: Thyroid nodules are a common clinical concern, with accurate diagnosis being critical for effective treatment and improved patient outcomes. Traditional ultrasound examinations rely heavily on the physician’s experience, which can lead to diagnostic variability. The integration of artificial intelligence (AI) into medical imaging offers a promising solution for enhancing diagnostic accuracy and efficiency. This study aimed to evaluate the effectiveness of the You Only Look Once v. 11 (YOLOv11) model in detecting and classifying thyroid nodules through ultrasound images, with the goal of supporting real-time clinical decision-making and improving diagnostic workflows.

Methods: We used the YOLOv11 model to analyze a dataset of 1,503 thyroid ultrasound images, divided into training (1,203 images), validation (150 images), and test (150 images) sets, comprising 742 benign and 778 malignant nodules. Advanced data augmentation and transfer learning techniques were applied to optimize model performance. Comparative analysis was conducted with other YOLO variants (YOLOv3 to YOLOv10) and residual network 50 (ResNet50) to assess their diagnostic capabilities.

Results: The YOLOv11 model exhibited superior performance in thyroid nodule detection as compared to other YOLO variants (from YOLOv3 to YOLOv10) and ResNet50. At an intersection over union (IoU) of 0.5, YOLOv11 achieved a precision (P) of 0.841 and recall (R) of 0.823, outperforming ResNet50’s P of 0.8333 and R of 0.8025. Among the YOLO variants, YOLOv11 consistently achieved the highest P and R values. For benign nodules, YOLOv11 obtained a P of 0.835 and R of 0.833, while for malignant nodules, it reached a P of 0.846 and a R of 0.813. Within the YOLOv11 model itself, performance varied across different IoU thresholds (0.25, 0.5, 0.7, and 0.9). Lower IoU thresholds generally resulted in better performance metrics, with P and R values decreasing as the IoU threshold increased.

Conclusions: YOLOv11 proved to be a powerful tool for thyroid nodule detection and malignancy classification, offering high P and real-time performance. These attributes are vital for dynamic ultrasound examinations and enhancing diagnostic efficiency. Future research will focus on expanding datasets and validating the model’s clinical utility in real-time settings.

Keywords: Thyroid nodule detection; You Only Look Once v. 11 (YOLOv11); ultrasound imaging; artificial intelligence (AI); malignancy classification


Submitted Feb 03, 2025. Accepted for publication Jun 17, 2025. Published online Aug 14, 2025.

doi: 10.21037/qims-2025-257


Introduction

The thyroid gland, a key organ of the human endocrine system, regulates metabolism through the secretion of thyroid hormones. Thyroid cancer, which arises from abnormal cell proliferation within the thyroid, has seen a significant increase in incidence rate globally, with this trend largely attributable to overdiagnosis driven by advancements in diagnostic imaging and tools. Thyroid nodules, a common clinical finding, are associated with factors such as genetic predisposition, immune system abnormalities, iodine intake, dietary habits, medication use, and emotional stress (1). Thyroid nodules are primarily divided into two categories: benign and malignant. Benign nodules typically include thyroid cysts and inflammatory nodules, which can mostly be treated effectively through medication or surgical removal. In contrast, malignant nodules account for 7% to 15% of all thyroid nodules (2). Their pathological types are more complex, mainly including primary thyroid cancer and metastatic thyroid cancer, which are primarily treated through surgical resection. Over the past 20 years, the prevalence of thyroid nodules has significantly increased, rising from 1–5% in the early 20th century to 33–68% today (3). According to the 2022 European Thyroid Association Guidelines, the management of pediatric thyroid nodules and differentiated thyroid carcinoma (DTC) requires expert care in experienced centers, emphasizing the importance of minimizing long-term side effects of treatment while maintaining the excellent prognosis of pediatric DTC. The guidelines also highlight the need for specific recommendations for children due to differences in clinical, molecular, and pathological characteristics as compared to adults (4). A study on hyalinizing trabecular tumors (HTTs) revealed that these rare thyroid neoplasms often pose diagnostic challenges due to their cytological similarities with malignant lesions such as papillary thyroid cancer and medullary thyroid cancer. It was further reported that 89% of HTTs had a negative concordant immunopanel and that 100% were wild-type BRAFV600E, underscoring the importance of ancillary tests in accurately diagnosing HTTs and avoiding unnecessary overtreatment (5).

The clinical diagnosis of thyroid nodules typically involves a variety of diagnostic methods, including ultrasound examination (6), fine-needle aspiration biopsy (7,8), blood analysis, genetic testing, computed tomography (CT), or magnetic resonance imaging (MRI) scans (9). Ultrasound examination, characterized by its safety and noninvasiveness has become the preferred imaging modality for thyroid nodules (2,10). Over the past few decades, the continuous advancement of high-frequency ultrasound technology and its related innovations has enabled the accurate detection of thyroid nodules with diameters of 2–3 mm. This detection capability holds significant clinical significance for the diagnosis of nodules. Early determination of the number and nature of thyroid nodules, especially the prediction of cervical lymph node metastasis in malignant nodules, is crucial for assisting clinicians in formulating precise treatment plans and improving patient outcomes (11). However, the accuracy of ultrasound diagnosis of thyroid nodules is highly dependent on the experience of the operating physician, and less experienced junior physicians may even miss diagnoses. Therefore, the development of an objective and effective intelligent diagnostic aid to assist physicians in making judgments is critical (12). In addition to traditional diagnostic methods, molecular diagnostic techniques can help clinicians to accurately determine the nature of thyroid nodules and guide treatment decisions by providing additional information. However, these methods have limitations such as their high cost, technical complexity, potential for false results, and variability in test performance across different platforms (13).

In recent years, the rapid development of artificial intelligence (AI) has engendered optimism in the field of medical diagnosis. AI-based computer-aided diagnostic (CAD) systems have exhibited considerable potential in the analysis of medical images, including those of thyroid nodules (14-16). These systems can transform a subjective diagnosis to an objective quantification, thereby improving diagnostic accuracy and reducing the bias caused by subjective factors (17-19). Numerous studies have demonstrated the effectiveness of AI in thyroid nodule detection and classification. For instance, a study by Kim et al. used a deep learning-based CAD system to analyze thyroid ultrasound images (20), achieving a high diagnostic accuracy rate. Other studies have examined the use of convolutional neural networks (CNNs) for thyroid nodule classification, highlighting the potential of AI in improving diagnostic efficiency (14,21,22).

In the diagnosis of thyroid nodules, real-time detection holds substantial clinical value. Real-time detection can significantly improve diagnostic efficiency, reduce patient waiting times, and optimize the diagnostic and treatment process. In dynamic ultrasound examinations, clinicians need to quickly and accurately identify nodules and arrive at preliminary judgments in order to promptly arrange further examinations or treatments (23). You Only Look Once (YOLO), a real-time object detection model, can swiftly process ultrasound images to offer clinicians instant nodule detection results, thereby facilitating rapid decision-making during examinations and enhancing diagnostic accuracy and efficiency. Additionally, YOLO’s unique strengths in thyroid ultrasound image detection lie in its single-stage detection mechanism. This study aims to develop a practical AI tool for real-time thyroid ultrasound by fine tuning You Only Look Once v. 11 (YOLOv11), the most recent single stage model that supports transfer learning, on labelled thyroid images so that nodules can be located and classified as benign or malignant instantly, thereby reducing reliance on operator experience and speeding up diagnosis. Unlike traditional two-stage methods, YOLO achieves faster detection speeds and lower computational complexity while being highly precise (24,25). We present this article in accordance with the TRIPOD+AI reporting checklist (available at https://qims.amegroups.com/article/view/10.21037/qims-2025-257/rc).


Methods

AI models

YOLO

The YOLO series has undergone a substantial evolution from YOLOv3 to YOLOv11, with each version introducing key improvements in architecture, performance, and efficiency. YOLOv3 introduced the Darknet-53 backbone and multiscale detection (26). YOLOv5 included cross-stage partial network (CSPNet) architecture and faster processing (27). YOLOv6 and YOLOv7 added lightweight designs, enhanced backbones, and anchor-free detection (28,29). YOLOv8 and YOLOv9 focused on lightweight structures and advanced training methods for real-time inference (25,30). YOLOv10 introduced large-kernel convolutions and partial self-attention modules (31). Most recently, YOLOv11 has integrated efficient C3K2 blocks and advanced attention mechanisms, achieving higher accuracy with fewer parameters and is optimized for real-time applications across various environments (32,33). The standard YOLOv11 network structures are illustrated in Figure 1.

Figure 1 YOLOv11 architecture. C2PSA, cross-stage partial with pyramid squeeze attention; Conv, convolutional; SPPF, spatial pyramid pooling fast; YOLOv11, You Only Look Once v. 11.

The YOLOv11 model is structured into three parts: backbone, neck, and head. It processes input images as follows: first, images are adaptively scaled to standardize dimensions. Then, the backbone uses an improved C3K2 module to extract hierarchical features (32). The neck employs a path aggregation network (PAN) structure with C3K2 for feature fusion across layers via upsampling and downsampling. Finally, the head uses a decoupled head, separating classification and detection, with depthwise separable convolutions in the classification branch to reduce parameters and computational load (33). These features render YOLOv11 highly suitable for the real-time detection of thyroid nodules, providing valuable support for medical diagnosis.

Residual network 50 (ResNet50)

Unlike YOLOv11, which is a lightweight object detection model designed for speed and efficiency, ResNet50 is a well-known deep CNN architecture that is part of the residual net (ResNet) family, developed by He et al. in 2016 (34). ResNet is often used as the backbone network in fast region-based CNNs (R-CNN). Its purpose is to address the common issues of vanishing and exploding gradients in deep network training through residual learning. ResNet50, which consists of 50 convolutional layers, is a prominent example of this architecture. Its key innovation is the use of residual blocks, in which each block creates a “residual map” by adding the input directly to the output via skip connections. This effectively mitigates the degradation problem often observed in deep networks.

Materials

Dataset

The dataset used in this study was obtained from Roboflow (San Francisco, CA, USA) and is licensed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license (35,36). It is specifically designed for object detection tasks and provides a comprehensive collection of annotated images suitable for training and evaluating YOLOv11 models. The dataset serves as a practical foundation for quickly validating model feasibility, facilitating preliminary technical route selection, and supporting future clinical research (35).

This study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. A total of 1,503 images from four projects were initially selected. All images and their corresponding labels were reviewed by a senior and a midlevel ultrasound physician to ensure accuracy and consistency. These reviewed images were then randomly shuffled and divided into three sets as follows: 1,203 images for the training set (80%), 150 images for the validation set (10%), and 150 images for the test set (10%). The dataset comprised 742 benign and 778 malignant thyroid nodule images.

Data augmentation

In our experiments, we used the “randaugment” option in YOLOv11, which applies a series of random data augmentation operations to increase the diversity of the training data and enhance the model’s robustness and generalization ability. These operations are stochastic, in that they introduce randomness to simulate a wide range of visual variations (37). Random augmentations through the “randaugment” option were applied to significantly enhance model performance by increasing data diversity, simulating real-world conditions, and improving generalization (38). Geometric augmentations included moderate scaling variations (scale variation =0.5) and selective left-right flipping (probability =0.5), while other geometric transformations such as rotation, shear, and perspective were disabled. Mosaic augmentation was fully enabled (probability =1.0) to introduce contextual diversity in the training data. The “Mixup” and “Copy-Paste” augmentations were disabled, ensuring that the focus remained on geometric and flipping-based enhancements. This configuration balanced complexity and effectiveness, providing sufficient variation to improve model generalization without introducing excessive computational overhead.

Experimental methods

We used a labeled dataset that was divided into three distinct sets: a training set, a validation set, and a test set. The training set was used to train the models, allowing them to learn the necessary features for accurate detection. The validation set was employed during the training process to fine-tune the models and prevent overfitting, ensuring that the models generalized well to unseen data. The test set was reserved for the final evaluation of model performance.

All models were trained for 100 epochs. Where applicable, transfer learning was employed to leverage pretrained weights and accelerate convergence. During each training epoch, the validation set was used to evaluate model performance. Metrics including as precision (P) and recall (R) were recorded to assess the models’ detection capabilities.

After training and validation, the YOLOv11 model demonstrated superior performance and was selected for further analysis. YOLO11n.pt carries pretrained weights provided by Ultralytics, derived from the Common Objects in Context (COCO) dataset. We evaluated YOLO11n.pt performance under different intersection over union (IoU) thresholds (0.25, 0.5, 0.7, and 0.9) to analyze how these thresholds impact detection P and R. This analysis was aimed at providing insights into the model’s robustness and adaptability to varying detection criteria.

We compared YOLOv11 to ResNet50, a well-established architecture that is widely used as a backbone network for one-stage detectors due to its strong feature extraction capabilities and efficient design. The results showed that YOLO11n.pt offers better performance.

The selected YOLO11n.pt model was tested on the reserved test set. We selected several images from the test set and compare the predicted bounding box coordinates with the ground truth labels. This comparison provided a straightforward assessment of the model’s performance in the test set.

Training environment

Table 1 outlines the hardware and software specifications for training the YOLOv11 model, including the central processing unit (CPU), graphics processing unit (GPU), system environment, and training platform details.

Table 1

Hardware and software specifications for model training

Category Configuration
CPU 12 vCPU Intel® Xeon® Platinum 8352V CPU @ 2.10 GHz
GPU RTX 3080 × 2 (20 GB) × 1
System environment PyTorch 2.5.1
Python 3.12 (Ubuntu 22.04)
CUDA 12.4
Training platform Ultralytics 8.3.49

CPU, central processing unit; CUDA, compute unified device architecture; GPU, graphics processing unit.

Transfer learning

In tasks involving the classification of thyroid nodules as benign or malignant, training a classification model from scratch presents several challenges. Models can be prone to overfitting, particularly when ultrasound image data are scarce, which subsequently leads to unsatisfactory recognition accuracy in practical applications. To address this, we opted to leverage the pretrained YOLO11n.pt model provided by Ultralytics (Frederick, MD, USA) and adapt it to the classification of thyroid nodules. This pretrained model, developed on a large-scale dataset, offered robust feature extraction capabilities that could significantly enhance the stability and generalizability of our model, effectively compensating for the limited training data.

In practice, the limited availability of ultrasound image data makes it difficult to train deep models directly. Such limitations not only increase training complexity but also often result in subpar predictive performance. However, through transfer learning, deep models can more efficiently extract key features of thyroid nodules, thereby significantly improving classification accuracy. Additionally, while deep models are characterized by numerous parameters and complex training processes that demand substantial resources, transfer learning allows for the fine-tuning of only the high-level features of the ultrasound images, thereby greatly reducing the demand for hardware resources and training time (39).

For ResNet50, by loading the pretrained weights of ResNet50 and fine-tuning the final fully connected layer, the model can achieve high-precision classification. This approach leverages the benefits of transfer learning, allowing the model to perform exceptionally well in specific tasks.

Overall, transfer learning offers an effective solution to the issue of insufficient data in thyroid ultrasound image classification. By transferring network parameters that have been well-trained in other image classification domains to our thyroid nodule classification task, we could significantly enhance model performance and mitigate the performance bottlenecks caused by limited data (39).

Parameters for model evaluation

The performance of the YOLOv11 model was evaluated using key metrics such as P, R, mean average precision (mAP), F1 curve, P curve, and R curve. Additionally, a confusion matrix was used to generate a detailed breakdown of the classification results (39,40). These metrics and tools collectively offered a comprehensive assessment of the model’s accuracy, robustness, and balance between false positives and false negatives, while the curves and confusion matrix provided insights into performance across different confidence thresholds and class-specific errors.

The formula for P is given as follows:

P=TPTP+FP

where TP (true positive) is the number of correctly predicted positive observations, and FP (false positive) is the number of incorrectly predicted positive observations.

The formula for R is given as follows:

R=TPFN+TP

where TP is the number of correctly predicted positive observations, and FN (false negative) is the number of incorrectly predicted negative observations.

The formula for mAP is given as follows:

mAP=1Ni=1NAPi

where APi is the average P for class i, and N is the total number of classes.

To evaluate model performance in object detection, the average precision when the IoU threshold is set to 0.5 (mAP@50) was used a metric. This indicates that detection is considered a TP if its IoU with the ground truth is at least 0.5 (41).

In contrast, mAP@50–95 is a more comprehensive and stringent metric that calculates the average precision across multiple IoU thresholds, typically ranging from 0.5 to 0.95 in increments of 0.05. This ensures the model performs well under various overlap requirements, providing a robust indicator of overall detection accuracy and reliability (41).

IoU is a key metric for evaluating the overlap between predicted and ground-truth bounding boxes. Defined as the ratio of their intersection area to the union area, IoU ranges from 0 to 1, with 1 indicating perfect overlap. It is crucial for assessing model performance in object detection, particularly in calculating metrics such as mAP at different thresholds. IoU measures both the model’s ability to identify objects and its P in localizing them, making it vital for applications requiring accurate object positioning (41).

In the context of the Ultralytics YOLO models, F1 curve refers to the plot of F1 scores at different confidence thresholds. The F1 score is the harmonic mean of P and R and offers a balanced trade-off between FPs and FNs. This curve helps in selecting the optimal confidence threshold in which the model achieves the highest F1 score, indicating the best balance between P and R. The P curve plots the P of the model across different confidence thresholds. It shows how the model’s accuracy in predicting true positives changes as the confidence threshold is adjusted. The R curve is a graphical representation showing how R values change across different confidence thresholds. This curve illustrates the model’s ability to identify all instances of objects in the validation dataset at various thresholds (40).

A confusion matrix is a table used to evaluate the performance of a classification model. It shows the counts of TPs, FPs, true negatives (TNs), and FNs. This matrix helps in understanding the model’s accuracy and errors, providing insights into its strengths and weaknesses.


Results

Table 2 offers a comprehensive evaluation of various YOLO models, focusing on their performance metrics across different classes. The validation was conducted on a dataset of 150 images and instances, and YOLO11n.pt emerged as the top performer with the highest P (0.841) and mAP@50 (0.874), demonstrating its superior accuracy in object detection. Despite a lower mAP@50–95 (0.482), which is typical for YOLO models due to the increased P required for higher IoU thresholds, YOLO11n.pt still outperformed other models, including YOLOv11.yaml, YOLOv10n.yaml, YOLOv9t.yaml, YOLOv8.yaml, YOLOv6.yaml, YOLOv5.yaml, and YOLOv3.yaml in terms of P and overall detection accuracy. These yaml files specify the network architectures for YOLOv3 through YOLOv11 but are shipped without any pretrained weights.

Table 2

Summary of the validation model

Model IoU Class Precision Recall mAP@50 mAP@50–95
yolo11n.pt 0.25 All 0.846 0.824 0.859 0.484
Benign 0.847 0.834 0.901 0.541
Malign 0.846 0.815 0.817 0.426
yolo11n.pt 0.5 All 0.841 0.823 0.874 0.482
Benign 0.835 0.833 0.909 0.54
Malign 0.846 0.813 0.839 0.425
yolo11n.pt 0.7 All 0.835 0.806 0.868 0.48
Benign 0.86 0.822 0.902 0.54
Malign 0.81 0.79 0.833 0.421
yolo11n.pt 0.9 All 0.808 0.73 0.816 0.464
Benign 0.89 0.781 0.872 0.525
Malign 0.725 0.679 0.76 0.403
yolov11.yaml 0.5 All 0.756 0.792 0.823 0.459
Benign 0.851 0.795 0.863 0.522
Malign 0.662 0.79 0.782 0.395
yolov10n.yaml 0.5 All 0.833 0.632 0.752 0.424
Benign 0.87 0.639 0.785 0.479
Malign 0.796 0.626 0.719 0.368
yolov9t.yaml 0.5 All 0.781 0.697 0.794 0.449
Benign 0.793 0.74 0.852 0.509
Malign 0.769 0.654 0.737 0.388
yolov8.yaml 0.5 All 0.798 0.728 0.808 0.456
Benign 0.839 0.74 0.851 0.504
Malign 0.757 0.716 0.764 0.408
yolov6.yaml 0.5 All 0.814 0.786 0.814 0.45
Benign 0.824 0.795 0.837 0.506
Malign 0.805 0.778 0.791 0.395
yolov5.yaml 0.5 All 0.71 0.797 0.795 0.418
Benign 0.78 0.767 0.838 0.491
Malign 0.639 0.827 0.751 0.345
yolov3.yaml 0.5 All 0.76 0.716 0.777 0.418
Benign 0.809 0.74 0.841 0.495
Malign 0.711 0.691 0.713 0.341

mAP@50: mean average precision when the IoU threshold is set to 0.5. mAP@50–95: mean average precision across multiple IoU thresholds, typically ranging from 0.5 to 0.95 in increments of 0.05. IoU, intersection over union; mAP, mean average precision.

YOLO11n.pt’s performance in detecting benign and malignant classes was particularly noteworthy. For benign cases, it achieved a P of 0.847 and an R of 0.834, while for malignant cases, the P was 0.846 and the R was 0.815. P reflects the proportion of TPs among all predicted positives, indicating the model’s ability to minimize FP. A high P is crucial in medical diagnostics, especially for malignant cases, to avoid unnecessary anxiety and treatment caused by false alarms. R, which is directly related to the FN rate (FNR), measures the proportion of TPs among all actual positives. A higher R corresponds to a lower FNR, meaning fewer actual malignant nodules are missed by the model. In clinical terms, a lower FNR is highly desirable as it reduces the risk of missing cancer cases. Based on the provided data, the performance of the YOLO11n.pt model across different IoU thresholds revealed a clear pattern: as the IoU threshold increased, the P remained relatively stable, while the R tended to decrease. This was particularly evident in both the benign and malign classes, in which the higher IoU thresholds led to stricter criteria for TPs, resulting in a lower R rate. For instance, in the benign class, the R decreased from 0.834 at IoU =0.25 to 0.781 at IoU =0.9, while the P slightly increased from 0.847 to 0.89. Similarly, in the malign class, the R decreased from 0.815 at IoU =0.25 to 0.679 at IoU =0.9, with P decreasing from 0.846 to 0.725. These findings highlight the trade-off between P and R as the IoU threshold increases, indicating that the model becomes more conservative in detecting TPs, which is crucial for applications in which high P is prioritized over R.

In the comparison of YOLO11n.pt and ResNet50, YOLO11n.pt demonstrated a slight advantage in P and R, with 0.5 IoU values of 0.841 and 0.823 compared to ResNet50’s 0.8333 and 0.8025, respectively. However, the key difference lay in their functionalities: YOLO11n.pt not only classified objects but also localized them by predicting bounding boxes, making it suitable for tasks requiring both identification and precise positioning. In contrast, ResNet50 focused solely on classification without providing localization information. Additionally, YOLO11n.pt was able to provide real-time deployment due to its ability to process images quickly and provide immediate feedback, which is crucial for applications such as surveillance or autonomous systems.

YOLO11n.pt had good performance in both P and R, effectively FPs and FNs. However, it is important to note that enhancing one metric may potentially impact the other. The trade-off between P and R across varying confidence thresholds is illustrated in Figure 2. YOLO11n.pt consistently achieved higher F1 scores across most confidence thresholds, indicating its ability to maintain high levels of both P and R. Accuracy and comprehensive detection hold a critical position in medical diagnostics. Figure 3 further highlights the models’ accuracy in identifying positive cases with minimal false alarms. The upward trend in the curves as confidence thresholds increased indicates that higher confidence levels corresponded to fewer FPs. YOLO11n.pt stood out with the highest P values across a broad range of confidence thresholds, underscoring its reliability in precise positive predictions. This is a significant advantage in applications in which FPs can lead to unnecessary costs or actions.

Figure 2 F1 score curves.
Figure 3 P curves. P, precision.

Figure 4 provides insight into each model’s ability to maintain high P while achieving broad R. This is particularly important in scenarios where both FP FNs have substantial consequences. YOLO11n.pt excelled in this aspect, maintaining high P at various R levels as compared to the other models. This underscores its robustness and effectiveness in real-world applications in which the cost of missed detections is high.

Figure 4 PR curves. PR, precision-recall.

Finally, we assessed how effectively each model captured all positive instances within the dataset (Figure 5). The curve representing YOLO11n.pt remained above others across most confidence thresholds, signifying its superior R performance. This indicates that the model is more adept at identifying TPs, even at lower confidence levels as compared to the other models. This capability is crucial for applications in which missing positive instances has severe implications.

Figure 5 R curves. R, recall.

The confusion matrix (Figure 6) indicated YOLO11n.pt’s strong classification performance across three classes: benign, malign, and background. For the benign class, it yielded 60 TPs, 5 FPs, and 8 FNs. In the malign class, the model yielded 69 TPs, 6 FPs, and 3 FNs. In the Figure 6, the intensity of the right color bar is directly proportional to the magnitude of the values; that is, the deeper the color, the greater the value at that location, while the lighter the color, the smaller the value at that location. However, for the background, some misclassification occurred, with 7 instances each of benign and malign being incorrectly labeled as background. Overall, YOLO11n.pt demonstrated high accuracy in classifying benign and malign cases, which is crucial for medical diagnostics but could be improved in terms of reducing background misclassifications.

Figure 6 Confusion matrix.

Figure 7 displays a comparison between labeled and predicted ultrasound images. On the left side, the labeled images show annotated regions of interest, with purple and red boxes indicating the positions of nodules. On the right side, the predicted images demonstrate the model’s ability to accurately locate these nodules, with blue (benign) and teal (malign) boxes indicating the positions. Additionally, the predictions include confidence scores (e.g., “benign 0.86”) that indicate the model’s certainty in classifying the nodules as benign or malignant. The alignment of the predicted boxes with the labeled regions highlights the model’s effectiveness in both nodule localization and classification.

Figure 7 Label images versus predicted images.

Discussion

The YOLO11n.pt model demonstrated superior performance in thyroid nodule detection and malignancy classification as compared to the other YOLO variants and ResNet50, achieving a high P of 0.841 and an mAP@50 of 0.874. These results are comparable to recent studies reporting nodule classification accuracies ranging from 0.83 to 0.87 for various CNN models trained on large datasets comprising thousands of cases (14). Our dataset, which exceeded 1,000 cases, further validates the effectiveness of the YOLO11n.pt model. Additionally, we conducted an experiment on a small dataset of 100 anal fistula orifices, yielding an accuracy of 0.92 and a R of 1. The high accuracy of this small dataset can be attributed to the alignment between the one-class nature of the pretrained dataset used for transfer learning and our small dataset, which enhanced the effectiveness of transfer learning. However, the generalization ability of the small dataset is not as robust as that of the large dataset. This means that large dataset system can stably provide relatively reliable diagnostic support under different clinical scenarios and different equipment conditions. By leveraging the stability of the large dataset system, we can significantly enhance the model’s generalization ability, ensuring more reliable and consistent diagnostic support across various clinical scenarios and equipment conditions, thereby improving the software’s utility and effectiveness in real-world applications.

The diagnostic accuracy of ultrasound examinations is influenced by a variety of factors, including the experience of the sonographer, the quality of the equipment, and the condition of the patient. It has been reported that the diagnostic accuracy rate of experienced sonographers is approximately 0.85 (42). When the system provides real-time assistance to sonographers, both the system and the physician make independent judgments, thereby reinforcing each other’s diagnostic capabilities in a bidirectional manner. This indicates that the real-time diagnostic assistance from the YOLOv11 system could significantly enhance the overall diagnostic accuracy, making it a highly effective tool in clinical practice.

However, the model’s R decreases as the IoU threshold increases, indicating a trade-off between P and R. This suggests that while the model becomes more conservative in detecting TPs with higher IoU thresholds, it may miss some actual nodules. This trade-off needs to be carefully considered in clinical applications.

Compared to ResNet50, YOLO11n.pt not only achieved better P and R but also provided real-time detection and localization through its single-stage detection mechanism. This makes it more suitable for dynamic ultrasound examinations, in which rapid and accurate decision-making is essential.

Despite these promising results, the dataset used in this study involved limitations that reduced the model’s clinical applicability. It is imperative to emphasize that at this stage, our YOLO11n.pt model exhibits a high rate of FNs, which renders it unsuitable for clinical application. The potential consequences of overlooking malignant thyroid nodules due to FNs are severe, as they may lead to delayed diagnosis and treatment, which can significantly impact patient health outcomes. Moreover, the lack of detailed patient demographics and morphological verification limits the dataset’s clinical relevance. Future work should focus on acquiring more comprehensive datasets with detailed patient information and verified diagnoses to enhance the model’s generalization ability. Additionally, the model’s performance should be evaluated in real-time clinical settings to assess its practical utility. Currently, we are in the process of collecting case information and validating video data for both liver and anal fistula orifices, building upon the findings of this study to conduct more in-depth AI research. These efforts are aimed at creating a more detailed and clinically relevant dataset, ultimately leading to the development of a reliable real-time CAD system.


Conclusions

The potential of the YOLOv11 model for thyroid nodule detection and malignancy classification based on ultrasound imaging was systematically evaluated in this study. YOLO11n.pt demonstrated superior performance relative to other YOLO variants and ResNet50, supporting its capacity to augment diagnostic accuracy and efficiency. The model’s real-time detection and localization capabilities, in conjunction with its high P and R, render it a valuable asset for medical professionals, with the potential to mitigate diagnostic bias and alleviate workload.

Future research should focus on obtaining more comprehensive datasets with verified diagnoses to enhance the model’s generalization ability. Moreover, assessing the model’s performance in real-time clinical settings is essential to determining its practical utility and to realizing its seamless integration into existing diagnostic workflows.


Acknowledgments

None.


Footnote

Reporting Checklist: The authors have completed the TRIPOD+AI reporting checklist. Available at https://qims.amegroups.com/article/view/10.21037/qims-2025-257/rc

Funding: This study received funding from Innovative Talents Program of Sichuan Provincial Health Commission (No. 24CXTD16).

Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://qims.amegroups.com/article/view/10.21037/qims-2025-257/coif). Z.L. is the General Manager of Glory Wireless Co. Ltd. The other authors have no conflicts of interest to declare.

Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. This study used publicly available datasets sourced from Roboflow. The dataset is licensed under CC BY 4.0. Given that the data within the dataset have been deidentified and do not contain private, personal, or sensitive information, ethical review by an ethics committee was not required. The use of the data complies with the dataset’s licensing agreements and terms of use. This study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments.

Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.


References

  1. Kitahara CM, Schneider AB. Epidemiology of Thyroid Cancer. Cancer Epidemiol Biomarkers Prev 2022;31:1284-97. [Crossref] [PubMed]
  2. Zheng J. Intelligent Diagnosis of Thyroid Nodules Based on Ultrasound Images. Tianjin University; 2021.
  3. Chen P, Feng C, Huang L, Chen H, Feng Y, Chang S. Exploring the research landscape of the past, present, and future of thyroid nodules. Front Med (Lausanne) 2022;9:831346. [Crossref] [PubMed]
  4. Lebbink CA, Links TP, Czarniecka A, Dias RP, Elisei R, Izatt L, Krude H, Lorenz K, Luster M, Newbold K, Piccardo A, Sobrinho-Simões M, Takano T, Paul van Trotsenburg AS, Verburg FA, van Santen HM. 2022 European Thyroid Association Guidelines for the management of pediatric thyroid nodules and differentiated thyroid carcinoma. Eur Thyroid J 2022;11:e220146. [Crossref] [PubMed]
  5. Dell’Aquila M, Gravina C, Cocomazzi A, Capodimonti S, Musarra T, Sfregola S, Fiorentino V, Revelli L, Martini M, Fadda G, Pantanowitz L, Larocca LM, Rossi ED. A large series of hyalinizing trabecular tumors: Cytomorphology and ancillary techniques on fine needle aspiration. Cancer Cytopathol 2019;127:390-8. [Crossref] [PubMed]
  6. Cong P, Wang XM, Zhang YF. Comparison of artificial intelligence, elastic imaging, and the thyroid imaging reporting and data system in the differential diagnosis of suspicious nodules. Quant Imaging Med Surg 2024;14:711-21. [Crossref] [PubMed]
  7. Fiorentino V. Dell’ Aquila M, Musarra T, Martini M, Capodimonti S, Fadda G, Curatolo M, Traini E, Raffaelli M, Lombardi CP, Pontecorvi A, Larocca LM, Pantanowitz L, Rossi ED. The Role of Cytology in the Diagnosis of Subcentimeter Thyroid Lesions. Diagnostics (Basel) 2021;11:1043. [Crossref] [PubMed]
  8. Wang CC, Friedman L, Kennedy GC, Wang H, Kebebew E, Steward DL, Zeiger MA, Westra WH, Wang Y, Khanafshar E, Fellegara G, Rosai J, Livolsi V, Lanman RB. A large multicenter correlation study of thyroid nodule cytopathology and histopathology. Thyroid 2011;21:243-51. [Crossref] [PubMed]
  9. Haugen BR, Alexander EK, Bible KC, Doherty GM, Mandel SJ, Nikiforov YE, Pacini F, Randolph GW, Sawka AM, Schlumberger M, Schuff KG, Sherman SI, Sosa JA, Steward DL, Tuttle RM, Wartofsky L. 2015 American Thyroid Association Management Guidelines for Adult Patients with Thyroid Nodules and Differentiated Thyroid Cancer: The American Thyroid Association Guidelines Task Force on Thyroid Nodules and Differentiated Thyroid Cancer. Thyroid 2016;26:1-133. [Crossref] [PubMed]
  10. Zheng TL, Yang N, Geng S, Zhao XY, Wang Y, Cheng DQ, Zhao L. An Improved Algorithm for Thyroid Nodule Detection in Ultrasound Images Based on Faster R-CNN. Journal of Sichuan University (Medical Science Edition) 2023;54:915-22. [Crossref] [PubMed]
  11. Takano T. Natural history of thyroid cancer [Review] Endocr J 2017;64:237-44.
  12. Wang B, Wan Z, Li C, Zhang M, Shi Y, Miao X, Jian Y, Luo Y, Yao J, Tian W. Identification of benign and malignant thyroid nodules based on dynamic AI ultrasound intelligent auxiliary diagnosis system. Front Endocrinol (Lausanne) 2022;13:1018321. [Crossref] [PubMed]
  13. Patel J, Klopper J, Cottrill EE. Molecular diagnostics in the evaluation of thyroid nodules: Current use and prospective opportunities. Front Endocrinol (Lausanne) 2023;14:1101410. [Crossref] [PubMed]
  14. Ma Y, Xu YC, Liu C, Zhang G, Kang WJ, Zhao D, Lin L, Wu SC. CNNs Ensemble Learning for Ultrasound Thyroid Nodule Diagnosis. Life Science Instruments 2021;19:52-7.
  15. Ren XQ. A Study on Thyroid Nodule Malignancy Prediction Model Based on Transformer. University of Electronic Science and Technology; 2024.
  16. Yang WT, Ma BY, Chen Y. A narrative review of deep learning in thyroid imaging: current progress and future prospects. Quant Imaging Med Surg 2024;14:2069-88. [Crossref] [PubMed]
  17. Sharifi Y, Amiri Tehranizadeh A, Danay Ashgzari M, Naseri Z. TIRADS-based artificial intelligence systems for ultrasound images of thyroid nodules: protocol for a systematic review. J Ultrasound 2025;28:151-8. [Crossref] [PubMed]
  18. Naganjaneyulu S, Lakshmi RS, NagulMeeraBee S, Reddy MSK. Thyroid Nodule Detection Using Deep Learning Strategies. 2024 International Conference on Emerging Systems and Intelligent Computing (ESIC). Bhubaneswar: IEEE; 2024:556-61.
  19. Kaur S, Singla J, Nkenyereye L, Jha S, Prashar D, Joshi GP, El-Sappagh S, Islam MS, Islam SMR. Medical Diagnostic Systems Using Artificial Intelligence (AI) Algorithms: Principles and Perspectives. IEEE Access 2020;8:228049-69.
  20. Kim YJ, Choi Y, Hur SJ, Park KS, Kim HJ, Seo M, Lee MK, Jung SL, Jung CK. Deep convolutional neural network for classification of thyroid nodules on ultrasound: Comparison of the diagnostic performance with that of radiologists. Eur J Radiol 2022;152:110335. [Crossref] [PubMed]
  21. Wang XQ, Yang F, Cao B, Liu J, Wei DJ, Cao H. Application of Convolutional Neural Networks in Thyroid Nodule Diagnosis. Laser & Optoelectronics Progress 2022;59:28-42.
  22. Liu T, Guo Q, Lian C, Ren X, Liang S, Yu J, Niu L, Sun W, Shen D. Automated detection and classification of thyroid nodules in ultrasound images using clinical-knowledge-guided convolutional neural networks. Med Image Anal 2019;58:101555. [Crossref] [PubMed]
  23. Luo HX. Development of an Automatic Recognition and Diagnosis System for Thyroid Nodules in Ultrasound Dynamic Videos Based on Deep Learning. 2022.
  24. Huang QN. A Study on Detection and Risk Stratification of Malignant Traits of Thyroid Nodules Based on C-TIRADS. Nanchang University; 2024.
  25. HussainM.YOLOv5, YOLOv8 and YOLOv10: The Go-To Detectors for Real-time Vision.2024. arXiv:2407.02988.
  26. RedmonJFarhadiA.YOLOv3: An Incremental Improvement.2018. arXiv:1804.02767.
  27. Wang Y, Zou X, Shi J, Liu M. YOLOv5-Based Dense Small Target Detection Algorithm for Aerial Images Using DIOU-NMS. Radioengineering 2024;33:12-22.
  28. LiCLiLJiangHWengKGengYLiLKeZLiQChengMNieWLiYZhangBLiangYZhouLXuXChuXWeiXWeiX.YOLOv6: A Single-Stage Object Detection Framework for Industrial Applications.2022. arXiv:2209.02976.
  29. Wang CY, Bochkovskiy A, Liao HYM. YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver: IEEE; 2023:7464-75.
  30. Akgül M, Kozan Hİ, Akyürek HA, Taşdemir Ş. Automated stenosis detection in coronary artery disease using yolov9c: Enhanced efficiency and accuracy in real-time applications. Journal of Real-Time Image Processing 2024;21:177.
  31. Wang A, Chen H, Liu L, Chen K, Lin Z, Han J. Yolov10: Real-time end-to-end object detection. Advances in Neural Information Processing Systems 2024;37:107984-8011.
  32. KhanamRHussainM.YOLOv11: An Overview of the Key Architectural Enhancements.2024. arXiv:2410.17725.
  33. JeghamNKohCYAbdelattiMHendawiA. Evaluating the Evolution of YOLO (You Only Look Once) Models: A Comprehensive Benchmark Study of YOLO11 and Its Predecessors.2024. arXiv:2411.00201.
  34. He K, Zhang X, Ren S, Sun J. Deep Residual Learning for Image Recognition. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Las Vegas, NV, USA: IEEE; 2016.
  35. Community R. tiroid Object Detection Dataset and Pre-Trained Model. Roboflow Universe. 2024. Available online: https://universe.roboflow.com/demo-k2z53/tiroid
  36. Community R. Thyroid_yolov8 Object Detection Dataset. Roboflow Universe. 2024. Available online: https://universe.roboflow.com/demo-k2z53/thyroid_yolov8/dataset/1
  37. Lin J, Hu G, Chen J. Mixed data augmentation and osprey search strategy for enhancing YOLO in tomato disease, pest, and weed detection. Expert Systems with Applications 2025;264:125737.
  38. Cubuk ED, Zoph B, Shlens J, Le QV. RandAugment: Practical automated data augmentation with a reduced search space. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). Seattle, WA, USA: IEEE; 2020.
  39. Liu T, Xie S, Yu J, Niu L, Sun W. Classification of thyroid nodules in ultrasound images using deep model based transfer learning and hybrid features. 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). New Orleans, LA, USA: IEEE; 2017.
  40. Sharma A, Kumar V, Longchamps L. Comparative performance of YOLOv8, YOLOv9, YOLOv10, YOLOv11 and Faster R-CNN models for detection of multiple weed species. Smart Agricultural Technology 2024;9:100648.
  41. Lin TY, Maire M, Belongie S, Hays J, Perona P, Ramanan D, Dollár P, Zitnick CL. Microsoft COCO: Common Objects in Context. Computer Vision–ECCV 2014:740-55.
  42. Huang P, Zheng B, Li M, Xu L, Rabbani S, Mayet AM, Chen C, Zhan B, Jun H. The Diagnostic Value of Artificial Intelligence Ultrasound S-Detect Technology for Thyroid Nodules. Comput Intell Neurosci 2022;2022:3656572. [Crossref] [PubMed]

(English Language Editor: J. Gray)

Cite this article as: Yang J, Luo Z, Wen Y, Zhang J. Artificial intelligence-enhanced ultrasound imaging for thyroid nodule detection and malignancy classification: a study on YOLOv11. Quant Imaging Med Surg 2025;15(9):7964-7976. doi: 10.21037/qims-2025-257

Download Citation