Deep learning-based computed tomography detection of early lymph node metastasis in head and neck cancer
Introduction
Head and neck squamous cell carcinoma (HNSCC) ranks among the most prevalent malignancies globally (1-4), with cervical lymph node metastasis (LNM) serving as a critical factor in determining disease staging, treatment decisions, and patient prognosis (5-8). Consequently, the precise preoperative identification of metastatic lymph nodes is essential for guiding surgical planning and optimizing clinical outcomes. Contrast-enhanced computed tomography (CT) is commonly employed for this purpose, owing to its widespread availability, standardized acquisition protocols, and cost-effectiveness (9-11). However, the reliable detection of LNM using CT remains a significant challenge in clinical practice.
Metastatic lymph nodes frequently exhibit small size, indistinct margins, and subtle enhancement patterns, particularly in early-stage disease, leading to the frequent oversight of smaller lesions. Diagnostic accuracy is heavily reliant on reader experience (12-16). These limitations contribute to substantial inter-observer variability and place a considerable burden on radiologists, underscoring the necessity for automated and objective diagnostic support.
Recent advancements in artificial intelligence (AI) have shown considerable promise in enhancing image-based cancer diagnosis through the facilitation of automated lesion detection and characterization (17,18). Early radiomics-based methodologies depended on handcrafted features extracted from predefined regions of interest (ROIs), thereby limiting their ability to capture intricate anatomical variations and subtle imaging cues associated with LNM (19,20). Conversely, deep learning models autonomously learn hierarchical feature representations directly from image data, achieving significant success across a broad spectrum of tumor detection and localization tasks.
In HNSCC imaging, deep learning techniques have been used for tumor classification, segmentation, and prognosis prediction across various modalities, including CT, magnetic resonance imaging (MRI), and positron emission tomography (PET) (21,22). However, the automated detection of LNM presents significant challenges. Methodologically, this task is characterized by two unique features: first, metastatic lymph nodes are generally small and situated within anatomically intricate backgrounds; second, the clinical implications of false-negative detections are considerably more severe than those of false-positive detections. These features inherently complicate the implementation of standard object detection frameworks.
Single-stage object detectors, which predict bounding boxes and class labels in a unified manner, frequently encounter challenges in achieving both high sensitivity and acceptable specificity under certain conditions. Optimizing aggressively for sensitivity may capture small metastatic nodes but tends to result in a high false-positive rate. Conversely, employing conservative detection thresholds can decrease false positives but at the risk of missing metastases. This trade-off between sensitivity and specificity is particularly problematic in the detection of LNM, where undetected lesions can lead to suboptimal surgical management.
Multi-stage and cascaded detection strategies present a promising alternative by breaking down the detection process into sequential sub-tasks, each with distinct objectives. Typically, an initial stage is employed to generate candidate regions with high sensitivity, which is then followed by a refinement stage that conducts more discriminative classification and localization. Although cascaded designs have demonstrated potential in medical image analysis, existing methodologies often depend on straightforward model stacking and lack explicit task-driven asymmetry tailored to the clinical characteristics of LNM.
We propose an asymmetric context-aware cascade detection (ACCD) framework for the automated identification of cervical LNM in contrast-enhanced CT images of patients with HNSCC. This framework reconceptualizes LNM detection as a two-stage asymmetric learning problem. The initial stage functions as a high-sensitivity candidate generation module, using an attention-enhanced You Only Look Once version 8 (YOLOv8)-based one-stage detector that is optimized to maximize recall and ensure comprehensive coverage of potentially suspicious lymph node regions. The subsequent stage involves context-aware fine-grained refinement within adaptively expanded candidate regions, employing a Faster region-based convolutional neural network (R-CNN)-based detector that explicitly learns to suppress false positives introduced by the recall-focused first stage. By separating the sensitivity-driven candidate generation from precision-oriented refinement, the ACCD framework seeks to enhance the detection of small and ambiguous metastatic lymph nodes while maintaining clinically acceptable false-positive rates. We present this article in accordance with the STARD-AI reporting checklist (available at https://qims.amegroups.com/article/view/10.21037/qims-2026-0585/rc).
Methods
Datasets description
In this study, a dataset comprising 6,860 contrast-enhanced CT scans was retrospectively assembled from 481 patients newly diagnosed with histopathologically confirmed HNSCC at Fujian Cancer Hospital between 2020 and 2024. All participants met the following inclusion criteria: (I) histopathological confirmation of malignant tumors in the region of the oral cavity, oropharynx, larynx, nasopharynx, or hypopharynx; (II) histopathological confirmation of metastatic lymph node involvement via percutaneous needle biopsy or surgical resection; (III) completion of CT imaging at Fujian Cancer Hospital; (IV) age between 18 and 80 years; and (V) provision of written informed consent. The exclusion criteria included the absence of informed consent and/or incomplete imaging data.
CT images were acquired using scanners from three leading manufacturers: Siemens, GE, and Philips. To ensure data independence, all images from an individual patient were assigned to the same dataset partition. Specifically, 4,802 images (70%) were allocated for training, 686 images (10%) for validation, and 1,372 images (20%) for testing. The distribution of the annotated CT scans across the different scanner manufacturers and dataset partitions is detailed in Table 1. Each scan was independently annotated by two senior radiologists, with any discrepancies resolved through adjudication by a third experienced radiologist.
Table 1
| Manufacturer and model | Training set | Validation set | Testing set | Total |
|---|---|---|---|---|
| Siemens Healthineers | 2,203 | 315 | 629 | 3,147 |
| SOMATOM force | 1,282 | 183 | 366 | 1,831 |
| SOMATOM definition | 921 | 132 | 263 | 1,316 |
| GE Healthcare | 1,637 | 234 | 467 | 2,338 |
| Revolution CT | 1,010 | 144 | 289 | 1,443 |
| Optima CT660 | 627 | 90 | 178 | 895 |
| Philips Healthcare | 962 | 137 | 276 | 1,375 |
| iQon Spectral CT | 509 | 73 | 145 | 727 |
| Brilliance iCT | 453 | 64 | 131 | 648 |
| Total scans (n) | 4,802 | 686 | 1,372 | 6,860 |
| Percentage (%) | ≈70.0 | ≈10.0 | ≈20.0 | 100 |
CT, computed tomography.
The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study was approved by the Research Ethics Committee of Fujian Cancer Hospital in March 2025 (No. SQ2025028), and written informed consent was obtained from each participant.
The enrolled cohort comprised both clinically node-negative (cN0) and clinically node-positive (cN+) HNSCC patients, thereby reflecting the full clinical spectrum encountered in routine preoperative imaging. Of the 481 enrolled patients, 174 (36.2%) were classified as cN0 and 307 (63.8%) as cN+, the latter of which were further subdivided into cN1 (n=127, 26.3%), cN2 (n=98, 20.4%), and cN3 (n=82, 17.1%). The detailed distribution of cN status, together with other baseline clinicopathological characteristics, is summarized in Table 2. Inclusion of both cN0 and cN+ subgroups was deemed essential to evaluate the diagnostic performance of the deep learning-based cervical lymph node metastasis detection model (DL-CervLNM) across patients with overt nodal disease as well as those at risk of harboring occult or sub-centimeter metastatic nodes, the latter representing a particularly challenging clinical scenario in which conventional radiological assessment is most prone to false-negative interpretations.
Table 2
| Variables | Values (n=481), n (%) |
|---|---|
| Age (years) | |
| <60 | 164 (34.1) |
| ≥60 | 317 (65.9) |
| Gender | |
| Female | 158 (32.8) |
| Male | 323 (67.2) |
| Tumor location | |
| Oral cavity | 207 (43.0) |
| Oropharynx | 76 (15.8) |
| Larynx | 65 (13.5) |
| Nasopharynx | 114 (23.7) |
| Hypopharynx | 19 (4.0) |
| cT stage | |
| cT1 | 13 (2.7) |
| cT2 | 31 (6.4) |
| cT3 | 228 (47.4) |
| cT4 | 209 (43.5) |
| cN status | |
| cN0 | 174 (36.2) |
| cN1 | 127 (26.3) |
| cN2 | 98 (20.4) |
| cN3 | 82 (17.1) |
cN, clinical nodal stage; cT, clinical primary tumor stage; HNSCC, head and neck squamous cell carcinoma.
Data preprocessing
To ensure the provision of high-quality inputs for subsequent deep learning analyses, a comprehensive preprocessing pipeline was implemented for all retrospectively collected high-resolution CT scans. To address the heterogeneity inherent in imaging data from multiple manufacturers, a hybrid denoising strategy was employed. This strategy involved the use of Gaussian filtering to attenuate high-frequency noise through spatial smoothing, complemented by median filtering to suppress impulsive (salt-and-pepper) noise, thereby preserving essential anatomical edges and texture details. Following noise reduction, all images were subjected to intensity normalization and Hounsfield unit (HU) windowing to standardize pixel value distributions across the cohort. This step is critical for minimizing model sensitivity to non-biological variance, thereby ensuring that the network focuses on task-relevant features. Finally, rigorous quality control measures were implemented, and scans exhibiting significant motion artifacts or incomplete fields of view were excluded to maintain the integrity of the dataset.
To establish a reliable ground truth, manual annotations of the primary HNSCC tumors and suspicious cervical lymph nodes were conducted to delineate the ROIs. Two senior radiologists, blinded to clinical outcomes, independently segmented the volumes using standardized software. Any discrepancies in boundary definitions or nodal involvement were resolved through adjudication by a third expert radiologist with more than 15 years of experience. This collaborative annotation process ensured high inter-observer agreement and the histopathological validity of the spatial labels, thus providing the “gold standard” for supervised training. Building on this ROI delineation, three anatomical confounding structures, including skeletal muscle, major blood vessels, and cortical bone, were further annotated using bounding boxes to improve discrimination between metastatic lymph nodes and surrounding tissues. The resulting dataset comprised multiclass bounding box annotations with consensus labels derived from the aforementioned multi-reader adjudication framework.
To address the challenge of limited medical imaging samples and enhance the model’s spatial invariance, a comprehensive data augmentation strategy was implemented. This involved applying random rotations (±15°) to emulate natural variations in patient positioning on the scanner gantry, as well as random scaling (ranging from 0.8 to 1.2 times the original size) to accommodate the diverse anatomical sizes of tumors and lymph nodes across the patient population. These techniques, in conjunction with horizontal flipping, effectively increased the labeled training set from 4,802 to 9,604 images, thereby significantly enhancing the network’s capacity to generalize across various tumor morphologies and orientations. The detailed workflow for data preprocessing and augmentation is depicted in Figure 1.
Reader study design
A total of 15 radiologists participated in the reader study, categorized as follows: (I) junior doctors (n=5): board-certified radiologists with ≤3 years of post-certification experience in head and neck imaging; (II) intermediate doctors (n=5): radiologists with 4–8 years of experience in head and neck imaging; and (III) senior doctors (n=5): radiologists with ≥9 years of experience specializing in head and neck oncology. All readers were blinded to the clinical outcomes and pathological reports. Each reader independently interpreted all cases in the unassisted condition. Following a washout period of at least 4 weeks to minimize recall bias, the readers repeated the interpretation of the same cases with DL-CervLNM assistance. Cases were presented in a randomized order that differed between the two reading sessions. Inter-reader agreement was evaluated using intraclass correlation coefficients.
Problem formulation and framework overview
Accurate detection of LNM in HNSCC is of paramount importance but presents significant challenges due to the subtle nature of CT features (e.g., small size and ambiguous boundaries), which frequently result in inter-observer variability, workflow inefficiencies, and a heightened risk of under-staging. Conventional single-stage detectors often encounter difficulties in achieving an optimal balance between sensitivity and specificity under these conditions. To address this issue, we reconceptualize LNM detection as an asymmetric cascade learning problem, wherein the detection process is divided into two sequential sub-tasks with distinct optimization objectives.
Specifically, we introduce an ACCD framework (Figure 2), which comprises a high-sensitivity candidate generation stage followed by a context-aware fine detection stage. The initial stage employs a YOLOv8-based detector optimized to enhance sensitivity and ensure comprehensive coverage of potential LNM regions. The subsequent stage uses a Faster R-CNN-based detector to reduce false positives and refine spatial localization by leveraging localized contextual information. The entire framework is evaluated as a unified detector within the original image space.
High-sensitivity candidate generation with attention-enhanced YOLO
In the initial phase of our methodology, a task-specific YOLOv8-based detector is used to identify potential lymph node regions within the input CT image. This approach diverges from traditional YOLO applications, where the detector is typically fine-tuned to deliver final predictions directly. Instead, our framework’s first stage is intentionally restructured to serve as a high-sensitivity candidate generator, emphasizing recall over precision to mitigate the likelihood of overlooking metastatic lymph nodes.
To improve the depiction of small and low-contrast lymph nodes, the standard YOLO architecture is augmented with a lightweight attention mechanism, which is incorporated into both the backbone and feature aggregation pathways. Specifically, channel-spatial attention modules are introduced following selected convolutional stages and feature pyramid outputs, allowing the network to dynamically highlight lymph node-relevant features while suppressing background interference from adjacent vessels, muscles, and glands. For a given intermediate feature map F, the attention-enhanced feature representation is defined as follows:
where and denote channel and spatial attention operations, respectively. This modification improves feature discrimination for subtle lymph node structures without introducing substantial computational overhead.
In addition to architectural enhancements, the YOLO detector is trained using a recall-oriented optimization strategy specifically designed to meet the clinical objective of preserving candidate regions. During the training phase, more lenient positive assignment criteria are employed to encourage the model to activate on lymph nodes with borderline appearances. Multi-scale training is used to enhance robustness against size variability across different patients and anatomical levels.
During inference, a deliberately low confidence threshold is applied to retain a diverse array of candidate regions, and non-maximum suppression is executed conservatively to preserve overlapping detections. Consequently, the first stage generates a dense set of candidate bounding boxes that intentionally includes both true metastatic lymph nodes and challenging false positives. These candidates are subsequently forwarded to the second stage for refinement and are not considered final predictions.
Formally, given a CT image I, the first stage outputs a candidate set as follows:
where bi denotes the bounding box coordinates in the original image space, and si represents the corresponding confidence score.
Context-aware candidate refinement
For each candidate region identified in the initial stage, an adaptive spatial expansion is implemented to incorporate the surrounding anatomical context. In clinical practice, radiologists evaluate lymph node malignancy by considering not only the node itself but also the perinodal tissue, adjacent vessels, and fat planes. Inspired by this practice, we propose a context-aware region cropping strategy that enlarges candidate regions by a predefined ratio relative to their size. The expanded region is then extracted from the original CT image and resized to a fixed resolution for further analysis. This strategy enables the refinement model to use contextual information while maintaining focus on the candidate lymph node.
To generate training samples for the refinement stage, candidate regions are aligned with ground-truth annotations using the intersection-over-union (IoU) metric. Candidate regions with substantial overlap with annotated metastatic lymph nodes are designated as positive samples, whereas those with minimal overlap are designated as negative samples. Ambiguous candidates are excluded to minimize label uncertainty.
This labeling strategy exhibits a distinct asymmetry between the two stages: the initial stage is designed to over-generate candidates to enhance sensitivity, while the subsequent stage is trained on a purposefully imbalanced set of challenging negatives that mirror the prevalent false-positive patterns identified by the first-stage detector. This asymmetric configuration allows the refinement stage to specifically learn to mitigate false positives resulting from the high-sensitivity candidate generation process.
Fine-grained detection with region-based refinement
In the second stage, a Faster R-CNN architecture is employed to conduct fine-grained detection within the contextually expanded candidate regions. This region-based framework, which incorporates ROI-based feature pooling along with localized classification and regression heads, is particularly adept at discerning subtle differences in appearance between metastatic lymph nodes and confounding anatomical structures.
During the training process, the refinement model is optimized through joint classification and bounding box regression objectives. To enhance the suppression of false positives, hard negative samples generated by the first-stage detector are assigned increased weights during training, thereby directing the model’s attention toward clinically challenging confounders.
The refinement stage subsequently produces both confidence scores and refined bounding boxes, aligned with the coordinate system of the cropped regions.
Inference and coordinate reprojection
During the inference phase, the input CT image undergoes initial processing by the high-sensitivity candidate generator. Subsequently, each candidate region is refined by the second-stage detector. The predicted bounding boxes are then reprojected onto the original image coordinate system, taking into account the spatial offsets of the cropped regions.
The final detection outcomes comprise refined bounding boxes and confidence scores, defined within the original CT image space, representing the output of the proposed ACCD framework. This framework is assessed as an integrated object detector. The final predictions are compared against ground-truth annotations in the original CT image space using standard object detection metrics, including mean average precision (mAP) at various IoU thresholds, recall, and precision. This evaluation protocol ensures a fair comparison with baseline detectors and addresses the clinical requirement for both high sensitivity and low false-positive rates.
Training configuration and statistical evaluation
The model was trained using NVIDIA RTX 4080SUPER graphics processing units, each with 16 GB of video random access memory, and the AdamW optimizer was employed. A carefully adjusted learning rate schedule was implemented to ensure stable convergence. The data augmentation strategy included random rotations within ±15 degrees and intensity scaling from 0.8 to 1.2 times the original values. To address overfitting, the training protocol incorporated an early stopping mechanism, with a patience threshold set at 20 epochs, based on monitoring the validation mAP.
The proposed framework was evaluated as a unified object detector, with final predictions compared against ground-truth annotations in the original CT image space using standard object detection criteria. This protocol ensures a fair comparison with baseline detectors while meeting clinical requirements for high sensitivity and low false-positive rates. The model’s performance was rigorously assessed using multiple metrics, including mAP at an mAP@0.5, precision, recall, F1-score, sensitivity, specificity, area under the receiver operating characteristic curve (AUC), and inference efficiency measured in frames per second (FPS).
The mAP@0.5 metric provides a comprehensive measure of model accuracy in both localization and classification, requiring a minimum overlap of 50% between predicted and ground-truth regions, and is a widely accepted standard in image detection tasks. Precision and recall are critical metrics that evaluate a model’s ability to minimize false positives and identify true positives, respectively, both of which are crucial for balancing unnecessary clinical interventions against the risk of missed diagnoses. The F1-score, defined as the harmonic mean of precision and recall, serves as a comprehensive metric for evaluating classifier performance. Sensitivity and specificity further elucidate diagnostic reliability, with sensitivity reflecting the accurate identification of true LNM cases, thereby minimizing false negatives, and specificity indicating the correct classification of non-metastatic nodes, thereby preventing unnecessary examinations. The AUC encapsulates discriminative performance across all classification thresholds, providing a robust measure that is independent of any particular cutoff. Additionally, inference speed, quantified in FPS, was included to evaluate computational efficiency, as rapid analysis is crucial for supporting real-world clinical workflows where timely image interpretation directly influences diagnosis and treatment planning. Collectively, these metrics encompass both the algorithmic efficacy and the clinical applicability of the proposed model.
To assess statistical robustness, we employed 1,000 bootstrap resamples to estimate 95% confidence intervals (CIs). Sensitivity and specificity comparisons were conducted using McNemar’s test, while DeLong’s test was used to analyze differences in the AUC values. All experiments were conducted using Python version 3.12.7, with PyTorch 2.5.1 and PyTorch-CUDA 12.4, to ensure reproducibility and scalability. These architectural and mechanistic optimizations enable DL-CervLNM to achieve improved detection performance while maintaining relatively low computational complexity and training efficiency, making it suitable for large-scale head and neck CT analysis in clinical and computer-aided diagnosis settings.
Results
Experimental results
For the evaluation of the proposed ACCD framework, an independent test set comprising 1,372 CT scans was used, with patient-level partitioning implemented to maintain data independence between the training and testing cohorts.
To comprehensively evaluate the detection performance of DL-CervLNM, a benchmarking study was conducted, comparing it against several widely used object detection models [i.e., Real-Time Detection Transformer (RT-DETR), You Only Look Once version 5 (YOLOv5), YOLOv8, and Faster R-CNN]. These models were chosen due to their established effectiveness and diverse architectural designs. All models were trained using identical dataset partitions and consistent preprocessing, augmentation, and optimization settings, ensuring that any performance differences were attributable to model architecture rather than experimental bias.
The evaluation of DL-CervLNM for LNM detection resulted in an mAP@0.5 of 94.9% and an F1-score of 91.2%, with an inference speed of 38.5 FPS. Comparative results with other models are detailed in Table 3. Notably, the model exhibited proficiency in detecting small-volume lesions. Moreover, DL-CervLNM demonstrated mAP@0.5 values exceeding 91% for various anatomical structures, including muscle, blood vessels, and bone, as detailed in Table 4.
Table 3
| Model | mAP@0.5 | Precision | Recall | F1-score | FPS |
|---|---|---|---|---|---|
| RT-DETR | 89.6% | 84.1% | 81% | 82.5% | 42.6 |
| YOLOv5 | 82.6% | 79.7% | 74.4% | 77% | 48.5 |
| YOLOv8 | 85.6% | 83.7% | 74.4% | 78.8% | 51.5 |
| Faster R-CNN | 90.9% | 87.5% | 85.1% | 86.3% | 34.7 |
| DL-CervLNM | 94.9% | 92.1% | 90.3% | 91.2% | 38.5 |
DL-CervLNM, deep learning-based cervical lymph node metastasis detection model; Faster R-CNN, Faster region-based convolutional neural network; FPS, frames per second; mAP@0.5, mean average precision at an intersection-over-union threshold of 0.5; RT-DETR, Real-Time Detection Transformer; YOLOv5, You Only Look Once version 5; YOLOv8, You Only Look Once version 8.
Table 4
| Class | mAP@0.5 | Precision | Recall |
|---|---|---|---|
| LNM | 94.9% | 92.1% | 90.3% |
| Muscle | 98.1% | 92% | 96.7% |
| Blood vessel | 91.3% | 88% | 86% |
| Bone | 98.7% | 94.1% | 99.2% |
DL-CervLNM, deep learning-based cervical lymph node metastasis detection model; LNM, lymph node metastasis; mAP@0.5, mean average precision at an intersection-over-union threshold of 0.5.
Precision-recall (PR) and F1-score analyses were conducted on the test set to evaluate and compare these models. As depicted in Figures 3A,3B, DL-CervLNM consistently achieved the highest precision across most recall intervals and maintained elevated F1 scores across nearly all confidence thresholds, with a peak F1-score exceeding 0.9. Conversely, Faster R-CNN and RT-DETR exhibited moderate performance, whereas YOLOv5 and YOLOv8 exhibited significant declines in precision at higher recall levels.
Receiver operating characteristic (ROC) curve analysis was conducted to assess model performance in detecting LNM. As depicted in Figure 3C, DL-CervLNM achieved an AUC of 0.980, with a 95% CI ranging from 0.968 to 0.992. Conversely, the AUC values for the YOLOv5, YOLOv8, Faster R-CNN, and RT-DETR models were 0.862, 0.909, 0.926, and 0.935, respectively. DeLong’s test confirmed the statistical significance of these differences (P<0.001). The ROC curve for DL-CervLNM consistently surpassed those of the other models across most false-positive rates, indicating a superior balance between sensitivity and specificity.
To comprehensively evaluate the robustness of DL-CervLNM across diverse imaging sources, a stratified analysis based on scanner type was performed on the independent test cohort. Performance metrics were assessed separately for CT images obtained from the Siemens, GE, and Philips scanners. DL-CervLNM consistently demonstrated high detection performance across all scanner manufacturers, with mAP@0.5 values exceeding 92% and AUC values exceeding 0.96, indicating minimal performance variability attributable to scanner manufacturer. These findings suggest that the proposed framework maintains stable detection capabilities despite variations in acquisition hardware and reconstruction parameters.
Visual comparison of LNM detection models
To further investigate the qualitative detection performance of different models, we analyzed their predictions on an identical axial CT slice featuring multiple metastatic lymph nodes. Conventional detectors, such as YOLOv5, YOLOv8, Faster R-CNN, and RT-DETR (Figure 4A), successfully identified several lymph node candidates. However, their predictions often exhibited spatial inaccuracies, incomplete lesion coverage, and missed detections, particularly for small or closely adjacent nodes. These limitations imply inadequate contextual modeling and diminished sensitivity to subtle morphological variations when compared to the ground truth.
In contrast, DL-CervLNM exhibited significantly enhanced localization accuracy, generating more compact and anatomically consistent bounding boxes with fewer false-positive responses. The model demonstrated notable robustness in detecting small-volume and densely clustered lymph nodes, which can be attributed to its multi-scale feature representation and cervical region-aware optimization strategy.
Consistent with these observations, the probability-based heatmap produced by DL-CervLNM demonstrated a highly concentrated response pattern, with activation localized to areas corresponding to clinically significant lymph node structures while largely suppressing background tissues (Figure 4B). Notably, increased confidence responses were detected in low-contrast regions associated with subtle metastatic involvement, indicating that the model effectively captured discriminative features beyond mere intensity cues. This behavior qualitatively suggests that DL-CervLNM enhances sensitivity and interpretability through context-aware feature learning, thereby substantiating its superior detection performance in the complex anatomy of the cervix.
Comparison of diagnostic performance
DL-CervLNM demonstrated significantly improved LNM detection compared with experienced radiologists on a test set comprising 1,372 CT images, including 327 LNM-negative CT images, with the remaining LNM-positive CT images containing a total of 1,228 annotated metastatic lymph nodes. At the lesion level, the model achieved a sensitivity of 90.6% (1,112/1,228). At the image level, the model achieved a specificity of 94.5% (309/327), representing absolute improvements of 2.2% and 4.3%, respectively, compared with the radiologists’ sensitivity and specificity of 88.4% and 90.2%. Additionally, the model’s AUC was 0.980, surpassing the radiologists’ consensus AUC of 0.884 by 0.096.
In terms of detecting micro-metastases, DL-CervLNM achieved an identification rate of 93.2%, which was significantly higher than the 76.4% achieved by radiologists, representing an improvement of 16.8% (P<0.01, McNemar’s test). This improvement corresponds to the detection of 17 additional micro-metastases per 100 cases. Notably, 23.6% of false-negative diagnoses in clinical practice were attributed to these small lesions, as detailed in Table 5.
Table 5
| Metric | DL-CervLNM (95% CI) | Radiologists (95% CI) | Difference (Δ) | P value |
|---|---|---|---|---|
| Sensitivity | 90.6% (89.3–95.1%) | 88.4% (83.9–91.8%) | +2.2% | 0.021* |
| Specificity | 94.5% (92.6–97.6%) | 90.2% (86.2–93.1%) | +4.3% | 0.008** |
| AUC | 0.980 (0.968–0.992) | 0.884 (0.852–0.912) | +0.096 | <0.001*** |
| Micro-metastasis detection rate | 93.2% (88.7–96.0%) | 76.4% (69.8–81.9%) | +16.8% | <0.001*** |
| PPV | 94.5% (91.2–96.6%) | 89.1% (84.7–92.4%) | +5.4% | 0.009** |
| NPV | 93.9% (90.1–96.2%) | 89.5% (85.1–92.8%) | +4.4% | 0.033* |
| F1-score | 0.912 | 0.866 | +0.046 | none |
Radiologist performance represents the average performance across 15 readers. *, P<0.05; **, P<0.01; ***, P<0.001. AUC, area under the receiver operating characteristic curve; CI, confidence interval; DL-CervLNM, deep learning-based cervical lymph node metastasis detection model; NPV, negative predictive value; PPV, positive predictive value.
Across varying levels of experience, the deployment of DL-CervLNM was associated with increased and less variable diagnostic AUC values for LNM detection compared to unassisted readings. As depicted in the violin plots (Figure 5), the AUC distributions for the junior, intermediate, and senior physicians exhibited an upward shift with the use of assistance, accompanied by reduced dispersion, thereby indicating enhanced consistency at the group level. The relative improvement was most pronounced among the junior readers and decreased with increasing experience, while the AUC values achieved with model assistance approached, but did not exceed, the model’s standalone performance (AUC =0.980). These patterns consistently demonstrate an assistive effect across all reader groups.
To assess the diagnostic efficacy of the proposed DL-CervLNM in clinical practice, a comparative analysis was conducted between the model-assisted workflow and traditional radiologist interpretation, as detailed in Table 6. For clinician evaluation, a stratified subsample of 572 CT scans was randomly selected from the 1,372-scan test set, with proportional representation maintained across scanner manufacturers (Siemens: 262; GE: 195; Philips: 115) and lesion-size strata (macro-metastases ≥10 mm: 341; micro-metastases <10 mm: 231).
Table 6
| Metric | Radiologists with DL-CervLNM | Radiologists | Improvement |
|---|---|---|---|
| Processing time per case | 4.2±0.8 min | 8.7±1.5 min | 51.7% |
| Micro-metastasis identification time | 0.3±0.1 s | 2.1±0.4 min | 99.8% |
| Report generation time | <10 s | 12.3±2.7 min | 98.6% |
Data are presented as mean ± standard deviation across 15 radiologists.
Thus, incorporating DL-CervLNM significantly enhanced overall workflow efficiency. Radiologists using the model completed case assessments and report generation considerably more rapidly than through manual interpretation alone. Further, the model significantly reduced the processing time for micro-metastasis detection, underscoring its potential to optimize diagnostic workflows and improve real-time clinical decision-making.
Discussion
Early detection and precise localization of cervical LNM are crucial for accurate staging and optimal therapeutic planning in patients with HNSCC. Experienced radiologists exhibit high diagnostic reliability in identifying macro-metastases (short-axis diameter >10 mm); however, their sensitivity significantly decreases when assessing sub-centimeter nodes. This diagnostic challenge is exacerbated by considerable inter-observer variability, as the interpretation of borderline nodes often remains subjective. Further, the manual evaluation of extensive imaging datasets is inherently labor-intensive and time-consuming, resulting in suboptimal clinical workflow efficiency. These limitations contribute to increased false-negative rates and a risk of under-staging. Consequently, enhancing the detection of micro-metastases is clinically imperative to refine the extent of neck dissection and guide the application of adjuvant therapy, ultimately improving patient prognosis (23-26).
This study focused on the initial and clinically crucial phase of the diagnostic workflow, specifically the automated localization of suspicious cervical lymph nodes on contrast-enhanced CT images (27). To achieve this, we developed and assessed an ACCD framework, which explicitly reconceptualizes LNM detection as a two-stage asymmetric learning problem. By separating high-sensitivity candidate generation from precision-oriented refinement, the proposed framework effectively addresses the inherent trade-off between sensitivity and specificity that constrains conventional single-stage detectors in tasks involving the detection of small targets.
The initial phase of the ACCD framework is meticulously optimized to enhance recall by employing attention-enhanced candidate generation, thereby ensuring thorough coverage of potential metastatic lymph nodes. Conversely, the subsequent phase is dedicated to minimizing false positives through context-aware region-based refinement. This task-oriented cascade architecture facilitates the efficient detection of subtle and ambiguous lymph nodes without adversely affecting workflow efficiency. An evaluation conducted on an independent patient-level test cohort revealed robust localization and detection performance, with the framework achieving a mAP of 94.9% and an AUC of 0.980.
The proposed framework demonstrated enhanced sensitivity for detecting micro-metastases compared to radiologists, underscoring its potential clinical utility in early-stage disease and challenging scenarios involving small targets. These findings indicate that asymmetric cascade detection offers a more appropriate methodological paradigm for LNM detection than monolithic detection architectures, especially in applications where missed detections have substantial clinical implications.
The automatic detection of small-volume lesions poses several distinct challenges, as elucidated in this study. Primarily, the dataset demonstrated a significant imbalance, characterized by a predominance of slices without LNM, which led to an increased rate of false positives by the classifier (28,29). This issue was particularly evident when small vessels or adjacent structures closely resembled LNM on individual axial slices. This challenge is not unique to algorithms; experts often need to review multiple slices or use multiplanar reconstruction to confirm suspicious findings, underscoring the limitations of relying exclusively on single-slice information (30-33). Second, although the integration of multi-scale feature encoding and attention modules was intended to improve visibility across lesions of different sizes, the model’s confidence decreased in cases of severe metal artifacts, anatomical distortions, or inconsistent contrast phases, sometimes leading to mislocalization. Third, the present study used a fully supervised learning methodology reliant on expert manual annotations. However, future research should investigate semi-supervised or weakly supervised approaches, such as employing high-confidence pseudo-labels, to alleviate the annotation burden associated with large-scale, multi-center datasets (34-37).
Numerous previous studies have underscored the potential applications of AI in imaging tasks associated with head and neck tumors (38-40). A recent systematic review and meta-analysis by Valizadeh et al. (2025) synthesized 23 eligible studies and reported a pooled AUC of 92% for the deep learning-based diagnosis of LNM in head and neck cancers (41), with PET/CT-based models yielding the highest sensitivity and specificity. The authors nonetheless noted that the majority of included studies relied solely on internal validation, which substantially constrains clinical translatability and external generalizability. Building on this evidence base, Song et al. (2025) provided a comprehensive review of AI applications across the HNSCC care continuum (42), encompassing radiologic detection, pathologic interpretation, and treatment response prediction. The authors emphasized that, despite the encouraging diagnostic performance reported in single-center cohorts, the field continues to be hampered by tumor heterogeneity, limited model interpretability, and a paucity of prospective multi-center validation. Similarly, Stawarz et al. (2025) conducted a focused systematic review on oropharyngeal cancer and demonstrated that radiomics- and deep learning-based models can preoperatively predict both LNM and extranodal extension with promising discriminative performance (43), while concurrently highlighting persistent shortcomings in Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis (TRIPOD)-compliant reporting and decision-curve analyses across the published literature.
Beyond these aggregate analyses, several recent original studies have advanced specific methodological approaches for LNM assessment, each with distinct strengths and limitations. In their 2025 study published in Head & Neck, Troise et al. employed radiomic features extracted from MRI in combination with a random forest classifier to predict occult cervical metastasis in a cohort of 75 patients with early-stage (clinical T1–2 tumor status and N0 nodal status) oral cavity carcinoma (44). The model demonstrated a test-set AUC of 0.80 and an overall accuracy of 80%. However, this method functions as a patient-level risk stratification tool limited to a single cancer subsite, relies on manual MRI segmentation, and is not feasible for real-time clinical application. In contrast, Liao et al. (2025) introduced a transfer-learning no-new U-Net framework in Scientific Reports (45), trained on a dataset of 11,013 annotated CT lymph nodes from 626 head and neck cancer patients across four hospitals. This model achieved segmentation accuracy comparable to that of experienced oncologists, with a Dice similarity coefficient ranging from 0.72 to 0.74. Nevertheless, its detection sensitivity on the internal test cohort was restricted to 54.6%, with a false-positive rate of 3.4 per volume, and its performance was particularly limited in identifying sub-centimeter micro-metastatic nodes.
Extending these efforts into a multi-vendor setting, Rinneburger et al. (2026) assessed a generic deep learning-based cervical lymph node segmentation model in 125 patients with head and neck cancer using CT data acquired from multiple scanner platforms (46). Although the model exhibited acceptable performance for typical lymph nodes, it systematically underperformed for nodes exhibiting central necrosis or atypically large dimensions, thereby reaffirming that off-the-shelf segmentation networks are inadequate for capturing the heterogeneous morphology of HNSCC metastases.
In a separate study, Qi et al. (2025) developed a deep learning–based model using dual-energy CT (47), using multi-sequence fused images, including iodine maps, fat maps, 70-keV monoenergetic images, and RHO/Z maps. This model underwent prospective validation in a cohort of 354 patients with oral squamous cell carcinoma, showing promising results in LNM discrimination. Nevertheless, the applicability of this approach is limited by its reliance on specialized dual-energy CT acquisition protocols, which are not available on standard single-energy scanners, thus hindering its generalizability in resource-constrained environments.
More recently, Dayan et al. (2026) developed an imaging-based AI model for extranodal extension detection and outcome prediction in human papillomavirus-positive oropharyngeal carcinoma (48), illustrating that AI can extend beyond metastatic node localization to downstream prognostic stratification. However, consistent with preceding work, Dayan et al.’s study was limited by its reliance on a single imaging modality and a relatively narrow disease subsite, collectively restricting the broader applicability of its findings.
In contrast to the approaches outlined above, the DL-CervLNM framework introduced in this study employs an asymmetric cascade detection architecture. Evaluated on a cohort of 481 HNSCC patients encompassing all primary subsites (test set: 1,372 CT slices), the framework achieved an AUC of 0.980 and a micro-metastasis detection rate of 93.2%, significantly exceeding the 76.4% detection rate of radiologists (P<0.001). Further, it supports real-time inference at 38.5 FPS, thereby facilitating seamless integration into Picture Archiving and Communication System (PACS) environments. These performance metrics compare favorably with those of all the reference studies cited above. Notably, DL-CervLNM is not intended to replace radiologist expertise but to function as an AI-assisted second reader. This approach demonstrably enhances diagnostic consistency across varying levels of radiologist experience and reduces interpretation time, thereby providing tangible incremental clinical value within existing oncological workflows.
The primary innovation of our proposed framework does not reside in incremental architectural modifications but rather in a task-driven reformulation of the detection problem. Specifically, LNM detection is decomposed into two asymmetric stages, each with distinct optimization objectives: a high-sensitivity candidate generation stage aimed at minimizing false-negative detections, and a context-aware refinement stage explicitly trained to suppress false positives introduced by the recall-oriented first stage. This asymmetric cascade design directly aligns with the clinical priority of avoiding missed metastatic lymph nodes during preoperative assessment and distinguishes our proposed framework from conventional single-stage or naively cascaded detectors.
From a clinical standpoint, the ACCD framework offers complementary value at various stages of the diagnostic workflow. Initially, it can operate as an automated pre-screening tool, rapidly identifying high-risk cases that require further examination by radiologists. Subsequently, during routine image interpretation, the system may serve as a concurrent computer-aided second reader, providing candidate regions and confidence indicators to reduce the risk of missing micro-metastases. Our findings indicate that such support can increase inter-reader consistency and reduce interpretation time, thereby facilitating its seamless integration into existing PACS environments or structured reporting workflows. However, optimal confidence thresholds may require task- and institution-specific calibration, and further validation across diverse clinical settings remains essential.
In this single-center retrospective study, the proposed ACCD framework exhibited strong performance in identifying cervical LNM on contrast-enhanced CT images. This was evidenced by consistent enhancements in localization accuracy, inter-reader agreement, and workflow efficiency. These findings highlight the potential clinical utility of the asymmetric cascade detection framework for detecting small lesions and underscore the need for multi-center prospective studies to establish its generalizability and real-world clinical impact.
Limitations
DL-CervLNM exhibited strong performance across various scanner types within this cohort; however, the study had a number of limitations. First, it was a single-center retrospective analysis with relatively standardized acquisition protocols. Consequently, the findings primarily underscore the applicability of the proposed framework in comparable clinical environments. To further substantiate its generalizability and enable large-scale clinical implementation, external validation using multi-center datasets with greater protocol diversity is essential. Second, the evaluation metrics were limited to image- and case-level detection and classification, without prospective validation against clinically significant endpoints such as staging concordance, alterations in treatment decisions, or follow-up outcomes. Third, performance benchmarks, such as maintaining an 85% specificity rate, may require calibration across various clinical environments to effectively balance sensitivity, positive predictive value, and the minimization of unnecessary biopsies.
Future research should focus on: (I) conducting multi-center, prospective trials to assess the impact of the DL-CervLNM framework on staging and treatment planning; (II) integrating cross-modality and multi-phase imaging techniques (e.g., ultrasound, MRI, or PET) to enhance the diagnostic robustness of the model; and (III) developing weakly supervised or point-annotation strategies for micro-lesions to alleviate the annotation burden and bolster the study’s findings.
Conclusions
In this diagnostic study, we developed and validated a deep learning framework, termed DL-CervLNM, specifically optimized for the detection of LNM in patients with HNSCC. Using a large-scale, meticulously annotated CT dataset, the model effectively addresses several critical challenges in current clinical practice, including low sensitivity for small-volume metastases, significant inter-reader variability, and suboptimal efficiency associated with manual radiological review. DL-CervLNM exhibited exceptional efficacy in detecting sub-centimeter lesions, which are often missed in traditional workflows, while also achieving a diagnostic throughput that considerably exceeded that of experienced radiologists. When employed as a clinical decision support tool, the model significantly enhanced interpretative consistency across varying levels of expertise, thereby minimizing subjective discrepancies. Despite these encouraging results, prospective multi-center trials and cross-platform validation are needed to confirm its broader applicability. Additionally, future research should aim to refine standardized thresholding strategies and develop seamless clinical integration protocols to enable its widespread adoption in oncological workflows.
Acknowledgments
None.
Footnote
Reporting Checklist: The authors have completed the STARD-AI reporting checklist. Available at https://qims.amegroups.com/article/view/10.21037/qims-2026-0585/rc
Data Sharing Statement: Available at https://qims.amegroups.com/article/view/10.21037/qims-2026-0585/dss
Funding: This work was supported in part by
Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://qims.amegroups.com/article/view/10.21037/qims-2026-0585/coif). All authors report that this work was supported in part by the Science and Technology Cooperation Program Project of CAS and Fujian Province (Nos. 2025T3030, 2026T3009, and 2026T3019); in part by the Outstanding Youth Project of “Forward Looking Leap Forward” of Haixi Institutes, CAS (No. CXZX-2024-JQ04); in part by the Major Science and Technology Project of Fuzhou (No. 2025-ZD-032); by the Joint Funds for the Innovation of Science and Technology, Fujian Province (No. 2024Y9604); and by the Natural Science Foundation of Fujian Province (No. 2025J011173). The authors have no other conflicts of interest to declare.
Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. This study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. This study was approved by the Research Ethics Committee of Fujian Cancer Hospital in March 2025 (No. SQ2025028). All imaging data were collected retrospectively and anonymized prior to analysis. Written informed consent was obtained from each participant before inclusion in the study.
Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.
References
- Chamoli A, Gosavi AS, Shirwadkar UP, Wangdale KV, Behera SK, Kurrey NK, Kalia K, Mandoli A. Overview of oral cavity squamous cell carcinoma: Risk factors, mechanisms, and diagnostics. Oral Oncol 2021;121:105451. [Crossref] [PubMed]
- Gong Y, Bao L, Xu T, Yi X, Chen J, Wang S, Pan Z, Huang P, Ge M. The tumor ecosystem in head and neck squamous cell carcinoma and advances in ecotherapy. Mol Cancer 2023;22:68. [Crossref] [PubMed]
- Barsouk A, Aluru JS, Rawla P, Saginala K, Barsouk A. Epidemiology, Risk Factors, and Prevention of Head and Neck Squamous Cell Carcinoma. Med Sci (Basel) 2023;11:42. [Crossref] [PubMed]
- Lee AM, Weaver AN, Acosta P, Harris L, Bowles DW. Review of Current and Future Medical Treatments in Head and Neck Squamous Cell Carcinoma. Cancers (Basel) 2024;16:3488. [Crossref] [PubMed]
- Ji H, Hu C, Yang X, Liu Y, Ji G, Ge S, Wang X, Wang M. Lymph node metastasis in cancer progression: molecular mechanisms, clinical significance and therapeutic interventions. Signal Transduct Target Ther 2023;8:367. [Crossref] [PubMed]
- Song B, Leroy A, Yang K, Khalighi S, Pandav K, Dam T, Lee J, Stock S, Li XT, Sonuga J, Fu P, Koyfman S, Saba NF, Patel MR, Madabhushi A. Deep Learning Model of Primary Tumor and Metastatic Cervical Lymph Nodes From CT for Outcome Predictions in Oropharyngeal Cancer. JAMA Netw Open 2025;8:e258094. [Crossref] [PubMed]
- Chen X, Zhang L, Lu H, Tan Y, Li B. Development and validation of a nomogram to predict cervical lymph node metastasis in head and neck squamous cell carcinoma. Front Oncol 2023;13:1174457. [Crossref] [PubMed]
- Payne K, Nenclares P, Schilling C. The impact of elective cervical lymph node treatment on the tumour immune response in head and neck squamous cell carcinoma: time for a change in treatment strategy? BJC Rep 2024;2:68. [Crossref] [PubMed]
- Struckmeier AK, Yekta E, Agaimy A, Kopp M, Buchbender M, Moest T, Lutz R, Kesting M. Diagnostic accuracy of contrast-enhanced computed tomography in assessing cervical lymph node status in patients with oral squamous cell carcinoma. J Cancer Res Clin Oncol 2023;149:17437-50. [Crossref] [PubMed]
- Luo YH, Mei XL, Liu QR, Jiang B, Zhang S, Zhang K, Wu X, Luo YM, Li YJ. Diagnosing cervical lymph node metastasis in oral squamous cell carcinoma based on third-generation dual-source, dual-energy computed tomography. Eur Radiol 2023;33:162-71. [Crossref] [PubMed]
- Stöcker C, Greve J, Beer M, Hosch B, Barth TFE, Hoffmann TK, von Witzleben A. Contrast-Enhanced Computed Tomography (CT) With Concordant Sonography as Sufficient Early Detection Tools for Recurrent and Persistent Cervical Metastases After (Chemo)radiotherapy (CRT). Head Neck 2025;47:999-1005. [Crossref] [PubMed]
- Yoon SY, Lee KS, Bezuidenhout AF, Kruskal JB. Spectrum of Cognitive Biases in Diagnostic Radiology. Radiographics 2024;44:e230059. [Crossref] [PubMed]
- Madsen CB, Rohde M, Gerke O, Godballe C, Sørensen JA. Diagnostic Accuracy of Up-Front PET/CT and MRI for Detecting Cervical Lymph Node Metastases in T1-T2 Oral Cavity Cancer-A Prospective Cohort Study. Diagnostics (Basel) 2023;13:3414. [Crossref] [PubMed]
- Liu J, Wang Z, Shao H, Qu D, Liu J, Yao L. Improving CT detection sensitivity for nodal metastases in oesophageal cancer with combination of smaller size and lymph node axial ratio. Eur Radiol 2018;28:188-95. [Crossref] [PubMed]
- Chung MS, Cheng KL, Choi YJ, Roh JL, Lee YS, Lee SS, Lee JH, Baek JH. Interobserver reproducibility of cervical lymph node measurements at CT in patients with head and neck squamous cell carcinoma. Clin Radiol 2016;71:1226-32. [Crossref] [PubMed]
- Alsibani A, Alqahtani A, Almohammadi R, Islam T, Alessa M, Aldhahri SF, Al-Qahtani KH. Comparing the Efficacy of CT, MRI, PET-CT, and US in the Detection of Cervical Lymph Node Metastases in Head and Neck Squamous Cell Carcinoma with Clinically Negative Neck Lymph Node: A Systematic Review and Meta-Analysis. J Clin Med 2024;13:7622. [Crossref] [PubMed]
- Dong D, Luo Z, Zheng Y, Liang Y, Zhao P, Feng L, Wang D, Cao Y, Zhao Z, Ma Y. Application of deep learning-based diagnostic systems in screening asymptomatic COVID-19 patients among oversea returnees. J Infect Dev Ctries 2022;16:1706-14. [Crossref] [PubMed]
- Chen X, Xu H, Qi Q, Sun C, Jin J, Zhao H, et al. AI-based chest CT semantic segmentation algorithm enables semi-automated lung cancer surgery planning by recognizing anatomical variants of pulmonary vessels. Front Oncol 2022;12:1021084. [Crossref] [PubMed]
- Bruixola G, Dualde-Beltrán D, Jimenez-Pastor A, Nogué A, Bellvís F, Fuster-Matanzo A, Alfaro-Cervelló C, Grimalt N, Salhab-Ibáñez N, Escorihuela V, Iglesias ME, Maroñas M, Alberich-Bayarri Á, Cervantes A, Tarazona N. CT-based clinical-radiomics model to predict progression and drive clinical applicability in locally advanced head and neck cancer. Eur Radiol 2025;35:4277-88. [Crossref] [PubMed]
- Scapicchio C, Gabelloni M, Barucci A, Cioni D, Saba L, Neri E. A deep look into radiomics. Radiol Med 2021;126:1296-311. [Crossref] [PubMed]
- Oreiller V, Andrearczyk V, Jreige M, Boughdad S, Elhalawani H, Castelli J, et al. Head and neck tumor segmentation in PET/CT: The HECKTOR challenge. Med Image Anal 2022;77:102336. [Crossref] [PubMed]
- Wang Y, Lombardo E, Huang L, Avanzo M, Fanetti G, Franchin G, Zschaeck S, Weingärtner J, Belka C, Riboldi M, Kurz C, Landry G. Comparison of deep learning networks for fully automated head and neck tumor delineation on multi-centric PET/CT images. Radiat Oncol 2024;19:3. [Crossref] [PubMed]
- Hoang JK, Vanka J, Ludwig BJ, Glastonbury CM. Evaluation of cervical lymph nodes in head and neck cancer with CT and MRI: tips, traps, and a systematic approach. AJR Am J Roentgenol 2013;200:W17-25. [Crossref] [PubMed]
- Lorusso G, Maggialetti N, Laugello F, Garofalo A, Villanova I, Greco S, Morelli C, Pignataro P, Lucarelli NM, Stabile Ianora AA. Diagnostic Performance of Magnetic Resonance Sequences in Staging Lymph Node Involvement and Extranodal Extension in Head and Neck Squamous Cell Carcinoma. Diagnostics (Basel) 2025;15:1251. [Crossref] [PubMed]
- Wang L, Liao W, Zhang S, Wang G. Head and Neck Tumor Segmentation of MRI from Pre- and Mid-Radiotherapy with Pre-Training, Data Augmentation and Dual Flow UNet. Head Neck Tumor Segm MR Guid Appl (2024) 2025;15273:75-86. [Crossref] [PubMed]
- Erdur AC, Rusche D, Scholz D, Kiechle J, Fischer S, Llorián-Salvador Ó, Buchner JA, Nguyen MQ, Etzel L, Weidner J, Metz MC, Wiestler B, Schnabel J, Rueckert D, Combs SE, Peeken JC. Deep learning for autosegmentation for radiotherapy treatment planning: State-of-the-art and novel perspectives. Strahlenther Onkol 2025;201:236-54. [Crossref] [PubMed]
- Roth HR, Lu L, Liu J, Yao J, Seff A, Cherry K, Kim L, Summers RM. Improving Computer-Aided Detection Using Convolutional Neural Networks and Random View Aggregation. IEEE Trans Med Imaging 2016;35:1170-81. [Crossref] [PubMed]
- Buda M, Maki A, Mazurowski MA. A systematic study of the class imbalance problem in convolutional neural networks. Neural Netw 2018;106:249-59. [Crossref] [PubMed]
- Wang J, Sun K, Cheng T, Jiang B, Deng C, Zhao Y, Liu D, Mu Y, Tan M, Wang X, Liu W, Xiao B. Deep High-Resolution Representation Learning for Visual Recognition. IEEE Trans Pattern Anal Mach Intell 2021;43:3349-64. [Crossref] [PubMed]
- van den Brekel MW, Castelijns JA, Stel HV, Golding RP, Meyer CJ, Snow GB. Modern imaging techniques and ultrasound-guided aspiration cytology for the assessment of neck node metastases: a prospective comparative study. Eur Arch Otorhinolaryngol 1993;250:11-7. [Crossref] [PubMed]
- Liao LJ, Lo WC, Hsu WL, Wang CT, Lai MS. Detection of cervical lymph node metastasis in head and neck cancer patients with clinically N0 neck-a meta-analysis comparing different imaging modalities. BMC Cancer 2012;12:236. [Crossref] [PubMed]
- Mukherjee S, Fischbein NJ, Baugnon KL, Policeni BA, Raghavan P. Contemporary Imaging and Reporting Strategies for Head and Neck Cancer: MRI, FDG PET/MRI, NI-RADS, and Carcinoma of Unknown Primary-AJR Expert Panel Narrative Review. AJR Am J Roentgenol 2023;220:160-72. [Crossref] [PubMed]
- Yuan N, Hassan MA, Ehrlich K, Weyers BW, Biddle G, Ivanovic V, Raslan OAA, Gui D, Abouyared M, Bewley AF, Birkeland AC, Farwell DG, Marcu L, Qi J. Early Detection of Lymph Node Metastasis Using Primary Head and Neck Cancer Computed Tomography and Fluorescence Lifetime Imaging. Diagnostics (Basel) 2024;14:2097. [Crossref] [PubMed]
- Sohn K, Berthelot D, Carlini N, Zhang Z, Zhang H, Raffel CA, Cubuk ED, Kurakin A, Li CL. FixMatch: simplifying semi-supervised learning with consistency and confidence. Adv Neural Inf Process Syst 2020;33:596-608.
- Oliver A, Odena A, Raffel C, Cubuk ED, Goodfellow I. Realistic evaluation of deep semi-supervised learning algorithms. Adv Neural Inf Process Syst 2018;31.
.Liu YC Ma CY He Z Kuo CW Chen K Zhang P Wu B Kira Z Vajda P Unbiased teacher for semi-supervised object detection. arXiv:2102.09480.- Su J, Luo Z, Lian S, Lin D, Li S. Mutual learning with reliable pseudo label for semi-supervised medical image segmentation. Med Image Anal 2024;94:103111. [Crossref] [PubMed]
- Mahmood H, Shaban M, Rajpoot N, Khurram SA. Artificial Intelligence-based methods in head and neck cancer diagnosis: an overview. Br J Cancer 2021;124:1934-40. [Crossref] [PubMed]
- Mäkitie AA, Alabi RO, Ng SP, Takes RP, Robbins KT, Ronen O, Shaha AR, Bradley PJ, Saba NF, Nuyts S, Triantafyllou A, Piazza C, Rinaldo A, Ferlito A. Artificial Intelligence in Head and Neck Cancer: A Systematic Review of Systematic Reviews. Adv Ther 2023;40:3360-80. [Crossref] [PubMed]
- Zhang MB, Meng ZL, Mao Y, Jiang X, Xu N, Xu QH, Tian J, Luo YK, Wang K. Cervical lymph node metastasis prediction from papillary thyroid carcinoma US videos: a prospective multicenter study. BMC Med 2024;22:153. [Crossref] [PubMed]
- Valizadeh P, Jannatdoust P, Pahlevan-Fallahy MT, Hassankhani A, Amoukhteh M, Bagherieh S, Ghadimi DJ, Gholamrezanezhad A. Diagnostic accuracy of radiomics and artificial intelligence models in diagnosing lymph node metastasis in head and neck cancers: a systematic review and meta-analysis. Neuroradiology 2025;67:449-67. [Crossref] [PubMed]
- Song B, Yadav I, Tsai JC, Madabhushi A, Kann BH. Artificial Intelligence for Head and Neck Squamous Cell Carcinoma: From Diagnosis to Treatment. Am Soc Clin Oncol Educ Book 2025;45:e472464. [Crossref] [PubMed]
- Stawarz K, Gorzelnik A, Klos W, Korzon J, Kissin F, Bieńkowska-Pluta K, Stawarz G, Rusetska N, Zwolinski J. Systematic review of artificial intelligence and radiomics for preoperative prediction of extranodal extension and lymph node metastasis in oropharyngeal cancer. Front Oncol 2025;15:1717641. [Crossref] [PubMed]
- Troise S, Ugga L, Esposito M, Positano M, Elefante A, Capasso S, Cuocolo R, Merola R, Committeri U, Abbate V, Bonavolontà P, Nocini R, Dell'Aversana Orabona G. The Role of Machine Learning to Detect Occult Neck Lymph Node Metastases in Early-Stage (T1-T2/N0) Oral Cavity Carcinomas. Head Neck 2025;47:2751-9. [Crossref] [PubMed]
- Liao W, Luo X, Li L, Xu J, He Y, Huang H, Zhang S. Automatic cervical lymph nodes detection and segmentation in heterogeneous computed tomography images using deep transfer learning. Sci Rep 2025;15:4250. [Crossref] [PubMed]
- Rinneburger M, Carolus H, Iuga AI, Weisthoff M, Lennartz S, Große Hokamp N, Caldeira LL, Jaiswal A, Maintz D, Laqua FC, Baeßler B, Klinder T, Persigehl T. Automated Lymph Node Localization and Segmentation in Patients with Head and Neck Cancer: Opportunities and Limitations of Using a Generic AI Model. Diagnostics (Basel) 2026;16:355. [Crossref] [PubMed]
- Qi YM, Zhang LJ, Wang Y, Duan XH, Li YJ, Xiao EH, Luo YH. Deep Learning Model Based on Dual-energy CT for Assessing Cervical Lymph Node Metastasis in Oral Squamous Cell Carcinoma. Acad Radiol 2025;32:6216-26. [Crossref] [PubMed]
- Dayan GS, Hénique G, Bahig H, Nelson K, Brodeur C, Christopoulos A, Filion E, Nguyen-Tan PF, O'Sullivan B, Ayad T, Bissada E, Tabet P, Guertin L, Desilets A, Kadoury S, Letourneau-Guillon L. Artificial Intelligence Model for Imaging-Based Extranodal Extension Detection and Outcome Prediction in Human Papillomavirus-Positive Oropharyngeal Cancer. JAMA Otolaryngol Head Neck Surg 2026;152:7-17. [Crossref] [PubMed]

