Text-guided multimodal deep learning in magnetic resonance imaging for spinal structures segmentation and lumbar abnormalities identification
Introduction
Lumbar abnormalities, such as vertebral degeneration, disc bulge, disc protrusion, disc extrusion, and Schmorl’s nodes, have demonstrated increasing prevalence and a trend toward earlier onset in recent years (1). Radiologically, these disorders present as vertebral osteophytes, facet joint hypertrophy, degenerative remodeling, and disc space narrowing accompanied by spinal canal stenosis. Patients commonly report chronic low back pain, with symptom severity and spatial distribution linked to the progression of degenerative changes. Lumbar magnetic resonance imaging (MRI) analysis, a primary non-invasive method for evaluating lumbar abnormalities, has integrated with deep learning (DL) to enhance quantitative imaging biomarkers (2,3), automate segmentation of vertebrae and discs (4,5), and enable multi-class disease classification (6-8).
MRI leverages strong magnetic fields and radiofrequency pulses to align proton spins, enabling high-resolution soft tissue visualization (9), including intervertebral discs, ligamentous structures, and neural elements. The functional spinal unit consists of two adjacent vertebrae, an intervertebral disc, spinal ligaments, and facet joints. Degenerative processes typically initiate in the nucleus pulposus, progressively affecting the intervertebral disc, annulus fibrosus, vertebral endplates, and bone marrow of adjacent vertebral bodies (10). These abnormalities commonly manifest as spondylolisthesis, loss of lumbar lordosis, degenerative scoliosis, and lateral osteophyte formation (11). Disc abnormalities in this study encompass bulges, protrusions, extrusions, and Schmorl’s nodes. Disc bulges correlate with bilateral lateral recess stenosis (12). Protrusions appear as focal central or paracentral herniations, whereas extrusions may compress the nerve roots (13). Schmorl’s nodes represent focal herniation of the nucleus pulposus through degenerated endplate channels (derived from regressed vascular foramina), with disc material protruding into the adjacent vertebral body and triggering reactive bone marrow changes (14).
Driven by recent advances in DL for lumbar spine imaging across X-ray, computed tomography (CT), and MRI modalities, convolutional neural networks (CNNs) and transformer-based models have achieved state-of-the-art performance in different medical image computing tasks (15). ConvNeXt V2, employed as an encoder (16), effectively captures multi-scale features for accurate vertebrae localization and analysis through its hierarchical architecture integrating fully convolutional masked autoencoder (FCMAE)-based self-supervised learning and a global response normalization (GRN) layer. The multi-stage depth-wise separable convolutions extract global contextual features reflecting spinal anatomical continuity while preserving fine-grained textures such as vertebral trabecular bone and intervertebral margins. Inspired by the remarkable cross-modal alignment demonstrated by Contrastive Language-Image Pretraining (CLIP), recent works have explored leveraging pre-trained vision-language models for pixel-level semantic understanding (17-19). Wang et al. (20) introduced a CLIP-driven image segmentation framework integrating multimodal representations from the CLIP architectures to achieve fine-grained text-to-pixel alignment. In the context of medical imaging, the text encoder of CLIP encodes medical terms into semantic vectors, which serve as guidance for the visual model to localize specific anatomical structures. The joint embedding space of CLIP enables a unified representation of text and images, facilitating cross-modal alignment and enabling visual models to better understand and respond to medical language descriptions. Cycle generative adversarial network (CycleGAN) (21) enables unpaired image translation while maintaining clinically critical anatomical features, empowering medical data augmentation strategies. Cross-modal medical imaging augmentation—encompassing diverse sequences within the same imaging modality or images from different devices—substantially enhances model inference robustness. Through the cycle-consistency constraints, CycleGAN enables unpaired image translation while maintaining anatomical features, empowering medical imaging augmentation strategies. For anatomical preservation, the dual mapping of CycleGAN ensures that synthetic images maintain original structural relationships, which is crucial for vertebral alignment in spinal imaging (22). Wang et al. (23) proposed a deformation-invariant CycleGAN (DicycleGAN), which integrates deformable convolutional layers to handle nonlinear domain shifts between MRI and CT modalities.
Early unimodal segmentation architectures, leveraging multi-scale feature extraction capabilities of CNN and Transformer designs (24-26), fused hierarchical features through direct concatenation to compute segmentation masks. However, these architectures are limited by their one-hot representation space, which fails to adequately capture semantic relationships among target categories. Another limitation is the partially labeled problem: in medical image benchmarks for organ and lesion segmentation, where high labor and expertise costs result in partial labeling, datasets often lack complete research-relevant annotations such as different target organs or lesion types or co-annotations of both. Xie et al. (27) mitigated the partially labeled problem by proposing TransDoDNet, which generates task-specific kernels via self-attention to model long-range organ dependencies, enabling flexible multi-task segmentation across multiple partially labeled datasets. Combining textual and visual information to address the limitations of unimodal approaches, Li et al. (28) proposed Language meets Vision Transformer (LViT), a text-augmented medical imaging segmentation model that integrates text annotations to alleviate high-quality labeled data scarcity, enabling improved pseudo label generation in semi-supervised learning.
In this study, we introduce a text-guided DL method for multi-class segmentation of vertebrae and intervertebral discs in sagittal T1-weighted imaging (T1WI) and T2-weighted imaging (T2WI) MRI, with concurrent identification of five lumbar abnormalities: vertebral degeneration, disc bulge, disc protrusion, disc extrusion, and Schmorl’s nodes. To mitigate the limitations of capturing semantic relationships between target categories, the text-guided module integrates a domain-adapted text encoder that generates image-aligned text embeddings, promoting semantic consistency between anatomical structures and radiological descriptions. Leveraging self-supervised pretraining, we employ ConvNeXt V2 as the backbone, pretrained on 1,975 unlabeled lumbar MRI scans to learn hierarchical anatomical representations and reduce dependency on extensive manual annotations. To further alleviate the limitations caused by partially labeled datasets, we first leverage CycleGAN-based augmentation to synthesize multi-sequence lumbar MRI from a single-source T2WI dataset, training an anatomical segmentation model to annotate vertebrae and intervertebral discs. This model is subsequently applied to a separate partially labeled dataset, where annotations are limited to abnormalities while vertebrae and discs remain unlabeled. The full task-specific annotations were completed through collaboration between radiologists and the anatomical segmentation model. Our study achieves robust multi-class segmentation and abnormality identification, exhibiting enhanced capability in capturing fine-grained anatomical details and modeling semantic relationships, with quantitative improvements observed across metrics relative to image-based unimodal approaches. We present this article in accordance with the CLEAR reporting checklist (available at https://qims.amegroups.com/article/view/10.21037/qims-2025-635/rc).
Methods
Study data
The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study was approved by the Ethics Committee of China-Japan Union Hospital of Jilin University (No. 2025KYYS036), and the requirement of informed consent was waived due to the retrospective nature of the study and the complete de-identification of all personally identifiable information.
Five centers (A to E) collectively contributed samples to support the study, each focusing on distinct research objectives. Figure 1 outlines the integrated workflow, where Figure 1A aligns with Center A, Figure 1B with Center B, Figure 1C with Center C, Figure 1D with Center E, and Figure 1E with Center D. Center A involved collection of 445 lumbar MRI cases to address model evaluation and sequence diversity, including 100 lumbar MRI studies collected from the China-Japan Union Hospital of Jilin University (2023.2 to 2024.11) and 345 cases (2022.1 to 2024.7) exclusively acquired to match the 5 sequences in Center D as shown in Figure 1A and Figure 1E. The 100-case subset comprised 50 lumbar-healthy cases and 50 lumbar-abnormal cases of 2 sequences based on T1WI and T2WI, with the abnormal samples selected based on radiological reports. All data included anonymized Digital Imaging and Communications in Medicine (DICOM) files, which were reviewed by two radiologists for clinical relevance and imaging accuracy to ensure that both healthy and abnormal cases met the criteria of our study. The inclusion criteria were as follows: (I) sagittal T1WI and T2WI sequences showing at least one abnormality related to vertebral degeneration, disc bulge, disc protrusion, disc extrusion, or Schmorl’s nodes; (II) scan coverage extending from the 12th thoracic vertebra (T12) to the first sacral vertebra (S1); (III) field of view (FOV) approximately 25–35 cm; and (IV) age 16–85 years. The exclusion criteria were as follows: (I) slice thickness exceeding 6 mm; (II) poor visualization of vertebral anatomy or lesion contrast; and (III) significant motion artifacts or other diagnostic confounders. Imaging was performed using a Siemens Skyra 3.0T MRI scanner (Siemens, Erlangen, Germany) with a 32-channel spine coil (gradient field strength: 45 mT/m, slew rate: 200 mT∙m−1∙ms−1). Following radiological review and dataset curation, the 100-case subset included 54 females, 46 males [age 16–82 years, mean ± standard deviation (SD): 42±17 years], utilized for final model evaluation and CycleGAN training. The 345 cases were used to address the sequence mismatch between the single T2WI sequence from Center B and the 5 sequences in Center D, as shown in Figure 1A,1B,1E, to achieve the CycleGAN training in this study.
Four public datasets were incorporated. Center B used the MRSpineSeg2021 dataset (25), containing 215 T2WI scans with pixel-level annotations for 19 spinal structures, to fine-tune vertebral and disc segmentation models. Center C adopted the RSNA 2024 Lumbar Spine Degenerative Classification dataset (29), comprising 1,975 T2WI and short TI inversion recovery (STIR) samples for self-supervised pretraining of the image encoder. Center D incorporated the Tianchi Spinal Disease dataset (30), which includes 201 sagittal MRI scans processed via DL segmentation, to fine-tune and evaluate the model performance in identifying lumbar abnormalities. Center E accessed the Lumbar Spine MRI dataset (31), consisting of 515 MRI studies with clinical descriptions of low back pain, to complete domain adaption for the text feature encoder, providing textual modality features. All four public datasets utilized in this study were used under valid copyright licenses, with all necessary permissions obtained in accordance with their respective licensing terms.
Center D contained partially labeled data that included annotations for 5 categories of lumbar abnormalities but lacked segmentation labels for vertebrae and intervertebral discs. To address this gap, data from multiple centers were employed. Center B provided a single T2WI-based sequence and served as the foundation for training a segmentation model, hereafter referred to as the “labeling model”, to label vertebrae and intervertebral discs. However, Center D comprised 5 imaging sequences based on T1WI and T2WI, whereas Center B included only one sequence, which created a modality mismatch. To bridge this mismatch, 345 cases acquired from Center A to match the 5 sequences in Center D were combined with 100 cases from Center A that involved the two most commonly scanned sequences based on T1WI and T2WI. The combination of these datasets aimed to expand sequence diversity, which enhanced adaptability of the labeling model to diverse imaging protocols (particularly sequence types) and facilitated robust extraction of image semantic features in lumbar MRI. Of the 100-case subset from Center A, 40 cases were selected as part of test set for our study, which included 20 healthy and 20 abnormal lumbar cases. CycleGAN was trained using these 445 Center A cases as the target sequence domain and Center B data as the source domain, generating multi-sequence augmented data to train the labeling model that collaborated with radiologists to finalize full annotations for Center D. Center C provided 1,975 unlabeled lumbar MRI scans for self-supervised pretraining of the image encoder, and Center E contributed 515 radiological reports for domain adaptation of the text encoder. The integration of these datasets across multiple centers formed the core of the methodology, which is illustrated in the workflow diagram shown in Figure 1.
Data processing
Data augmentation
To address the mismatch between the single T2WI sequence of Center B and 5 sequences of T1WI and T2WI from Center D, we trained the CycleGAN (32) to synthesize MRI data of multiple sequences, enhancing the adaptability of the labeling model. The training process used Center B as the source domain and Center A as the target domain. The target domain from Center A included 345 cases matched to 5 sequences of Center D: (I) “T2-weighted turbo spin echo fat-suppressed Dixon technique sagittal in-phase imaging” (t2_tse_fs-dixon_sag_in); (II) “T2-weighted turbo spin echo fat-suppressed Dixon technique sagittal out-of-phase imaging” (t2_tse_fs-dixon_sag_opp); (III) “T2-weighted turbo spin echo fat-suppressed Dixon technique sagittal fat-only imaging” (t2_tse_fs-dixon_sag_F); (IV) “T2-weighted turbo spin echo fat-suppressed Dixon technique sagittal water-only imaging” (t2_tse_fs-dixon_sag_W); and (V) “T1-weighted turbo spin echo sagittal imaging” (t1_tse_sag), along with 100 cases involving 2 sequences specific to Center A: (I) “T2-weighted turbo spin echo sagittal assisted compressed sensing imaging” (t2_tse_sag_ACS); and (II) “T1-weighted turbo spin echo sagittal assisted compressed sensing imaging” (t1_tse_sag_ACS). Visualization examples are shown in Figure 2.
The CycleGAN was trained individually for each of the 7 target domain sequences over 200 epochs on an NVIDIA GeForce RTX 3090 graphics processing unit (GPU) with 24 GB of video random access memory (VRAM) (NVIDIA, Santa Clara, CA, USA). Data augmentation included random horizontal flipping, minor brightness and contrast variations, and limited affine transformations to diversify training samples while preserving anatomical consistency. Identical augmentation strategies were applied to both source and target domains. The adversarial training framework utilized a least squares-based loss [least squares generative adversarial networks (LSGAN)] for the discriminators, and the generators were optimized to minimize adversarial objectives alongside cycle consistency loss weighted by λcycle =10.0. The Adam optimizer was employed with a learning rate set to 2e−4, where the exponential decay rates for the first and second moment estimates were β1 =0.5 and β2 =0.999, respectively. During inference on GPU, the model processed individual slices in 0.051 seconds on average.
The source data from Center B included 2,037 slices. The total number of slices used to train the labeling model, which combines source and synthetic data, amounted to 15,812, as calculated in Table 1. The synthetic slices, generated by CycleGAN from source data in Center B, were reconstructed into corresponding image volumes using the spatial information recorded in the scans from Center B. The synthetic data underwent rigid registration using the SlicerElastix (33) plugin in 3D Slicer software (https://www.slicer.org/) and was subsequently validated by two radiologists to ensure the anatomical clarity of vertebral and intervertebral disc structures before integration into the dataset. The preprocessed source and synthetic data were then used to train the labeling model.
Table 1
| Target sequence | Adopted original slices | Adopted synthesized slices |
|---|---|---|
| t2_tse_fs-dixon_sag_in | 1,596 | 1,904 |
| t2_tse_fs-dixon_sag_opp | 1,523 | 2,002 |
| t2_tse_fs-dixon_sag_F | 1,361 | 1,898 |
| t2_tse_fs-dixon_sag_W | 1,422 | 2,021 |
| t1_tse_sag | 1,635 | 1,977 |
| t2_tse_sag_ACS | 1,389 | 2,017 |
| t1_tse_sag_ACS | 1,446 | 1,956 |
CycleGAN, cycle generative adversarial network.
To overcome the partially labeled limitation of Center D, which is annotated for 5 abnormalities but lacks vertebral and intervertebral disc segmentation labels, we developed the labeling model using ConvNeXt-V2 (34) as the backbone to segment 19 spinal structures, including 10 vertebral categories (T9–L5, S1) and 9 intervertebral disc categories (T11/12–L5/S1). The encoder of the model uses parameters initialized from our self-supervised pretraining results, derived from processing on 1,975 unlabeled lumbar MRI scans from Center C through a dual-task pretraining that combined image reconstruction and degeneration severity classification to capture hierarchical anatomical features, as shown in Figure 3. The labeling model was trained on the original and CycleGAN-augmented data from Center B for 50 epochs using the Adam optimizer (learning rate =1e−4, β1 =0.5, β2 =0.999) for 90 epochs on an NVIDIA GeForce RTX 3090 GPU equipped with 24 GB of VRAM. Data augmentation included random cropping of regions covering 60–80% of the slice area and centered on the midline of the lumbar spinal column, brightness, and contrast adjustments within –20% to 20% intensity variation, bidirectional random flipping, and Gaussian blur with σ∈[0.5,1.5].
Imaging data annotations
Leveraging the trained labeling model, annotations of 19 spinal structures, which include 10 vertebral categories (T9–L5, S1) and 9 intervertebral disc categories (T11/12–L5/S1), were completed for the 40-case test set from Center A and all data in Center D. Specifically for Center D, the abnormality localization annotations inherent in the data (5 lumbar abnormalities and their corresponding anatomical sites) were used to map abnormalities to precise pixel coordinates in segmentation masks. This process expanded the annotations into 24-channel binary masks, including the previously annotated 10 vertebrae (T9–L5, S1) and 9 intervertebral discs (T11/12–L5/S1), along with 5 lumbar abnormalities (1 type of vertebral degeneration and 4 intervertebral disc abnormality types), with each category allocated to a separate channel. The test set for the text-guided DL method is illustrated in Table 2.
Table 2
| Data center | Annotations | Cases | Slices | Testing tasks |
|---|---|---|---|---|
| Center A | 19 spinal structures | 40 | 732 | Segmentation |
| Center D | 19 spinal structures + 5 lumbar abnormalities | 31 | 701 | Segmentation; abnormality identification |
All annotations were validated and refined through collaboration with two radiologists. The full annotated data of Center D were divided into training, validation, and test sets for the text-guided DL method as detailed in Table 3.
Table 3
| Dataset | Cases | Abnormal lumbar vertebrae | Abnormal lumbar discs | Sex ratio (M/F) | Age (years) [M (SD); (Q1, Q3)] |
|---|---|---|---|---|---|
| Train | 150 | 592 | 482 | 1.38 (87/63) | 52.06 (18.96); (44, 65) |
| Valid | 20 | 82 | 69 | 0.25 (4/16) | 61.55 (8.49); (55, 68) |
| Test | 31 | 137 | 119 | 0.82 (14/17) | 55.74 (18.35); (49, 68) |
M, mean; M/F, male/female; SD, standard deviation.
This study enhanced segmentation performance by integrating textual features. Leveraging clinical reports of lumbar MRI scans from Center E, we performed self-supervised domain adaptation to develop a text encoder. Structured textual prompts were created for each vertebra, intervertebral disc, and abnormality category. Anatomical prompts for vertebrae and intervertebral discs were generated using the template: “This is a lumbar MRI slice showing [anatomical category]”, where anatomical category ∈ {vertebrae (10 classes), discs (9 classes)}. For lumbar abnormalities, prompts followed the format: “[anatomical category] presents with [normal/abnormal findings]”, where abnormalities ∈ {vertebral degeneration, disc bulge, disc protrusion, disc extrusion, Schmorl’s nodes}.
Methodology framework
Self-supervised pretraining
The self-supervised pretraining framework employed ConvNeXt V2-P (34) as the backbone to learn image representations from unannotated stack of two-dimensional (2D) slices from Center C via a dual-task optimization strategy, illustrated in Figure 3A. During the masking phase, 60% of each slice was randomly masked. The encoder processed masked inputs to produce hierarchical latent representations through cascaded convolutional blocks, the parameters of which were iteratively updated to capture contextual relationships between visible and masked regions. A lightweight decoder reconstructed the masked regions through layer-wise feature upsampling, with reconstruction accuracy evaluated using mean squared error (MSE) loss between recovered and original images. Concurrently, image representations were flattened via a multi-layer perceptron (MLP) module and classified into three categories (normal, moderate degeneration, severe degeneration) using anomaly grading labels from Center C, optimized by cross-entropy loss. The framework was trained for 300 epochs on an NVIDIA GeForce RTX 3090 GPU equipped with 24 GB of VRAM, using the AdamW optimizer with an initial learning rate of 3e−4 (β1 =0.9, β2 =0.998), and a weight decay of 1e−3 applied to all non-bias parameters. The learning rate schedule adhered to a cosine annealing strategy with a 40-epoch warmup phase, linearly ramping from 0 to the initial value during warmup and decaying to 3e−6 over the remaining 260 epochs. Data augmentation encompassed random in-plane rotations of –15° to 15°, 35% probability of horizontal flipping, center-biased random cropping (60–80% of original size), and Gaussian blur with .
To derive representations in the textual modality, we employed the MedCLIP (35) text encoder as the backbone, initialized with pre-trained weights from MedCLIP. For each textual description in Center E, 40% of the content was masked to support masked language model (MLM) pretraining. Adjacent text descriptions were explicitly labeled as positive pairs, whereas non-adjacent or semantically unrelated descriptions were randomly sampled as negative controls for next sentence prediction (NSP) pretraining. As illustrated in Figure 3B, domain adaptation for text encoder using radiological report descriptions was subsequently performed on an NVIDIA GeForce RTX 3090 GPU with 24 GB of VRAM for 90 epochs. The training framework incorporated two loss functions with equal weighting: cross-entropy loss for MLM to optimize semantic reconstruction of masked tokens and binary cross-entropy (BCE) loss for NSP to model contextual coherence between sentence pairs. The AdamW optimizer was applied with a learning rate of 5e−5 (β1 =0.9, β2 =0.999) and a weight decay of 0.01. The learning rate was initially set to 5e−5 and adjusted using a cosine annealing schedule with a linear warmup phase over the first 10 epochs, starting from 0 and reaching the initial value at the end of the warmup. Subsequently, the learning rate decayed to 5e−7 over the remaining 80 epochs following the cosine annealing strategy. This dual-task pretraining framework, which combines MLM-driven semantic reconstruction and NSP-based contextual coherence learning, enabled the encoder to capture domain-specific linguistic patterns while preserving clinical semantic relationships inherent in radiological reports.
Text-guided segmentation model
Figure 4 illustrates the overall framework of the proposed method. The frozen domain-adapted text encoder converted textual descriptions of 24 categories into 256-dimensional embeddings, which were concatenated with 256-dimensional globally average-pooled image features derived from the frozen image encoder. The fused vector was processed through an MLP module, aligning its dimension with the output channels of the image decoder and producing the parameter matrix μ. During training, 2D MRI slices were normalized using min-max scaling and fed into the model as tensors with dimensions [4, 256, 256, 3], where 4 indicates the batch size, 256×256 denotes the slice resolution, and 3 represents the input channels. The image encoder was initialized with weights from ConvNeXt V2-P (ImageNet-1K FCMAE pre-trained weights). To match the 3-channel input requirement of these pre-trained weights, single-channel MRI slices were replicated across all three channels, forming a 3-channel tensor structure. Following self-supervised pretraining as shown in Figure 3A, the image encoder parameters were frozen to exclusively perform forward propagation for hierarchical feature extraction, without participating in backpropagation. A symmetrical decoder was constructed to mirror the downsampling architecture of the image encoder. At each upsampling stage, the decoder outputs underwent attention computation (37) with corresponding encoder features from equivalent downsampling levels. These features were then fused through skip connections via channel-wise concatenation, followed by progressive upsampling until reaching the spatial resolution matching the initial downsampling stage. Before final segmentation, the parameter matrix μ was reshaped into a four-dimensional (4D) tensor (24×N×1×1), where N is the channel dimension of the image decoder output, to parameterize a group convolution (GC) operation with 4 groups. The GC operation directly modulated the decoder output features via 1×1 convolutions. The GC output was then activated by a channel-wise Sigmoid to produce segmentation probabilities.
The segmentation output comprised 24 categorical channels, with each channel dedicated to a specific anatomical structure or lumbar abnormality. The first 19 channels represented vertebrae and intervertebral discs, whereas the remaining 5 channels included five types of lumbar abnormalities associated with these structures. Each channel represented a 256×256-pixel 2D binary mask, where foreground pixels were distinguished from background by thresholding the output probability at 0.5. To balance convergence speed, improve adaptation to small anatomical targets, and enhance segmentation accuracy, the training objective combined weighted BCE loss (38) and Dice coefficient loss (39) as: . In the training dataset of Center D, the model underwent 50 fine-tuning epochs. During the initial 25 epochs, α and β were fixed at 0.25 and 0.75, respectively. For the subsequent 25 epochs, these weights were progressively adjusted to α=0.75 and β=0.25 through linear scheduling. The model was trained on an NVIDIA GeForce RTX 3090 GPU equipped with 24 GB of VRAM. For inference, the processing time averaged 0.33 seconds per individual slice on the GPU. Data augmentation involved random cropping of regions spanning 60–90% in the slice area, with each crop centered on the midline of the lumbar spinal column. It also included brightness and contrast adjustments within a –15% to 15% intensity range, bidirectional random flipping with a 50% probability for each axis, and Gaussian blur with .
Metrics
Segmentation performance was evaluated using mean Intersection over Union (mIoU) based on annotations verified by radiologists, which were used for the overall assessment of vertebrae and intervertebral discs. For subclass-specific evaluation, Dice coefficient, recall, and precision were adopted as metrics. To assess stability of segmentation performance, each region of interest (ROI) was bisected along the midline of its bounding rectangle into upper and lower subregions based on sagittal anatomical morphology, with corresponding Dice scores calculated as upper-Dice and lower-Dice. Additionally, recall, false-positive rate (FPR), and precision were employed to measure the identification of five lumbar abnormalities.
Statistical analyses were conducted using SciPy.Stats library (NumFOCUS, Austin, Texas, USA) from Python. Normality of the upper-Dice, lower-Dice, and mIoU distributions was assessed using Shapiro-Wilk and D’Agostino-Pearson tests (P>0.05 for all metrics). Paired t-tests with Cohen’s d effect size were conducted to assess: (I) intra-ROI differences between upper-Dice and lower-Dice values (P>0.05 indicating no significant differences); (II) inter-model mIoU differences (P<0.05 indicates significant differences). For lumbar abnormality identification, model performance was validated through 1,000 bootstrap resampling iterations. Statistical significance was set at P<0.05 after confirming normality with Shapiro-Wilk testing.
Results
Segmentation performance of vertebrae and discs
When evaluated on the test set (Center D: 31 cases, 701 slices; Center A: 40 cases, 732 slices), our proposed method achieved a mIoU of 0.823±0.053 for multi-class segmentation of 19 vertebral/disc structures, with the lumbar region demonstrating a mIoU of 0.816±0.073 for lumbar vertebrae (L1–L5) and 0.817±0.053 for lumbar discs. Paired t-test analysis (P=0.744) and Cohen’s d effect size (d=0.01) revealed no statistically significant differences between upper-Dice and lower-Dice metrics (P=0.744 >0.5), with d<0.2 indicating negligible practical significance. As detailed in Table 4, the model demonstrated high stability and accuracy across thoracic, lumbar, and sacral regions.
Table 4
| ROIs | Upper-Dice | Lower-Dice | mIoU |
|---|---|---|---|
| Thoracic (T9–T12) | 0.880±0.031 | 0.873±0.038 | 0.837±0.061 |
| Lumbar (L1–L5) | 0.850±0.044 | 0.849±0.032 | 0.816±0.073 |
| Thoracic disc | 0.858±0.042 | 0.862±0.043 | 0.821±0.054 |
| Lumbar disc | 0.844±0.029 | 0.846±0.034 | 0.817±0.053 |
| Sacrum | 0.899±0.017 | 0.900±0.025 | 0.840±0.044 |
Data are expressed as mean ± SD. mIoU, mean Intersection over Union; ROIs, regions of interest; SD, standard deviation.
The segmentation results for 19 vertebral/disc categories across the T9-to-sacrum anatomical region (spanning thoracic vertebrae T9–T12, lumbar vertebrae L1–L5, and 9 corresponding intervertebral discs) were binarized using a 0.5 probability threshold and subjected to connected-component analysis. As shown in Figure 5, the Dice coefficient distributions exhibited high consistency between upper and lower subregions of ROIs, with minimal variation in mIoU values across the five anatomical regions, indicating robust spatial stability of the segmentation model.
The results presented in Table 5 indicate that the proposed method (Ours) demonstrated superior performance across all metrics in the test set, achieving an upper-Dice score of 0.859±0.040, lower-Dice of 0.858±0.038, and mIoU of 0.823±0.053. Notably, Ours-VisualOnly (derived from Ours by removing the text-guided module) exhibited performance comparable to nnU-Net (40), with upper-Dice values of 0.840 and 0.842 for Ours-VisualOnly and nnU-Net, respectively, and lower-Dice scores of 0.834 and 0.833, respectively. However, both models underperformed relative to the complete text-guided method proposed in this study (Ours). Paired t-test analysis and Cohen’s d effect size further validated significant performance differences: Ours outperformed Ours-VisualOnly (P<0.01, d=0.38), nnU-Net (P<0.01, d=0.26), and MT-U-Net (41) (P<0.01, d=0.61) in mIoU.
Table 5
| Model | Upper-Dice | Lower-Dice | mIoU |
|---|---|---|---|
| nnU-Net | 0.842±0.043 | 0.833±0.037 | 0.806±0.054 |
| MT-U-Net | 0.821±0.064 | 0.801±0.079 | 0.766±0.073 |
| Ours-VisualOnly | 0.840±0.058 | 0.834±0.061 | 0.791±0.067 |
| Ours | 0.859±0.040 | 0.858±0.038 | 0.823±0.053 |
Data are expressed as mean ± SD. mIoU, mean Intersection over Union; SD, standard deviation.
As illustrated in Figure 6, the visualization examples demonstrate that the proposed method maintained robust segmentation performance across diverse imaging conditions. Notably, it preserved anatomical segmentation accuracy in cases with localization line obstructions while effectively discriminating vertebrae from intervertebral discs in low-resolution interpolated images, even when confronted with blurred structural boundaries or suboptimal contrast conditions. These findings collectively indicate the inherent robustness to variations in image quality.
Identification for five lumbar abnormalities
The output channels from 20 to 24 correspond to five lumbar abnormalities: vertebral degeneration, disc bulge, disc protrusion, disc extrusion, and Schmorl’s nodes. Each channel’s foreground pixels represent candidate abnormal regions, with expanded ROIs defined by 6-pixel dilation from the minimum bounding rectangles of connected components. All models (nnU-Net, MT-U-Net, Ours-VisualOnly, and Ours) employed the same strategy of selecting the class with the highest pixel count in connected components to determine lumbar abnormalities. Using 0.5 as the predictive probability threshold and after 1,000 bootstrap iterations, the average recall, average FPR, and average precision of nnU-Net, MT-U-Net, Ours-VisualOnly, and Ours for the five abnormalities are presented in Table 6.
Table 6
| Model | Recall | FPR | Precision |
|---|---|---|---|
| nnU-Net | 0.842±0.029** | 0.096±0.020* | 0.858±0.030* |
| MT-U-Net | 0.817±0.035** | 0.111±0.019* | 0.845±0.021** |
| Ours-VisualOnly | 0.839±0.032** | 0.090±0.023* | 0.855±0.026* |
| Ours | 0.867±0.027 | 0.079±0.015 | 0.893±0.028 |
Data are expressed as mean ± SD. *, P<0.05; **, P<0.01 (vs. Ours). FPR, false positive rate; SD, standard deviation.
As shown in Figure 7A, the AUC values of our model for each lumbar abnormality outperform the other models presented in Figure 7B-7D. Specifically, for vertebral degeneration and Schmorl’s nodes, our model exhibited a significant advantage over the other three models. For the disc extrusion, although our model still demonstrated an advantage upon comprehensive consideration, this edge is relatively less prominent than in other categories.
Figure 8 demonstrates the visualization of the proposed method “Ours” for identifying five lumbar abnormalities. Three representative cases were selected, including normal lumbar vertebrae/intervertebral discs, vertebral degeneration, disc bulge, disc protrusion, disc extrusion, and Schmorl’s nodes. Using annotations reviewed by radiologists as the ground truth, the visualization shows recognition results for each lumbar abnormality type. Notably, the probability for the L2–L3 disc bulge region was 0.43, which was below the decision threshold of 0.5. All other identification of lumbar abnormalities maintained full consistency with the annotation standards. Furthermore, the model demonstrated high consistency in visualization between T1WI and T2WI samples of the same case [Figure 8 (3a-3d)], exhibiting robust performance across different MRI sequences.
Discussion
This study proposed a text-guided multimodal DL method to address challenges in the lumbar MRI analysis process. Our method focused on segmenting 19 spinal structures mainly in the lumbar region, along with the identification of 5 lumbar abnormalities. To mitigate the limitations of image-only unimodal methods that struggled to model the contextual correlations latent in research-relevant categories, we introduced domain-specific text embeddings generated by a CLIP-based text encoder. These embeddings encoded the relationship representations among radiological terms of relevant categories in this study, improving the semantic relationship modeling in the text modality. We first retained the main architecture of the image-only unimodal models and applied self-supervised pretraining to ConvNeXt V2-P as the image encoder. Then, text embeddings were concatenated with image features and fed into an MLP module to generate a fused representation, which parameterized a GC for direct modulation of decoder output features. By integrating text-based categorical relationship modeling, we enhanced the feature extraction performance, enabling more precise contextual understanding and semantic alignment compared to image-only methods. In medical image analysis tasks, two common challenges manifest as incomplete dataset annotations and cross-dataset sequence mismatches. To address these, we first employed CycleGAN to synthesize multi-sequence samples for mismatched imaging sequences, subsequently training a labeling model on a combination of original samples and synthetic outputs. By collaborating with radiologists, we leveraged the labeling model to propagate vertebral and intervertebral disc segmentation results and supplemented the missing category labels in the incomplete annotation dataset. Our study enhanced segmentation performance for 19 spinal structures and identification of 5 lumbar abnormalities in multi-sequence lumbar MRI, outperforming image-only baseline models in different evaluation metrics.
Lumbar abnormalities exhibit complex MRI manifestations. Previous studies (25,29,30,31) have predominantly focused on task-specific goals for specific MRI sequences. These studies encompassed ROI segmentation of vertebrae and intervertebral discs, detection and grading of degenerative abnormalities, and multimodal contrastive learning for aligning imaging findings with radiological reports. To address sequence mismatch limitations in cross-dataset scenarios, multi-sequence lumbar MRI scans from Center A underwent CycleGAN-based style transfer for data augmentation, which aligns with emerging trends in medical image synthesis where cross-sequence conversions are increasingly utilized for such purposes. Wu et al. (42) developed an interpretable generative adversarial network integrating spatiotemporal features to synthesize coronary CT angiography from cardiac CT perfusion data. Furthermore, through a systematic comparison, Koch et al. (43) demonstrated the superior performance of diffusion models in converting magnetic resonance angiography (MRA) to CT angiography (CTA). Herein, we optimized the lumbar MRI study data within this fixed-task paradigm, enhancing feature diversity and model generalizability for targeted spinal structure analysis.
Self-supervised learning has demonstrated that pretraining models using different pretext tasks (34,44) can yield highly generalizable visual representations. Given the limited annotations for spinal structures and lumbar abnormalities, this study utilized extensive unlabeled lumbar MRI data (1,975 imaging scans) for pretraining of the image encoder. These scans included both normal findings and clinical cases exhibiting neuroforaminal stenosis, subarticular recess narrowing, and spinal canal narrowing, with data sourced from 8 global regions across 5 continents. During the pretraining phase, through the joint optimization of image reconstruction and degeneration grading, our method was enabled to learn the comprehensive patterns of vertebral anatomy and the subtle mechanisms underlying imaging abnormalities. Our method integrates clinical abnormality grading with image reconstruction into the pretraining pipeline, whereas prior studies (15,16)—even those adopting multi-task methods—focused exclusively on image-level objectives, enabling comprehensive representation learning. Building on multimodal models such as CLIP (45), the text encoder offers textual embeddings, which substantially boost the capabilities of abnormality characterization for lumbar spinal structures. In our study, we domain-adapted the text encoder using lumbar MRI radiology reports sourced from Center E. These reports contain structured clinical descriptions of disc bulges, thecal sac compression, annular fissures, and spondylolisthesis. Considering the strong domain-relatedness between the pretraining data and the target task, the optimized text encoder efficiently converts standardized medical prompts into refined semantic embeddings. These cross-modal representations, derived from the pretrained image encoder and the domain-adapted text encoder, offer complementary information for model inference in two mechanisms. First, the representations establish explicit correlations by fusing image features with text embeddings, enhancing contextual interpretation of imaging details. Second, they model semantic relationships among research categories, mitigating the feature sparsity inherent in categorical labeling.
In related research, multi-class segmentation models have been widely adopted for precise lesion localization and classification. Xu et al. (46) optimized the Swin-Transformer to achieve segmentation of 3 categories of gliomas across 4 MRI sequences, demonstrating significant performance gains in boundary-ambiguous regions. Markhali et al. (47) implemented U-Net-based automated segmentation of intervertebral discs with varying degeneration grades in age-diverse T1WI lumbar MRI scans. Compared to two-stage methods that sequentially perform ROI segmentation followed by feature analysis and classification, multi-class segmentation strategies leverage global contextual information during classification decisions, enabling more abundant feature interdependencies. By integrating clinical prompts, text embeddings for each category were computed and fused with imaging features to establish inter-class semantic relationships, effectively tackling the feature sparsity inherent in traditional one-hot labeling. Yan et al. (48) presented a semantics-guided training framework which enhances pseudo-label quality through CLIP-based semantic alignment and improves segmentation robustness by minimizing tri-view discrepancy. In our study, lumbar abnormality annotations were transformed into structured medical prompts containing anatomical locations and abnormality types. These prompts undergo cross-modal semantic computation with image features to obtain comprehensive inter-class semantic associations. The fused vectors are then used to guide the segmentation decoder.
In clinical implementation, the proposed method demonstrates potential for integration into certified DICOM viewers as a modular component, enabling intelligent processing of raw MRI data streams. Anatomical segmentation masks and abnormality annotations, presented as interactive overlays on imaging interfaces, exhibit clinical potential to support radiological review within the Picture Archiving and Communication System (PACS) workflows.
This study has several limitations. First, supervised training faced data shortages compared to self-supervised pretraining, which used unlabeled data. Our future research will refine model pseudo-labels through manual editing to build better segmentation annotations. This iterative process is expected to enhance model performance while establishing more robust validation benchmarks for novel methods. Regarding multimodal research, feature fusion aims to investigate latent correlations between textual reports and medical images, with the goal of achieving precise cross-modal alignment and enhancing diagnostic efficacy. Additionally, future studies will aim to leverage CLIP mechanisms in an alternative sequential process to further align and fuse image-text embeddings, enabling the model to generate radiology reports corresponding to lumbar MRI scans. Furthermore, we intend to explore context integration by stacking three adjacent slices as channel dimensions, a strategy that allows direct parameter initialization from our current model to facilitate fine-tuning for spatial context modeling.
Conclusions
In summary, this study proposed a text-guided DL method for segmenting 19 spinal structures and identifying 5 lumbar abnormalities across multi-sequence MRI. The model demonstrated robust performance in integrating cross-modal features to enhance imaging detail segmentation and optimize abnormality identification, highlighting its clinical potential for assisting lumbar spine diagnosis through reducing radiologist workload via automated analysis workflows.
Acknowledgments
None.
Footnote
Reporting Checklist: The authors have completed the CLEAR reporting checklist. Available at https://qims.amegroups.com/article/view/10.21037/qims-2025-635/rc
Data Sharing Statement: Available at https://qims.amegroups.com/article/view/10.21037/qims-2025-635/dss
Funding: This work was supported by
Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://qims.amegroups.com/article/view/10.21037/qims-2025-635/coif). The authors have no conflicts of interest to declare.
Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study was approved by the Ethics Committee of China-Japan Union Hospital of Jilin University (No. 2025KYYS036), and the requirement of informed consent was waived due to the retrospective nature of the study and the complete de-identification of all personally identifiable information.
Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.
References
- Suthar P, Patel R, Mehta C, Patel N. MRI evaluation of lumbar disc degenerative disease. J Clin Diagn Res 2015;9:TC04-9. [Crossref] [PubMed]
- Gaonkar B, Villaroman D, Beckett J, Ahn C, Attiah M, Babayan D, Villablanca JP, Salamon N, Bui A, Macyszyn L. Quantitative Analysis of Spinal Canal Areas in the Lumbar Spine: An Imaging Informatics and Machine Learning Study. AJNR Am J Neuroradiol 2019;40:1586-91. [Crossref] [PubMed]
- Mannil M, Burgstaller JM, Thanabalasingam A, Winklhofer S, Betz M, Held U, Guggenberger R. Texture analysis of paraspinal musculature in MRI of the lumbar spine: analysis of the lumbar stenosis outcome study (LSOS) data. Skeletal Radiol 2018;47:947-54. [Crossref] [PubMed]
- Chen T, Su ZH, Liu Z, Wang M, Cui ZF, Zhao L, Yang LJ, Zhang WC, Liu X, Liu J, Tan SY, Li SL, Feng QJ, Pang SM, Lu H. Automated Magnetic Resonance Image Segmentation of Spinal Structures at the L4-5 Level with Deep Learning: 3D Reconstruction of Lumbar Intervertebral Foramen. Orthop Surg 2022;14:2256-64. [Crossref] [PubMed]
- Masood RF, Taj IA, Khan MB, Qureshi MA, Hassan T. Deep learning based vertebral body segmentation with extraction of spinal measurements and disorder disease classification. Biomed Signal Process Control 2022;71:103230.
- Cheung JPY, Kuang X, Lai MKL, Cheung KM, Karppinen J, Samartzis D, Wu H, Zhao F, Zheng Z, Zhang T. Learning-based fully automated prediction of lumbar disc degeneration progression with specified clinical parameters and preliminary validation. Eur Spine J 2022;31:1960-8. [Crossref] [PubMed]
- Lehnen NC, Haase R, Faber J, Rüber T, Vatter H, Radbruch A, Schmeel FC. Detection of Degenerative Changes on MR Images of the Lumbar Spine with a Convolutional Neural Network: A Feasibility Study. Diagnostics (Basel) 2021.
- Liawrungrueang W, Kim P, Kotheeranurak V, Jitpakdee K, Sarasombath P. Automatic Detection, Classification, and Grading of Lumbar Intervertebral Disc Degeneration Using an Artificial Neural Network Model. Diagnostics (Basel) 2023.
- Grover VP, Tognarelli JM, Crossey MM, Cox IJ, Taylor-Robinson SD, McPhail MJ. Magnetic Resonance Imaging: Principles and Techniques: Lessons for Clinicians. J Clin Exp Hepatol 2015;5:246-55. [Crossref] [PubMed]
- Kushchayev SV, Glushko T, Jarraya M, Schuleri KH, Preul MC, Brooks ML, Teytelboym OM. ABCs of the degenerative spine. Insights Imaging 2018;9:253-74. [Crossref] [PubMed]
- Malfair D, Beall DP. Imaging the degenerative diseases of the lumbar spine. Magn Reson Imaging Clin N Am 2007;15:221-38. vi. [Crossref] [PubMed]
- Adams A, Roche O, Mazumder A, Davagnanam I, Mankad K. Imaging of degenerative lumbar intervertebral discs; linking anatomy, pathology and imaging. Postgrad Med J 2014;90:511-9. [Crossref] [PubMed]
- Fadda A, Lang J, Forterre F. Far lateral lumbar disc extrusion: MRI findings and surgical treatment. Vet Comp Orthop Traumatol 2013;26:318-22. [Crossref] [PubMed]
- Mok FP, Samartzis D, Karppinen J, Luk KD, Fong DY, Cheung KM. ISSLS prize winner: prevalence, determinants, and association of Schmorl nodes of the lumbar spine with disc degeneration: a population-based study of 2449 individuals. Spine (Phila Pa 1976) 2010;35:1944-52. [Crossref] [PubMed]
- Zhang R. A state-of-the-art survey of deep learning for lumbar spine image analysis: X-ray, CT, and MRI. AI Medicine 2024;1:3.
- Deng S, Yang Y, Wang J, Li A, Li Z. Efficient SpineUNetX for X-ray: A spine segmentation network based on ConvNeXt and UNet. J Vis Commun Image Represent 2024;103:104245.
- Kakkar M, Shanbhag D, Aladahalli C, Language Augmentation M GR. in CLIP for Improved Anatomy Detection on Multi-modal Medical Images. Annu Int Conf IEEE Eng Med Biol Soc 2024;2024:1-4. [Crossref] [PubMed]
- Zhou Z, Lei Y, Zhang B, Liu L, Liu Y. ZegCLIP: Towards Adapting CLIP for Zero-shot Semantic Segmentation. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE; 2023:11175-85.
- Liu B, Lu D, Wei D, Wu X, Wang Y, Zhang Y, Zheng Y. Improving Medical Vision-Language Contrastive Pretraining With Semantics-Aware Triage. IEEE Trans Med Imaging 2023;42:3579-89. [Crossref] [PubMed]
- Wang Z, Lu Y, Li Q, Tao X, Guo Y, Gong M, Liu T. CRIS: CLIP-Driven Referring Image Segmentation. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA. IEEE; 2022:11676-85.
- Zhu JY, Park T, Isola P, Efros AA. Unpaired image-to-image translation using cycle-consistent adversarial networks. 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy. IEEE: 2017:2242-51.
- Vrettos K, Koltsakis E, Zibis AH, Karantanas AH, Klontzas ME. Generative adversarial networks for spine imaging: A critical review of current applications. Eur J Radiol 2024;171:111313. [Crossref] [PubMed]
- Wang C, Macnaught G, Papanastasiou G, MacGillivray T, Newby D. Unsupervised Learning for Cross-Domain Medical Image Synthesis Using Deformation Invariant Cycle Consistency Networks. In: Simulation and Synthesis in Medical Imaging: Third International Workshop, SASHIMI 2018, Held in Conjunction with MICCAI 2018, Granada, Spain. Springer; 2018;3:52-60.
- Lu H, Li M, Yu K, Zhang Y, Yu L. Lumbar spine segmentation method based on deep learning. J Appl Clin Med Phys 2023;24:e13996. [Crossref] [PubMed]
- Pang S, Pang C, Zhao L, Chen Y, Su Z, Zhou Y, Huang M, Yang W, Lu H, Feng Q. SpineParseNet: Spine Parsing for Volumetric MR Image by a Two-Stage Segmentation Framework With Semantic Image Representation. IEEE Trans Med Imaging 2021;40:262-73. [Crossref] [PubMed]
- Tao R, Liu W, Zheng G. Spine-transformers: Vertebra labeling and segmentation in arbitrary field-of-view spine CTs via 3D transformers. Med Image Anal 2022;75:102258. [Crossref] [PubMed]
- Xie Y, Zhang J, Xia Y, Shen C. Learning From Partially Labeled Data for Multi-Organ and Tumor Segmentation. IEEE Trans Pattern Anal Mach Intell 2023;45:14905-19. [Crossref] [PubMed]
- Li Z, Li Y, Li Q, Wang P, Guo D, Lu L, Jin D, Zhang Y, Hong Q. LViT: Language Meets Vision Transformer in Medical Image Segmentation. IEEE Trans Med Imaging 2024;43:96-107. [Crossref] [PubMed]
- Tyler R, Jason T, Robyn B, Errol C, Adam F, Felipe K, Mongan J, Prevedello L, Vazirabad M. RSNA 2024 Lumbar Spine Degenerative Classification. Available online: https://kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification, 2024, Kaggle.
- Tianchi. Tianchi Spinal Disease Dataset. 2020. Available online: https://tianchi.aliyun.com/dataset/dataDetail?dataId=79463
- Natalia F, Meidia H, Afriliana N, Al-Kafri AS, Sudirman S, Simpson A, Sophian A, Al-Jumaily M, Al-Rashdan W, Bashtawi M. Development of Ground Truth Data for Automatic Lumbar Spine MRI Image Segmentation. 2018 IEEE 20th International Conference on High Performance Computing and Communications; IEEE 16th International Conference on Smart City; IEEE 4th International Conference on Data Science and Systems (HPCC/SmartCity/DSS), Exeter, UK. IEEE; 2018:1449-54.
- Wang J, Wu QMJ, Pourpanah F. DC-cycleGAN: Bidirectional CT-to-MR synthesis from unpaired data. Comput Med Imaging Graph 2023;108:102249. [Crossref] [PubMed]
- Klein S, Staring M, Murphy K, Viergever MA, Pluim JP. elastix: a toolbox for intensity-based medical image registration. IEEE Trans Med Imaging 2010;29:196-205. [Crossref] [PubMed]
- Woo S, Debnath S, Hu R, Chen X, Liu Z, Xie S, Kweon IS. Convnext v2: Co-designing and scaling convnets with masked autoencoders. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada. IEEE; 2023:16133-42.
- Wang Z, Wu Z, Agarwal D, Sun J. MedCLIP: Contrastive Learning from Unpaired Medical Images and Text. Proc Conf Empir Methods Nat Lang Process 2022;2022:3876-87.
- Liu J, Zhang Y, Chen JN, Xiao J, Lu Y, Landman BA, Yuan Y, Yuille A, Tang Y, Zhou Z. Clip-driven universal model for organ segmentation and tumor detection. 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France. IEEE; 2023:21095-107.
- Islam M, Vibashan VS, Jose VJM, Wijethilake N, Utkarsh U, Ren H. Brain tumor segmentation and survival prediction using 3D attention UNet. Brainlesion: Glioma, Multiple Sclerosis, Stroke and Traumatic Brain Injuries: 5th International Workshop, BrainLes 2019, Held in Conjunction with MICCAI 2019, Shenzhen, China, October 17, 2019, Revised Selected Papers, Part I 5. Springer International Publishing; 2020:262-72.
- Ruby U, Yendapalli V. Binary cross entropy with deep learning technique for image classification. International Journal of Advanced Trends in Computer Science and Engineering 2020;9: [Crossref]
- Ma J, Chen J, Ng M, Huang R, Li Y, Li C, Yang X, Martel AL. Loss odyssey in medical image segmentation. Med Image Anal 2021;71:102035. [Crossref] [PubMed]
- Isensee F, Jaeger PF, Kohl SAA, Petersen J, Maier-Hein KH. nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation. Nat Methods 2021;18:203-11. [Crossref] [PubMed]
- Wang H, Xie S, Lin L, Iwamoto Y, Han XH, Chen, YW, Tong R. Mixed transformer u-net for medical image segmentation. ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); 23-27 May 2022; Singapore, Singapore. IEEE; 2022:2390-4.
- Wu C, Zhang H, Chen J, Gao Z, Zhang P, Muhammad K, Del Ser J. Vessel-GAN: Angiographic reconstructions from myocardial CT perfusion with explainable generative adversarial networks. Future Generation Computer Systems 2022;130:128-39.
- Koch A, Aydin OU, Hilbert A, Rieger J, Tanioka S, Ishida F, Frey D. Cross-modality image synthesis from TOF-MRA to CTA using diffusion-based models. Med Image Anal 2025; Epub ahead of print. [Crossref]
- He K, Chen X, Xie S, Li Y, Dollár P, Girshick R. Masked autoencoders are scalable vision learners. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 18-24 June 2022; New Orleans, LA, USA. IEEE; 2022:15979-88.
- Radford A, Kim J W, Hallacy C, Ramesh A, Goh G, Agarwal S, Sastry G, Askell A, Mishkin P, Clark J, Krueger G, Sutskever I. Learning transferable visual models from natural language supervision. International Conference on Machine Learning, PMLR 2021;139:8748-63.
- Xu Y, Yu K, Qi G, Gong Y, Qu X, Yin L, Yang P. Brain tumour segmentation framework with deep nuanced reasoning and Swin‐T. IET Image Process 2024;18:1550-64.
- Markhali MI, Peloquin JM, Meadows KD, Newman HR, Elliott DM. Neural network segmentation of disc volume from magnetic resonance images and the effect of degeneration and spinal level. JOR Spine 2024;7:e70000. [Crossref] [PubMed]
- Yan K, Cai Q, Zhang F, Cao Z, Liu Z. SGTC: Semantic-Guided Triplet Co-training for Sparsely Annotated Semi-Supervised Medical Image Segmentation. Proceedings of the AAAI Conference on Artificial Intelligence 2025;39:9112-20.

