Mamba-based brain tumor segmentation of incomplete multi-modal MR images
Introduction
Background
Accurate evaluation of brain tumors is not only essential for understanding the progression of various types of encephalopathy and tumor-related systemic diseases but also plays a critical role in clinical diagnosis, prognosis, and treatment planning. Among available imaging modalities, magnetic resonance imaging (MRI) stands out as a widely used, non-invasive structural imaging technique capable of providing detailed and high-resolution visualizations of intracranial anatomy. In particular, multi-modal MRI—which includes sequences such as fluid-attenuated inversion recovery (FLAIR), T1-weighted (T1), contrast-enhanced T1 (T1c), and T2-weighted (T2)—offers complementary information by capturing different tissue characteristics, contrast levels, and temporal dynamics (1). However, in real-world clinical settings, complete multi-modal magnetic resonance (MR) images are often not available for every patient case. This limitation poses significant challenges for automated medical image analysis and diagnostic algorithms that rely on full multi-modal input. These practical challenges underscore the need for robust and adaptable methods capable of handling incomplete multi-modal data while still maintaining reliable diagnostic performance (2). In addition, brain tumors in MRI scans often visualize heterogeneous structures, meaning different sub-regions of the tumor show distinct biological and radiological characteristics. MRI sequences (e.g., T1, T1c, T2, FLAIR) capture different biological properties for brain tumor evaluation (3).
Existing approaches based on convolutional neural networks (CNNs) have played a central role in medical image segmentation, since automating the tasks of detecting, segmenting, and classifying tissue or lesions can free up the doctors’ time for higher value tasks and reduce errors due to fatigue and subjectivity (4,5). U-Net architectures are often developed, such as a Sharp U-Net architecture, which is designed with new connections between the encoder and decoder subnetworks using a depthwise convolution of the encoder feature maps with a sharpening spatial filter to address the semantic gap issue between the encoder and decoder features (4,5). However, due to the inherent locality of convolution operations, CNN-based models often struggle to capture global contextual features, which limits their ability to fully exploit the complementary information across multiple modalities. While CNNs excel at capturing local, hierarchical features (edges, textures, small patterns), their fixed and often limited receptive fields struggle to model long-range dependencies. In medical imaging, understanding the global context of an organ or lesion (e.g., the relationship between different parts of a large tumor, or how a kidney relates to the spine) is crucial for accurate segmentation. Without sufficient global context, CNNs can produce fragmented segmentation or miss large, diffuse structures. CNNs focus heavily on local neighborhoods, which can lead to misclassifications when global information is required to disambiguate similar local patterns.
More recently, vision transformers (ViTs) have been successfully applied to medical image segmentation, thanks to their self-attention mechanism, which is particularly effective at modeling long-range dependencies. By dividing input images into patches or tokens, transformers are able to compute pairwise interactions between all patches, enabling a more holistic understanding of the image. This ability has led to the integration of transformer modules in multi-modal MRI segmentation frameworks (6-9), improving the capture of global features and mitigating the effects of absent or incomplete modalities. However, this comes at a significant computational cost, scaling quadratically with the input sequence length. For high-resolution two-dimensional (2D) medical images or even three-dimensional (3D) volumetric data, this becomes computationally prohibitive and memory-intensive, limiting their practical application.
State-space models (SSMs) like Mamba have demonstrated potential on effectively reducing computational complexity to a linear level while capturing global features. This motivates us to develop a new framework for multi-modal medical image segmentation task. Mamba’s selective SSM design allows it to model long-range dependencies with computational complexity that scales linearly with the sequence length. This enables it to order the magnitude more efficient than transformers for high-resolution images, unlocking the possibility of processing entire medical scans (even 3D volumes) without downsampling them excessively or resorting to patch-based inference that compromises global context. Despite the linear complexity, Mamba has demonstrated the ability to effectively capture global contextual information, similar to transformers. It does this through its selective scan mechanism, which allows it to dynamically filter and propagate information along the sequence, effectively focusing on relevant long-range dependencies.
In this work, a new framework is designed to seamlessly integrate Mamba’s global context modeling with CNNs’ robust local feature extraction capabilities. This allows the model to simultaneously extract the fine details (e.g., tumor boundaries, subtle texture changes) and the broader anatomical context (e.g., the tumor’s relationship to surrounding organs, its overall shape and extent). Different MRI modalities (T1, T1c, T2, FLAIR) provide distinct, complementary information. A new framework can incorporate separate CNN encoders for each modality to extract modality-specific features. These features can then be fused before or within Mamba blocks to allow the SSM to learn global correlations across modalities. This is a powerful way to leverage all available modality features. Cross-level fusion Mamba (CFM) blocks are designed to capture global features of the low-level feature through its contextual learning mechanism. Cross-level feature fusion can be extended to fuse features not only across different levels within a single modality’s processing stream but also across different modalities at various levels of abstraction. This allows the model to leverage the complementary information from all available MRI sequences. By integrating information from diverse modalities, the model can create a richer representation that better distinguishes tumor subregions [e.g., enhancing tumor (ET), necrosis, edema] from healthy tissue, even in cases where one modality alone might be ambiguous, and therefore, cross-level fusion can potentially improve boundaries of tumor subregions. In addition, a cross-level uncertainty (CU) constraint is imposed on each class of the final predicted tumor. This strategy integrates modal information via lesion uncertainty and employs a weight mechanism, guiding the model to learn more discriminative regional features. Cross-level features imply that uncertainty is considered not just at the final output layer but potentially at different feature levels within the network. High-level features might capture the overall presence of a tumor but lack precise boundary information, leading to uncertainty at fine scales. Low-level features might have sharp edges but lack global context, leading to uncertainty in segmentation. By imposing a constraint across these levels, the model is forced to ensure that its uncertainty estimations are consistent and informed by both coarse and fine details. For example, if a high-level feature map is uncertain about a tumor region, the low-level features contributing to that region should also reflect a corresponding level of uncertainty regarding their local details. Uncertainty is often highest at object boundaries or inconsistent regions among different modalities. By specifically constraining uncertainty at these critical regions across different feature levels, the model is compelled to learn more precise and confident delineations. It forces the network to refine the features that contribute to distinguishing the tumor from its surroundings.
Related work
Brain tumor segmentation in multi-modal MR images is a critical task for diagnosis, treatment planning, and prognosis. Recent research has heavily focused on leveraging deep learning techniques, particularly CNNs and their variants, due to their remarkable performance in image analysis. Multi-modal MRI (e.g., T1, T1c, T2, FLAIR sequences) provides complementary information about the tumor, its subregions (ET, necrotic/cystic core, and peritumoral edema), and surrounding healthy tissue, which is crucial for accurate delineation.
Multi-modality MR image for brain disease diagnosis
Deep learning is increasingly utilized in medical imaging for various applications, including cancer detection, neuroimaging for cognitive impairment such as Alzheimer’s disease (AD), and real-time imaging for surgical guidance (10). Deep learning models, such as CNNs, have been widely applied to MRI data for AD detection. The Dual-3DM3AD approach, which uses a mixed transformer-based semantic segmentation and triplet pre-processing method, represents a significant advancement in the early multi-class diagnosis of AD (11).
Deep learning models can classify brain tumors into different categories, such as glioma, meningioma, and pituitary tumors, based on MRI features (12). Ensemble deep learning techniques have been employed to improve the accuracy and robustness of brain tumor classification (13). Deep learning models can differentiate between high- and low-grade glioblastoma multiforme while varying feature correlation limits (14).
Deep learning techniques can predict survival in patients with brain tumors, aiding in treatment planning and patient management (15). Predicting the survival of patients with brain tumors is a critical task that can inform treatment decisions (16). Deep learning models can integrate clinical factors, imaging features, and genomic data to provide personalized survival predictions (17). Dynamic segmentation networks integrate adversarial learning, dynamic CNNs, and attention mechanisms to enhance performance (18). Radiomics models, based on multi-view multiscale analysis, have been used to predict overall survival (19). Context-aware deep learning models can also predict survival by considering the uncertainty of tumor location in radiological images (20).
Multi-modality MR image segmentation
The U-Net architecture remains a cornerstone for medical image segmentation due to its ability to capture both local and global contextual information. Integrating attention gates or modules (e.g., convolutional block attention module, coordinate attention) (21,22) helps the model focus on salient features, particularly in the tumor regions, and suppress irrelevant information, leading to better segmentation accuracy and boundary delineation. Researchers are exploring multi-encoder U-Nets (23) that can process different MRI modalities separately and then fuse the extracted features at various levels. This allows for better leveraging of the distinct information from each modality. For volumetric MRI data, 3D U-Nets or extensions like V-Net (24) are commonly used to capture 3D spatial relationships and hierarchical features, providing better volumetric segmentation. However, they often come with increased computational complexity. Transformers, originally popular in natural language processing, are gaining traction in medical imaging due to their ability to model long-range dependencies. UNETR (Transformers for 3D Medical Image Segmentation) (25) is an example of this trend. While promising, integrating them effectively with CNNs (hybrid models) is a current area of research. While primarily an object detection model, improved versions of YOLO are being adapted for segmentation tasks in brain MRI, incorporating modules like Atrous Spatial Pyramid Pooling (ASPP) and attention mechanisms to enhance multi-scale feature representation and focus on critical tumor features (26). Models are designed to dynamically focus on salient features across modalities by jointly learning interdependencies between imaging sequences, leading to more precise boundary delineations, especially for challenging cases like low-grade astrocytomas (27). Quantifying the uncertainty of segmentation predictions is becoming increasingly important for clinical adoption (28). This helps clinicians identify areas where the model might be less confident, guiding manual review and building trust in CNNs. Approaches include Monte Carlo dropout and test-time augmentation. While complex models often achieve higher accuracy, computational intensity is a concern. Brain tumors exhibit high variability in shape, size, location, and intensity characteristics, making segmentation challenging. Models need to be robust to this variability (29). Ensemble networks that incorporate multi-view attention mechanisms are gaining prominence in brain tumor segmentation (30). These networks leverage information from multiple MRI modalities, such as T1, T2, and T1c images, to improve segmentation accuracy (31). Multi-view attention mechanisms allow the model to focus on the most relevant features in each modality, while the ensemble approach combines the predictions of multiple models to enhance robustness and reduce errors (32). Integration of multi-modal images across multiple scales, and effectively eliminating noise interference in brain tumor images, remains a major limitation in segmentation tasks (33).
Absent modality synthesis
The field of incomplete multi-modal MR image segmentation has seen the emergence of generative methods in addressing the issues of absent modalities by synthesizing the absent data from available modalities. U-HVED (34) was a hetero-modal variational encoder-decoder for tumor segmentation of absent modalities. Multi-modal variational auto-encoders were introduced to create the common latent variables, and a mixture sampling procedure was introduced to evaluate the reconstruction error on all the modalities. Sharma et al. (35) proposed a variant of generative adversarial network (GAN) capable of leveraging redundant information contained within multiple available sequences in order to generate one or more absent sequences for a patient scan. Hyper-GAE (36) was a unified and adaptive multi-modal MR image synthesis network for tumor segmentation with absent modalities. The feature-level and image-level completion was achieved by a shared hyper-encoder for embedding each available modality into the feature space, a graph-attention-based fusion block to aggregate the features of available modalities to the fused features, and a shared hyper-decoder for image reconstruction. In addition, a hypernet-based modulation module to adaptively utilize the real and synthetic modalities. Meng et al. (37) proposed a unified multi-modal modality masked diffusion network to synthesize arbitrary absent modalities of brain MRI within a single network. Based on the learned common latent space, absent modalities can be drawn from Gaussian noise via a diffusion model. Zhang et al. (38) proposed a unified multi-modal image synthesis method for absent modality imputation. A commonality and discrepancy-sensitive encoder for the generator was designed to exploit both modality-invariant and specific information contained in input modalities. A dynamic feature unification module was designed to hardly and softly integrate information from a varying number of available modalities. A multi-input, multi-output network was proposed to combine information from all the available pulse sequences and synthesize the absent ones in a single forward pass.
Incomplete modality fusion and interaction
HeMIS (39) was a respective convolutional pipeline to independently process each modality for computing mapwise statistics such as the mean and the variance. The mean and variance feature maps of available modalities were concatenated and fed into a final set of convolutional stages. Chen et al. (40) used feature disentanglement to decompose the input modalities into the modality-specific appearance code for each modality and the modality-invariant content code for the segmentation task. The disentangled content code from each modality was fused into a shared representation for incomplete modalities. RFNet (41) was designed with a region-aware fusion module to fuse modal feature from available image modalities according to different regions. The region-aware fusion was conducted on the divided features in each region, and modal-wise attention weights were learned individually in different regions to aggregate the corresponding features. SMU-Net (42) was a style matching U-Net (SMU-Net) for brain tumor segmentation on MR images. A content and style-matching mechanism was proposed to distill the informative features from the full-modality network into an absent modality network. mmFormer (6) consisted of the hybrid modality-specific encoders, an inter-modal transformer, and a decoder. The hybrid modality-specific encoder was used to extract both local and global context information within a specific modality by bridging a convolutional encoder and an intramodal transformer. The auxiliary regularizer in the convolutional decoder was used to force the decoder to generate segmentation even when certain modalities were absent. Qiu et al. (43) froze the model with incomplete modalities and incorporate a lightweight modal-aware adapter during finetuning. The features generated by the modality state classifier was used as prompts that were conscious of the absence of modalities. M2FTrans (7) introduced learnable fusion tokens and masked self-attention to stably build long range dependency across modalities, while being more flexible to learn from incomplete modalities and modality-specific features were further re-weighted through spatial weight attention and channel-wise fusion transformers for feature redundancy reduction and modality rebalancing. IMS2Trans (8) was a Swin Transformer network by utilizing a single encoder to extract latent feature maps from all available modalities. A lightweight shifted multi-layer perceptron coupled with a masking bottleneck was designed. A feature distillation strategy between individual modalities and the entire set of modalities was used to by comparing the features of each modality with the averaged features computed across all modalities. MIFPN (9) was designed to introduce modality invariant feature prompts in modal information interaction, guiding the model to learn the incomplete modality information and facilitating integration. Modality-aware masks and modality selection strategies were used to merge shallow features of the encoder.
Methods
To mitigate reduced model performance from absent modalities, our novel network is elaborated, specifically engineered for the segmentation of brain tumors in MRI scans with arbitrary MRI modalities absent, as shown in Figure 1A. During the interaction stage, introducing SSMs allows for capturing deep knowledge of absent modalities from cross-modal correlations, thereby effectively learning inter-modal information and guiding the model to learn information about the absent modality. As shown in Figure 1B, the Mamba block is proposed for each modality encoded feature with single-axis scan along width for efficiency. As shown in Figure 1C, the Mamba fusion (MF) block is proposed to foster the learning of absent modality features, leading to a more comprehensive representation of multi-modal MR images for tumor segmentation. CFM blocks are designed to capture global features of the low-level feature through its contextual learning mechanism, as shown in Figure 1D. A CU constraint is imposed on each class of the final predicted tumor. This strategy integrates modal information via lesion uncertainty and employs a weight mechanism, guiding the model to learn more discriminative regional features. Through these components, the model significantly improves brain tumor image segmentation accuracy, especially in scenarios with absent modalities.
Datasets
Two brain tumor segmentation datasets are used to evaluate the proposed method. The two datasets were distributed in the multi-modal Brain Tumor Segmentation Challenge (BraTS) (1): BraTS2018 and BraTS2020. These datasets comprise four MRI modalities: FLAIR, T1, T1c, and T2. Following previous works (6-9,40), we have implemented several preprocessing steps, including the removal of non-brain black background and normalization of each modality to zero mean and unit variance. For consistency and fair comparison (6-9,40), the same grouping strategy as the above previous works is strictly utilized. The BraTS2018 dataset is divided into 199 training samples, 29 validation samples, and 57 test samples, with evaluation conducted using a three-fold cross-validation approach. Similarly, the BraTS2020 dataset is divided into 219 training, 50 validation, and 100 test samples. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments.
MF block
An MF block is proposed to efficiently capture global contextual relationships across multi-modal MR images. Its core strength lies in using SSMs to effectively learn and integrate features from absent modalities during interactions, which significantly improves the completeness of the fused features. Given that fused features inherently possess extensive cross-modal correlations, vision Mamba (ViM) plays a crucial role in fostering the learning of these absent modality features. This leads to a more comprehensive representation of multi-modal MR images, ultimately ET segmentation. As illustrated in Figure 1B, our Mamba-based fusion block is specifically designed to fuse modality-specific features through these cross-modal interactions.
In our task, one of the multi-modal MR images can be defined as , and represents ground truth of the tumor annotated by experienced doctors. To better extract modality-specific feature, the modality features are independently extracted. For each modality, a modality-specific encoder is constructed to obtain the corresponding high-level feature as
where respectively denotes the channel, depth, height, and width of fm, Em denotes the modality-specific encoder of the mth modality MR image Im and θm denotes the learnable parameter of Em. In our framework, Em is constructed as the encoder of 3D-U-Net, which comprises five feature extraction blocks in succession. Each block comprises a 3D convolutional layer with a 3×3×3 kernel size, instance normalization, and a LeakyReLU activation function. A convolutional layer with a stride of 2 between consecutive levels is used to downsample the feature maps. The number of channels is sequentially set to 16, 32, 64, 128, and 256 at each level.
To efficiently capture intra-modality contextual relations, we propose a mamba block for each modality-specific feature fm. Since Mamba can only encode sequential data, the vanilla 2D ViM block in (44) has four cross-scans horizontally and vertically across spatial dimensions to generate four spatial vision sequences, which are then processed by four separate S6 models to extract spatial features. However, such a scanning strategy is inefficient for extracting spatial features from 3D volume images. Adding scans in each additional direction increases the number of sequences in parallel and actually leads to an increase in the memory consumption, the computational complexity, and the number of parameters. Therefore, we propose the spatial unidirectional scan strategy to efficiently extract global spatial features. As shown in Figure 1B, for the input feature of our 3D ViM block, the channel is permuted to the last dimension and layer normalization is performed to produce the feature . Two linear transformations are parallel applied to to change the same channel as . The sigmoid linear unit (SiLU) activation is applied to the resulted feature in the bottom path as
where is the output feature of the bottom path, and denotes a linear transformation for changing the channel from to . Similarly, a linear transformation is applied to to change the same channel as in the top path and the channel is permuted. CIL is used to extract local features, where CIL denotes a 3D convolutional layer with a 3×3×3 kernel size, instance normalization, and a LeakyReLU activation function. A voxel-by-voxel scan along the width of feature is performed to generate a spatial vision sequence. The voxel-by-voxel scan is implemented by voxel flattening and dimension permuting and the spatial vision sequence can be obtained as ft. This sequence ft is then fed into an S6 model for state-space evolution and SSM within Mamba can be defined as follows:
where , and are all the training parameters of the S6 model and denotes the Mamba hidden state dimension. After efficient spatial Mamba processing, a layer normalization and a reshaping operation are performed to obtain a feature with , and a linear projection are performed to original embedded dimension . Finally, the Mamba feature is obtained by dimension permuting as the original input as follows
where denotes the Hadamard product operation.
To simulate modality unavailability in real scenarios, a Bernoulli indicator is introduced for modality-specific feature zeroing, i.e., modality-specific feature of a absent modality is replaced by a zero tensor. As shown in Figure 1C, for the input zeroed and concatenated feature of our 3D ViM block, the channel is permuted to the last dimension and layer normalization is performed to produce the feature . Two linear transformations are parallel applied to to change the same channel as . The SiLU activation is applied to the resulted feature in the bottom path as
where is the output feature of the bottom path. Similarly, a linear transformation is applied to to change the same channel as in the top path and the channel is permuted to obtain . CIL and voxel-by-voxel scan along the width of feature are performed to generate a spatial vision sequence. The spatial vision sequence can be obtained as . This sequence is then fed into an S6 model within Mamba can be defined as follows
where , , and are all the training parameters of the S6 model. After efficient spatial Mamba processing, a layer normalization and a reshaping operation are performed to obtain a feature with , and a linear projection are performed to original embedded dimension . Finally, the Mamba feature is obtained by dimension permuting as the original input as
The fused Mamba feature is further added to the input zeroed and concatenated feature with a CIL operation as
For the input concatenated feature of our 3D ViM block, the channel is permuted to the last dimension and layer normalization is performed to produce the feature . Two linear transformations are parallel applied to to change the same channel as . CIL and voxel-by-voxel scan along the width of feature are performed to generate a spatial vision sequence. The spatial vision sequence can be obtained as . This sequence is then fed into an S6 model within Mamba can be defined as
The fused feature with the proposed Mamba block of available modalities is aligned with that of full modalities, such that the proposed Mamba block can extract the similar feature of full modalities. Mahalanobis distance is used to measure the difference between fused features of available modalities and full modalities as
where is the covariance matrix of the features. Unlike Euclidean distance, Mahalanobis distance accounts for the correlations between fused features according to their variance. It measures distance in terms of standard deviations from the mean of a distribution.
CFM block
Global context is critical: tumors are not isolated; their shape, location, and relationship to surrounding structures (e.g., ventricles, white matter) determine their segmentation. To fuse the fine details and the macro features, CFM blocks are designed as the blocks of U-Net decoder, as shown in Figure 1C. In this architecture, the high-level feature is upsampled and fed into the Mamba branch. The Mamba layer captures global features of the low-level feature through its contextual learning mechanism. Mamba’s contextual learning mechanism, through selective state updates and long-range propagation, efficiently models these relationships. By combining upsampled high-level features (semantic guidance) with low-level features (fine details), Mamba ensures that the global context it captures is both semantically meaningful high-level knowledge and low-level details—ultimately improving the model’s ability to segment complex, heterogeneous structures like brain tumors. This mechanism distinguishes Mamba from attention-based or purely local methods, making it well-suited for our architecture’s goal of fusing multi-scale features while preserving global contextual coherence.
To be similar to Figure 1B, the first branch upsamples and projects the input features linearly (linear), and processes the sequence data through 3D convolution operations to enhance feature representation. These processed features are then fed into the SSM. The SSM enables global modeling of the sequence, allowing effective information propagation across different voxels and capturing the sequence’s global characteristics of the high-level feature. In addition, the SSM focuses on the most important of the sequence while ignoring less relevant voxels to improve both the efficiency and effectiveness of high-level feature extraction. For the input feature of our CF-Mamba block, upsampling is performed and the channel is permuted to the last dimension and layer normalization is performed to produce the feature . Two linear transformations are parallel applied to to change the same channel as . The SiLU activation is applied to the resulted feature in the bottom path as
where is the output feature of the bottom path of the l−1th level fused feature, and denotes a linear transformation for changing the channel from to . Similarly, a linear transformation is applied to to change the same channel as in the top path and the channel is permuted of the l−1th level fused feature. CIL is used to extract local features. The voxel-by-voxel scan is implemented by voxel flattening and dimension permuting and the spatial vision sequence can be obtained as . This sequence is then fed into an S6 model for state-space evolution and SSM within Mamba as
where , and are all the training parameters of the S6 model. After efficient spatial Mamba processing, a layer normalization and a reshaping operation are performed to obtain a feature with a size of . The output features of the l−1 level features undergo the Hadamard product operation as
The available and zeroed modality-specific features are concatenated and fed into a CIL block to fuse the skip connections. Similarly, the output feature is then processed as the top path to capture long-range dependencies of the fine details in the lower-level feature, and the Hadamard product operation is used to fuse the two-level features as
where denotes the Mamba feature of the skip connections in the lth level as Eq. [12]. and are then concatenated and fed into a linear transformation layer and a CIL block to further fuse the two-level features. The decoder comprises four CFM blocks and a 3D convolutional layer to predict the tumor.
CU
Due to the limitation of the model, multiple-level features do not have a fixed contour or shape of a lesion or an organ, and even the neighboring-level features may vary largely from each other for the same image. The shape or size of a lesion or an organ may vary largely from each other for different patients or different time. This probably leads to the degradation of deep neural network models for variable shape or size of a lesion or an organ in multiple-level features. Therefore, it is necessary to measure the uncertain lesion area sizes.
To achieve more accurate segmentation of brain tumors in the presence of missing modalities, we adopt a deep supervision segmentation strategy. For the lth level fused feature , a 2l−1 upsampling operation is applied to restore the fused feature to the original resolution, and then the output feature is fed into a 3D convolutional layer to predict the tumor. The loss functions utilized in the model training process consist of Dice loss and cross-entropy (CE) loss. These two loss functions are widely used in medical image segmentation tasks due to their effectiveness in handling class imbalance and capturing fine details in segmentation boundaries. The Dice loss is particularly advantageous for improving the overlap between predicted and annotated tumors, and the CE loss provides a measure for pixel-wise classification. Compared with the ground truth to compute the loss can be defined by
where denote the lth level predicted tumor, denotes a 2l−1 upsampling operation, Conv3D denotes a 3D convolutional layer. denotes the vth voxel of the clsth class of the label, and denotes the vth voxel of the clsth class of the lth level predicted tumor. Nv denotes the number of voxels and Ncls the number of classes.
Considering that different levels of features contribute differently to the final segmentation result, we design a level consistency mechanism to measure the relationships between different levels of features and make more effective use of different levels of features. This strategy maximizes the utilization of different scales of features and semantic information at each level, enhances our understanding of brain tumors, and enables more accurate segmentation. The CU of the clsth class can be defined as
where denotes the number of levels. The CU can be used to constrain the final segmentation prediction as
where denotes the vth voxel of the clsth class of the final predicted tumor.
The overall loss includes the deep supervision loss Ll and the final segmentation loss Lf, written as
where λ is the coefficient of the Mahalanobis distance LM.
Results
Evaluation metrics
To comprehensively analyze the segmentation performance of the proposed method, the dice similarity coefficient (DSC) and Hausdorff distance (HD) are used as evaluation metrics. The DSC measures the volume overlap between the predicted tumors and the manually annotated labels by experienced doctors, with higher DSC scores indicating better segmentation accuracy. The 95% HD (HD95) is used to measure the surface distances between the predicted tumors and the manually annotated labels, with lower values indicating more approximate surface as that of ground truth.
Implementation details
Our method is implemented in PyTorch. Our method and baselines are trained and tested by utilizing a Nvidia GeForce RTX 4090. During the training process, the batch size is set to 2, and the AdamW optimizer is used to optimize model parameters, with an initial learning rate of 0.0002 and a weight decay of 0.0001. Considering memory consumption, each modality is randomly cropped to a subvolume with 80×80×80 voxels, and various data augmentation techniques such as random rotations, intensity variations, and mirror flipping are used. In addition, a warm-up learning rate adjustment strategy and a polynomial decay strategy is also used. To simulate absent modality scenarios, 15 different cases of modality absent are constructed as FLAIR, T1, T1c, T2, T1c + T2, T1 + T1c, FLAIR + T1, T1 + T2, FLAIR + T2, FLAIR + T1c, FLAIR + T1 + T1c, FLAIR + T1 + T2, FLAIR + T1c + T2, T1 + T1c + T2, FLAIR + T1 + T1c + T2. The 15 patterns are sampled patterns uniformly at training time and the test set is evaluated all 15 patterns evenly. The folds of the BraTS2018 dataset and the BraTS2020 dataset are followed the previous works. In the inference stage, sliding-window inference with overlap covers the entire volume and no bias patches toward lesion voxels during training. As previous works (6-9,40,41), random factor with uniform distribution is generated to randomly select one of these scenarios to simulate absent modalities.
Comparison with state-of-the-art method
Based on the availability of source codes and data splits, six state-of-the-art methods, adopting the publicly available data splits for incomplete multi-modal brain tumor segmentation, were included for comparison, including RobustSeg (41), RFNet (41), mmFormer (6), M2FTrans (7), IMS2Trans (8), and MIFPN (9).
Comparison on BraTS2018
Table 1 summarized the quantitative comparison results across all 15 multi-modal combinations on the BraTS2018 dataset. RobustSeg achieved mean DSC scores of 83.25, 72.72, and 49.95 for whole tumor (WT), tumor core (TC), and ET, respectively. Among the comparison methods, MIFPN demonstrated the best overall segmentation performance for WT, TC, and ET, surpassing RobustSeg with a mean increase of 3.11, 5.64, and 14.53 in DSC scores, respectively. Our proposed method, however, achieved the overall best segmentation performance for WT, TC, and ET, outperforming MIFPN with an increase of 2.05, 2.86, and 7.05 in the mean DSC scores, respectively.
Table 1
| Method | Class | FLAIR | T1 | T1c | T2 | T1c + T2 | T1 + T1c | FLAIR + T1 | T1 + T2 | FLAIR + T2 | FLAIR + T1c | FLAIR + T1 + T1c | FLAIR + T1 + T2 | FLAIR + T1c + T2 | T1 + T1c + T2 | FLAIR + T1 + T1c + T2 | Mean |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| RobustSeg | WT | 81.21±10.71 | 70.09±18.54 | 63.83±21.70 | 79.48±13.13 | 87.16±7.06 | 73.64±16.08 | 85.41±8.61 | 85.74±7.99 | 87.33±7.48 | 87.77±6.97 | 89.19±5.95 | 88.96±6.29 | 90.20±6.27 | 87.85±6.68 | 90.82±5.60 | 83.25±10.55 |
| TC | 58.42±22.04 | 47.62±24.62 | 53.36±23.79 | 56.82±22.89 | 84.79±7.30 | 81.64±9.36 | 67.88±14.45 | 67.22±16.72 | 75.33±12.34 | 83.43±8.95 | 85.67±7.59 | 70.89±14.56 | 85.15±7.87 | 86.14±7.62 | 86.37±6.13 | 72.72±13.09 | |
| ET | 30.55±29.17 | 34.23±26.97 | 18.48±39.13 | 28.64±32.83 | 68.82±14.65 | 68.01±16.00 | 33.44±31.95 | 34.29±26.94 | 51.29±23.87 | 67.54±14.61 | 68.82±12.47 | 37.56±27.47 | 68.33±15.84 | 70.05±12.28 | 69.17±15.11 | 49.95±21.52 | |
| RFNet | WT | 84.86±8.33 | 71.29±18.66 | 68.95±18.01 | 83.67±9.80 | 86.71±7.31 | 74.81±15.11 | 88.12±7.01 | 87.74±7.48 | 88.73±6.76 | 87.71±6.76 | 89.36±6.17 | 90.59±5.55 | 90.32±5.81 | 87.33±7.60 | 90.81±5.05 | 84.73±8.55 |
| TC | 60.85±19.58 | 58.36±20.40 | 54.28±23.77 | 56.88±22.85 | 83.87±7.42 | 81.31±9.91 | 69.74±14.83 | 68.65±14.73 | 75.96±11.30 | 84.03±8.46 | 85.26±8.11 | 71.95±13.18 | 85.01±7.94 | 84.93±7.69 | 85.89±7.20 | 73.80±12.58 | |
| ET | 36.21±31.90 | 39.52±24.80 | 26.97±36.52 | 34.27±30.89 | 73.62±10.82 | 72.38±12.98 | 39.74±24.71 | 40.48±27.38 | 57.08±20.17 | 73.23±11.78 | 73.84±10.73 | 43.41±28.30 | 73.58±11.62 | 74.35±10.26 | 74.06±11.67 | 55.52±18.24 | |
| mmFormer | WT | 84.01±10.71 | 77.49±18.54 | 78.37±21.70 | 86.24±13.13 | 85.51±7.06 | 79.97±16.08 | 87.71±8.61 | 85.61±7.99 | 87.85±7.48 | 87.85±6.97 | 88.21±5.95 | 88.31±6.29 | 88.44±6.27 | 86.19±6.68 | 88.62±5.60 | 85.36±10.55 |
| TC | 68.37±16.13 | 61.22±18.61 | 61.60±19.20 | 64.24±16.45 | 82.39±8.98 | 77.59±10.08 | 73.43±12.22 | 73.08±12.11 | 78.12±10.50 | 82.89±8.04 | 84.66±7.36 | 75.67±11.19 | 85.61±6.62 | 84.43±8.56 | 86.16±6.37 | 75.96±12.98 | |
| ET | 42.41±23.04 | 39.99±26.40 | 33.33±33.34 | 34.94±32.53 | 73.07±11.58 | 72.58±12.06 | 39.19±26.15 | 44.82±22.62 | 58.48±17.02 | 72.62±12.05 | 73.91±10.96 | 46.02±24.29 | 73.67±11.32 | 73.78±13.11 | 73.96±13.02 | 56.85±18.99 | |
| M2FTrans | WT | 84.47±9.47 | 75.76±13.57 | 74.43±16.11 | 87.42±7.80 | 86.51±8.23 | 78.89±12.03 | 89.06±6.45 | 86.03±8.94 | 89.46±5.80 | 89.48±6.31 | 89.61±6.55 | 89.74±6.46 | 90.17±5.41 | 86.93±8.50 | 90.17±5.41 | 85.88±9.04 |
| TC | 69.18±13.87 | 63.68±18.52 | 66.19±15.55 | 67.73±16.78 | 86.32±6.43 | 84.51±6.97 | 73.75±13.65 | 73.89±14.10 | 79.67±10.77 | 86.08±6.40 | 86.53±6.60 | 75.91±11.08 | 86.64±6.81 | 86.59±6.71 | 86.91±6.41 | 78.24±10.88 | |
| ET | 47.03±23.84 | 34.02±31.67 | 38.77±26.94 | 39.49±24.20 | 74.41±10.49 | 74.91±11.29 | 44.53±27.74 | 49.16±20.34 | 61.73±19.14 | 74.79±12.61 | 76.05±10.30 | 49.99±25.01 | 75.21±10.41 | 75.57±10.50 | 76.11±9.56 | 59.45±19.87 | |
| IMS2Trans | WT | 85.12±9.23 | 72.62±16.70 | 68.05±19.81 | 84.43±9.96 | 88.66±7.37 | 75.71±13.85 | 90.09±6.44 | 88.33±7.24 | 90.74±5.74 | 90.82±5.78 | 91.61±4.95 | 90.94±4.98 | 91.32±5.64 | 88.98±6.83 | 92.67±4.03 | 86.01±7.97 |
| TC | 70.61±14.99 | 71.67±13.88 | 72.24±14.71 | 71.15±13.56 | 83.12±8.61 | 82.71±8.47 | 76.11±13.14 | 74.79±13.61 | 78.46±10.12 | 82.98±8.00 | 83.78±8.43 | 76.43±10.84 | 83.41±8.79 | 83.73±8.62 | 83.92±7.40 | 78.35±11.91 | |
| ET | 47.12±24.85 | 42.29±28.86 | 46.11±22.63 | 48.52±22.14 | 72.77±11.98 | 72.87±11.67 | 50.97±23.53 | 49.74±21.11 | 61.61±19.20 | 72.74±12.54 | 74.17±11.62 | 51.32±19.96 | 73.74±12.60 | 73.38±11.71 | 74.91±12.55 | 60.82±16.46 | |
| MIFPN | WT | 85.01±8.24 | 78.49±12.69 | 79.37±11.35 | 87.24±8.04 | 86.51±7.42 | 80.97±10.66 | 88.71±6.21 | 86.61±8.70 | 88.85±6.36 | 88.85±6.24 | 89.21±6.15 | 89.31±6.95 | 89.44±5.81 | 87.19±7.17 | 89.62±6.02 | 86.36±8.46 |
| TC | 70.81±15.47 | 71.67±13.03 | 72.24±12.77 | 71.15±14.14 | 83.12±8.44 | 82.71±9.16 | 76.11±11.71 | 74.79±13.87 | 78.46±10.99 | 83.02±8.32 | 83.78±7.95 | 76.44±11.54 | 83.41±8.79 | 83.73±8.62 | 83.92±8.52 | 78.36±9.95 | |
| ET | 50.12±24.94 | 55.29±21.01 | 49.11±22.90 | 51.52±24.24 | 75.77±11.87 | 75.87±12.07 | 53.97±22.55 | 52.74±23.16 | 64.61±16.63 | 75.74±9.70 | 77.17±10.73 | 54.32±20.10 | 76.74±11.63 | 76.38±10.87 | 77.91±10.38 | 64.48±17.05 | |
| Ours | WT | 87.77±6.60 | 79.29±11.18 | 80.71±8.68 | 88.37±5.82 | 89.11±5.66 | 81.71±8.60 | 90.37±5.30 | 88.56±6.29 | 91.05±4.74 | 91.43±4.54 | 91.83±3.76 | 91.54±4.15 | 91.87±4.23 | 89.32±5.45 | 93.16±3.49 | 88.41±5.22 |
| TC | 72.38±13.81 | 69.68±13.64 | 74.96±10.27 | 73.17±12.34 | 87.48±5.76 | 87.92±5.68 | 77.25±10.24 | 76.78±9.75 | 80.78±9.03 | 87.31±6.35 | 88.28±4.92 | 78.53±10.52 | 87.16±6.42 | 88.25±5.76 | 88.37±5.00 | 81.22±9.20 | |
| ET | 61.39±11.58 | 60.79±12.55 | 63.29±11.01 | 61.61±13.05 | 80.32±7.87 | 80.35±7.47 | 60.31±13.89 | 63.37±13.92 | 66.31±12.47 | 81.59±5.71 | 82.44±6.85 | 67.68±12.60 | 80.63±6.78 | 81.02±7.02 | 81.92±7.05 | 71.53±9.40 |
Data are presented as mean ± standard deviation. BraTS, Brain Tumor Segmentation Challenge; DSC, dice similarity coefficient; ET, enhancing tumor; FLAIR, fluid-attenuated inversion recovery; T1, T1-weighted; T1c, contrast-enhanced T1; T2, T2-weighted; TC, tumor core; WT, whole tumor.
To validate that our method offers improvements over others, we used a Wilcoxon matched-pairs signed-ranks two-tailed test for the mean DSC and HD95 for WT, TC, and ET to calculate the P values. The P values for the scores to verify whether our method significantly improved segmentation results under these conditions. As shown in Table 2, P<0.05 of Wilcoxon matched-pairs signed-ranks two-tailed test for the mean DSC scores for WT, TC, and ET confirmed that our method statistically significantly outperformed previous methods.
Table 2
| Method | Class | BraTS2018 | BraTS2020 | |||
|---|---|---|---|---|---|---|
| DSC | HD95 | DSC | HD95 | |||
| RobustSeg | WT | 5.16E−12 | 5.16E−18 | 6.6E−14 | 6.6E−28 | |
| TC | 7.86E−18 | 7.86E−26 | 2.47E−26 | 2.47E−31 | ||
| ET | 5.17E−10 | 5.17E−18 | 3.89E−26 | 3.89E−26 | ||
| RFNet | WT | 0.0000672 | 6.72E−23 | 6.11E−14 | 6.11E−27 | |
| TC | 2.02E−16 | 2.02E−33 | 2.1E−22 | 2.1E−33 | ||
| ET | 6E−34 | 6E−21 | 3.98E−40 | 3.98E−21 | ||
| mmFormer | WT | 0.0000287 | 2.87E−11 | 4.45E−10 | 4.45E−13 | |
| TC | 2.41E−12 | 2.41E−10 | 8.03E−18 | 8.03E−12 | ||
| ET | 3.1E−30 | 3.1E−12 | 3.43E−36 | 3.43E−13 | ||
| M2FTrans | WT | 0.000183 | 0.00183 | 0.0000314 | 3.14E−11 | |
| TC | 0.000264 | 2.64E−10 | 7.69E−16 | 7.69E−10 | ||
| ET | 2.5E−26 | 2.5E−10 | 2.75E−28 | 2.75E−11 | ||
| IMS2Trans | WT | 0.0131 | 0.00131 | 0.000262 | 2.62E−10 | |
| TC | 0.0369 | 3.69E−14 | 7.61E−16 | 7.61E−12 | ||
| ET | 2.26E−18 | 2.26E−09 | 7.24E−26 | 7.24E−12 | ||
| MIFPN | WT | 0.015 | 0.0103 | 0.0253 | 0.0131 | |
| TC | 0.012 | 0.0248 | 0.0135 | 0.0258 | ||
| ET | 9.9E−16 | 0.0142 | 0.0186 | 0.0186 | ||
BraTS, Brain Tumor Segmentation Challenge; DSC, dice similarity coefficient; ET, enhancing tumor; HD95, 95% Hausdorff distance; TC, tumor core; WT, whole tumor.
In the three missing-modality patterns on the BraTS2018, the proposed method improved the mean DSC scores for WT, TC, and ET of 1.5, 1.1, and 10.3 over MIFPN and the mean HD95 for WT, TC, and ET of 0.9, 1.76, and 1.41 over MIFPN. In the two missing-modality patterns on the BraTS2018, the proposed method improved the mean DSC scores for WT, TC, and ET of 1.96, 3.2, and 5.5 over MIFPN and the mean HD95 for WT, TC, and ET of 1.16, 1.26, and 1.01 over MIFPN. In the one missing-modality patterns on the BraTS2018, the proposed method improved the mean DSC scores for WT, TC, and ET of 2.35, 3.72, and 6.79 over MIFPN and the mean HD95 for WT, TC, and ET of 0.92, 0.26, and 0.17 over MIFPN. In clinical reality, certain modality absences (e.g., missing T1c due to contrast allergy) were more common and clinically consequential than others. For the pattern of missing only T1c on the BraTS2018 dataset, the proposed method surpassed MIFPN with mean DSC score increases of 2.23, 2.09, and 13.36 for WT, TC, and ET. For the pattern of missing both T1 and T1c on the BraTS2018 dataset, the proposed method surpassed MIFPN with mean DSC score increases of 2.2, 2.32, and 1.7 for WT, TC, and ET. These results showed that less missing-modality patterns and larger gains in the mean DSC scores of the BraTS2018.
Table 3 summarized the HD95 for all fifteen multi-modal combinations on the BraTS2018 dataset. RFNet recorded the highest mean HD95 scores across all combinations, with values of 12.52, 15.3, and 10.36 for WT, TC, and ET, respectively. Among the comparison methods, MIFPN achieved the lowest mean HD95 scores at 6.84 (WT), 6.41 (TC), and 5.34 (ET). Our proposed method, however, attained the overall lowest mean HD95 for WT, TC, and ET, surpassing MIFPN with a decrease of 1.06, 1.2, and 0.98 in the mean HD95 scores, respectively. P<0.05 of Wilcoxon matched-pairs signed-ranks two-tailed test for the mean HD95 for WT, TC, and ET confirmed that our method statistically significantly outperformed previous methods.
Table 3
| Method | Class | FLAIR | T1 | T1c | T2 | T1c + T2 | T1 + T1c | FLAIR + T1 | T1 + T2 | FLAIR + T2 | FLAIR + T1c | FLAIR + T1 + T1c | FLAIR + T1 + T2 | FLAIR + T1c + T2 | T1 + T1c + T2 | FLAIR + T1 + T1c + T2 | Mean |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| RobustSeg | WT | 17.18±16.84 | 20.64±19.81 | 22.13±21.24 | 17.69±17.34 | 7.33±6.89 | 12.26±11.16 | 11.73±11.50 | 7.42±7.35 | 9.18±8.35 | 8.41±8.16 | 6.71±6.64 | 6.65±5.99 | 5.47±5.36 | 5.92±5.39 | 5.51±5.45 | 10.95±10.40 |
| TC | 15.87±14.92 | 26.41±24.03 | 26.36±25.31 | 20.83±20.21 | 6.93±6.65 | 13.87±13.73 | 15.29±13.91 | 10.04±9.54 | 12.62±11.86 | 11.35±10.44 | 6.36±5.79 | 9.44±9.06 | 7.59±7.21 | 7.96±7.40 | 5.18±4.66 | 13.07±12.94 | |
| ET | 10.49±10.39 | 19.34±18.76 | 15.92±15.76 | 15.97±15.17 | 4.92±4.87 | 10.79±10.36 | 11.03±10.81 | 7.07±6.58 | 8.53±7.76 | 7.41±7.04 | 4.64±4.59 | 8.66±8.57 | 6.61±6.02 | 7.12±7.05 | 4.51±4.06 | 9.53±9.34 | |
| RFNet | WT | 10.15±9.34 | 36.19±32.93 | 29.88±27.49 | 13.81±12.29 | 11.99±11.03 | 22.84±20.78 | 8.78±7.99 | 6.39±5.62 | 6.33±5.51 | 9.03±8.13 | 6.09±5.42 | 6.89±5.86 | 5.81±5.00 | 7.79±7.32 | 5.79±4.92 | 12.51±11.01 |
| TC | 8.60±7.57 | 54.07±51.37 | 34.56±29.38 | 21.30±18.74 | 18.02±15.32 | 28.64±24.63 | 5.31±4.51 | 5.16±4.75 | 7.01±6.66 | 11.66±10.61 | 8.54±7.77 | 7.36±6.40 | 5.94±5.52 | 8.12±6.90 | 5.28±4.54 | 15.30±13.16 | |
| ET | 14.13±12.15 | 16.06±14.45 | 19.07±17.54 | 16.44±14.96 | 7.56±6.58 | 8.89±8.45 | 13.65±12.97 | 7.11±6.75 | 10.33±9.81 | 10.39±9.45 | 8.21±7.64 | 9.06±8.52 | 6.43±5.72 | 3.92±3.41 | 4.16±3.79 | 10.36±9.84 | |
| mmFormer | WT | 9.34±16.84 | 11.36±19.81 | 11.39±21.24 | 10.61±17.34 | 8.93±6.89 | 8.93±11.16 | 7.81±11.50 | 7.27±7.35 | 8.21±8.35 | 8.09±8.16 | 7.26±6.64 | 7.44±5.99 | 8.18±5.36 | 7.62±5.39 | 7.41±5.45 | 8.66±10.40 |
| TC | 9.75±8.97 | 8.97±8.43 | 12.71±11.82 | 13.56±11.93 | 4.71±4.00 | 5.69±5.35 | 8.76±7.80 | 8.39±7.13 | 7.61±7.15 | 6.59±6.26 | 5.27±4.48 | 8.09±7.28 | 4.96±4.22 | 4.51±4.10 | 4.73±4.40 | 7.62±6.93 | |
| ET | 9.23±8.21 | 6.29±5.47 | 9.73±8.56 | 10.17±8.75 | 5.92±5.15 | 5.24±4.51 | 9.18±7.80 | 7.93±7.37 | 8.38±7.63 | 6.83±6.22 | 6.16±5.73 | 8.57±7.63 | 6.78±6.03 | 4.99±4.69 | 6.51±5.53 | 7.46±6.56 | |
| M2FTrans | WT | 8.76±7.80 | 12.20±10.98 | 13.56±11.93 | 10.79±9.82 | 6.61±5.62 | 9.72±8.65 | 6.42±5.59 | 6.46±5.75 | 5.98±5.14 | 6.01±5.47 | 5.59±5.09 | 5.25±4.78 | 5.59±5.09 | 6.28±5.84 | 5.13±4.36 | 7.62±6.48 |
| TC | 9.05±8.24 | 9.72±8.94 | 13.32±11.59 | 12.44±10.95 | 5.88±5.12 | 6.25±5.81 | 8.48±7.55 | 8.04±7.56 | 7.47±6.65 | 6.49±5.71 | 5.91±5.26 | 7.42±6.90 | 6.36±5.85 | 5.43±5.16 | 5.49±5.00 | 7.85±7.14 | |
| ET | 9.25±8.42 | 5.37±4.94 | 11.01±10.24 | 12.84±11.56 | 3.95±3.63 | 4.11±3.82 | 9.67±8.90 | 8.98±8.53 | 7.48±7.03 | 5.62±5.00 | 4.39±4.08 | 8.78±7.64 | 4.43±3.99 | 3.53±3.21 | 3.56±3.28 | 6.86±6.52 | |
| IMS2Trans | WT | 8.55±7.44 | 12.26±10.30 | 11.83±10.17 | 11.51±10.36 | 6.12±5.14 | 11.71±9.60 | 5.04±4.54 | 6.17±5.00 | 5.88±4.94 | 5.39±4.85 | 4.11±3.62 | 4.88±4.29 | 4.24±3.56 | 4.83±3.91 | 4.02±3.50 | 7.10±6.18 |
| TC | 10.29±9.16 | 9.05±7.78 | 9.88±8.60 | 11.21±9.98 | 9.78±8.02 | 6.32±5.18 | 8.83±7.68 | 8.67±7.11 | 9.07±7.44 | 8.99±8.09 | 6.64±5.51 | 8.38±7.46 | 9.61±8.26 | 7.92±6.65 | 8.92±7.40 | 8.90±7.83 | |
| ET | 8.56±7.28 | 6.32±5.37 | 10.75±9.25 | 11.68±10.40 | 3.86±3.13 | 4.37±3.67 | 9.99±8.59 | 7.97±7.01 | 6.88±5.92 | 5.18±4.35 | 4.21±3.49 | 7.84±6.27 | 4.43±3.94 | 3.32±2.72 | 3.94±3.15 | 6.62±5.36 | |
| MIFPN | WT | 6.90±5.66 | 10.94±9.41 | 12.02±9.62 | 7.77±6.22 | 5.95±5.24 | 9.15±7.69 | 6.19±5.26 | 6.39±5.75 | 5.61±4.66 | 6.44±5.60 | 5.51±4.63 | 4.63±3.70 | 4.99±4.14 | 5.62±4.66 | 4.53±3.76 | 6.84±6.02 |
| TC | 8.21±7.39 | 11.86±10.56 | 10.81±9.40 | 10.97±9.21 | 5.53±4.42 | 10.12±8.10 | 4.67±3.83 | 5.74±4.71 | 5.09±4.38 | 4.26±3.54 | 3.61±3.18 | 4.37±3.76 | 3.18±2.77 | 3.76±3.35 | 3.96±3.33 | 6.41±5.26 | |
| ET | 7.16±6.16 | 10.59±9.21 | 9.16±7.51 | 9.28±7.70 | 4.25±3.61 | 6.81±5.86 | 3.74±3.14 | 5.49±4.78 | 4.34±3.60 | 3.79±3.07 | 2.61±2.24 | 3.49±2.83 | 2.48±2.18 | 3.16±2.62 | 3.83±3.41 | 5.35±4.60 | |
| Ours | WT | 5.38±4.36 | 9.64±8.10 | 9.23±8.31 | 10.54±9.49 | 4.29±3.65 | 8.72±6.98 | 5.01±4.01 | 4.78±3.87 | 5.05±4.44 | 4.89±4.35 | 4.09±3.27 | 4.59±3.86 | 4.04±3.64 | 4.35±3.83 | 2.24±1.90 | 5.79±4.98 |
| TC | 7.58±6.75 | 7.43±6.02 | 9.21±7.64 | 10.14±8.42 | 4.24±3.48 | 5.74±4.59 | 4.35±3.74 | 4.52±3.80 | 5.03±4.17 | 3.96±3.45 | 3.31±2.71 | 4.19±3.52 | 3.02±2.45 | 3.36±2.92 | 2.08±1.85 | 5.21±4.64 | |
| ET | 6.51±5.66 | 4.26±3.45 | 9.44±7.93 | 9.94±8.65 | 3.25±2.67 | 4.36±3.71 | 3.32±2.66 | 4.11±3.53 | 4.15±3.40 | 3.15±2.61 | 2.43±2.11 | 3.36±2.69 | 2.39±2.08 | 2.88±2.48 | 1.91±1.62 | 4.36±3.53 |
Data are presented as mean ± standard deviation. BraTS, Brain Tumor Segmentation Challenge; ET, enhancing tumor; FLAIR, fluid-attenuated inversion recovery; HD95, 95% Hausdorff distance; T1, T1-weighted; T1c, contrast-enhanced T1; T2, T2-weighted; TC, tumor core; WT, whole tumor.
Comparison on BraTS2020
To further analyze the impact of the number of absent modalities on our model’s segmentation performance, we conducted comparative experiments against state-of-the-art methods using the BraTS2020 dataset.
Similar observations regarding performance trends are found on the BraTS2020 dataset, as summarized in Table 4. RobustSeg recorded the lowest mean DSC scores across all fifteen multi-modal combinations for WT and TC, at 82.72 and 70.38, respectively. For ET, RobustSeg’s mean DSC score was slightly higher (0.47) than RFNet. Among the comparison methods, MIFPN achieved the highest mean DSC scores for WT, TC, and ET, outperforming RobustSeg with mean increases of 5.28, 9.77, and 17.57, respectively. Our proposed method, however, attained the best overall segmentation performance for WT, TC, and ET, surpassing MIFPN with mean DSC score increases of 1.32, 2.57, and 1.86, respectively. P<0.05 of Wilcoxon matched-pairs signed-ranks two-tailed test for the mean DSC scores for WT, TC, and ET confirmed that our method statistically significantly outperformed previous methods.
Table 4
| Method | Class | FLAIR | T1 | T1c | T2 | T1c + T2 | T1 + T1c | FLAIR + T1 | T1 + T2 | FLAIR + T2 | FLAIR + T1c | FLAIR + T1 + T1c | FLAIR + T1 + T2 | FLAIR + T1c + T2 | T1 + T1c + T2 | FLAIR + T1 + T1c + T2 | Mean |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| RobustSeg | WT | 80.36±10.80 | 68.18±17.50 | 63.15±21.74 | 80.52±11.88 | 86.22±8.82 | 72.77±16.34 | 85.97±8.56 | 85.48±8.71 | 87.58±6.96 | 87.54±7.35 | 88.51±7.24 | 87.13±7.21 | 90.03±5.78 | 87.01±7.66 | 90.41±5.85 | 82.72±11.23 |
| TC | 57.72±22.41 | 44.04±26.86 | 51.15±26.87 | 53.24±21.51 | 82.21±8.36 | 79.29±10.56 | 65.48±16.22 | 65.38±18.35 | 71.14±14.14 | 81.61±9.75 | 83.81±8.58 | 69.09±14.22 | 83.29±7.69 | 83.88±7.58 | 84.41±7.64 | 70.38±13.33 | |
| ET | 34.09±30.98 | 26.59±33.77 | 21.01±36.34 | 30.82±28.36 | 72.68±10.93 | 71.09±13.88 | 36.72±29.74 | 37.78±26.13 | 50.54±22.26 | 72.21±12.23 | 73.47±10.61 | 41.64±26.85 | 72.91±12.73 | 74.34±10.26 | 73.92±11.74 | 52.65±18.94 | |
| RFNet | WT | 83.95±9.15 | 67.77±17.73 | 69.97±18.02 | 81.24±11.07 | 85.39±8.18 | 73.95±16.67 | 86.21±8.96 | 86.64±7.62 | 87.48±7.64 | 85.71±8.72 | 87.41±7.05 | 88.81±7.05 | 88.63±6.25 | 86.07±8.08 | 88.86±6.80 | 83.21±9.91 |
| TC | 65.21±16.35 | 55.41±20.51 | 59.59±22.23 | 60.23±19.89 | 78.86±10.15 | 76.16±13.11 | 69.25±13.84 | 69.49±16.78 | 72.81±14.41 | 78.62±11.33 | 81.31±8.97 | 71.65±13.89 | 81.44±8.91 | 81.16±9.99 | 81.95±8.84 | 72.21±13.62 | |
| ET | 38.77±30.62 | 35.46±28.40 | 28.23±33.73 | 30.11±28.65 | 68.98±13.03 | 68.18±13.68 | 34.55±29.45 | 40.74±29.04 | 50.14±20.44 | 67.96±13.46 | 69.15±13.88 | 41.24±27.62 | 69.52±14.94 | 69.76±14.82 | 69.92±14.44 | 52.18±22.48 | |
| mmFormer | WT | 84.67±10.80 | 74.67±17.50 | 74.26±21.74 | 82.89±11.88 | 86.35±8.82 | 78.66±16.34 | 80.71±8.56 | 86.14±8.71 | 89.26±6.96 | 89.16±7.35 | 89.44±7.24 | 89.75±7.21 | 90.06±5.78 | 86.84±7.66 | 90.18±5.85 | 84.87±11.23 |
| TC | 60.53±20.92 | 50.88±26.03 | 60.75±17.66 | 60.37±19.42 | 83.79±8.92 | 82.93±7.68 | 82.05±8.98 | 70.85±16.03 | 75.22±12.39 | 83.21±7.89 | 83.89±8.06 | 73.72±12.09 | 83.79±7.29 | 84.26±7.24 | 84.31±7.85 | 74.70±12.40 | |
| ET | 39.98±24.01 | 38.53±28.89 | 31.01±33.12 | 34.79±26.74 | 70.46±11.82 | 69.93±13.83 | 49.44±24.77 | 41.32±28.17 | 52.15±22.49 | 69.44±15.28 | 71.24±12.37 | 43.82±27.53 | 70.79±13.44 | 70.76±11.70 | 70.57±12.07 | 54.95±20.27 | |
| M2FTrans | WT | 84.92±8.75 | 73.11±17.21 | 74.69±15.69 | 87.92±7.25 | 87.06±7.76 | 79.43±11.31 | 89.74±6.57 | 86.59±7.91 | 89.94±5.53 | 89.92±6.25 | 90.21±5.78 | 90.31±5.62 | 90.67±5.97 | 87.46±7.90 | 90.71±5.11 | 86.18±7.88 |
| TC | 60.59±19.31 | 59.98±21.61 | 64.41±16.02 | 67.19±16.41 | 84.34±7.52 | 80.11±9.95 | 73.01±14.30 | 72.31±12.46 | 71.16±15.29 | 84.48±8.23 | 84.84±7.58 | 68.35±17.41 | 84.89±8.16 | 84.78±8.22 | 85.14±7.73 | 75.04±11.73 | |
| ET | 42.95±28.53 | 41.84±27.34 | 43.65±28.18 | 42.44±27.63 | 70.75±14.04 | 70.47±12.40 | 50.51±23.26 | 45.09±23.06 | 53.62±19.02 | 70.93±13.66 | 71.99±11.76 | 55.58±19.10 | 71.19±13.54 | 71.77±11.86 | 71.99±11.76 | 58.32±16.67 | |
| IMS2Trans | WT | 85.62±7.91 | 73.04±16.72 | 68.59±18.22 | 85.51±8.11 | 89.46±6.32 | 76.25±13.06 | 90.41±6.14 | 89.11±6.64 | 91.17±5.56 | 91.19±5.20 | 91.88±5.12 | 92.41±4.55 | 92.78±4.12 | 89.91±5.95 | 93.11±4.34 | 86.70±8.65 |
| TC | 61.86±20.60 | 72.68±12.29 | 54.79±21.70 | 58.38±19.15 | 84.62±7.54 | 81.91±9.95 | 70.39±14.51 | 68.74±15.32 | 73.91±12.52 | 84.24±8.04 | 85.61±7.20 | 72.11±15.06 | 85.38±7.75 | 85.89±7.34 | 86.27±6.45 | 75.12±11.20 | |
| ET | 52.44±21.40 | 51.19±20.50 | 47.76±22.99 | 45.29±26.81 | 75.72±9.95 | 73.91±13.05 | 60.32±19.84 | 56.13±20.62 | 61.74±18.75 | 75.18±11.42 | 75.78±11.14 | 69.07±14.23 | 75.56±10.51 | 76.51±9.63 | 76.05±11.26 | 64.84±17.23 | |
| MIFPN | WT | 86.43±7.60 | 80.02±12.79 | 81.02±11.58 | 87.99±6.97 | 88.17±7.10 | 82.71±10.72 | 90.64±5.34 | 88.27±6.57 | 90.61±5.63 | 90.55±5.86 | 91.07±4.91 | 91.09±5.26 | 91.18±5.64 | 88.94±6.97 | 91.42±4.89 | 88.01±6.95 |
| TC | 73.45±12.48 | 70.23±13.99 | 74.44±14.06 | 73.95±12.76 | 85.34±7.18 | 81.13±9.81 | 78.56±11.36 | 77.06±10.55 | 79.64±9.16 | 85.59±7.93 | 86.14±7.21 | 78.91±10.33 | 85.77±7.26 | 85.91±6.62 | 86.17±6.64 | 80.15±10.32 | |
| ET | 54.56±19.54 | 51.22±24.39 | 53.58±23.21 | 54.89±22.10 | 79.95±8.62 | 78.51±10.32 | 67.41±13.69 | 56.96±21.52 | 65.06±15.02 | 79.51±8.40 | 80.85±9.00 | 78.76±9.77 | 84.22±7.73 | 84.53±6.19 | 83.32±8.01 | 70.22±13.40 | |
| Ours | WT | 89.01±5.28 | 80.23±9.49 | 81.18±8.85 | 88.81±5.60 | 90.21±5.09 | 83.07±8.97 | 91.65±4.59 | 89.91±4.94 | 91.49±4.51 | 91.76±3.79 | 92.87±3.57 | 92.83±3.59 | 93.22±3.53 | 90.12±5.04 | 93.43±3.35 | 89.32±4.91 |
| TC | 83.91±7.72 | 76.65±10.27 | 78.06±9.21 | 81.61±8.28 | 88.32±5.84 | 74.49±12.76 | 77.22±10.93 | 77.52±8.99 | 80.15±8.34 | 88.71±4.52 | 89.67±4.86 | 78.24±9.36 | 88.26±5.64 | 89.18±5.30 | 88.95±4.42 | 82.73±7.43 | |
| ET | 56.95±15.50 | 53.91±16.59 | 54.03±15.63 | 55.31±13.41 | 85.06±5.98 | 79.49±6.97 | 73.21±8.57 | 65.15±11.85 | 67.34±11.43 | 85.81±5.53 | 86.19±4.70 | 59.76±14.49 | 85.84±5.52 | 87.06±4.14 | 86.11±4.58 | 72.08±9.21 |
Data are presented as mean ± standard deviation. BraTS, Brain Tumor Segmentation Challenge; DSC, dice similarity coefficient; ET, enhancing tumor; FLAIR, fluid-attenuated inversion recovery; T1, T1-weighted; T1c, contrast-enhanced T1; T2, T2-weighted; TC, tumor core; WT, whole tumor.
In the three missing-modality patterns on the BraTS2020, the proposed method improved the mean DSC scores for WT, TC, and ET of 0.94, 7.04, and 1.48 over MIFPN and the mean HD95 for WT, TC, and ET of 0.49, 0.97, and 0.06 over MIFPN. In the two missing-modality patterns on the BraTS2020, the proposed method improved the mean DSC scores for WT, TC, and ET of 1.19, 0.15, and 4.7 over MIFPN and the mean HD95 for WT, TC, and ET of 0.67, 0.97, and 0.55 over MIFPN. In the one missing-modality patterns on the BraTS2020, the proposed method improved the mean DSC scores for WT, TC, and ET of 1.69, 2.15, and 2.37 over MIFPN and the mean HD95 for WT, TC, and ET of 0.68, 0.41, and 1.21 over MIFPN. These results showed that less missing-modality patterns and larger gains in the mean DSC scores of the BraTS2020. For the pattern of missing only T1c on the BraTS2020 dataset, the proposed method surpassed MIFPN with mean DSC score increases of 1.74 for WT. For the pattern of missing both T1 and T1c on the BraTS2020 dataset, the proposed method surpassed MIFPN with mean DSC score increases of 0.88, 0.51, and 2.28 for WT, TC, and ET. These results also showed that less missing-modality patterns and larger gains in the mean DSC scores of the BraTS2020.
Taking into account the issue of multiple testing, we applied the Holm-Bonferroni correction, with the corrected significance threshold set at 0.05. Tables 1-5 demonstrated that our method showed statistically significant performance improvements on two datasets, with most P values being below 0.05. This confirmed the stability of our approach under various modality settings.
Table 5
| Method | Class | FLAIR | T1 | T1c | T2 | T1c + T2 | T1 + T1c | FLAIR + T1 | T1 + T2 | FLAIR + T2 | FLAIR + T1c | FLAIR + T1 + T1c | FLAIR + T1 + T2 | FLAIR + T1c + T2 | T1 + T1c + T2 | FLAIR + T1 + T1c + T2 | Mean |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| RobustSeg | WT | 19.50±17.75 | 28.29±26.31 | 25.29±23.27 | 19.53±18.94 | 10.05±9.85 | 18.72±17.41 | 13.64±12.69 | 9.95±9.85 | 12.04±11.20 | 11.97±10.77 | 10.84±10.73 | 8.96±8.51 | 7.70±7.01 | 7.02±6.95 | 6.66±6.53 | 14.01±12.75 |
| TC | 21.87±20.12 | 34.63±31.86 | 26.74±25.94 | 20.18±18.16 | 9.17±8.71 | 20.75±20.54 | 13.72±12.90 | 11.95±10.99 | 12.64±11.63 | 10.59±10.17 | 7.66±7.35 | 9.13±8.67 | 7.88±7.80 | 6.17±5.98 | 5.62±5.34 | 14.58±13.85 | |
| ET | 15.67±15.20 | 27.13±25.50 | 20.14±19.94 | 15.91±15.59 | 8.09±7.69 | 17.55±17.02 | 12.18±11.69 | 9.65±8.78 | 9.67±9.57 | 8.85±8.58 | 6.63±6.36 | 8.83±8.48 | 7.39±7.09 | 5.17±5.12 | 5.75±5.58 | 11.91±10.96 | |
| RFNet | WT | 19.32±17.00 | 43.42±38.64 | 19.79±17.22 | 15.55±13.22 | 13.62±11.85 | 23.82±22.63 | 7.29±6.71 | 8.35±7.85 | 6.24±5.87 | 12.48±11.61 | 7.31±6.87 | 8.09±7.04 | 5.91±5.44 | 7.69±6.84 | 6.48±5.70 | 13.69±12.73 |
| TC | 16.86±14.67 | 59.72±51.36 | 16.34±14.22 | 22.68±19.96 | 21.79±20.26 | 24.13±22.20 | 6.49±5.71 | 6.71±5.70 | 7.74±7.12 | 17.06±15.01 | 6.67±6.14 | 7.31±6.65 | 5.82±5.47 | 7.35±6.76 | 5.03±4.48 | 15.45±13.44 | |
| ET | 10.32±9.60 | 17.16±15.44 | 19.83±17.45 | 17.76±16.69 | 7.15±6.22 | 8.37±7.11 | 12.46±10.72 | 6.38±5.55 | 9.56±9.08 | 12.03±10.35 | 8.08±6.95 | 7.87±7.16 | 5.52±5.24 | 5.74±5.40 | 4.43±4.03 | 10.18±9.57 | |
| mmFormer | WT | 9.75±17.75 | 17.98±26.31 | 16.25±23.27 | 14.17±18.94 | 7.17±9.85 | 13.98±17.41 | 6.61±12.69 | 7.26±9.85 | 6.08±11.20 | 7.39±10.77 | 5.86±10.73 | 6.81±8.51 | 5.37±7.01 | 5.47±6.95 | 4.68±6.53 | 8.99±12.75 |
| TC | 9.75±8.97 | 8.87±7.81 | 12.81±11.02 | 14.08±11.97 | 5.51±5.01 | 10.44±9.50 | 8.94±8.49 | 7.92±6.81 | 8.32±7.82 | 7.58±6.90 | 6.02±5.12 | 8.07±7.51 | 5.74±5.45 | 4.43±4.21 | 5.48±4.93 | 8.26±7.52 | |
| ET | 8.75±7.61 | 7.44±6.40 | 10.87±10.33 | 11.90±10.47 | 4.11±3.86 | 8.92±8.21 | 9.48±8.53 | 8.04±7.40 | 7.25±6.16 | 6.27±5.89 | 5.26±4.52 | 8.13±7.15 | 5.12±4.71 | 3.43±2.92 | 4.34±4.12 | 7.29±6.63 | |
| M2FTrans | WT | 9.27±8.34 | 16.67±15.67 | 15.13±13.62 | 13.77±12.81 | 6.53±5.75 | 11.24±10.57 | 6.44±6.12 | 6.24±5.87 | 6.17±5.37 | 6.11±5.68 | 5.43±5.00 | 5.19±4.83 | 6.13±5.27 | 5.86±5.27 | 5.03±4.73 | 8.35±7.60 |
| TC | 9.34±8.50 | 9.11±7.83 | 11.42±10.62 | 11.87±10.33 | 5.72±4.98 | 9.54±8.20 | 8.25±7.43 | 7.86±6.84 | 7.75±7.21 | 6.35±5.65 | 5.93±5.16 | 7.48±6.96 | 6.73±6.06 | 5.16±4.70 | 5.54±4.71 | 7.87±7.16 | |
| ET | 8.12±7.39 | 6.03±5.13 | 9.97±9.37 | 10.76±10.22 | 4.63±4.03 | 8.85±8.32 | 9.33±8.86 | 7.86±6.92 | 7.01±6.10 | 5.21±4.43 | 4.57±3.98 | 7.84±6.82 | 4.69±4.31 | 4.95±4.55 | 4.12±3.54 | 6.93±5.96 | |
| IMS2Trans | WT | 8.16±6.53 | 17.05±14.49 | 17.33±14.38 | 11.55±9.93 | 5.58±4.69 | 11.74±10.33 | 5.89±4.77 | 5.34±4.59 | 6.11±4.89 | 5.69±4.89 | 5.27±4.58 | 4.85±3.88 | 5.91±4.79 | 5.22±4.44 | 4.45±3.83 | 8.01±6.41 |
| TC | 9.14±7.49 | 18.71±16.84 | 17.97±15.45 | 14.86±12.48 | 4.87±4.04 | 10.51±9.04 | 8.25±6.60 | 5.92±5.27 | 7.35±6.03 | 4.99±4.09 | 4.92±4.38 | 5.68±4.60 | 4.88±4.05 | 4.52±3.71 | 4.66±4.01 | 8.48±6.78 | |
| ET | 8.26±7.10 | 15.67±14.10 | 14.88±12.20 | 10.93±9.62 | 4.14±3.73 | 8.25±7.01 | 8.45±6.84 | 5.42±4.72 | 6.23±5.48 | 3.62±3.08 | 2.93±2.49 | 6.26±5.57 | 3.20±2.85 | 3.11±2.55 | 2.96±2.37 | 6.95±5.84 | |
| MIFPN | WT | 8.30±6.72 | 9.01±7.57 | 8.60±7.40 | 7.25±6.09 | 7.31±6.43 | 5.59±4.64 | 5.01±4.06 | 4.62±3.74 | 5.78±4.86 | 5.14±4.11 | 4.30±3.70 | 4.72±4.25 | 5.47±4.49 | 4.36±3.71 | 4.45±3.87 | 5.99±4.85 |
| TC | 7.63±6.41 | 8.32±6.66 | 7.72±6.95 | 7.46±6.49 | 7.49±6.74 | 5.36±4.50 | 5.08±4.47 | 6.43±5.47 | 5.54±4.43 | 5.21±4.64 | 3.17±2.60 | 5.58±4.85 | 6.28±5.46 | 4.59±3.99 | 5.46±4.75 | 6.09±5.36 | |
| ET | 6.31±5.55 | 6.97±5.99 | 6.01±4.93 | 6.53±5.81 | 3.59±3.09 | 2.36±1.91 | 5.68±5.11 | 5.51±4.79 | 3.37±3.03 | 3.84±3.15 | 3.03±2.64 | 5.24±4.61 | 3.70±3.00 | 3.63±3.12 | 3.20±2.56 | 4.60±4.00 | |
| Ours | WT | 4.87±3.90 | 9.83±7.96 | 9.65±8.40 | 6.87±6.18 | 4.02±3.30 | 7.13±6.27 | 4.39±3.91 | 4.28±3.51 | 5.12±4.15 | 4.47±3.75 | 3.43±2.98 | 4.31±3.88 | 4.95±4.41 | 3.42±2.87 | 3.81±3.16 | 5.37±4.73 |
| TC | 7.25±6.38 | 6.04±5.01 | 7.51±6.08 | 6.45±5.61 | 4.42±3.71 | 4.32±3.63 | 4.86±3.94 | 5.15±4.38 | 5.53±4.70 | 5.01±4.21 | 3.05±2.50 | 6.83±5.46 | 4.78±4.16 | 3.34±2.74 | 3.86±3.32 | 5.23±4.50 | |
| ET | 6.15±5.54 | 5.95±5.12 | 8.40±7.31 | 5.58±4.91 | 3.47±3.09 | 2.05±1.82 | 4.33±3.51 | 4.88±4.05 | 3.26±2.87 | 3.03±2.48 | 2.85±2.42 | 4.07±3.42 | 1.81±1.47 | 2.01±1.73 | 1.46±1.17 | 3.95±3.28 |
Data are presented as mean ± standard deviation. BraTS, Brain Tumor Segmentation Challenge; ET, enhancing tumor; FLAIR, fluid-attenuated inversion recovery; HD95, 95% Hausdorff distance; T1, T1-weighted; T1c, contrast-enhanced T1; T2, T2-weighted; TC, tumor core; WT, whole tumor.
Computational efficiency
We have provided a comparison of computational metrics such as inference time, training time, graphics processing unit (GPU) memory footprint, or number of parameters against previous methods on the BraTS2020 dataset. The details were shown in Table 6. RFNet had the smallest number of trainable parameters, but MIFPN has the largest number of trainable parameters. Compared to the best transformer-based method MIFPN, our achieved better performance with fewer model parameters and peak GPU memory, due to the inherent property of transformers involving more model parameters. In our implementation, selective SSM/Mamba scales linearly with sequence length. This aligned with ViM’s resource consumption with progressively increasing resolutions as reported in previous work (45). IMS2Trans was Swin Transformer network which consumes the most peak GPU memory. However, the Mamba-based design offered a superior accuracy with moderate training time and least inference time, which was critical for clinical deployment.
Table 6
| Method | DSC, % | HD95, voxels | Parameters, M | Peak GPU memory, Gb | Training time (per epoch), s | Inference time (per volume), s | |||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| WT | TC | ET | WT | TC | ET | ||||||
| RobustSeg | 82.72±11.23 | 70.38±13.33 | 52.65±18.94 | 14.01±12.75 | 14.58±13.85 | 11.91±10.96 | 37.58 | 5.19 | 83.46 | 9.37 | |
| RFNet | 83.21±9.91 | 72.21±13.62 | 52.18±22.48 | 13.69±12.73 | 15.45±13.44 | 10.18±9.57 | 8.98 | 15.89 | 111.53 | 11.74 | |
| mmFormer | 84.87±11.23 | 74.70±12.40 | 54.95±20.27 | 8.99±12.75 | 8.26±7.52 | 7.29±6.63 | 35.99 | 6.83 | 100.83 | 23.53 | |
| M2FTrans | 86.18±7.88 | 75.04±11.73 | 58.32±16.67 | 8.35±7.60 | 7.87±7.16 | 6.93±5.96 | 45.42 | 13.53 | 124.74 | 37.36 | |
| IMS2Trans | 86.70±8.65 | 75.12±11.20 | 64.84±17.23 | 8.01±6.41 | 8.48±6.78 | 6.95±5.84 | 4.86 | 78.45 | 189.08 | 48.38 | |
| MIFPN | 88.01±6.95 | 80.15±10.32 | 70.22±13.40 | 5.99±4.85 | 6.09±5.36 | 4.60±4.00 | 107.72 | 21.86 | 109.24 | 6.52 | |
| Ours | 89.32±4.91 | 82.73±7.43 | 72.08±9.21 | 5.37±4.73 | 5.23±4.50 | 3.95±3.28 | 60.67 | 13.08 | 101.64 | 4.38 | |
Data are presented as mean ± standard deviation. BraTS, Brain Tumor Segmentation Challenge; DSC, dice similarity coefficient; ET, enhancing tumor; GPU, graphics processing unit; HD95, 95% Hausdorff distance; M, million; TC, tumor core; WT, whole tumor.
Visualization
As illustrated in Figure 2, we tested our method to visually demonstrate its performance in scenarios with absent modalities, showcasing segmentation results across 15 configurations. Our results were optimal when all modalities were present and remained competitive even with some modalities absent. The visualizations were presented in axial, sagittal, and coronal views, as shown in Figure 2, where it clearly demonstrated that our method outperformed MIFPN in the majority of the 15 combinations with a superior segmentation performance.
Ablation study
In this study, we have proposed a method consisting of four components: Mamba block, MF block, CFM block, and CU. In this section, we analyzed the effects of each component on the BraTS2020 dataset using the mean DSC score as the evaluation metric to demonstrate their effectiveness, as in Table 7. 3D-CNN with convolution fusion of the input zeroed and concatenated feature was used as the baseline.
Table 7
| Components (Mamba/MF/LM/CFM/CU) | DSC, % | |||
|---|---|---|---|---|
| ET | TC | WT | Mean | |
| ×/×/×/×/× | 60.34±19.46 | 76.56±13.29 | 85.24±9.52 | 74.05±38.42 |
| √/×/×/×/× | 63.57±18.55 | 78.15±11.64 | 86.38±8.51 | 76.03±35.16 |
| √/√/×/×/× | 65.06±17.71 | 79.31±10.36 | 86.84±7.64 | 77.07±34.03 |
| √/√/√/×/× | 67.84±16.82 | 80.63±9.51 | 87.55±7.15 | 78.67±33.62 |
| √/√/√/√/× | 70.71±14.56 | 81.45±8.74 | 88.07±5.81 | 80.08±30.68 |
| √/√/√/×/√ | 69.23±14.24 | 81.08±9.03 | 87.93±7.86 | 79.41±31.64 |
| √/√/√/√/√ | 72.08±9.21 | 82.73±7.43 | 89.32±4.91 | 81.38±28.68 |
Data are presented as mean ± standard deviation. BraTS, Brain Tumor Segmentation Challenge; CFM, cross-level fusion Mamba; CU, cross-level uncertainty; DSC, dice similarity coefficient; ET, enhancing tumor; LM, Mahalanobis distance in Eq. [18]; MF, Mamba fusion; TC, tumor core; WT, whole tumor.
Mamba block
To assess the contribution of the proposed Mamba block for feature extraction, we conducted an ablation study. As shown in Table 7, our baseline model achieved a mean Dice score of 74.05 for brain tumor segmentation. When integrating the last encoder blocks with our Mamba block, the mean Dice score increased to 76.03. Gains on ET are larger than TC and WT because edema benefits from improved discrimination by Mamba. This improvement demonstrates the effectiveness of the Mamba block in capturing multi-modality dependencies crucial for accurate tumor segmentation.
To evaluate the effectiveness of our model architecture, we have tried single unidirectional scan along height, depth, depth, width and inversed width, width and height, width, height and depth on the BraTS2020 dataset in Table 8. Contrary to expectations, the single unidirectional scan along width did not exhibit a notable decline in performance on the BraTS2020 dataset and multi-axis or bidirectional scans did not yield significant improvements. In fact, its slower convergence rate resulted in inferior performance metrics compared to the baseline under identical training settings.
Table 8
| Scan axis | DSC, % | HD95, voxels | Parameters, M | Peak GPU memory, Gb | Training time (per epoch), s | Inference time (per volume), s | |||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| WT | TC | ET | WT | TC | ET | ||||||
| Width | 89.32±4.91 | 82.73±7.43 | 72.08±9.21 | 5.37±4.73 | 5.23±4.50 | 3.95±3.28 | 60.67 | 13.08 | 101.64 | 4.38 | |
| Height | 89.41±4.86 | 82.82±7.39 | 72.17±9.25 | 5.30±4.70 | 5.16±4.52 | 3.88±3.30 | 60.67 | 13.08 | 101.34 | 4.45 | |
| Depth | 89.25±4.95 | 82.65±7.48 | 72.01±9.16 | 5.45±4.76 | 5.31±4.48 | 4.03±3.25 | 60.67 | 13.08 | 102.67 | 4.78 | |
| Width and inversed width | 89.58±4.82 | 82.99±7.49 | 72.26±9.07 | 5.26±4.70 | 5.12±4.53 | 3.84±3.30 | 61.03 | 13.95 | 107.01 | 5.38 | |
| Width and height | 89.62±4.79 | 83.01±7.29 | 72.33±9.32 | 5.23±4.75 | 5.09±4.48 | 3.81±3.26 | 61.47 | 14.81 | 222.74 | 17.81 | |
| Width, height, and depth | 89.65±4.87 | 83.08±7.51 | 72.39±9.15 | 5.21±4.69 | 5.07±4.52 | 3.79±3.31 | 61.93 | 15.55 | 458.19 | 33.63 | |
Data are presented as mean ± standard deviation. BraTS, Brain Tumor Segmentation Challenge; DSC, dice similarity coefficient; ET, enhancing tumor; GPU, graphics processing unit; HD95, 95% Hausdorff distance; M, million; TC, tumor core; WT, whole tumor.
MF block
The efficacy of the MF block was rigorously evaluated, focusing exclusively on its contribution to feature integration. Our findings confirmed that the MF block facilitates the effective fusion and interaction of potentially incomplete modal features.
To clearly show the contribution of this specific loss term to the model’s ability to align features from incomplete and complete modalities, we have included an ablation within Table 7 that specifically evaluated the contribution of the LM loss term by comparing the MF module with and without it, to demonstrate its necessity for robust performance with missing modalities. Quantitative analysis, presented in Table 7, revealed improvements in segmentation performance: the mean DSC scores for the ET, TC, and WT regions increased by 2.87, 0.82, and 0.52, respectively, when incorporating the LM loss term over the baseline + Mamba block + MF block without LM. The mean DSC scores for the ET, TC, and WT regions increased by 4.27, 2.48, and 1.17, respectively, when incorporating the MF block over the baseline + Mamba block + LM. This culminated in an average increase of 2.64 in the mean DSC across all three regions. These results strongly suggested that the MF block offered a superior mechanism for efficiently learning and consolidating multi-modal features during the fusion and interaction phases, particularly when contrasted with more simplistic fusion methodologies.
CFM block
We evaluated the effectiveness of the CFM block, a critical component designed to enhance multi-modal medical image segmentation. This module harnessed the global modeling strengths of SSM to ensure efficient information propagation across all voxels. This dual benefit allowed the CFM block to not only improve the representation quality of individual modalities but also to facilitate the contextual inference of absent information and the robust fusion of diverse multi-modal features. Quantitative results, presented in Table 7, demonstrated the CFM block’s individual contribution: mean DSC scores surpassed the Baseline by 2.87, 0.82, and 0.52 in the ET, TC, and WT regions, respectively, yielding an average regional improvement of 1.41.
More strikingly, the synergistic combination of the Mamba block, the MF block, and the CFM block led to even greater performance enhancements. Table 5 further showed that this combined approach increased mean DSC scores by 10.37, 4.89, and 2.83 for ET, TC, and WT, respectively, culminating in an overall average increase of 6.03 across the three regions. This compelling evidence highlighted how the incorporation of SSM modules enables our method to effectively synthesize relatively complete modal features, even in scenarios with absent input modalities.
CU
Quantitative results in Table 7 revealed that the integration of CU led to respective increases of 1.39, 0.45, and 0.38 in segmentation DSC for the ET, TC, and WT regions when compared against the ‘Mamba + MF’. These gains validated the CU strategy’s contribution to integrating multi-level features and achieving more comprehensive and robust modal representations.
To more intuitively demonstrate the effectiveness of our CU strategy, we have visualized the uncertainty maps, as shown in Figure 3. As shown in Figure 3C, the uncertainty map can highlight non-ET, and its boundary is also consistent with the contour of non-ET, indicating that our approach can focus on this region. As shown in Figure 3D,3E, the uncertainty maps also can highlight peritumoral edema and ET.
Hyperparameter selection
Deep learning models learn complex, often correlated, features from different modalities. For example, a feature indicating “contrast enhancement” from T1c might be correlated with a feature indicating “edema” from FLAIR in tumor regions. Euclidean distance treats all dimensions equally and independently. Two feature vectors might look “far” in Euclidean space due to high variance in a single feature, but Mahalanobis distance would correctly show them as “close” if their variation is consistent with the learned correlations within a class. This leads to more precise boundary delineation. Mahalanobis distance, by incorporating the inverse covariance matrix, inherently decorrelates the features and scales them according to their variance. This means it can accurately measure the true statistical distance in a high-dimensional fused feature space, even if individual features are highly correlated. This is crucial because it allows the model to better align fused feature of incomplete modalities to that of complete modalities.
The Mahalanobis distance weight (denoted as λ) is a critical hyperparameter in our framework, as it balances the loss term measuring the difference between fused features of available modalities and full modalities. To evaluate its impact on model performance, we conducted a systematic sensitivity analysis on the BraTS2020 dataset. We tested λ across a reasonable interval: {2.0, 4.0, 8.0, 12.0, 14.0}. This range was chosen to cover scenarios where the Mahalanobis distance term has minimal (e.g., λ=2.0) to dominant (e.g., λ=12.0) influence on the total loss. All other hyperparameters (e.g., learning rate, network depth) were fixed across experiments. For each λ, we performed 5 independent repeated experiments to account for randomness in training initialization, reporting mean ± standard deviation of key metrics in Table 9. The low standard deviation (<0.5) across 5 repetitions confirmed the consistency of these trends, reinforcing that the optimal λ range was robust to random initialization. This structure highlighted the rigor of our repeated experiments, connected the hyperparameter’s impact to model behavior, and tied the findings back to the broader utility of our method.
Table 9
| λ | DSC, % | |||
|---|---|---|---|---|
| ET | TC | WT | Mean | |
| 2 | 71.39±0.32 | 81.70±0.28 | 87.19±0.37 | 79.90±0.25 |
| 4 | 72.40±0.35 | 82.86±0.31 | 89.34±0.33 | 81.00±0.29 |
| 8 | 72.08±0.29 | 82.73±0.36 | 89.32±0.30 | 81.38±0.34 |
| 12 | 72.68±0.33 | 82.49±0.27 | 88.22±0.38 | 80.89±0.31 |
| 14 | 70.03±0.39 | 82.18±0.26 | 87.71±0.35 | 79.54±0.22 |
Data are presented as mean ± standard deviation. DSC, dice similarity coefficient; ET, enhancing tumor; TC, tumor core; WT, whole tumor.
Discussion
Although the proposed method demonstrates superior performance compared to existing approaches on both the BraTS2018 and BraTS2020 datasets, it still suffers from the following two limitations: (I) the missingness patterns in the data of this study were artificially simulated and may not accurately represent clinically realistic, not-missing-at-random (NMAR) scenarios. In research on Mamba-based SSMs, scenarios involving missing data in incomplete multi-modal MRI are typically constructed through simulation methods (e.g., randomly removing a specific modality). The design logic of such simulations often assumes that the missing process is independent of the pathological information in MR images and the patient’s condition. However, missing data in clinical multi-modal MRI generally exhibits NMAR characteristics. Simulated missingness fails to replicate the complex relationship—where the cause of missingness is deeply tied to the clinical information in images. Consequently, when Mamba-based SSMs process real-world incomplete clinical data, they tend to suffer from biases in modal feature fusion (e.g., over-reliance on complete modalities and neglect of potential associated information from missing modalities). This ultimately leads to issues such as blurred boundary segmentation and missed detection of small lesions, limiting the model’s clinical generalization ability. As shown in Figure 4A,4B, multiple modalities without T1c where ET does not exist; however, as shown in Figure 4C, the proposed method incorrectly segments non-enhancing TC as ET and surrounding normal tissues incorrectly segments as edema. (II) The single-axis scanning mechanism of Mamba blocks processes data along one dimension, which can lead to anisotropic feature representations and reduced spatial coherence. When Mamba-based SSMs extract features from multi-modal MR images, Mamba blocks typically adopt a single-axis scanning strategy (e.g., processing 3D MR images through sequential slicing only along the z-axis, i.e., the body’s long axis). This scanning method directly induces anisotropy—an imbalance in resolution and information density across different spatial dimensions of the MR image: for example, under single-axis scanning, the slice interval along the z-axis may reach 2 mm (low resolution), while the pixel spacing in the x-/y-axis (axial plane) is only 0.5 mm (high resolution). The negative impact of anisotropy on image segmentation manifests in two key ways: first, critical anatomical details in low-resolution dimensions (e.g., z-axis) are blurred, preventing Mamba blocks from effectively capturing these weak features. Second, resolution differences across dimensions cause the Mamba-based SSM to overemphasize information from high-resolution dimensions during feature fusion, resulting in imbalanced feature weighting. This further reduces the spatial consistency of segmentation results (e.g., “discontinuities” in segmented lesion regions along the z-axis). As shown in Figure 4D,4E, multiple modalities without T1c where edema is large and small ETs scatter in the non-enhancing TC; however, as shown in Figure 4C,4F, the green arrows point to surrounding normal tissues, which are incorrectly segmented as edema by the proposed method; and the blue arrows point to small ETs, which are incorrectly segmented by the proposed method.
To address the aforementioned issues, future research plans to optimize the feature extraction process of Mamba blocks using a multi-axis scanning strategy: By scanning multi-modal MR images along the three spatial axes (x, y, and z) respectively, feature data with balanced resolution across all dimensions will be considered. This approach not only supplements critical anatomical details in low-resolution dimensions (a shortcoming of single-axis scanning) and mitigates the interference of anisotropy on feature extraction but also provides a reference for real-world clinical missing scenarios through the complete data acquired via multi-axis scanning. This will support the development of a missing data generation mechanism that aligns with NMAR patterns, thereby enhancing the clinical applicability of the model in incomplete multi-modal MR image segmentation.
The integration of deep learning into clinical radiology workflows holds significant promise for improving brain disease diagnostic accuracy and efficiency (46-48). Integrating deep learning algorithms for brain tumor segmentation into existing clinical workflows typically requires incorporating them seamlessly into the Picture Archiving and Communication Systems (PACS) (46,47,49). Future research plans to integrate the proposed method for efficient radiology operations of brain tumors. Crucially, this integration must also support the display of uncertainty maps alongside the final tumor segmentation. These maps highlight regions where the deep learning model is less confident in its prediction of the tumor boundary, providing the radiologist with vital information to identify potential segmentation errors or areas requiring further human scrutiny before finalizing the diagnostic report.
Data privacy is a significant concern when dealing with sensitive medical data such as brain MR images. Federated learning (FL) offers a promising approach to train models on decentralized data sources without directly sharing the data, thus preserving privacy (45,50,51). In a FL framework, local models are trained on individual client devices or institutions, and only the model updates are shared with a central server for aggregation (52). Future research plans to combine FL with privacy-enhancing technologies like homomorphic encryption to further protect the sensitive data of patients with brain tumors (53).
Conclusions
In this study, we have introduced an approach for interacting with incomplete multi-modal information with state-space models. During this process, the MF block is proposed to foster the learning of absent modality features, leading to a more comprehensive representation of multi-modal MR images for tumor segmentation, mitigating the challenges associated with feature incompleteness due to absent modalities and enhancing the model’s capability to navigate these complex situations. CFM blocks are designed to capture global features of the low-level feature through its contextual learning mechanism. A CU constraint is imposed on each class of the final predicted tumor. This strategy integrates modal feature via lesion uncertainty and employs a weight mechanism, guiding the model to learn more discriminative regional features. Extensive experiments on BraTS2018 and BraTS2020 show that our method exhibits superior performance under various incomplete multi-modal settings compared to existing state-of-the-art methods.
Acknowledgments
None.
Footnote
Funding: This work was supported by
Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://qims.amegroups.com/article/view/10.21037/qims-2025-1913/coif). The authors have no conflicts of interest to declare.
Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. This study used publicly available datasets sourced from BraTS. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments.
Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.
References
- Menze BH, Jakab A, Bauer S, Kalpathy-Cramer J, Farahani K, Kirby J, et al. The Multimodal Brain Tumor Image Segmentation Benchmark (BRATS). IEEE Trans Med Imaging 2015;34:1993-2024. [Crossref] [PubMed]
- Ghaffari M, Sowmya A, Oliver R. Automated Brain Tumor Segmentation Using Multimodal Brain Scans: A Survey Based on Models Submitted to the BraTS 2012-2018 Challenges. IEEE Rev Biomed Eng 2020;13:156-68. [Crossref] [PubMed]
- Aggarwal M, Tiwari AK, Sarathi MP. Comparative analysis of deep learning models on brain tumor segmentation datasets: BraTS 2015-2020 datasets. Rev Intell Artif 2022;36:863-71.
- Zunair H, Ben Hamza A. Sharp U-Net: Depthwise convolutional network for biomedical image segmentation. Comput Biol Med 2021;136:104699. [Crossref] [PubMed]
- Verma R, Kumar N, Patil A, Kurian NC, Rane S, Graham S, et al. MoNuSAC2020: A Multi-Organ Nuclei Segmentation and Classification Challenge. IEEE Trans Med Imaging 2021;40:3413-23. [Crossref] [PubMed]
- Zhang Y, He N, Yang J, Li Y, Wei D, Huang Y, Zhang Y, He Z, Zheng Y. mmFormer: Multimodal medical transformer for incomplete multimodal learning of brain tumor segmentation. In: International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer;2022:107-117.
- Shi J, Yu L, Cheng Q, Yang X, Cheng KT, Yan Z. Mftrans: Modality-masked fusion transformer for incomplete multi-modality brain tumor segmentation. IEEE J Biomed Health Inform 2023;28:379-90. [Crossref] [PubMed]
- Zhang D, Wang C, Chen T, Chen W, Shen Y. Scalable Swin Transformer network for brain tumor segmentation from incomplete MRI modalities. Artif Intell Med 2024;149:102788. [Crossref] [PubMed]
- Diao YQ, Fang HH, Yu HY, Li F, Xu YW. Multimodal invariant feature prompt network for brain tumor segmentation with missing modalities. Neurocomputing 2025;616:128847.
- Chen XL, Xie HR, Tao XH, Wang FL, Leng MM, Lei BY. Artificial intelligence and multimodal data fusion for smart healthcare: topic modeling and bibliometrics. Artif Intell Rev 2024;57:91.
- Khan AA, Mahendran RK, Perumal K, Faheem M. Dual-3DM3AD: Mixed Transformer Based Semantic Segmentation and Triplet Pre-Processing for Early Multi-Class Alzheimer’s Diagnosis. IEEE Trans Neural Syst Rehabil Eng 2024;32:696-707. [Crossref] [PubMed]
- Jader RF, Kareem S, Awla H. Ensemble deep learning technique for detecting MRI brain tumors. Appl Comput Intell Soft Comput 2024;2024:6615468.
- Islam N, Azam S, Islam S, Kanchan MH, Parvez AS, Islam M. An improved deep learning-based hybrid model with ensemble techniques for brain tumor detection from MR image. Informatics Med Unlocked 2024;47:101483.
- Wankhede DS, Selvarani R. Dynamic architecture-based deep learning approach for glioblastoma brain tumor survival prediction. Neurosci Inform 2022;2:100062.
- Hashmi A, Osman AH. Brain tumor classification using conditional segmentation with residual network and attention approach by extreme gradient boost. Appl Sci 2022;12:10791.
- Yogananda CGB, Wagner B, Nalawade SS, Murugesan GK, Pinho MC, Fei B, Madhuranthakam AJ, Maldjian JA. Fully automated brain tumor segmentation and survival prediction of gliomas using deep learning and MRI. In: International MICCAI Brainlesion Workshop. Cham: Springer; 2020:99-112.
- Khalighi S, Reddy K, Midya A, Pandav KB, Madabhushi A, Abedalthagafi M. Artificial intelligence in neuro-oncology: advances and challenges in brain tumor diagnosis, prognosis, and precision treatment. NPJ Precis Oncol 2024;8:80. [Crossref] [PubMed]
- Yeafi A, Islam M, Yusuf SU. A deep learning framework for 3D brain tumor segmentation and survival prediction. Healthc Anal 2025;8:100418.
- Fiaz K, Madni TM, Anwar F, Janjua UI, Rafi A, Abid MMN, Sultana N. Brain tumor segmentation and multiview multiscale-based radiomic model for patient’s overall survival prediction. Int J Imaging Sys Tech 2022;32:982-99.
- Pei L, Vidyaratne L, Rahman MM, Iftekharuddin KM. Context aware deep learning for brain tumor segmentation, subtype classification, and survival prediction using radiology images. Sci Rep 2020;10:19726. [Crossref] [PubMed]
- Xun SY, Zhang Y, Duan SX, Wang MW, Chen JG, Tong T, Gao QQ. LAM CT, Hu MH, Tan T. ARGA-UNet: Advanced U-Net segmentation model using residual grouped convolution and attention mechanism for brain tumor MR image segmentation. Virtual Reality Intell Hardware 2024;6:203-16.
- Yang T, Lu X, Yang L, Yang M, Chen J, Zhao H. Application of MRI image segmentation algorithm for brain tumors based on improved YOLO. Front Neurosci 2025;18:1510175. [Crossref] [PubMed]
- Xue RZ, Zhang ZF, Zhao YX, Zhang Q, Liang MH. MMEFU-Net: A Mamba-guided multi-encoder fusion U-Net for tumor segmentation in CT images. IEEE Access 2025;13:76257-70.
- Prathipati SC, Satpathy SK. Transforming 3D brain tumour image segmentation: an enhanced V-Net approach for precise diagnosis and treatment planning. In: 2024 International Conference on Advances in Computing, Communication and Applied Informatics (ACCAI). IEEE; 2024:1-6.
- Hatamizadeh A, Tang Y, Nath V, Yang D, Myronenko A, Landman B, Roth HR, Xu D. Unetr: Transformers for 3d medical image segmentation. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. IEEE; 2022:574-84.
- Almufareh MF, Imran M, Khan A, Humayun M, Asim M. Automated brain tumor segmentation and classification in MRI using YOLO-based deep learning. IEEE Access 2024;12:16189-207.
- Seshimo H, Rashed EA. Segmentation of Low-Grade Brain Tumors Using Mutual Attention Multimodal MRI. Sensors (Basel) 2024;24:7576. [Crossref] [PubMed]
- Zhao J, Xing Z, Chen Z, Wan L, Han T, Fu H, Zhu L. Uncertainty-Aware Multi-Dimensional Mutual Learning for Brain and Brain Tumor Segmentation. IEEE J Biomed Health Inform 2023;27:4362-72. [Crossref] [PubMed]
- Xue H, Yao Y, Teng Y. Multi-modal tumor segmentation methods based on deep learning: a narrative review. Quant Imaging Med Surg 2024;14:1122-40. [Crossref] [PubMed]
- Manjunatha B, Manohar G, Anand LV, Manikandan M, Kavitha B, David DB. Automated brain tumor segmentation with deep learning. Latin Am Appl Res Int J 2025;55:309-14.
- Thenmoezhi N, Perumal B, Lakshmi A. Multi-view image fusion using ensemble deep learning algorithm for MRI and CT images. ACM Trans Asian Low-Resour Lang Inf Process 2024;23:1-24.
- Lu JC, Huang YH, Zhang YW, Ding H. MSA-net: An efficient attention-aware 3D network for brain tumor segmentation in MRI. In: Proceedings of the 3rd International Conference on Computer, Artificial Intelligence and Control Engineering; 2024:131-6.
- Uppal D, Prakash S. Ms2adm-bts: Multi-scale dual attention-guided diffusion model for volumetric brain tumor segmentation. Pattern Recognit Lett 2025;198:115-22.
- Dorent R, Joutard S, Modat M, Ourselin S, Vercauteren T. Hetero-modal variational encoder-decoder for joint modality completion and segmentation. In: International Conference on Medical Image Computing and Computer-Assisted Intervention. Cham: Springer International Publishing; 2019:74-82.
- Sharma A, Hamarneh G. Missing MRI Pulse Sequence Synthesis Using Multi-Modal Generative Adversarial Network. IEEE Trans Med Imaging 2020;39:1170-83. [Crossref] [PubMed]
- Yang H, Sun J, Xu Z. Learning Unified Hyper-Network for Multi-Modal MR Image Synthesis and Tumor Segmentation With Missing Modalities. IEEE Trans Med Imaging 2023;42:3678-89. [Crossref] [PubMed]
- Meng X, Sun K, Xu J, He X, Shen D. Multi-Modal Modality-Masked Diffusion Network for Brain MRI Synthesis With Random Modality Missing. IEEE Trans Med Imaging 2024;43:2587-98. [Crossref] [PubMed]
- Zhang Y, Peng C, Wang Q, Song D, Li K, Kevin Zhou S. Unified Multi-Modal Image Synthesis for Missing Modality Imputation. IEEE Trans Med Imaging 2025;44:4-18. [Crossref] [PubMed]
- Havaei M, Guizard N, Chapados N, Bengio Y. Hemis: Hetero-modal image segmentation. In: International Conference on Medical Image Computing and Computer-Assisted Intervention. Cham: Springer International Publishing; 2016:469-77.
- Chen C, Dou Q, Jin Y, Chen H, Qin J, Heng PA. Robust multimodal brain tumor segmentation via feature disentanglement and gated fusion. In: International Conference on Medical Image Computing and Computer-Assisted Intervention. Cham: Springer International Publishing; 2019:447-56.
- Ding Y, Yu X, Yang Y. RFNet: Region-aware fusion network for incomplete multi-modal brain tumor segmentation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision; 2021:3975-84.
- Azad R, Khosravi N, Merhof D. SMU-net: Style matching U-net for brain tumor segmentation with missing modalities. In: International Conference on Medical Imaging with Deep Learning. PMLR; 2022:48-62.
- Qiu YS, Zhao ZY, Yao HD, Chen DL, Wang Z. Modal-aware visual prompting for incomplete multi-modal brain tumor segmentation. In: Proceedings of the 31st ACM International Conference on Multimedia; 2023:3228-39.
- Liu Y, Tian Y, Zhao Y, Yu H, Xie L, Wang Y, Ye Q, Jiao J, Liu Y. VMamba: Visual state space model. Adv Neural Inf Process Syst 2024;37:103031-63.
- Sohn JH, Chillakuru YR, Lee S, Lee AY, Kelil T, Hess CP, Seo Y, Vu T, Joe BN. An Open-Source, Vender Agnostic Hardware and Software Pipeline for Integration of Artificial Intelligence in Radiology Workflow. J Digit Imaging 2020;33:1041-6. [Crossref] [PubMed]
- Blezek DJ, Olson-Williams L, Missert A, Korfiatis P. AI Integration in the Clinical Workflow. J Digit Imaging 2021;34:1435-46. [Crossref] [PubMed]
- Canguit JA, Sison M. Diagnostic efficiency, challenges and job satisfaction of radiologic technologists in picture archiving and communication system (PACS). Int J Acad Res Prog Educ Dev 2025;14:2393-408.
- Zhang L, LaBelle W, Unberath M, Chen H, Hu J, Li G, Dreizin D. A vendor-agnostic, PACS integrated, and DICOM-compatible software-server pipeline for testing segmentation algorithms within the clinical radiology workflow. Front Med (Lausanne) 2023;10:1241570. [Crossref] [PubMed]
- Albalawi E, Mahesh TR, Thakur A, Kumar VV, Gupta M, Khan SB, Almusharraf A. Integrated approach of federated learning with transfer learning for classification and diagnosis of brain tumor. BMC Med Imaging 2024;24:110. Erratum in: BMC Med Imaging 2024;24:161. [Crossref] [PubMed]
- Zhang W, Jin W, Rho S, Jiang F, Yang C. A federated learning framework for brain tumor segmentation without sharing patient data. Int J Imaging Syst Technol 2024;34:e23147.
- Joshi M, Pal A, Sankarasubbu M. Federated learning for healthcare domain-pipeline, applications and challenges. ACM Trans Comput Healthc 2022;3:1-36.
- Mazher M, Razzak I, Qayyum A, Tanveer M, Beier S, Khan T, Niederer SA. Self-supervised spatial-temporal transformer fusion-based federated framework for 4D cardiovascular image segmentation. Inf Fusion 2024;106:102256.
- Wang B, Li H, Guo Y, Wang J. PPFLHE: A privacy-preserving federated learning scheme with homomorphic encryption for healthcare data. Appl Soft Comput 2023;146:110677.

