Mamba-AE: pixel-wise anomaly detection auto-encoder in chemical exchange saturation transfer magnetic resonance imaging
Original Article

Mamba-AE: pixel-wise anomaly detection auto-encoder in chemical exchange saturation transfer magnetic resonance imaging

Yingcheng Zhao1 ORCID logo, Zhenzhen You1, Lanfen Chen2, Zhenghao Shi1, Lin Wang1, Yuanlin Zhang1, Lu Sun1, Xiaowei He3, Xiaoli Wang2

1School of Computer Science and Engineering, Xi’an University of Technology, Xi’an, China; 2School of Medical Imaging, Shandong Second Medical University, Weifang, China; 3Xi’an Key Lab of Radiomics and Intelligent Perception, School of Information Sciences and Technology, Northwest University, Xi’an, China

Contributions: (I) Conception and design: Y Zhao; (II) Administrative support: X Wang, X He; (III) Provision of study materials or patients: X Wang; (IV) Collection and assembly of data: All authors; (V) Data analysis and interpretation: Y Zhao, Z You, L Chen, Z Shi; (VI) Manuscript writing: All authors; (VII) Final approval of manuscript: All authors.

Correspondence to: Xiaoli Wang, PhD. School of Medical Imaging, Shandong Second Medical University, No. 7166 Baotong West Street, Weifang 261053, China. Email: wxlpine@163.com; Xiaowei He, PhD. Xi’an Key Lab of Radiomics and Intelligent Perception, School of Information Sciences and Technology, Northwest University, No. 1 Xuefu Avenue, Xi’an 710127, China. Email: hexw@nwu.edu.cn.

Background: Chemical exchange saturation transfer (CEST) magnetic resonance imaging (MRI) demonstrates significant potential for early disease detection, yet encounters challenges in lesion analysis due to interference from concomitant effects during imaging. Although machine learning-based unsupervised anomaly detection (UAD) methods offer feasible solutions for identifying subtle lesions in contaminated images, current UAD approaches primarily dependent on convolutional neural network (CNN) or Transformer architectures to process image-level data from conventional imaging modalities, exhibit limitations when applied to CEST data characterized by spectral information. This study aimed to propose a novel UAD framework tailored for CEST spectral data to overcome these limitations, enabling precise pixel-wise anomaly detection for early-stage lesions.

Methods: The proposed framework named Mamba-AE employs a multi-layers encoder-decoder architecture with stacked Mamba blocks as the core component. Each Mamba block integrates three key modules: selective state-space models (SSMs) to dynamically adjust parameters based on input sequences, enabling adaptive long-range dependency modeling of CEST spectral data; Gated Multi-Layer Perceptron with Sigmoid Linear Unit (SiLU) activation and residual connections to enhance non-linear feature learning and stability; Multi-Scale Feature Alignment that aligns hierarchical encoder-decoder features via cosine similarity, preserving physiological semantics across scales. The framework adopts a dual-domain reconstruction strategy: data-space reconstruction minimizes pixel-wise spectral errors through Huber loss, while feature-space reconstruction enforces consistency between paired encoder-decoder layer features via cosine similarity loss. Anomaly scores are generated by combining normalized residuals from data-space errors and discrepancies in feature-space alignment, enabling precise pixel-level lesion localization in early-stage CEST data. For quantitative evaluation of lesion detection performance, the area under the curve (AUC) and Dice similarity coefficient are employed as core metrics. The framework was trained exclusively on CEST spectra from healthy rat brain tissue and validated on an independent dataset from a rat model of transient focal cerebral ischemia induced by middle cerebral artery occlusion (MCAO) with reperfusion.

Results: The proposed method has been validated on ischemic stroke rat datasets at 2, 6, and 24 hours post-occlusion. Mamba-AE demonstrated superior performance across all stages. At the critical 2-hour time point, it achieved the highest performance among all methods, with an AUC of 95.91% and a Dice score of 88.43% (sensitivity >89%, specificity >92%). The advantage remained substantial at 6 hours (AUC: 92.19%, Dice: 85.23%) and 24 hours (AUC: 94.64%, Dice: 87.79%), consistently outperforming other approaches. Furthermore, anomaly heatmaps revealed precise lesion localization capabilities, with spatial accuracy correlating strongly with histopathological ground truth measurements.

Conclusions: Mamba-AE provides a computationally efficient and robust solution for early disease detection in CEST MRI. Its integration of spectral feature learning and structural preservation highlights its potential for clinical applications requiring high sensitivity to subtle pathological changes.

Keywords: Chemical exchange saturation transfer imaging (CEST imaging); unsupervised anomaly detection (UAD); Mamba; dual-domain reconstruction; ischemic stroke


Submitted Sep 09, 2025. Accepted for publication Jan 19, 2026. Published online Feb 11, 2026.

doi: 10.21037/qims-2025-1952


Introduction

Chemical exchange saturation transfer (CEST) magnetic resonance imaging (MRI) is a novel molecular imaging technique (1). As shown in Figure 1A, different from the traditional medical scans that image only with grayscale or red, green, blue (RGB) channels, CEST covers a much broader spectral range over the resonance frequencies of multiple types of protons (2). Therefore, this imaging method can specifically detect low-concentration endogenous metabolic substances such as glucose (3), glutamate (4), creatine (5), and utilize them as non-invasive biomarkers. Due to this characteristic, numerous studies have demonstrated the capability of CEST imaging in detecting various diseases, including tumor (6), Alzheimer’s disease (7), and stroke (8), particularly in their early stages. Thus, the technique holds promising clinical application prospects.

Figure 1 Schematic diagram of CEST imaging results. (A) A typical set of early-stage CEST ischemic stroke data, raw CEST spectra from individual pixels in different tissues, along with the corresponding quantitative APT spectra obtained using MCMC-based inverse Z-spectrum analysis; (B) a comparison of images obtained from CEST and conventional MRI modalities during the early stage of ischemic stroke. Red arrow: indication of the lesion. APT, Amide Proton Transfer; CEST, chemical exchange saturation transfer; MCMC, Markov Chain Monte Carlo; MRI, magnetic resonance imaging; T2W, T2-weighted.

Although the CEST effect of exchangeable protons such as amine and amide gives the abovementioned metabolites the ability to indicate diseased tissues, the scanning process is susceptible to interference from a variety of concomitant signals that are insensitive to lesion manifestation owing to the intrinsic mechanism of CEST imaging, such as the direct saturation (DS) effect and the magnetization transfer (MT) effect from semi-solid macromolecules (9). These effects dilute the key diagnostic information of interest, severely weakening the contrast of the focal tissue in the image, especially in the early stages of disease where manifestations are not yet prominent. Consequently, quantitative analysis methods are typically required to eliminate these interfering effects, these methods can be roughly categorized into model-free approaches and model-based approaches. Model-free methods, such as MT ratio asymmetry (MTRasym) (10) and inverse Z-spectrum analysis (MTRRex) (9), feature simple computational processes but relatively limited precision due to their reliance on simplified Z-spectrum analysis without accounting for the influence of Nuclear Overhauser Effect (NOE), MT and other concomitant effects. In contrast, model-based methods, including multi-pool Lorentzian fitting (11) and Markov Chain Monte Carlo (MCMC) quantification (12), achieve higher quantitative accuracy (ACC) by leveraging the theoretical model, Bloch-McConnell equation to fit and separate overlapping effects, but often involve cumbersome computational workflows, which pose a significant challenge to clinical diagnosis. As shown in Figure 1, for the early ischemic at 2 hours post-occlusion, conventional weighted imaging modalities such as T2-weighted (T2W) imaging, diffusion-weighted imaging (DWI), raw CEST images [at 3.5 ppm where the Amide Proton Transfer (APT) effect is strongest], and the corresponding Z-spectra all fail to exhibit significant pathological manifestations. In contrast, the quantitative spectra and images of APT, obtained by quantifying APT through sophisticated MCMC parameter quantification combined with inverse Z-spectrum analysis, successfully capture the occurrence of ischemia.

Detecting lesions manually from disturbed images is undoubtedly very difficult and inefficient, fortunately, deep neural networks have now become a popular tool for analyzing medical images and can help researchers extract pathological information from mixed signals. Although supervised deep learning methods have been shown to achieve satisfactory performance in disease detection and lesion segmentation tasks, the collection and annotation of high-quality pathology data is still a challenging task (13), especially for CEST imaging, which has not yet been widely applied in clinical settings. It is also difficult to incorporate the various abnormalities that occur during the disease process when constructing the training dataset. Compared with supervised methods, unsupervised anomaly detection (UAD) aims to identify abnormal information by learning the pattern of normal samples (14-16), thus eliminating the requirement to annotate pathological data and greatly facilitating the construction of datasets. Therefore, UAD has attracted much attention in the field of medical image analysis (17,18).

Due to the reliance on normal samples only, most current UAD methods typically attempt to learn the distribution paradigm of images, and build detection strategies upon this foundation (19-21). These algorithms can be broadly classified into anomaly synthesis-based methods and reconstruction-based methods. The former try to generate pseudo-anomalies in normal data to cover all possible anomaly types, thereby modeling UAD as a fully supervised binary classification problem and usually constructing discriminators to differentiate them (22-24); whereas the latter assume that a network trained on normal data patterns can reconstruct normal tissue images but fails to reconstruct unseen abnormal tissue images (25-27). Data reconstruction methods often leverage the auto-encoder (28,29) or generative adversarial network models (30,31) for training, aiming to minimize the discrepancy between the original input and the reconstructed data, known as the reconstruction error. The trained model provides a corresponding reconstructed image for the input image and calculates an anomaly score based on the distance between the two. While feature reconstruction methods map images into a feature space for reconstruction (32-34). Compared to image reconstruction, features may offer more discriminative and low-noise representations (18,35).

The CEST spectrum contains multiple exchange signals from substances sensitive to lesions. By benefiting from the abundant spectral information, pixel-wise analysis of spectra can better uncover the response patterns of signals to pathologies. Currently, reconstruction-based methods typically employ convolutional neural networks (CNNs) as their backbone models (25,35). Although CNNs can effectively learn the semantic information of local contexts, they are limited by their receptive fields and lack the ability for long-distance modeling, making it challenging to handle CEST spectrum sequences. Models based on recurrent neural networks (RNNs) (36-38) and long short-term memory networks (LSTMs) (39-41) have also been introduced into UAD tasks, demonstrating feasibility in time series. These models learn point-by-point representations of sequences to evaluate errors. However, this approach has difficulties in providing a comprehensive description of the sequence context, and for complex patterns, the information provided by point-by-point representations is relatively limited, which may result in the learning process being dominated by normal data points, making it difficult to distinguish anomalies. The Transformer model has also garnered attention in the UAD field due to its excellent long-range dependency modeling capabilities (42-44), but the quadratic complexity brought by its self-attention mechanism constrains the model’s performance.

Recently, state-space models (SSMs) have demonstrated exceptional performance in the analysis of continuous long sequence data (45-47). As a type of SSMs with a selection mechanism, Mamba can select relevant information in an input-dependent manner, thereby possessing stronger long-range modeling capabilities while maintaining linear computational complexity (48). Consequently, it has garnered extensive attention in various fields (49-51).

Therefore, to address these limitations, this study aims to develop a novel UAD framework. Specifically, we propose Mamba-AE, a pixel-wise framework that applies the Mamba model to capture long-range dependencies in CEST Z-spectra, thereby enabling precise early lesion detection with a focus on ischemic stroke. Specifically, we use multiple Mamba blocks to capture the long-range dependencies in the CEST spectrum and learn the normally organized signal patterns. The model takes the spectrum as input and outputs a reconstruction of it for end-to-end training. In addition, we consider the role of reconstruction on both data and feature dimensions, and introduce reconstruction loss to jointly constrain the learning process in two domains, data space and potential feature space, respectively.

The main contributions of this method can be summarized as follows:

  • An auto-encoder framework based on the Mamba backbone is proposed for CEST data anomaly detection. This framework effectively leverages Mamba’s powerful long-range feature extraction capabilities to model long-distance CEST spectra;
  • Joint constraints are established in both data and feature spaces to simultaneously learn the distribution patterns of normal samples from both domains;
  • We have constructed three rat brain ischemia datasets at different stages and evaluated the effectiveness of the Mamba-AE model on each. Besides, comparative experiment results with multiple methods have proven the superiority of proposed method in pixel-wise anomaly detection for CEST.

We present this article in accordance with the TRIPOD reporting checklist (available at https://qims.amegroups.com/article/view/10.21037/qims-2025-1952/rc).


Methods

Animal model and dataset

The data employed in this study were derived from a cerebral ischemia dataset in rats, developed by our research team. The animal models consisted of 24 healthy, 9 to 10-week-old male Sprague-Dawley rats with a body weight of 240 grams. Specifically, 7 healthy rats were used as the training set, while the test sets comprised 7, 6, and 4 rats at 2, 6, and 24 hours post-ischemia, respectively. A subset of these rats underwent middle cerebral artery occlusion (MCAO) (52) via intraluminal suture. Subsequently, reperfusion was induced by removing the nylon suture two hours after MCAO, thereby establishing a robust model of focal cerebral ischemia. All animal experiments were performed under a project license (No. 2019SDL092) granted by the Institutional Ethics Committee of Shandong Second Medical University, in compliance with the institutional guidelines for the care and use of animals.

CEST data acquisition

Data acquisition was performed on a 7-Tesla horizontal bore scanner (Bruker Biospec, Ettlingen, Germany) using a transmit-receiver volume coil (diameter =40 mm). Following the acquisition of T2W images for anatomical reference across multiple slices, CEST MRI was acquired at a coronal slice positioned at the center of the striatum. The CEST sequence utilized a continuous wave pre-saturation pulse (saturation time, Tsat =2,500 ms), followed by a rapid acquisition with relaxation enhancement (RARE) readout [RARE factor =32; repetition time/echo time (TR/TE) =5,000 ms/4 ms]. A total of 135 Z-spectra images with saturation offsets (∆ω) ranging from −10 to 10 ppm were collected sequentially at saturation radio frequency (RF) amplitudes (B1-sat) of 0.7, 1, and 2 µT, respectively. A B0 inhomogeneity map was obtained using the water saturation shift referencing (WASSR) method, which involved acquiring saturation images with a weak saturation pulse (B1-sat =0.5 µT; Tsat =500 ms) swept from −1 to 1 ppm in 0.1 ppm increments. An S0 image (without a saturation pulse) was also acquired for signal normalization. The following parameters were consistent across all the above-mentioned sequences: slice thickness =1 mm, field of view (FOV) =34×28 mm2, and matrix size =94×64. Based on these meticulously defined scanning parameters, imaging sessions were conducted at four critical time points throughout the post-ischemic disease progression: pre-modeling, 2, 6, and 24 hours post-reperfusion.

Based on the abovementioned experimental setup, we obtained a dataset containing 603 frequencies and a total of 9,934 pixels. Among these pixels, 3,739 are from scans of non-ischemic rats, serving as the training dataset within the study, while the remainder are derived from scans of the three post-ischemic stages, which constitute the test dataset.

Data pre-processing

We utilized the Matlab platform (MathWorks, Natick, USA) to complete the data pre-processing operations, and implemented our network model with the PyTorch framework, all runs were conducted on an NVIDIA GeForce RTX 4090 GPU (NVIDIA, Santa Clara, CA, USA). The Adam optimizer is adopted to optimize the model parameters, with a learning rate of 0.0005, and other parameters are set as β1 =0.5, β2 =0.999. The model was trained for 200 epochs with a batch size of 64. During the visualization presentation of testing stage only, the image scale is resized to 256×256 to improve the display performance. Please note that this post-processing step did not affect the anomaly score calculation, which was performed entirely on the original data.

Given a training set Dtrain= {x1, x2, …, xN} with N normal CEST samples and a test set Dtest = {(xt1, yt1), (xt2, yt2), …, (xtM, ytM)} with M labeled CEST samples, where a CEST image xiH×W×f, a pixel-wise label yiH×W. H, W, and f signify the height, width, and the frequencies per pixel of an image. Our goal is to train a model to detect anomalous pixels (lesion regions) from Dtest.

We preprocess the raw CEST data to facilitate better training. First, the CEST images were corrected for field inhomogeneity (53), and then normalized using the unsaturated CEST images. Next, we performed dimensional transformation on the dataset to convert it into a collection of spectral data for all pixels. The prepared data are acquired under B1 irradiation pulse intensities of 0.7, 1, and 2 µT, with these spectra serving as inputs for three channels. The resulting training set Dtrainn×f×c, in which n represents the total number of pixels, and c=3, represents the number of channels.

Mamba-AE network architecture

An overview of the proposed Mamba-AE is presented in Figure 2. The framework is based on a conventional auto-encoder, comprising an encoder and a decoder, each composed of 4 layers. Each layer consists of several stacked Mamba blocks, which serve as the core component for spectral feature extraction and reconstruction.

Figure 2 Overview of the proposed Mamba-AE approach. The Mamba-AE employs a multi-layered encoder-decoder framework, that: (I) extracts and reconstructs CEST spectra through multiple stacked Mamba blocks; (II) constructs a global loss function/anomaly score by leveraging multi-scale distances between pairwise layers. CEST, chemical exchange saturation transfer; RMS, root mean squared; SiLU, Sigmoid Linear Unit; SSM, selective state-space model.

The Mamba block is built upon a selective SSM, a novel sequence modeling architecture that combines the high performance of Transformers with the computational efficiency of recurrent models. Its key innovation is a selection mechanism that allows the model to dynamically adjust its parameters based on the input, enabling it to focus on or ignore specific parts of the lengthy CEST spectrum. This overcomes the limitations of CNNs, which have a limited receptive field, and Transformers, which suffer from quadratic computational complexity when processing long sequences. For our task, this translates to efficient and robust modeling of global dependencies in CEST data.

The fundamental SSM maps a 1D input sequence x(t) to a hidden state h(t) and an output y(t). The selective mechanism of Mamba dynamically generates the parameters B and C, as well as the time step Δ, as functions of the input x(t). This input-dependent configuration is crucial for adapting to the complex patterns in CEST Z-spectra.

To enable digital computation, the continuous system is discretized using a zero-order hold (ZOH) rule with the input-dependent step size Δ, transforming the continuous parameters into discrete parameters A¯ and B¯. This yields the core recurrent computation of the SSM at each time step:

ht=A¯ht1+B¯xt

yt=Cht+Dxt

Beyond the core SSM, a complete Mamba block, as illustrated in the lower part of Figure 2, integrates several key components for stable and effective training: root mean squared norm (RMS norm) for normalization, a Gated Multi-Layer Perceptron (MLP) with Sigmoid Linear Unit (SiLU) activation for enhanced non-linear feature transformation, and residual connections to facilitate the training of deep networks. The processing of an input feature fi1 in the i-th Mamba block (MBi) can be summarized as:

fi=MBi(fi1)

fi1=SSM(Silu(Conv(Linear(RN(fi1)))))

fi=fi1+Linear(Silu(Linear(RN(fi1)))fi1)

where Conv, RMS normalization(RN), Linear, and SiLU denote a 1×1 convolutional layer, an RMS norm layer, a linear projection layer, and the SiLU activation function, respectively, and represents element-wise multiplication.

Loss function

In a multi-layer auto-encoder, each layer has corresponding feature inputs and outputs, and these sub-encoder-decoder structures are essentially feature reconstruction processes. The features learned in these processes provide multi-scale and low-noise data representations. Therefore, beyond the traditional data reconstruction commonly used in UAD tasks, our research also focuses on global feature reconstruction to fully leveraging the feature extraction capabilities of the encoder. We incorporate the cosine similarity between paired encoder-decoder layers as part of the model’s loss. For our proposed four-layer structure, except for the outermost layers, which directly process the input and output data, we utilize the input features of the k-th layer encoder and the corresponding output features of the k-th layer decoder at different scales. By calculating the cosine similarity, we measure the distance between features pixel by pixel, which serves as our feature reconstruction loss. This component can be expressed as follows:

Lfeature=k=24(1FEKTFDKFEKFDK)

The encoder/decoder features for the k-th layer are denoted as FEK, FDKb×lk×d, where lk=f2k1 represent the corresponding sequence length, and denotes the L2 norm.

Relying solely on feature reconstruction loss is insufficient for our reconstruction network to learn the contextual information of the input data. It is necessary to constrain the generative model by measuring the distance between the original sequence and the generated sequence. Therefore, we introduce a reconstruction loss for the data. Among the commonly used losses in UAD studies, the L1 loss function exhibits high robustness, while the L2 loss is more sensitive to outliers but has the advantage of faster convergence. Therefore, we adopt the Huber loss, which is a compromise between L1 and L2 losses. This loss function is flexible and can adaptively select a computation strategy based on the comparison of the difference between the reconstructed value and the actual value and a preset threshold. Thus, it can efficiently handle non-Gaussian noise while maintaining efficiency, which can be formulated as:

Ldata(y,y˜)={12(yy˜)2,if|yy˜|δδ|yy˜|12δ2,if|yy˜|>δ

Here, y,y˜b×f×c represent the original data values and the predicted values, respectively, while the parameter δ serves as a weight to select the calculation formula. When the prediction error is less than δ, the mean squared error (MSE), also known as L2, is used. Conversely, when the prediction error exceeds δ, a linear error, which approximates L1, is adopted. In our research, tests have shown that the performance benefits the most when δ=1.

Based on the aforementioned two loss functions, the overall objective used to optimize the entire model can be formulated as:

L=αLfeature+βLdata

Here, α and β are the loss weights, which have been adjusted through testing and set to 1 and 0.4 respectively in our study to achieve optimal performance.

Testing process and anomaly map

During the testing stage, a set of test data y can be input into the trained model to obtain the feature maps FEK, FDK at various layers and the reconstructed data y˜. We then evaluate the distance between them, similar to the way we calculate loss, the anomaly score for a single pixel is composed of feature-level and data-level anomaly scores, which can be computed as follows:

A(h,w)=k=24[1FEK(h,w)TFDK(h,w)FEK(h,w)FDK(h,w)]+[y(h,w)y˜(h,w)]2

Wherein, a pair of (h,w) locates a spatial point in the original image, and a larger A value indicates a higher anomaly likelihood at the corresponding position. Compared to loss calculation, since the difference between predicted and true values is generally small after optimization is completed, here we adopt the MSE metric for computing spectrum-level scores, and remove the constraint weights on this part, allowing anomalies among the data to serve as the dominant factor. The final anomaly map ManoH×W is obtained by describing each pixel with A, and a global normalization is performed on this map.

Evaluation metrics

To evaluate the performance of our study, we select the area under the curve (AUC), Dice score (Dice), average classification ACC, sensitivity (SEN), and specificity (SPE) as the evaluation metrics. Below are the calculation formulas for these metrics, as well as for the intermediate term precision (PRE):

SEN=TPTP+FN

SPE=TNFP+TN

ACC=TP+TNTP+TN+FP+FN

PRE=TPTP+FP

Dice=2×|PredGT||Pred|+|GT|=2×PRE×SENPRE+SEN

where the TP, TN, FP, FN represent the pixel counts for true positive, true negative, false positive, and false negative, and the Pred and GT represent the predicted and ground truth segmentation sets, respectively. AUC refers to the area under the receiver operating characteristic (ROC) curve, which is plotted based on the true positive rate (TPR) against the false positive rate (FPR) across various classification thresholds. The thresholds we use to calculate these metrics are determined based on the optimal Dice score.

Experimental setup

In our experiments, apart from utilizing the basic autoencoder (AE) and variational autoencoder (VAE), which rely on data reconstruction error for anomaly detection, we also incorporated several state-of-the-art anomaly detection methods for comparison: f-AnoGAN (27), GANomaly (28), self-supervised and translation-consistent features for anomaly detection (SALAD) (52), RD4AD (22), and encoder-decoder contrast (EDC) (15). We reproduced these models based on the publicly available codes provided by the authors and conducted model training and testing using the dataset described in this paper. It is worth noting that, to maintain consistency in the research objects of this study, for most of these models, we modified some of the layers that originally processed two-dimensional data to instead process one-dimensional data, enabling them to accommodate our spectral-level inputs, while keeping their overall structures unchanged. However, RD4AD and EDC are exceptions as they still process image-level data, because their methods leverage pre-trained models on ImageNet and cannot be arbitrarily altered. To accommodate this requirement, we preprocessed the CEST data specifically for these two models by using the basic S0-normalized image at each saturation frequency offset as an independent single-channel 2D input. Each frequency’s image was processed separately by the models. This approach allowed us to leverage their pre-trained architectures while conforming to the image-based input requirement.

To further evaluate the effectiveness and contributions of the proposed components, we designed comprehensive ablation studies. These primarily include a comparison of the effectiveness of different network models and our adopted Mamba model as the backbone of the auto-encoder, as well as a comparison of the impact on performance when combining the data-level and feature-level reconstruction loss functions versus using them separately. For the backbone comparison, we maintained the same auto-encoder architecture while systematically replacing the core sequence modeling component. The Mamba blocks were substituted with fundamental layers including MLP, RNN, and LSTM layers, respectively. For the loss function comparison, we evaluated the performance under the constraints of data-level loss based on MSE, feature-level loss based on cosine similarity, and the proposed combined loss, with the model using data-level loss taken as the baseline. In all ablation experiments, the anomaly score calculation method and other default settings remained consistent with those described in the main model.


Results

Overall performance comparison

We evaluated the lesion detection performance of the proposed method and the current state-of-the-art methods introduced in the Methods section at different stages after ischemic stroke, conducting experiments on datasets collected at 2, 6, and 24 hours post-ischemia. Both the proposed method and the state-of-the-art methods are trained 10 times using different random seeds. Table 1 presents the average results of all methods. Overall, our proposed method significantly outperformed the comparison methods on all datasets across various metrics. Specifically, lesion detection on the 2-hour dataset is particularly challenging because early ischemia typically lacks prominent pathological manifestations. Among the comparison methods, some utilizing simple models such as AE and VAE exhibited poor performance, insufficient for effectively identifying lesion pixels. Methods incorporating strategies such as adversarial learning or feature-level constraints yielding results with pathological significance, yet there is room for improvement in terms of reliability. In addition, methods utilizing pre-trained models (RD4AD and EDC) exhibit even better performance. Compared to these methods, our approach achieved more ideal performance across all metrics on this dataset, demonstrating its ability to stably detect strokes in the early stages, with AUC, Dice score, ACC, and SPE of 95.91%, 88.43%, 91.14%, and 91.97%, exceeding previous SOTA methods 4.42%, 0.31%, 5.17%, and 12.63%, respectively.

Table 1

Anomaly detection performance (%) yielded by different methods on 2, 6, and 24 hours datasets

Datasets Metrics AE VAE f-AnoGAN GANomaly SALAD RD4AD EDC Ours
2 hours AUC 66.07 59.37 78.17 83.63 84.64 83.57 91.49 95.91
F1-score 58.60 62.82 69.69 73.54 72.82 72.08 88.12 88.43
ACC 61.22 50.45 71.20 78.41 76.70 74.39 85.97 91.14
SEN 71.18 95.85 85.38 78.11 80.70 82.40 91.63 89.76
SPE 54.98 15.26 62.22 78.60 74.17 79.34 78.53 91.97
6 hours AUC 73.55 69.56 67.50 84.02 88.94 82.97 90.91 92.19
F1-score 54.17 57.46 49.14 77.29 75.21 69.09 76.07 85.23
ACC 65.70 69.60 68.53 82.04 84.00 81.11 85.03 90.64
SEN 73.39 62.10 61.29 72.15 85.05 70.37 78.76 81.45
SPE 62.77 73.31 70.92 89.30 83.58 85.71 87.74 95.20
24 hours AUC 71.93 78.44 80.38 90.01 89.05 93.41 92.98 94.64
F1-score 56.40 65.90 69.54 78.64 76.97 79.44 83.48 87.79
ACC 69.26 75.71 76.94 85.86 85.10 84.90 88.37 92.43
SEN 62.18 73.25 82.17 80.89 77.71 89.94 91.72 88.42
SPE 72.59 76.88 74.47 88.22 88.59 82.48 86.79 96.77

ACC, accuracy; AE, autoencoder; AUC, area under the curve; EDC; SALAD; SEN, sensitivity; SPE, specificity; VAE, variational autoencoder.

For the 6-hour ischemia dataset, pathological manifestations in tissues typically become more evident, making it less difficult to detect abnormalities compared to the early stage, but accurately identifying specific pixels remains challenging. On this dataset, the performance of various comparison methods improved, but the results of AE and VAE were still unsatisfactory (Dice score <60%, ACC <70%), while GANomaly, SALAD and EDC performed well, with metrics close to the optimal values. Among these methods, our method achieved the best results across all metrics, with AUC and Dice score of 92.19% and 85.23%, respectively, and an ACC of 90.64%, surpassing the best results by 1.28%, 7.94%, and 5.61%, respectively.

The detection task on the 24-hour dataset is relatively straightforward because pathological changes in ischemic tissues are usually pronounced at this stage. Therefore, the performance of various comparison methods on this dataset was acceptable, with almost all methods achieving AUC and ACC values exceeding 80%. However, f-AnoGAN’s SPE and SALAD’s SEN were lower, indicating the bias of these models. Our method maintained balanced performance across all metrics, with AUC, Dice score, ACC, SEN, and SPE of 94.64%, 87.79%, 92.43%, 88.42%, and 96.77%, respectively. Among them, AUC, Dice score, ACC, and SPE exceeded the highest values by 1.23%, 4.31%, 4.06%, and 8.18%, respectively.

It is worth mentioning that, despite conducting image-level detection rather than pixel-level detection, the RD4AD and EDC methods exhibit considerable performance close to our approach, demonstrating the potential of pre-trained models. Compared to RD4AD, EDC further improves performance, indicating the effectiveness of its improvements aimed at addressing the shortcomings of pre-trained models in medical imaging. However, as the detection difficulty of the dataset decreases, their performance is no longer sufficiently outstanding. According to subsequent abnormal heatmap experiments, it can be observed that such methods struggle to detect the local details of abnormal regions, reflecting their applicability issues in pixel-wise UAD.

Heatmap visualization comparison

Using pixel-wise anomaly score calculations, we generated interpretable heatmaps for the detection results of various comparison methods and the proposed method on these three datasets. Higher anomaly values are reflected as hotter regions in the images, as shown in Figure 3, which now additionally includes co-registered T2-weighted imaging and DWI for anatomical reference and clinical context. The displayed CEST images used for reference are derived from manually selected 3.5-ppm scans that exhibit good lesion contrast effects. For the 2-hour dataset, basic AE and VAE can only roughly reflect some discrete pixel points within the ischemic core region, but they also incorrectly show many hot regions outside the ischemic tissue. f-AnoGAN demonstrates the general ischemic region, but the hotspots are relatively discrete, exhibiting a very low signal-to-noise ratio. GANomaly and SALAD exhibit clearer hot region contours, but the localization of lesion tissue is not accurate enough. Notably, at this early stage, the CEST-based anomaly maps reveal pathological regions that show limited manifestation in the corresponding T2-weighted imaging and DWI images, supporting our method’s SEN to early metabolic alterations. For the 6- and 24-hour datasets, almost all methods can roughly show the ischemic core, but simple models overdetect many anomaly points and still struggle to clearly delineate the boundaries of the anomaly regions. Methods such as f-AnoGAN, GANomaly, and SALAD exhibit results close to our method, but they do not fully reflect the lesion areas and have slightly inaccurate contours. As disease progresses, the anomaly regions detected by our method show increasing spatial correspondence with the abnormalities visible in T2-weighted imaging and DWI, while maintaining superior boundary definition. Due to differences in detection mechanisms, we specifically note the results of RD4AD and EDC. Unlike pixel-wise anomaly scoring, these methods use image-level detection, which results in smoother and complete regions. However, it struggles to present the true extent and local details of ischemic tissue with complex shapes. It can be observed that RD4AD focuses on displaying the ischemic core and the overall range but fails to provide contours that closely approximate the true lesion area, while EDC has seen significant improvement, although not a perfect match, it can roughly reflect the local manifold of the lesion. Compared to these methods, our proposed method effectively covers the ischemic regions on various datasets and leverages the Mamba model and hierarchical feature comparison to thoroughly learn the characteristics of normal pixel CEST spectra sequences, thereby suppressing misjudgments arising from inherent physiological differences between tissues.

Figure 3 Visualization of anomaly heatmaps for datasets at various stages post-ischemia, where higher heat values (red regions) indicate higher anomaly scores, suggesting the potential presence of pathology at those locations. The first three columns display the original CEST images (normalized Z-spectrum image at 3.5 ppm), co-registered T2-weighted images, and DWI of the same subject for anatomical reference. The fourth column shows pixel-level anomaly annotations, and subsequent columns exhibit comparative heatmaps generated by several techniques, namely AE, VAE, f-AnoGAN, GANomaly, SALAD, RD4AD, EDC, and our method, respectively. Rows correspond to 2, 6, and 24 hours post-ischemia time points. AE, autoencoder; CEST, chemical exchange saturation transfer; DWI, diffusion-weighted imaging; EDC, SALAD, T2W, T2-weighted; VAE, variational autoencoder.

The capability of our method to delineate lesions, especially in the critical 2-hour post-ischemia phase, stems from its direct SEN to the underlying pathophysiology. The anomalies detected correspond to regions of tissue acidosis, where a drop in pH inhibits the amide proton exchange rate (ksw), a metabolic event that precedes the structural damage visible on DWI or T2W imaging. This provides a clear physiological basis for the observed hotspots in our heatmaps, which often highlight areas that are not yet apparent on conventional MRI. Consequently, our approach offers more than just pixel-level ACC; it provides an early metabolic signature of ischemia. This has significant translational potential, as it could aid in identifying the ischemic penumbra—salvageable tissue that is a primary target for early therapy, thereby refining diagnosis and treatment decisions in acute stroke.

Ablation study

In this section, we further evaluate the effectiveness and contributions of the improvement strategies proposed in our method using an ablation study, following the experimental design outlined in the Methods section. The ablation experiments involve both qualitative and quantitative evaluations. For qualitative comparison, we still use heatmaps for visualization, while for quantitative comparison, we use metrics including AUC, Dice score, ACC, SEN, and SPE.

We compared the performance of various feature extraction networks in UAD, specifically testing the three models, the conventional MLP, RNN, and LSTM as the backbone for the encoder/decoder, the MLP-based model was used as the baseline. In all experiments, the combined loss function described earlier was used for training, while the anomaly score calculation method and default settings remained consistent. The visualization results are shown in Figure 4. It can be seen that the baseline model exhibits many falsely detected hot regions, but compared to the performance of the simplest AE model (Figure 3), it improves the completeness of the detected ischemic region on the early ischemic dataset, indirectly reflecting the effectiveness of the combined loss function adopted in the proposed method. Models using RNN and LSTM show some improvement in over-detection issues, but the challenge of completeness and connectivity regarding the ischemic region has not been perfectly resolved yet. In contrast, the proposed Mamba-based model provides clearer lesion boundaries and a more uniform internal structure of the ischemic region.

Figure 4 Visualization of anomaly heatmaps using different AE backbone networks, where higher heat values (red regions) signify higher anomaly scores, indicating the potential presence of pathology at those locations. The first column from the left displays the original CEST images (3.5 ppm) for reference, the second column shows pixel-level anomaly annotations, and subsequent columns exhibit the results using MLP, RNN, LSTM as the backbone networks, and our method, respectively. AE, autoencoder; CEST, chemical exchange saturation transfer; LSTM, long short-term memory network; MLP, Multi-Layer Perceptron; RNN, recurrent neural network.

The corresponding quantitative results are shown in Table 2. It can be seen that, with other training strategies remaining consistent, the Mamba-based model achieves better performance on all datasets compared to traditional sequential models. In particular, on the more challenging 2-hour dataset, its AUC, Dice score, ACC, SEN, and SPE are improved by 12.52%, 9.79%, 12.5%, 12.18%, and 12.25% respectively compared to the baseline model, and by 8.37%, 12.66%, 10.91%, 4.82%, and 12.77% respectively compared to the best performance among the other two sequential models. In the 6-hour dataset, our method achieved the best results across all metrics, with AUC and Dice score of 92.19% and 85.23%, respectively, and an ACC of 90.64%, surpassing the best results of other backbones by significant margins. For the 24-hour dataset, our method maintained balanced performance across all metrics, with AUC, Dice score, ACC, SEN, and SPE of 94.64%, 87.79%, 92.43%, 88.42%, and 96.77%, respectively, demonstrating comprehensive leadership. The complete quantitative comparisons are detailed in Table 2. These consistent results across different stages of pathology fully demonstrate the effectiveness of our model for sequential data.

Table 2

Anomaly detection performance (%) yielded of our framework using different backbone networks on 2, 6, and 24 hours datasets

Datasets Metrics MLP RNN LSTM Ours
2 hours AUC 75.05 87.54 82.32 95.91
F1-score 73.60 75.77 73.63 88.43
ACC 71.14 80.23 77.05 91.14
SEN 79.37 81.93 84.94 89.76
SPE 62.67 79.20 72.26 91.97
6 hours AUC 81.97 89.98 86.87 92.19
F1-score 71.03 77.45 72.73 85.23
ACC 75.07 85.79 80.70 90.64
SEN 82.01 73.39 77.42 81.45
SPE 70.94 91.97 82.33 95.20
24 hours AUC 81.59 92.14 92.19 94.64
F1-score 74.85 82.24 80.65 87.79
ACC 74.33 88.32 87.70 92.43
SEN 81.94 79.62 79.62 88.42
SPE 67.69 90.33 91.54 96.77

ACC, accuracy; AUC, area under the curve; LSTM, long short-term memory network; MLP, Multi-Layer Perceptron; RNN, recurrent neural network; SEN, sensitivity; SPE, specificity.

We also evaluated the performance of the Mamba-AE framework under the constraints of different loss functions in data space and feature space. Specifically, these included data-level loss based on MSE, feature-level loss based on cosine similarity, and the proposed combined loss. The model using data-level loss was taken as the baseline. In all experiments, the Mamba network was used as the backbone for training, while the anomaly score calculation method and default settings remained consistent. The visualization results are shown in Figure 5. It can be observed that neither the data-level loss nor the feature-level loss alone provides comprehensive coverage of the lesion area. While the data-level loss produces relatively complete lesion contours, it fails to capture some fine details. Similarly, the feature-level loss shows limited ability to fully cover the lesion regions. In contrast, our proposed method using the combined loss function provides clearer local textures and fewer false detections compared to the other two.

Figure 5 Visualization of anomaly heatmaps utilizing various loss functions and anomaly scores, where higher heat values (red regions) signify higher anomaly scores, indicating a potential presence of pathology at those locations. The first column from the left displays the original CEST images (3.5 ppm) for reference, the second column shows pixel-level anomaly annotations, while the subsequent columns exhibit the results obtained using data loss and anomaly score, feature loss and anomaly score, and our method, respectively. CEST, chemical exchange saturation transfer.

In terms of quantitative evaluation, the experimental results are shown in Table 3. It can be seen that using only feature-level loss as a constraint is insufficient. The baseline model performs better than the model trained using only feature-level loss on all three datasets. However, our model using the combined loss function achieves significant improvements, with increases of 10.74%, 16.04%, and 14.5% in AUC, Dice score, and ACC respectively compared to the baseline model on the 2-hour dataset; increases of 6.72%, 10.92%, and 5.65% on the 6-hour dataset; and increases of 8.21%, 8.06%, and 4.93% on the 24-hour dataset. These results indicate that combining data-level loss and feature-level loss enables the model to better learn the semantic feature patterns of normal samples and tissue structures, thereby improving the performance of UAD in terms of data reconstruction.

Table 3

Anomaly detection performance (%) yielded of our Mamba-AE framework with different loss functions and different anomaly scores in the data and feature spaces on 2, 6, and 24 hours datasets

Datasets Metrics, loss, anomaly score AdataLdata AfeatureLfeature Ours
2 hours AUC 85.17 77.33 95.91
F1-score 72.39 66.86 88.43
ACC 76.64 73.35 91.14
SEN 82.32 71.08 89.76
SPE 73.29 74.73 91.97
6 hours AUC 85.47 75.92 92.19
F1-score 74.31 61.29 85.23
ACC 84.99 74.26 90.64
SEN 65.32 61.79 81.45
SPE 94.78 80.40 95.20
24 hours AUC 86.43 79.26 94.64
F1-score 79.73 62.72 87.79
ACC 87.50 78.07 92.43
SEN 76.92 57.32 88.42
SPE 92.47 87.92 96.77

ACC, accuracy; AUC, area under the curve; SEN, sensitivity; SPE, specificity.


Discussion

The proposed Mamba-AE framework demonstrates significant advancements in UAD for CEST MRI, particularly in addressing challenges associated with early-stage disease identification. Through comparative experiments with current mainstream UAD methods and ablation studies on the framework’s components, the results fully demonstrate the necessity of developing models tailored to CEST MRI data characteristics and the effectiveness of the proposed approach.

This capability is grounded in the unique contrast mechanism of CEST. The early anomalies detected by our model primarily reflect regions of tissue acidosis, where a decreased pH inhibits the amide proton exchange rate (ksw), a metabolic alteration that occurs prior to the structural tissue damage typically captured by DWI or T2W imaging. Consequently, our method provides a complementary, metabolism-sensitive perspective for identifying ischemic injury, potentially revealing pathological changes that are not yet visible on conventional MRI.

The multi-frequency scanning mechanism of CEST generates Z-spectra, which are critical for analyzing pathological biomarker information. However, the overlapping resonance peaks caused by concomitant effects lead to a “dilution” effect, resulting in high inter-frequency image heterogeneity and low tissue region discriminability. Consequently, feature extraction based on 2D images often fails to capture the most lesion-relevant semantic information. This is further evidenced by the performance gap between our method and image-based pre-trained models (RD4AD, EDC), which can be attributed to the fundamental domain shift between natural images and medical CEST data. This limitation underscores the necessity of developing specialized architectures tailored to spectral characteristics. The selective state-space mechanism of the Mamba model enables efficient and robust long-range dependency modeling, avoiding scalability limitations of sequential context integration or visual deformation, thus offering unique advantages in processing long-sequence data. It captures subtle pathological deviations in complex Z-spectra, while the hierarchical feature reconstruction strategy leverages multi-scale encoder-decoder outputs to provide low-noise representations, enhancing anomaly discrimination. Ablation experiments on the backbone demonstrate that the Mamba model overcomes the narrow receptive field of traditional sequence processing models, enabling the detection of abnormal changes by learning the distribution patterns of broad resonance peaks in normal tissue Z-spectra.

During encoding-decoding, deep features may deviate from normal sample distributions due to model capacity limitations or noise interference. Constraining feature alignment between intermediate encoder-decoder layers through feature-space reconstruction forces the model to maintain multi-scale semantic consistency in the latent space. Analysis of experimental results suggests that normal brain anatomical structures exhibit stable spatial patterns in deep features, while abnormal regions disrupt these patterns. Feature-space reconstruction compels the model to preserve critical structural information during compression-reconstruction via cross-layer alignment, thereby highlighting pattern collapse caused by lesions. Additionally, feature-space reconstruction distinguishes noise-induced errors (e.g., magnetic field inhomogeneity or motion artifacts) from true pathological changes through high-level semantic constraints. Complementing this, data-space reconstruction ensures global input-output consistency, effectively capturing prominent anomaly patterns. The dual-domain strategy synergizes data-feature co-optimization, balancing global consistency and local SEN, and jointly captures anomaly signals from pixel-level to semantic-level granularity.

The primary clinical implication of this work lies in the potential for ultra-early stroke assessment. By detecting metabolic dysfunction directly, our approach could aid in identifying the ischemic penumbra tissue that is metabolically compromised but structurally salvageable, thereby providing critical information for treatment decisions in the acute phase.

Despite the promising results, this study has certain limitations. The primary limitation is the constrained number of animal subjects, a common challenge in pre-clinical CEST-MRI studies due to the extensive acquisition time required for high-resolution Z-spectra. However, as noted in the Methods section, the pixel-wise learning paradigm effectively amplifies the sample size for the anomaly detection task, with the model training on thousands of individual spectral sequences. This approach has demonstrated robust performance, particularly in the critical early stages of ischemia. Nevertheless, the generalizability of the model would benefit from validation on a larger and more diverse cohort.

Future research should focus on expanding Mamba-AE’s capabilities through four key directions: First, multi-disease validation across diverse CEST datasets to establish broader clinical applicability and robustness. Second, integrating Mamba’s spectral analysis strengths with spatial attention mechanisms for comprehensive spatio-spectral feature learning, potentially enhancing holistic tissue characterization. Third, developing adaptive preprocessing modules capable of dynamically adjusting to heterogeneous CEST acquisition protocols, improving model flexibility across imaging platforms. Finally, translating these advancements into clinical practice through prospective trials that rigorously evaluate real-world diagnostic performance, workflow integration, and practical utility in routine healthcare settings.


Conclusions

In this paper, we propose Mamba-AE, a Mamba network-based anomaly detection model for CEST MRI. This model offers robust capability in capturing long-range dependencies and efficiently models CEST spectrum sequences with low computational complexity. Additionally, we introduce a joint loss function that combines a similarity loss between the encoder and decoder with a reconstruction loss in the data space. This design facilitates the model to learn semantically meaningful representations of normal samples, thereby enhancing the discrimination of pathological deviations. We have validated our method on three datasets of rats with cerebral ischemia at different stages, which exhibit significant variations in pathological manifestations, allowing for a comprehensive assessment of the robustness of the proposed method. The results demonstrate that our method achieves the best performance both qualitatively and quantitatively across all datasets, confirming its effectiveness for early disease detection in CEST MRI.


Acknowledgments

None.


Footnote

Reporting Checklist: The authors have completed the TRIPOD reporting checklist. Available at https://qims.amegroups.com/article/view/10.21037/qims-2025-1952/rc

Data Sharing Statement: Available at https://qims.amegroups.com/article/view/10.21037/qims-2025-1952/dss

Funding: This research work was supported by the National Natural Science Foundation of China (NSFC) (grant No. 62201472) and Natural Science Basic Research Program of Shaanxi Province (grant No. 2025JC-YBQN-934).

Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://qims.amegroups.com/article/view/10.21037/qims-2025-1952/coif). The authors have no conflicts of interest to declare.

Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. All animal experiments were performed under a project license (No. 2019SDL092) granted by the Institutional Ethics Committee of Shandong Second Medical University, in compliance with the institutional guidelines for the care and use of animals.

Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.


References

  1. Ward KM, Aletras AH, Balaban RS. A new class of contrast agents for MRI based on proton chemical exchange dependent saturation transfer (CEST). J Magn Reson 2000;143:79-87. [Crossref] [PubMed]
  2. van Zijl PC, Yadav NN. Chemical exchange saturation transfer (CEST): what is in a name and what isn't? Magn Reson Med 2011;65:927-48. [Crossref] [PubMed]
  3. Xu X, Chan KW, Knutsson L, Artemov D, Xu J, Liu G, Kato Y, Lal B, Laterra J, McMahon MT, van Zijl PC. Dynamic glucose enhanced (DGE) MRI for combined imaging of blood-brain barrier break down and increased blood volume in brain cancer. Magn Reson Med 2015;74:1556-63. [Crossref] [PubMed]
  4. Cai K, Singh A, Roalf DR, Nanga RP, Haris M, Hariharan H, Gur R, Reddy R. Mapping glutamate in subcortical brain structures using high-resolution GluCEST MRI. NMR Biomed 2013;26:1278-84. [Crossref] [PubMed]
  5. Cai K, Singh A, Poptani H, Li W, Yang S, Lu Y, Hariharan H, Zhou XJ, Reddy R. CEST signal at 2ppm (CEST@2ppm) from Z-spectral fitting correlates with creatine distribution in brain tumor. NMR Biomed 2015;28:1-8. [Crossref] [PubMed]
  6. Krikken E, Khlebnikov V, Zaiss M, Jibodh RA, van Diest PJ, Luijten PR, Klomp DWJ, van Laarhoven HWM, Wijnen JP. Amide chemical exchange saturation transfer at 7 T: a possible biomarker for detecting early response to neoadjuvant chemotherapy in breast cancer patients. Breast Cancer Res 2018;20:51. [Crossref] [PubMed]
  7. Chen L, van Zijl PCM, Wei Z, Lu H, Duan W, Wong PC, Li T, Xu J. Early detection of Alzheimer's disease using creatine chemical exchange saturation transfer magnetic resonance imaging. Neuroimage 2021;236:118071. [Crossref] [PubMed]
  8. Yu L, Chen Y, Chen M, Luo X, Jiang S, Zhang Y, Chen H, Gong T, Zhou J, Li C. Amide Proton Transfer MRI Signal as a Surrogate Biomarker of Ischemic Stroke Recovery in Patients With Supportive Treatment. Front Neurol 2019;10:104. [Crossref] [PubMed]
  9. Zaiss M, Xu J, Goerke S, Khan IS, Singer RJ, Gore JC, Gochberg DF, Bachert P. Inverse Z-spectrum analysis for spillover-, MT-, and T1 -corrected steady-state pulsed CEST-MRI--application to pH-weighted MRI of acute stroke. NMR Biomed 2014;27:240-52. [Crossref] [PubMed]
  10. Zhou J, Payen JF, Wilson DA, Traystman RJ, van Zijl PC. Using the amide proton signals of intracellular proteins and peptides to detect pH effects in MRI. Nat Med 2003;9:1085-90. [Crossref] [PubMed]
  11. Zaiss M, Schmitt B, Bachert P. Quantitative separation of CEST effect from magnetization transfer and spillover effects by Lorentzian-line-fit analysis of z-spectra. J Magn Reson 2011;211:149-55. [Crossref] [PubMed]
  12. Zhao Y, Wang X, Wang Y, Wang B, Zhang L, Wei X, He X. Application of a Markov chain Monte Carlo method for robust quantification in chemical exchange saturation transfer magnetic resonance imaging. Quant Imaging Med Surg 2022;12:5140-55. [Crossref] [PubMed]
  13. Wang D, Zhang Y, Zhang K, Wang L. Focalmix: Semi-supervised learning for 3d medical image detection. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 13-19 June 2020; Seattle, WA, USA. IEEE; 2020:3950-9. doi: 10.1109/CVPR42600.2020.00401.
  14. Gao HH, Qiu BY, Barroso RJD, Hussain W, Xu YS, Wang XH. TSMAE: A Novel Anomaly Detection Approach for Internet of Things Time Series Data Using Memory-Augmented Autoencoder. IEEE Transactions on Network Science and Engineering 2023;10:2978-90.
  15. Kim J, Kang H, Kang P. Time-series anomaly detection with stacked Transformer representations and 1D convolutional network. Engineering Applications of Artificial Intelligence 2023;120:105964.
  16. Zhang YX, Chen YQ, Wang JD, Pan ZW. Unsupervised Deep Anomaly Detection for Multi-Sensor Time-Series Signals. IEEE Transactions on Knowledge and Data Engineering 2023;35:2118-32.
  17. Zhao H, Li Y, He N, Ma K, Fang L, Li H, Zheng Y. Anomaly Detection for Medical Images Using Self-Supervised and Translation-Consistent Features. IEEE Trans Med Imaging 2021;40:3641-51. [Crossref] [PubMed]
  18. Guo J, Lu S, Jia L, Zhang W, Li H. Encoder-Decoder Contrast for Unsupervised Anomaly Detection in Medical Images. IEEE Trans Med Imaging 2024;43:1102-12. [Crossref] [PubMed]
  19. Jiang A, Huang C, Cao Q, Wu S, Zeng Z, Chen K, Zhang Y, Wang Y. Multi-scale cross-restoration framework for electrocardiogram anomaly detection. International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI); Vancouver, Canada. Springer; 2023. doi:10.1007/978-3-031-43907-0_9.
  20. Bao J, Sun H, Deng H, He Y, Zhang Z, Li X. BMAD: Benchmarks for medical anomaly detection. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW); 17-18 June 2024; Seattle, WA, USA. IEEE; 2024:4042-53. doi: 10.1109/CVPRW63382.2024.00408.
  21. Cai Y, Chen H, Yang X, Zhou Y, Cheng KT. Dual-distribution discrepancy with self-supervised refinement for anomaly detection in medical images. Med Image Anal 2023;86:102794. [Crossref] [PubMed]
  22. Zavrtanik V, Kristan M, Skočaj D. DRAEM-A discriminatively trained reconstruction embedding for surface anomaly detection. IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada. IEEE; 2021:8310-9. doi: 10.1109/ICCV48922.2021.00822.
  23. Li CL, Sohn K, Yoon J, Pfister T. Cutpaste: Self-supervised learning for anomaly detection and localization. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 20-25 June 2021; Nashville, TN, USA. IEEE; 2021:9659-69. doi: 10.1109/CVPR46437.2021.00954.
  24. Hu T, Zhang J, Yi R, Du Y, Chen X, Liu L, Wang Y, Wang C. Anomalydiffusion: Few-shot anomaly image generation with diffusion model. Proceedings of the AAAI Conference on Artificial Intelligence 2024;38:8526-34.
  25. Deng H, Li X. Anomaly detection via reverse distillation from one-class embedding. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 18-24 June 2022; New Orleans, LA, USA. 2022:9727-36. doi: 10.1109/CVPR52688.2022.00951.
  26. Liang Y, Zhang J, Zhao S, Wu R, Liu Y, Pan S. Omni-Frequency Channel-Selection Representations for Unsupervised Anomaly Detection. IEEE Trans Image Process 2023;32:4327-40. [Crossref] [PubMed]
  27. He H, Zhang J, Chen H, Chen X, Li Z, Chen X, Wang Y, Wang C, Xie L. A diffusion-based framework for multi-class anomaly detection. Proceedings of the AAAI Conference on Artificial Intelligence. 2024;38:8472-80.
  28. Marimont SN, Tarroni G. Anomaly detection through latent space restoration using vector quantized variational autoencoders. 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI), Nice, France. IEEE; 2021:1764-7. doi: 10.1109/ISBI48211.2021.9433778.
  29. Baur C, Wiestler B, Albarqouni S, Navab N. Deep autoencoding models for unsupervised anomaly segmentation in brain MR images. In: Crimi A, Bakas S, Kuijf H, Keyvan F, Reyes M, van Walsum T. editors. Brainlesion: Glioma, Multiple Sclerosis, Stroke and Traumatic Brain Injuries. 4th International Workshop, Granada, Spain. Springer; 2018. doi: 10.1007/978-3-030-11723-8_16.
  30. Schlegl T, Seeböck P, Waldstein SM, Langs G, Schmidt-Erfurth U. f-AnoGAN: Fast unsupervised anomaly detection with generative adversarial networks. Med Image Anal 2019;54:30-44. [Crossref] [PubMed]
  31. Akcay S, Atapour-Abarghouei A, Breckon TP. GANomaly: Semi-supervised anomaly detection via adversarial training. Computer Vision–ACCV 2018: 14th Asian Conference on Computer Vision (ACCV), Perth, Australia. Springer; 2018:622-37. doi: 10.1007/978-3-030-20893-6_39.
  32. You Z, Yang K, Luo W, Cui L, Zheng Y, Le X. ADTR: Anomaly detection transformer with feature reconstruction. In: Tanveer M, Agarwal S, Ozawa S, Ekbal A, Jatowt A. editors. International Conference on Neural Information Processing. Springer; 2022. doi:10.1007/978-3-031-30111-7_26.
  33. Meissen F, Paetzold J, Kaissis G, Rueckert D. Unsupervised anomaly localization with structural feature-autoencoders. In: Bakas S, Crimi A, Baid U, Malec S, Pytlarz M, Baheti B, Zenk M, Dorent R. editors. International MICCAI Brainlesion Workshop. 2022. doi: 10.1007/978-3-031-33842-7_2.
  34. Shi Y, Yang J, Qi Z. Unsupervised anomaly segmentation via deep feature reconstruction. Neurocomputing 2020;424:9-22.
  35. You Z, Cui L, Shen Y, Yang K, Lu X, Zheng Y, Le X. A unified model for multi-class anomaly detection. Advances in Neural Information Processing Systems. 2022;35:4571-84.
  36. Su Y, Zhao Y, Niu C, Liu R, Sun W, Pei D. Robust anomaly detection for multivariate time series through stochastic recurrent neural network. Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2019:2828-37. doi: 10.1145/3292500.3330672.
  37. Shen L, Li Z, Kwok J. Timeseries anomaly detection using temporal hierarchical one-class network. Advances in Neural Information Processing Systems 2020;33:13016-26.
  38. Li Z, Zhao Y, Han J, Su Y, Jiao R, Wen X, Pei D. Multivariate time series anomaly detection and interpretation using hierarchical inter-metric and temporal embedding. Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining. 2021:3220-30. doi: 10.1145/3447548.3467075.
  39. Park D, Hoshi Y, Kemp CC. A multimodal anomaly detector for robot-assisted feeding using an LSTM-based variational autoencoder. IEEE Robotics and Automation Letters 2018;3:1544-51.
  40. Hundman K, Constantinou V, Laporte C, Colwell I, Soderstrom T. Detecting spacecraft anomalies using lstms and nonparametric dynamic thresholding. Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. 2018:387-95. doi: 10.1145/3219819.3219845.
  41. Tariq S, Lee S, Shin Y, Lee MS, Jung O, Chung D, Woo SS. Detecting anomalies in space using multivariate convolutional LSTM with mixtures of probabilistic PCA. Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. 2019:2123-33. doi: 10.1145/3292500.3330776.
  42. Kang H, Kang P. Transformer-based multivariate time series anomaly detection using inter-variable attention mechanism. Knowledge-Based Systems 2024;290:111507.
  43. Xu J, Wu H, Wang J, Long M. Anomaly transformer: Time series anomaly detection with association discrepancy. arXiv 2021. doi:10.48550/arXiv.2110.02642.
  44. Tuli S, Casale G, Jennings NR. TranAD: Deep transformer networks for anomaly detection in multivariate time series data. arXiv 2022. doi:10.48550/arXiv.2201.07284.
  45. Gu A, Goel K, Ré C. Efficiently modeling long sequences with structured state spaces. arXiv 2021. doi:10.48550/arXiv:2111.00396.
  46. Smith JT, Warrington A, Linderman SW. Simplified state space layers for sequence modeling. arXiv 2022. doi:10.48550/arXiv:2208.04933.
  47. Mehta H, Gupta A, Cutkosky A, Neyshabur B. Long range language modeling via gated state spaces. arXiv 2022. doi:10.48550/arXiv:2206.13947.
  48. Gu A, Dao T. Mamba: Linear-time sequence modeling with selective state spaces. arXiv 2023. doi:arXiv:10.48550/2312.00752.
  49. Zhu L, Liao B, Zhang Q, Wang X, Liu W, Wang X. Vision mamba: Efficient visual representation learning with bidirectional state space model. arXiv 2024. doi:10.48550/arXiv:2401.09417.
  50. Ruan J, Li J, Xiang S. VM-UNet: Vision mamba unet for medical image segmentation. arXiv 2024. doi:10.48550/arXiv:2402.02491.
  51. Lieber O, Lenz B, Bata H, Cohen G, Osin J, Dalmedigos I, Safahi E, Meirom S, Belinkov Y, Shalev-Shwartz S. Jamba: A hybrid transformer-mamba language model. arXiv 2024. doi:10.48550/arXiv:2403.19887.
  52. Longa EZ, Weinstein PR, Carlson S, Cummins R. Reversible middle cerebral artery occlusion without craniectomy in rats. Stroke 1989;20:84-91. [Crossref] [PubMed]
  53. Kim M, Gillen J, Landman BA, Zhou J, van Zijl PC. Water saturation shift referencing (WASSR) for chemical exchange saturation transfer (CEST) experiments. Magn Reson Med 2009;61:1441-50. [Crossref] [PubMed]
Cite this article as: Zhao Y, You Z, Chen L, Shi Z, Wang L, Zhang Y, Sun L, He X, Wang X. Mamba-AE: pixel-wise anomaly detection auto-encoder in chemical exchange saturation transfer magnetic resonance imaging. Quant Imaging Med Surg 2026;16(3):218. doi: 10.21037/qims-2025-1952

Download Citation