CMMIQA: a prompt-driven cross-modality multi-organ medical image quality assessment model
Introduction
Three-dimensional (3D) medical imaging technology plays an important role in clinical diagnosis. Computed tomography (CT) and magnetic resonance imaging (MRI) are two of the most common 3D medical imaging techniques (1). With the volume of CT and MRI scans, maintaining the quality of these images is essential for effective diagnosis and equipment optimization. More complete anatomical information is provided by excellent images, which aid medical professionals in making precise diagnosis and treatment plans. Inadequate image quality can cause missed or incorrect diagnoses. Consequently, assessing the quality of medical images is quite important.
Medical image quality assessment (IQA) refers to the process of quantitative or qualitative assessment of physical characteristics (noise, spatial resolution, contrast, etc.) and diagnostic value (body position, lesion visibility, etc.) of medical images through subjective expert assessment or objective algorithm analysis. Its core goal is to ensure that the image can clearly and accurately present the anatomical or pathological features to meet the needs of clinical diagnosis, treatment planning or scientific analysis. IQA for CT and MRI generally depends on subjective assessment by medical imaging specialists or physicians. However, dealing with large-scale dataset can be challenging because subjective assessment takes a lot of time. Now, scientists are using computer-assisted instruments or software to assess image quality in an objective, measurable method. To get objective assessment results, measurement metrics like Structural Similarity, peak-signal-to-noise ratio (2), and contrast signal to noise ratio (3) are proposed. According to the availability of high-quality images as a reference, objective assessment can be divided into full reference IQA (FR-IQA), reduced reference IQA (RR-IQA), and no-reference IQA (NR-IQA) (4). The development of FR-IQA techniques is further restricted by the difficulty of obtaining a significant number of high-dose images as a reference and, given the clinical hazards and expenses associated with radiation. Consequently, medical image assessment’s study is focused on NR-IQA.
With deep learning’s ascent in recent years, academics have started automating IQA with deep learning techniques. Lei et al. conducted medical image IQA of various modalities using deep learning and machine learning techniques (5,6). Mason et al. (7) examined 10 FR-IQA methods. Furthermore, Kastryulin et al. (8) assessed the use of 35 distinct IQA techniques in MRI. These techniques included FR-IQA methods and NR-IQA methods. However, due to the model’s weak generalization and robustness, it cannot be used for medical IQA with various modalities and organs. At the same time, the lack of training data is also a major constraint to the development of artificial intelligence (AI) IQA models. On the one hand, manually labeled datasets are expensive; on the other hand, most of the existing models use synthetic two-dimensional (2D) data for inferencing, while clinical collection is mostly 3D sequences, and the accuracy remains to be verified.
To address these challenges, we built the first 3D Cross-Modality Multi-Organ Medical Image Quality Assessment Database, named CMMIQA-DB. The data set construction is mainly based on conventional data, without obvious artifacts or major defects, and has diagnostic value. Based on this, we proposed a hybrid-convolution (Hybrid-conv) model with aware prompts, which focuses on assessing the physical properties and underlying features of images, and provides an accurate scheme for assessing the quality of medical images in different modalities, named CMMIQA. Figure 1 shows an overview of CMMIQA-DB and CMMIQA. Our contributions are summarized below:
- The first 3D cross-modalities multi-organ medical IQA dataset was established, named CMMIQA-DB. The dataset consists of 2,400 brain MRI sequences and breast MRI sequences with different magnetic field strengths, as well as chest CT sequences with different radiation doses. This dataset is created based on real rather than synthetic data.
- We proposed a Hybrid-conv model with aware prompts to provide an accurate scheme for assessing the quality of medical images in different modalities and organs, named CMMIQA. The model extracts image characteristics by Hybrid-conv network, and provides guidance to the model through carefully designed prompts. Finally, the quality scores of each sequence are obtained through multi-layer perceptron (MLP) quality regression.
- Compared with the state-of-the-art (SOTA) models, the proposed model performs well in terms of performance and efficiency, while being able to transfer to other 2D and 3D medical image datasets as a pre-trained model. We expect that CMMIQA will become a fundamental tool for medical IQA.
Methods
Dataset
In the field of medical images, there are still no complete data sets to assess the quality of medical images in different modalities. To fill this gap, we leveraged CT data collected by healthcare institutions and MRI data from public databases to form our cross-modality medical image assessment database, named CMMIQA-DB. The data set composition and related parameters are shown in detail in Table 1. Figure 2 shows sample images from CMMIQA-DB.
Table 1
| Modality | Position | Type | Label | Case | Image number (average) | Slice thickness (mm) | Size (pixels) | Source | Parameters | ||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Exposure (mAs) | Peak kilovoltage (kVp) | Magnetic field strength (T) | Repetition time/echo time (ms) | Parallel imaging | |||||||||
| CT | Chest | Lung window | 0 | 200 | 314 | 1 | 512×512 | In-house | 30 | 120 | – | – | – |
| 1 | 200 | 225 | 1.5 | 512×512 | 50 | 120 | – | – | – | ||||
| 2 | 200 | 278 | 2 | 512×512 | 150 | 120 | – | – | – | ||||
| 3 | 200 | 330 | 1 | 512×512 | 200 | 120 | – | – | – | ||||
| Soft tissue window | 0 | 200 | 64 | 5 | 512×512 | 30 | 120 | – | – | – | |||
| 1 | 200 | 61 | 5 | 512×512 | 50 | 120 | – | – | – | ||||
| 2 | 200 | 63 | 5 | 512×512 | 150 | 120 | – | – | – | ||||
| 3 | 200 | 62 | 5 | 512×512 | 200 | 120 | – | – | – | ||||
| MRI | Brain | T1 | 0 | 200 | 170 | 1.2 | 256×256 | ADNI | – | – | 1.5 | 8.6/4 | √ |
| 1 | 200 | 170 | 1.2 | 256×256 | – | – | 3 | 6.8/3.2 | √ | ||||
| Breast | T1 | 0 | 200 | 66 | 2; 3 | 320×320; 512×512 | Duke | – | – | 1.5 | 450–837/7.856–14.58 | × | |
| 1 | 200 | 60 | 2; 3 | 256×256; 512×512 | – | – | 3 | 500–781/7.9–10.136 | × | ||||
√, parallel imaging; ×, non-parallel imaging. ADNI, Alzheimer’s Disease Neuroimaging Initiative; CMMIQA-DB, Cross-Modality Multi-Organ Medical Image Quality Assessment Database; CT, computed tomography; MRI, magnetic resonance imaging; T, Tesla.
First, we selected 1,600 3D chest CT data from Shenzhen People’s Hospital (in-house dataset). The images collected in the data set strictly followed the CT examination operation procedure, and were collected according to the scanning position (supine position), range (thoracic entrance-posterior costal septum angle, field of view is about 320 mm) and scanning conditions (120 kV, automatic exposure control) as required by the standardized scanning protocol of routine chest scan. In the process of CT scanning, the higher the radiation dose, the lower CT image noise, and the easier it is to perceive low-contrast structures, that is, the higher the image quality. The radiation dose is determined by the tube’s peak kilovolt (kvp), slice scanning time, and tube current. In it, the tube current and slice scanning time together act as mAs related to radiation dose and image quality. Increasing mAs increases the dose proportionally, so the greater the mAs and the higher the dose, the higher the corresponding image quality (9). The mAs levels we selected (30/50/150/200 mAs) covered the clinical range from low-dose screening to diagnostic protocols, fixing 120 kVp to control mAs effects. Data are derived from different scan instances (not within the same scanning object). The optimal dose is selected according to the patient’s body size and tissue density of the scan site through automatic exposure control to reduce the impact of patient’s weight and anatomical structure variability on image quality. The image exposure within the same sequence is consistent. In the process of CT image reconstruction, clear images with different contrasts can be obtained by adjusting window width (WW), window level (WL) and post-processing function to further help diagnosis. The CT value is often called the Hounsfield unit (HU) and reflects the degree of X-ray absorption by the tissue. We selected lung window sequence and soft tissue window sequence, two commonly used sequences in clinical diagnosis. The lung window sequence (WW 1,500–2,000 HU, WL 450–600 HU) optimized and sharpening of image edges for pulmonary parenchyma visualization, enhancing contrast between airways and alveoli structures. Soft tissue window sequences (WW 300–500 HU, WL 40–60 HU) designed for mediastinal analysis, emphasizing subtle density differences in cardiac/vascular structure. For every type and dose, there were 200 sequences.
For MRI data, the magnetic field strength is an accurate measure of MRI image quality, usually expressed in Tesla (T). It has a major impact on the signal-to-noise ratio and spatial resolution, and the higher the magnetic field strength, the higher the image quality (10). In the clinic, 1.5T and 3T magnetic fields strength are most frequently utilized. As a result, we chose 3D MRI sequences of the breast from the Duke-Breast-Cancer-MRI (11) data set and 3D MRI sequences of the brain from the Alzheimer’s Disease Neuroimaging Initiative (ADNI) MRI (12) data set, all of which had magnetic field strengths of 1.5T and 3T. All data were T1 sequences, and only non-contrast sequences were selected to eliminate signal changes caused by contrast agents. All scanning parameters follow the scanning protocol developed by ADNI and Duke. For each type and magnetic field strength, there are 200 sequences.
We use radiation dose and magnetic field strength as labels for model training. On the one hand, expert annotations are scarce, and there are no publicly available real datasets with manual annotations. We train the model using these two labels, as a pre-training step, which can be finetuned with expert annotation datasets for downstream tasks. “Cross-dataset validation” section also demonstrates the validity of this training strategy. On the other hand, because of the settings the dose and magnetic field strength might be missing in the Digital Imaging and Communications in Medicine (DICOM) tags during CT and MRI scans. At the same time, in existing public datasets, such information is often missing because of the process of data anonymization. In addition, many other medical image formats, such as ‘Nifti’, do not have provide the relevant information. In certain cases, this prediction is still be useful.
By integrating CT and MRI data, we created a comprehensive dataset of 2,400 samples. It is of great significance to develop good medical IQA method. To the best of our knowledge, this is the first quality assessment dataset constructed from real rather than synthetic 3D medical image data. This study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study was approved by Shenzhen Municipal People’s Hospital Clinical Research Ethics Committee (No. EC-20200226-1018) and individual consent for this retrospective analysis was waived.
Model
The model extracts the features of images in various modalities using the Hybrid-conv network. It then includes the image type information of various modalities to direct the network during quality regression, and ultimately delivers the results of the image quality prediction. The network structure is depicted in Figure 3. The feature extraction module is specifically divided into five parts, with the four primary layers being the convolution layer, BatchNorm layer, Nonlinear layer, and Maximum pooling layer included in each portion. Of them, the final two sections use hyper-convolution (Hyper-conv), whereas the first three partial convolution layers use traditional 3D convolution layers. We create an MLP module for quality prediction outcomes after feature extraction. The MLP module includes three fully connected (FC) layers, ReLU activation layers and Dropout layers.
Prompt strategy
We use three different prompts to inform the model: modality, position, and type. Each prompt provides special insights into the model, which will effectively improve the explainability of the model. Among them, the modality (CT/MRI) prompts the image modality to help the model understand the different gray features of CT and MRI images. Position (chest/brain/breast) indicates the anatomical position to which the image belongs and helps the model understand the context of the image based on the anatomical structure. Type (lung window/soft tissue window/T1) indicates the type of the image and helps the model understand the semantic features of different types of images.
The prompt strategy consists of several key steps: first, we take three corresponding prompts based on the input image, and obtain the different prompt codes using One-Hot Encoding (13). After concatenation these prompt codes, use the FC layer to project into the same dimensions as the Hyper-conv layer, ensuring compatibility. At this point, the prompt codes will be converted to information code and processed by each subsequent Hyper-conv layer. The implementation of information code can be expressed as:
where info represents information code, we expect to use it to control the output of the convolution layer.
Hybrid-conv
In feature extraction module, we use traditional convolution layer and Hyper-conv superposition to extract features. The goal of Hyper-conv is to dynamically generate convolution kernel through information code to realize shared parameter learning and lightweight adaptation of different image sequences. We further refined its implementation, motivated by Han et al. (14), to obtain a more compact module Hyper-conv layer, as shown in Figure 3. Its core ideas include:
- Shared weight bank. All sequences share a trainable weight bank (param) that stores all possible convolution kernel base vectors, and the convolution kernel for a specific sequence is generated through subsequent steps. Avoiding storing convolution kernels independently for each sequence can significantly reduce the number of parameters.
where Cin and Cout are convolution layer of input and output channel number, K as the convolution kernel size, Cw is weight bank dimension, represents the capacity of the weight bank. - Mapping function. The information code (info) obtained by the prompt strategy learns its association with the weight bank through the MLP, and predicts a mapping function f for info, maps info as a hidden space vector of the weight bank, and uses it to combine the weight bank.
where WMLP is the linear layer weight and bMLP is the bias. - Dynamic convolution kernel generation. Through the mapping function f predicted by info, each base kernel in the weight bank is weighted and summed with f to generate a specific kernel of the target sequence.
where is the matrix product. Finally, the implementation of the Hyper-conv layer is as follows:
where x is the input feature, y is the output feature.
Hyper-conv can effectively improve the accuracy and generalization speed of the model, and accelerate the convergence speed of the model. Despite the obvious advantages of Hyper-conv, we still chose a combination of traditional convolution and Hyper-conv. This is because using Hyper-conv too early will prevents the model from fully learning its underlying features, which reduce the effectiveness of the model. The relevant experiments will be introduced in “Hybrid-conv” section. The design significantly reduces the number of parameters in the model while maintaining flexibility for different image sequences, making the model suitable for cross-modality multi-organ IQA tasks.
Quality regression
We use an MLP composed of three FC layers to regression the extracted image features to the image quality score. MLP maps the sequence sharing features and specific features extracted by the Hybrid-conv module to the unified quality scoring space to solve the problem of consistency in quality assessment across modality images.
Among them, the first FC layer (FC1, 4,096 neurons) maps high-dimensional inputs to 4,096-dimensional space for extracting global statistical features (noise, texture, etc.), and the higher number of neurons enables the preservation of critical information while avoiding information loss. The second FC layer (FC2, 1,024 neurons) is further compressed to 1,024 dimensions, removing redundant features by dimensionality reduction, preserving quality-related features (edge sharpness, contrast, etc.), and enhancing the robustness of the model. The third FC layer (FC3, 1 neuron) maps the features to a single quality score, synthesizing all the information to generate the final prediction. MLP uses the ReLU activation function to solve the nonlinear mapping problem between high-dimensional features and quality scores, helping the model to fit complex quality cases. At the same time, the Dropout layer (drop rate of 0.5) is added to prevent the model from overfitting. The regression process of the predicted score is given by the following formula:
where x is the input feature, , , , , , , d is the input feature dimension, represents the ReLU activation function, and represent Dropout.
It is worth noting that even though our label values are discrete, we still chose the regression task. On the one hand, the kinds of quality labels are not the same for images of different modalities (for example, CT data contains four labels and MRI data contains only two labels). The design of the classification task will cause the redundancy of multiple classification heads and reduce the efficiency of the model. The regression task allows the same output header to be shared across modalities, avoiding the increase in model complexity caused by modality differences. On the other hand, although dose and magnetic field strength are positively correlated with image quality, these parameters cannot fully characterize image quality. Regression scores rather than classification grades were chosen as the model task to distinguish images of different quality at the same dose level and to more sensitively capture quality differences in subclasses (e.g., 0.2 points indicate quality characteristics between classes 1–2). Hence, we uniformly use quality regression to get the final predicted score, rather than a simple classification of ordinals.
At the same time, if the future introduction of physician subjective score (continuous), regression framework can be directly compatible with our cross-validation data sets (“Cross-dataset validation” section proves this point), and the classification task needs to redraw the threshold. In addition, the model can be classified by adjusting the threshold value according to the demand after the output of continuous values (for example, ≤0.2 points can be regarded as a poor quality category), which ensures the flexibility of the model.
Loss function
The mean square error (MSE) loss function is used to learn this model. The MSE loss function helps to optimize the model by punishing the MSE between the predicted value and the true value. The smaller the value of MSE, the smaller the difference between the predicted value and the true value of the model, and the better the performance of the model. Our labels are ordered quality levels (0 to 3 indicates low to high quality), and MSE helps the model learn the relative distance between different quality levels, the calculation formula is as follows:
where n is the number of samples, and is the true value and the predicted value of the model.
Training procedure
Implement details
When training and testing on CMMIQA-DB, we follow the common practice of dataset splitting by leaving out 80% for training and 20% for testing. To eliminate the bias in one single split, we use 5-fold cross validation.
The model is built on the Python3.9 environment and the PyTorch framework. We use Adam optimizer initialized by a learning rate of 0.0001. We train CMMIQA for 50 epochs under a batch size of 1 on a server with an Intel(R) Xeon(R) W2245 CPU @ 3.90 GHz, NVIDIA RTX A6000 computer.
Evaluation metrics
We employed the Spearman rank correlation coefficient (SRCC), Pearson linear correlation coefficient (PLCC), and root MSE (RMSE), which are the three metrics most frequently used to quantify the IQA model, for our model evaluation (15). SRCC runs a linear correlation analysis on the rank sizes of the two target arrays to assess the monotonicity of the prediction made by the IQA algorithm. PLCC quantifies the precision of the predictions made by the IQA algorithm and characterizes the linear correlation between two sets of data. RMSE computes a measure of the absolute error between the two sets of data to forecast the expected consistency of two sets of data. The calculation formulas are as follows:
[9]]
where and are the reference quality scores and predicted quality scores respectively, and are the average values of the reference quality scores and predicted quality scores respectively, N represents the number of samples, and represents the difference between the reference quality scores and predicted quality scores of the i-th image.
Results
Benchmark performance
The proposed model was trained and tested on CMMIQA-DB as well as six sub-test sets (the CT dataset, MRI dataset, lung window dataset, soft tissue window dataset, brain dataset, and breast dataset). The sub-test sets were randomly sampled stratified by dose/magnetic field strength to ensure a balanced sample for all quality levels, with each subset containing 100 test samples. Table 2 displays the results of the experiment. Figure 4 shows the visualized results of the model. At the same time, we tested subsets and set 0.5 as the threshold, with test results above 0.5 classified as “Good” and below 0.5 classified as “Bad”, hoping to show our experimental results more clearly. The confusion matrix diagram is shown in Figure 5. As shown in the figure, the result of confusion matrix is consistent with the result of correlation analysis. The experimental results show that the model has strong correlation and low prediction error on the whole test set. In the sub-test set, MRI data results were significantly better than CT data, brain MRI data performed best, and lung window data performed relatively poor. This shows that the model performance difference is related to the task complexity.
Table 2
| Dataset | SRCC | PLCC | RMSE |
|---|---|---|---|
| Test | 0.7476 | 0.7153 | 0.3345 |
| CT | 0.7095 | 0.7023 | 0.3327 |
| MRI | 0.8154 | 0.7949 | 0.3104 |
| CT-lung | 0.7010 | 0.6977 | 0.3447 |
| CT-soft | 0.7168 | 0.7474 | 0.3181 |
| MRI-breast | 0.7878 | 0.7647 | 0.3416 |
| MRI-brain | 0.8344† | 0.8287† | 0.2923† |
†, the best-performing model. CMMIQA, Cross-Modality Multi-Organ Medical Image Quality Assessment; CT, computed tomography; MRI, magnetic resonance imaging; SRCC, Spearman rank correlation coefficient; PLCC, Pearson linear correlation coefficient; RMSE, root mean square error.
Ablation studies
To verify the effectiveness of each module in CMMIQA, we conducted a comprehensive ablation experiment, including Hybrid-conv and prompt strategy. The results of these ablation experiments are shown in Table 3. All results were averaged by five-fold cross-validation. Model 1 is constructed using only stacked convolution layers and does not add Hyper-conv and prompt as our baseline. From Model 2 to Model 6, 1–5 Hyper-conv layers are added respectively, which is used to verify the effectiveness of Hybrid-conv network. The larger the number, the more Hyper-conv layers are added. Model 7 adds a prompt strategy on Model 1. Model 8 adds a prompt strategy on Model 3. The final model, CMMIQA, is the complete model we proposed, where we combine the best results from Hybrid-conv and prompt strategy. Based on the results of these 9 models, we can analyze the contribution of each module in CMMIQA.
Table 3
| Model | Hybrid-conv (%) | Prompts | SRCC | PLCC | RMSE |
|---|---|---|---|---|---|
| 1 | 0 | × | 0.5178 | 0.5506 | 0.3740 |
| 2 | 20 | × | 0.6146 | 0.5607 | 0.3850 |
| 3 | 40 | × | 0.7207 | 0.7039 | 0.3682 |
| 4 | 60 | × | 0.6814 | 0.6686 | 0.3754 |
| 5 | 80 | × | 0.6025 | 0.6011 | 0.3563 |
| 6 | 100 | × | 0.5498 | 0.5315 | 0.4074 |
| 7 | 0 | √ | 0.6238 | 0.6197 | 0.3885 |
| 8 | 100 | √ | 0.7254 | 0.6576 | 0.5235 |
| CMMIQA | 40 | √ | 0.7476† | 0.7153† | 0.3345† |
†, the best-performing model. √, add prompts; ×, no additional prompts. CMMIQA-DB, Cross-Modality Multi-Organ Medical Image Quality Assessment Database; Hybrid-conv, hybrid-convolution; SRCC, Spearman rank correlation coefficient; PLCC, Pearson linear correlation coefficient; RMSE, root mean square error.
Hybrid-conv
For CMMIQA, a significant improvement lies in the method of feature extraction. First, we tested the effect of Hybrid-conv on the model. We added different numbers of Hyper-conv layers in different locations, and Model 1 to Model 6 evaluated the effectiveness of Hybrid-conv. It should be noted that in order to avoid redundant data presentation, only the experimental results of the corresponding number of Hyper-conv layers in the optimal position are shown in Table 3. The experimental results show that the addition of Hyper-conv layers can improve the performance of the model. However, in the process of the experiment, we found that if the Hyper-conv layer is applied too early, the neural network will have difficulty discovering the basic features of the input image. Therefore, the best choice is to combine the Hyper-conv layer with the last two feature extraction layers.
Prompt strategy
At the same time, we study the influence of prompt strategy on the model. We added prompt strategy to Model 7, Model 8, and CMMIQA. The experimental results show that the addition of prompts can effectively enhance the ability of the network to detect different image types, and enhance the prediction effect of the model on the quality of cross-modality multi-organ medical images.
It is noteworthy that with the same experimental setting and hyper-parameters, the feature extraction module with the Hyper-conv layer converges faster in the training and inferencing process, saving a third of the training process.
Performance comparison
Reference algorithms
To verify the validity of the CMMIQA, BRISQUE (16), DBCNN (17), CNNIQA (18), PaQ2PiQ (19), CLIP (20), HyperIQA (21) six NR-IQA methods were used as reference algorithms for performance comparison. BRISQUE is a no-reference spatial IQA algorithm. DBCNN uses deep bilinear convolutional neural networks for blind IQA. CNNIQA, as its name suggests, uses convolutional neural network for NR-IQA of images. PaQ2PiQ provides a comprehensive assessment of global and local images. CLIP assessed the quality perception and abstract perception of images in a zero-shot manner and, likewise, incorporated effective prompts. Hyper-IQA propose a self-adaptive hyper network architecture to blind assess image quality. We trained all the models using the same training strategy and ensured their convergence.
Results analysis
Figure 6 shows a comparison of the performance and efficiency of CMMIQA and other SOTA models in the full test set. We compared SRCC, PLCC and RMSE, and calculated the inference time of an average single 3D medical image sequence. The results show that CMMIQA achieves the best performance and the highest efficiency. In addition, in the process of the experiment, we found that the results of the neural network-based IQA method are generally better than the traditional IQA method. This is further evidence of how beneficial deep learning techniques can be for assessing the quality of medical images.
Cross-dataset validation
Most of the existing medical image IQA methods are based on synthetic 2D data training. To further verify the validity of our model, we retrained the network using a publicly available synthetic dataset (22). The dataset consisted of 1,000 abdominal soft tissue window images labeled by five professional radiologists. In order to accept 2D image as input, the max-pool layer in CMMIQA is adjusted synchronously with the network depth modification. For effective comparison, we chose those medical image IQA methods that use the same type of image for training (23-25). Table 4 shows the experimental results. Since the code and training data for the comparison method are not open source, we strictly use the metric values reported in their original paper for comparison. Metrics not covered or with inconsistent scales were marked with “N/A” to avoid cross-literature statistical bias. Our model adopted two training strategies: (I) retraining with random initial weights (CMMIQA-1); (II) re-training with weights of CMMIQA-DB pre-training (CMMIQA-2). The experimental results show that our model is also competitive in synthesizing 2D dataset. The results of CMMIQA-2 are also better than those of CMMIQA-1. This suggests that pre-trained CMMIQA can also be efficiently transfer to other 2D or 3D medical image datasets as a foundation model.
Table 4
| Model | SRCC | PLCC | RMSE |
|---|---|---|---|
| Ref1 (23) | 0.8302 | 0.8770† | N/A |
| Ref2 (24) | 0.8040 | N/A | N/A |
| Ref3 (25) | 0.7480 | 0.7540 | N/A |
| CMMIQA-1 | 0.8523‡ | 0.8430 | 0.1622‡ |
| CMMIQA-2 | 0.8618† | 0.8705‡ | 0.1589† |
“N/A” indicates that the metric is not reported in the original paper or that the quality scale used is inconsistent with this study. †, the best-performing model; ‡, the second-best performing model. CMMIQA, Cross-Modality Multi-Organ Medical Image Quality Assessment; Ref, reference; SRCC, Spearman rank correlation coefficient; PLCC, Pearson linear correlation coefficient; RMSE, root mean square error.
Discussion
As AI technology advances, methods based on deep learning to assess image quality are starting to be used in the medical domain (5,6). Researchers have synthesized 2D medical images of varying quality for neural network training by adding noise to real data, and they have seen some success (23-25). Nonetheless, the synthesized data’s image quality is easy to distinguish, and the majority of the medical data utilized in the medical institutions consists of 3D image sequences. Simultaneously, the majority of the current models are designed to assess images in a single modality, making them incapable of adjusting to the intricate context of several modalities and multiple scanning locations found in clinical settings.
Our proposed CMMIQA-DB and CMMIQA provide a promising direction for cross-modality multi-organ 3D medical IQA. We construct a 3D cross-modality multi-organ medical IQA dataset based on real rather than synthetic data for the first time. Based on this, a Hybrid-conv model with aware prompt is proposed. The model dynamically generates convolution kernel for images of different modalities, organs and types by means of prompt strategy and Hyper-conv layer, which can simultaneously process multiple types of data and avoid the limitation of the traditional single-task model requiring repeated training. The experimental results show that the model is stable in many kinds of data sets, and can realize efficient assessment under the premise of ensuring accuracy. At the same time, the model can be transferred to other 2D and 3D medical image data sets as a pre-training model, showing good cross-modality generalization ability and multi-organ adaptability.
However, our dataset and model are not without flaws. Firstly, our dataset only included CT and MRI scans, and the generalization of ultrasound, X-ray and other modalities needs to be further verified. In the future, we plan to increase our cross-modality multi-organ image collection. Furthermore, even though there is a strong positive correlation between dose and magnetic field strength and image quality, other parameters like artifact and body position must also be taken into account in clinical practice. Therefore, the benchmark for network training should be a quality score that is assessed subjectively by a physician or imaging professional. This will grow to be a significant area of research for us in the future.
In clinical work, CMMIQA can be integrated into the medical system to help screen low-quality images and remind them to re-scan, reducing the cost of manual re-examination. The model’s cross-modality, lightweight design facilitates deployment in different environments, such as multi-device joint platforms, driving standardized quality assessment. CMMIQA can be used as a pre-processing module to embed other frameworks for data pre-screening, thereby helping to improve the performance of downstream task models such as segmentation and classification. Similarly, it can also be used as a post-processing module to evaluate the model performance of downstream tasks such as reconstruction and denoising. We expect CMMIQA to become a basic tool for medical IQA.
Conclusions
In this paper, we are committed to assess the quality of real 3D cross-modality multi-organ medical image. For this purpose, we built a dataset consist of 2,400 medical sequences named CMMIQA-DB. The dataset contains medical sequences of different modalities, organs and types. Based on CMMIQA-DB, a scheme to accurately assess the quality of medical images of different modalities is proposed. The core of this scheme is to use Hybrid-conv to extract image features under the guidance of prompt strategy to regression prediction scores. The experimental results show that CMMIQA has great potential in generalization, accuracy and inference efficiency compared with existing methods, while being able to transfer to other 2D and 3D medical image datasets as a pre-training model. We envision that CMMIQA will become a powerful tool for medical IQA.
Acknowledgments
None.
Footnote
Funding: This work was supported by
Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://qims.amegroups.com/article/view/10.21037/qims-2025-127/coif). The authors have no conflicts of interest to declare.
Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. This study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study was approved by Shenzhen Municipal People’s Hospital Clinical Research Ethics Committee (No. EC-20200226-1018) and individual consent for this retrospective analysis was waived.
Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.
References
- van Beek EJ, Hoffman EA. Functional imaging: CT and MRI. Clin Chest Med 2008;29:195-216. vii. [Crossref] [PubMed]
- Hore A, Ziou D. Image quality metrics: PSNR vs. SSIM. In: 2010 20th international conference on pattern recognition. IEEE; 2010:2366-9.
- Yao S, Lin W, Ong E, Lu Z. Contrast signal-to-noise ratio for image quality assessment. In: IEEE International Conference on Image Processing 2005. IEEE; 2005;1:397-400.
- Chow LS, Paramesran R. Review of medical image quality assessment. Biomedical Signal Processing and Control 2016;27:145-54.
- Lei K, Syed AB, Zhu X, Pauly JM, Vasanawala SS. Artifact- and content-specific quality assessment for MRI with image rulers. Med Image Anal 2022;77:102344. [Crossref] [PubMed]
- Liu S, Thung KH, Lin W, Shen D, Yap PT. Hierarchical Nonlocal Residual Networks for Image Quality Assessment of Pediatric Diffusion MRI With Limited and Noisy Annotations. IEEE Trans Med Imaging 2020;39:3691-702. [Crossref] [PubMed]
- Mason A, Rioux J, Clarke SE, Costa A, Schmidt M, Keough V, Huynh T, Beyea S. Comparison of Objective Image Quality Metrics to Expert Radiologists' Scoring of Diagnostic Quality of MR Images. IEEE Trans Med Imaging 2020;39:1064-72. [Crossref] [PubMed]
- Kastryulin S, Zakirov J, Pezzotti N, Dylov DV. Image quality assessment for magnetic resonance imaging. IEEE Access 2023;11:14154-68.
- Payne JT. CT radiation dose and image quality. Radiol Clin North Am 2005;43:953-62. vii. [Crossref] [PubMed]
- Wu CW, Chuang KH, Wai YY, Wan YL, Chen JH, Liu HL. Vascular space occupancy-dependent functional MRI by tissue suppression. J Magn Reson Imaging 2008;28:219-26. [Crossref] [PubMed]
- Chitalia R, Pati S, Bhalerao M, Thakur SP, Jahani N, Belenky V, McDonald ES, Gibbs J, Newitt DC, Hylton NM, Kontos D, Bakas S. Expert tumor annotations and radiomics for locally advanced breast cancer in DCE-MRI for ACRIN 6657/I-SPY1. Sci Data 2022;9:440. [Crossref] [PubMed]
- Wyman BT, Harvey DJ, Crawford K, Bernstein MA, Carmichael O, Cole PE, et al. Standardization of analysis sets for reporting results from ADNI MRI data. Alzheimers Dement 2013;9:332-7. [Crossref] [PubMed]
- Rodríguez P, Bautista MA, Gonzalez J, Escalera S. Beyond one-hot encoding: Lower dimensional target embedding. Image and Vision Computing 2018;75:21-31.
- Han L, Tan T, Zhang T, Huang Y, Wang X, Gao Y, Teuwen J, Mann R. Synthesis-based imaging-differentiation representation learning for multi-sequence 3D/4D MRI. Med Image Anal 2024;92:103044. [Crossref] [PubMed]
- Hu B, Li L, Wu J, Qian J. Subjective and objective quality assessment for image restoration: A critical survey. Signal Processing: Image Communication 2020;85:115839.
- Mittal A, Moorthy AK, Bovik AC. No-reference image quality assessment in the spatial domain. IEEE Trans Image Process 2012;21:4695-708. [Crossref] [PubMed]
- Wang J, Chan KC, Loy CC. Exploring clip for assessing the look and feel of images. Proceedings of the AAAI Conference on Artificial Intelligence 2023;37:2555-63.
- Khmag A, Kamarudin N. Natural image deblurring using recursive deep convolutional neural network (R-DbCNN) and second-generation wavelets. In: 2019 IEEE International Conference on signal and image processing applications (ICSIPA). IEEE; 2019:285-90.
- Su S, Yan Q, Zhu Y, Zhang C, Ge X, Sun J, Zhang Y. Blindly assess image quality in the wild guided by a self-adaptive hyper network. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2020:3667-76.
- Ying Z, Niu H, Gupta P, Mahajan D, Ghadiyaram D, Bovik A. From patches to pictures (PaQ-2-PiQ): Mapping the perceptual space of picture quality. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2020:3575-85.
- Li R, Yang H, Yu T, Pan Z. CNN model for screen content image quality assessment based on region difference. In: 2019 IEEE 4th International Conference on Signal and Image Processing (ICSIP). IEEE; 2019:1010-4.
- Lee W, Wagner F, Maier A, Wang A, Baek J, Hsieh SS, Choi JH. Low-dose Computed Tomography Perceptual Image Quality Assessment Grand Challenge Dataset. In: Medical Image Computing and Computer Assisted Intervention; 2023.
- Gao Q, Zhu M, Li D, Bian Z, Ma J. CT image quality assessment based on prior information of pre-restored images. Nan Fang Yi Ke Da Xue Xue Bao 2021;41:230-7. [Crossref] [PubMed]
- Duan J, Cai J, Zhi S, Mou X. Blind CT Image Quality Assessment Model Based on CT Image Statistics. In: The Fourth International Symposium on Image Computing and Digital Medicine; 2020:201-5.
- Chen Z, Hu B, Niu C, Chen T, Li Y, Shan H, Wang G. IQAGPT: Image Quality Assessment with Vision-language and ChatGPT Models. arXiv 2023. arXiv:2312.15663.


