Quantitative generation of diffusion-weighted imaging from non-contrast CT in acute ischemic stroke: a multi-scanner deep-learning model with external validation
Introduction
Stroke, an acute cerebrovascular disease, is the second leading cause of disability and death on a global scale. It has waged an alarming $34 billion economic toll each year. The majority of strokes are of an ischemic nature, accounting for over 87% of all cases, while the remainder are hemorrhagic (1). When an acute ischemic stroke (AIS) occurs, immediate and accurate diagnosis is essential not to miss the best time to treat these patients. Early reperfusion therapies, including intravenous thrombolysis (IVT) and endovascular treatment (EVT), have been established as the two most effective interventions in AIS (2,3). IVT is known to substantially decrease the long-term risk of disability after an AIS. Nevertheless, the IVT window is confined to 4.5 hours (extendable to 6 hours in highly sophisticated centers). Thrombolytic therapy cannot be administered if the delay exceeds the specified therapeutic window (4,5). It is worth noting that a delay of one minute in the door-to-needle time (DNT) has a strong relationship with the 90-day survival rate, the rate of intracranial hemorrhage (ICH) within 36 hours, and functional outcomes at the third month, such as activities of daily living (ADL), domicile, and motor power (6).
Neuroimaging, including computed tomography (CT) and magnetic resonance imaging (MRI), helps assess pre-thrombolysis acute stroke (7). However, it is essential to note that these imaging methods are not without limitations. The artifacts observed in CT and MRI images have been also found to simulate stroke lesion intensity and morphology, making the diagnosis exceptionally challenging for neuroradiologists (8,9). Non-contrast CT (NCCT) is an easily accessible and fast imaging technique, but it has low contrast sensitivity, which prevents the visualization of 3–9% of AIS (10,11). The early ischemic changes, often subtle and represented by localized brain edema and hypodensity, are challenging to identify even with CT scans and may be masked by normal physiological processes or stable old lesions (12). On the other hand, diffusion-weighted imaging (DWI) is increasingly recognized as the gold standard for AIS diagnosis. The device has a high sensitivity range of 73% to 92% for detecting hyperacute ischemic strokes within the first 3 hours from symptom onset, and this efficacy is maintained beyond six hours (13,14). Yet, DWI may not always be readily available in emergent clinical settings. A DWI scan, which is not acquired at the time of first availability, delays treatment and worsens patient outcomes. Therefore, the ability to recapitulate AIS findings on DWI from CT images would provide a great deal of utility in early patient management.
Recent advancements in artificial intelligence, particularly deep learning, have revolutionized medical imaging technologies (15). In addition to improving diagnostic accuracy, deep learning methods have shown promise in speeding up clinical workflows and decision-making (16-20).
A line of deep learning approaches has gained attention in stroke imaging, helping to improve diagnosis, prognosis, and lesion delineation. Numerous studies have been conducted on the detection and prediction of ischemic stroke outcome using multimodal MRI and radiomic features. For instance, Yang et al. reported an MRI-based deep learning biomarker of 3-month functional outcomes in the cohort of AIS patients with an area under the receiver operating characteristic curve (AUC) similar to conventional clinical risk scores with further performance gains when these scores were combined (21).
A segment of this research has recently begun to focus on stroke detection, primarily through the use of artificial intelligence to enhance vascular imaging and DWI segmentation (22-29). Kuang et al. propose a hybrid convolutional neural network-Transformer (CNN-Transformer) model to segment AIS lesions on NCCT, effectively capturing both local and global features and outperforming state-of-the-art methods (30). Gheibi et al. developed a CNN-Res framework for multimodal MRI lesion segmentation in which they demonstrate that a well-rounded pairwise architecture can jointly model local as well as global information without adding to network complexity (31). Kuang et al. in their work presented the I2PC-Net, utilizing multiscale convolution and symmetry-enhancement mechanisms that enable robust ischemic lesion segmentation from non-enhanced CT scans (32).
Beyond diagnostic and segmentation tasks, deep learning has also been extensively studied for medical image translation between modalities. A variety of generative adversarial network (GAN)-based and transformer-based frameworks have been proposed to synthesize images across modalities for clinical analysis, data harmonization, and interpretation. For example, transformer-enhanced GANs such as MMTrans have been developed to exploit global context while translating multi-modal medical images, achieving improved synthesis quality on both aligned and unpaired datasets compared with conventional models (33).
In the context of CT and MRI translation specifically, several studies have begun to explore direct cross-modality synthesis. A modified CycleGAN with attention guidance and gradient consistency regularization was proposed to translate non-contrast CT to fluid-attenuated inversion recovery (FLAIR) MRI in unpaired stroke image datasets, explicitly preserving lesion morphology during translation and demonstrating effective harmonization between CT and MRI modalities (34). Another study has employed deep learning models such as CycleGAN variants to establish bidirectional mappings between CT and MRI images, aiming to generate realistic MRI images from CT data for improved diagnostic support when MRI is unavailable. Song et al. proposed a generative model-based framework that synthesizes follow-up DWI from CT perfusion data to improve ischemic stroke lesion segmentation, achieving a Dice score of 62.4% by leveraging the generated DWI for downstream analysis (35). However, their approach relies on CT perfusion imaging and focuses primarily on segmentation performance, whereas the present study investigates direct quantitative DWI synthesis from widely available non-contrast CT.
These works illustrate the growing interest in CT-to-MRI translation; however, most previous studies have focused on synthesizing structural or anatomical contrasts rather than diffusion-weighted contrasts particularly relevant to AIS. Direct end-to-end synthesis of NCCT to diffusion-weighted MRI remains underexplored. Therefore, in this study, we propose a modified ControlNet model to enable more faithful diffusion signal reconstruction from NCCT and systematically evaluate its performance against existing DWI synthesis methods.
Nonetheless, the potential benefit of translating images from NCCT to DWI end-to-end directly for AIS patients has not been well studied. This study employs the modified ControlNet model to account for it and allow more faithful reconstruction of diffusion-weighted magnetic resonance imaging data. A rule-based model aims to utilize the fast imaging of NCCT to provide a rapid and high-accuracy diagnostic pathway, while using the high contrast provided by DWI. The proposed framework has been systematically evaluated using comprehensive image-quality metrics on whole-brain and lesion-specific scans. Moreover, extensive comparisons were made with other deep learning-based DWI synthesis methods. These in vivo results are consistent with the notion that this new model has high potential to profoundly revolutionize AIS management by significantly reducing time-to-intervention and improving lesion identification precision.
In this paper, we aim to investigate a problem that involves converting NCCT images of cerebral ischemic stroke patients into DWI images using deep learning models, thereby facilitating rapid and early screening for AIS. We present this article in accordance with the TRIPOD+AI reporting checklist (available at https://qims.amegroups.com/article/view/10.21037/qims-2025-2000/rc).
Methods
Datasets
The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study was approved by the Ethics Committee of Longgang Central Hospital of Shenzhen (No. 2023ECPJ077), and individual consent for this retrospective analysis was waived. We gathered patient data from the Longgang Central Hospital of Shenzhen between 2013 and 2022, encompassing 224 patients aged from 37 to 89 years. All patients underwent CT scan within 24 hours after onset of stroke symptoms, followed by IVT or pharmacotherapy. DWI of brain was performed within 3 days after CT. The median National Institutes of Health Stroke Scale (NIHSS) score at admission was 4 [interquartile range (IQR) 2, 6], and the median stroke lesion volume was 4.44 mL (IQR, 1.75, 26.75 mL). The dataset was divided into training, validation, and testing cohorts, as summarized in Table 1. The table presents the baseline demographic and clinical characteristics of patients in each cohort, including age, sex distribution, NIHSS score, and modified Rankin Scale (mRS) score.
Table 1
| Characteristic | Training (n=180) | Validation (n=20) | Testing (n=24) | P value |
|---|---|---|---|---|
| Age (years) | 58.26±12.01 | 64.25±12.10 | 59.19±12.37 | 0.110 |
| Sex | <0.001 | |||
| Male | 131 (72.8) | 17 (85.0) | 12 (50.0) | |
| Female | 49 (27.2) | 3 (15.0) | 12 (50.0) | |
| NIHSS score | 4.84±3.59 | 5.75±7.95 | 6.38±5.88 | 0.746 |
| mRS score | 1.82±1.28 | 2.60±1.43 | 2.86±1.39 | 0.001 |
Data are presented as mean ± standard deviation or n (%). mRS, modified Rankin Scale; NIHSS, National Institutes of Health Stroke Scale.
NCCT images were obtained using a Philips Brilliance 16 scanner with the following parameters: tube voltage of 140 kVp, slice thickness of 6 mm, and a matrix size of 512×512 mm2. DWI scans were acquired using a 1.5 T Philips Ingenia scanner and a 3.0 T Siemens Prisma scanner. The 1.5 T DWI parameters were as follows: repetition time (TR)/echo time (TE) =3,330 ms/104 ms; slice thickness =6 mm; matrix =164×110; field of view (FOV) =230×230; b value =1,000. The parameters of the 3.0 T DWI scans were as follows: TR/TE =3,300 ms/54 ms; slice thickness =5 mm; matrix =140×140; FOV =220×220; b value =1,000.
Image preprocessing
All images have been resized to an isotropic voxel size of 0.85×0.85×1 mm3. Fine-tuning the CT window width and level to 80 and 35, respectively, brings about a more dramatic display with greater emphasis on brain infarct lesions. After histogram standardization, which was performed using the Nyúl method implemented in the intensity-normalization Python package, all MR images were N4 bias field corrected to reduce intensity inhomogeneities. Multimodal registration between NECT and diffusion MRI was performed using a two-stage approach consisting of rigid registration followed by affine registration, implemented with the SimpleElastix Python package.
At the beginning of the skull-stripping process, FSL-BET was used to extract the brain region from DWI images. A binary mask generated from these extracted regions was applied to CT images for skull stripping with high precision.
Model and framework
We have modified our Multi-Task ControlNet to run on 2D slices using FiLM-style affine modulation. The model comprises four components, as illustrated in Figure 1: a shared encoder, a segmentation decoder, a synthesis decoder, and a FiLM affine modulation module.
As shown in Figure 1, the original U-Net encoder path is used in this paper. Our model consists of four stages of down-sampling, each with three layers: a 3×3 convolution with a stride of 2, GroupNorm, and ReLU. The encoder takes an input NCCT slice and yields feature maps , where is the deepest bottleneck representation. All weights of the encoder are frozen during training due to its pre-trained NCCT feature extraction.
The segmentation branch uses four up-sampling blocks to recreate the encoder structure. Each block consists of a transposed convolution with a stride of 2, followed by a 3×3 network, GroupNorm, and ReLU. The result at each scale is an anatomical prior map , where this encoding includes tissue regions.
The synthesis branch structure clearly mirrors the decoder of segmentation, with four pieces for resolution increasing. Starting from , maps that will be modulated by weight coefficients are reconstructed . At each up-sampling stage in the synthesis decoder, the corresponding anatomical prior map is fed through a 1×1 convolutional layer; here we refer to it as the FiLM generator. This outputs a tensor of shape . The tensor is then split channel-wise into two equal parts: these are the scale parameters and shift parameters .
In Formula 1, denotes the intermediate feature activations from the synthesis decoder at the i-th decoding stage, which are adaptively modulated by FiLM-generated scale and shift parameters. Element-wise transformation denotes multiplication is used in conjunction with Hadamard, which is shown in the following formula, meaning that all the individual parts are multiplied at once.
This FiLM-based affine modulation follows the standard formulation introduced by Perez et al. (36), after affine modulation, goes through standard up-sampling and convolution within the decoder channels. The FiLM generator structure is illustrated in Figure 2.
It is worth noting that all of these convolutions are initialized using Xavier’s method. This means that during the start of training, becomes close to 1 and approaches 0. Because before the network has learned to use physical information—especially for anatomical cues—faithless paths should always exist. When we train this end-to-end process, what these FiLM layers in fact do is channel-wise and spatially adaptive tuning. They allow the model to look for specific areas in DWI synthesis to amplify or weaken, based on localization-specific tissue structures coded by .
Loss function
Training minimizes a composite loss:
The loss is a voxel-wise reconstruction loss that penalizes absolute intensity differences between the ground truth DWI slice and the synthesized slice , which can be expressed as:
With regard to structural similarity loss,
Structural similarity index measure (SSIM) is computed over local 11×11 windows as follows:
with and denoting local means and variances, the covariance, and , small positive constants.
In this formula, and are known as the local means and variances. In addition, the symbol represents the covariance, while and represent two small stabilising constants.
About the issue of Dice loss for anatomical consistency, compares the segmentation map derived from the input NCCT against predicted from the synthesised DWI, thereby enforcing structural alignment. Here, the Dice loss is applied to latent segmentation representations rather than directly to DWI intensity images, serving as a structural consistency constraint across modalities.
The weighting coefficients of the loss terms were empirically tuned on the validation set to balance reconstruction accuracy and structural consistency.
Training details
The proposed network was implemented in Python (version 3.9) using the PyTorch deep learning framework (version 2.0). Model training was conducted on a workstation equipped with an NVIDIA RTX 4090 GPU, using CUDA 11.8. Training took approximately 10 hours. Optimization was performed using the Adam optimizer, with an initial learning rate of 1×10−4 and default momentum parameters. The batch size was set to 8. The network was trained for a maximum of 200 epochs. To improve model robustness and reduce overfitting, data augmentation was applied during training, including random in-plane rotations and horizontal flipping. Specifically, early stopping was employed if the validation loss did not improve for 20 consecutive epochs, and the model with the lowest validation loss was selected for final evaluation.
Model inference was conducted in a slice-wise manner on the same GPU. For each axial slice, the average inference time was 2.46 seconds, which produced an overall running time of approximately 5–7 minutes to produce a complete synthesized 3D DWI volume for a typical brain scan consisting of 120–160 slices.
Quantitative assessment indicators
Several assessment indicators are used to evaluate model performance. These include peak signal-to-noise ratio (PSNR), SSIM, and normalized mean squared error (NMSE). These metrics, under consideration, provide a comprehensive assessment of the model’s performance.
The mean squared error (MSE) is a fundamental metric in image processing and is often used to measure the average of the squared differences between pixel values for two given images, the MSE was calculated as following formula:
NMSE can be expressed as follows:
Capitalizing on the MSE, the PSNR is a metric used to gauge the quality of a reconstructed image in decibels (dB). The term is defined as follows:
In the context of two image ground truths and synthetic DWI, the SSIM is a measure that focuses on pixel-wise differences and is defined as:
Statistical analysis
In order to assess consistency and variation in the synthesized DWI compared to the reference DWI, the Bland-Altman analysis described was conducted to evaluate the agreement of the synthesized DWI images with the relevant ground-truth DWI. For each voxel in the brain mask, the intensity difference between the synthesized DWI and the reference DWI was calculated and plotted vs. the mean intensity of the two images according to Bland-Altman method. The mean difference was interpreted as the systematic bias between the two images, and the 95% limits of agreement were defined as mean difference ±1.96 standard deviations. This analysis was conducted to assess consistency and variation in the synthesized DWI compared to the reference DWI, instead of performing statistical hypothesis testing.
Clinical scoring metrics
The integrated DWI was systematically assessed by qualitative clinical evaluation in the current study for clinical applicability and quality. Medical professionals of our affiliated hospital reviewed this study. Researchers prioritized accuracy, crispness of anatomical structures, and definition of pathological features.
The evaluation consisted of artifact reduction, noise suppression, contrast retention, lesion discrimination, and overall image quality assessments for DWI. The criteria were scored on a scale of 1 through 5, with “excellent” indicating the highest value and “no information available” the lowest.
Results
Visual results and quantitative analysis
As shown in Figure 3, the figure presents a selection of examples of synthetic DWI generated by the network, alongside their corresponding reference DWI and associated error maps. As evident from the figure, the generated results demonstrate significant disparities in lesion areas compared to normal tissue.
As shown in Figure 4, the DWI images generated by the proposed algorithm are presented beside those generated by alternative algorithms to enable comparison. Furthermore, the same applies to the actual DWI images. A careful comparison reveals that the DWI images generated by our modified ControlNet, as well as those produced by the U-net and Pix2pix, can all clearly display high signal intensity areas affected by a stroke. This difference in output is particularly striking compared with other networks, such as 2D-CycleGAN, 3D-CycleGAN, and DualGAN, which completely abandon authentic DWI’s signal intensity profile when depicting stroke areas.
Images produced by the proposed approach not only consistently outperform those generated by other networks but also show no significant difference internally between and around tissue structures in both focal and non-focal areas. Rather than producing authentic DWI images, they come from within a few tissues, if not a single tissue, of them.
We assess the model’s performance using various metrics, including PSNR, SSIM, and NMSE. These metrics help give a comprehensive overview of the model’s overall performance. The last column quantitatively demonstrates that the precision of our synthetic DWI images exceeds that of both authentic DWI images and images generated by other networks. Although different algorithms can generate various characteristics of images, our method outstrips all other networks in terms of image generation. In Figure 5, our model can achieve a PSNR of 29.32±1.84 dB, as well as SSIM and NMSE scores of 0.867±0.032 and 0.293±0.041, respectively, thus proving its preeminent ability in this field.
To further illustrate the spatial characteristics of reconstruction errors, an ROI-based analysis was performed on a representative testing subject. As is shown in Figure 6, three regions of interest were manually selected to cover infarct core, peri-infarct tissue, and non-infarct brain regions. This analysis was intended to provide an intuitive, qualitative visualization of model behavior across different tissue types, rather than a statistical comparison across the full cohort.
Firstly, as far as reconstruction errors are concerned, there is a consistent finding in each of the 3 ROIs that for NMSE, the proposed method is always the best (0.355 in ROI 1, 0.259 in ROI 2, 0.183 in ROI 3), followed by Pix2Pix (0.385/0.277/0.193), DualGAN (0.387/0.292/0.204), U-Net (0.379/0.301/0.195), 2D-CycleGAN (0.382/0.307/0.202), and 3D-CycleGAN (0.395/0.321/0.236). This shows that ControlNet can more accurately align the new DWI to its ground truth than any of the other methods researched until now.
Secondly, in terms of resistance to noise, it can be seen that our network has improved PSNR in all three ROIs over the nearest competitor by a rate of about 1 to 3 dB, demonstrating an advantage in its ability to suppress reconstruction artifacts.
Finally, the highest SSIM is achieved by our network with 0.933 in ROI 2, and it maintains an admirable position over all three ROIs—0.881 for ROI 1, greater than DualGAN’s 0.892, and 0.937 for ROI 3, where the difference between Pix2Pix and the second-place finisher in terms of SSIM is less than 0.01. This indicates that in anatomically diverse regions, our proposed method exhibits greater reliability compared with other methods of automatic data synthesis.
By contrast, both U-Net baseline models and traditional GAN-based methods show much less consistency and become increasingly worse with the contrast required (ROI 1 and ROI 2).
Statistical analysis
Our original DWI images are seen at the right lower corner. The figures in this row are Bland-Altman plots of how artificial data compared with their corresponding real information; ours is in the second column, and they are followed by outputs from four other processes, respectively. As Figure 7 indicates, however, the average difference between our synthesized DWI images and true ones is −0.288. This result lies within a 95% confidence interval of −13.06 to 13.06, compared to competing methods, the proposed ControlNet-based approach shows smaller mean bias and narrower limits of agreement than the DWI reference, implying that it is much more consistent with the reference DWI.
Stroke core discrimination analysis
In addition, voxel-wise ROC-AUC was computed to evaluate stroke core discrimination based on synthesized DWI. For each patient in the testing cohort, synthesized DWI intensity was used as a continuous score to distinguish infarct core voxels from non-core voxels with reference to the ground-truth core masks, and the AUC values were summarized as mean ± standard deviation across patients. As shown in Table 2, the proposed ControlNet achieved the highest AUC (0.764±0.026) among all compared methods, indicating superior capability in distinguishing infarct core from non-core tissue based on synthesized diffusion contrast.
Table 2
| Method | AUC |
|---|---|
| U-Net | 0.687±0.055 |
| Pix2Pix | 0.711±0.039 |
| 2D CycleGAN | 0.542±0.067 |
| 3D CycleGAN | 0.490±0.102 |
| DualGAN | 0.603±0.082 |
| ControlNet | 0.764±0.026 |
Data are presented as mean ± standard deviation. AUC, area under the receiver operating characteristic curve; DWI, diffusion-weighted imaging; GAN, generative adversarial network.
The case-level analysis was conducted on the independent testing cohort to identify false-positive (FP) and false-negative (FN) lesion depiction on synthesized DWI. An FP case was defined as visually apparent lesion-like hyperintensity on synthesized DWI without a corresponding lesion on reference DWI, while an FN case was defined as a lesion present on reference DWI mask that was not visible or was substantially underestimated on synthesized DWI. Within the 24 test cases, the ControlNet model showed one FP case (4.2%) and four FN cases (16.7%).
Clinical scoring
A total of four radiologists independently assessed the synthetic DWI images produced by various methods on seven patients. For each patient, the generated DWI images were scored on a 1–5 scale based on each image quality criterion (overall image quality, noise suppression, lesion conspicuity, artifact reduction, and contrast recovery). The final scores were obtained by averaging the ratings across the seven patients. The aggregated results and individual radiologist scores are shown in Figure 8.
General image quality was judged to be of a high standard, achieving an average of 4.78 points. Among noise reduction techniques, only the application of noise suppression received a score over 4 points, namely 4.55. The rating score for the discrimination of lesions in a clinical context was 4.58. Artifact reduction received a high score of 4.08. Contrast retention did even better, with an average of 4.83.
Discussion
This paper introduces an enhanced ControlNet model for converting NCCT images of stroke patients to DWI. Our approach offers a practical way of obtaining DWI from NCCT in an emergency setting, improving the efficiency and ease of diagnosis of acute ischaemic stroke. Furthermore, this process saves valuable time for obtaining the MRI, leading to earlier treatment and better outcomes overall.
By incorporating a frozen, pre-trained encoder to extract NCCT features and a parallel segmentation decoder that produces meaningful prior maps that possess similar structure information at each decoding stage, our network seeks to combine the best of both worlds. This FiLM-style affine modulation is designed to make the network brighten or darken exactly where it needs to for DWI synthesis. This is done without disrupting the stable representations achieved by our backbone model. The unique aspect of our FiLM method is the fine-grained channel- and location-specific control it gives to the DWI contours. Thus, compared with a naive concatenation or addition operation, we get clearer DWI contours and more faithful structures. Moreover, only the segmentation branch and FiLM generators undergo training, so our method retains the efficiency and generalization characteristics of a frozen U-Net core when adapting to different anatomies. Lastly, this multitask framework encourages synergy between segmentation and synthesis objectives, rendering synthesised DWI volumes that are in voxel-dimension intensity and closely adhere to true anatomical boundaries.
While the findings of this study are promising, some limitations should be acknowledged. The very small number of cases considered for testing was another main limitation of the current study. There was also the use of a separate test set to assess performance of the model, but the small number of trials may make an assessment on the external validity of a proposed approach in different clinical conditions, stroke subtypes, and imaging protocols challenging. The research also adds that infrequent presentations and rare anatomic or pathological deviations may be underrepresented. Also, it should be emphasized that the above FP and FN prevalence estimates originated from a relatively small testing cohort and must be interpreted with caution. The small sample size may especially be insufficient to reflect the variability in lesion appearance seen in routine clinical practice. However, these exploratory findings shed light on potential failure modes of synthesized DWI, with false negatives more often detected in cases with subtle lesions on NCCT. More extensive multicenter studies are needed to define error prevalence more reliably and assess clinical risk. Our findings will be generalized through broader multicenter collaborations and validated externally across institutions to increase the robustness and clinical utility of the model.
Secondly, the developed model was constructed and tested on non-enhanced CT images collected during the initial part of an AIS (within 6 hours of the onset of symptoms). Since ischemic lesions undergo progressive changes in size, the imaging features in later-stage stroke may vary widely from early phase images, and how the model performs beyond the early time window is a matter to be addressed. Furthermore, often, DWI in routine clinical practice is obtained after therapeutic treatment including IVT or drug therapy. Indeed, the combined CT imaging of the brain and the MRI analysis in this study might represent different pathological states of the brain depending on the type, timing, and efficacy of the intervening therapy. This mismatch, due to time, may compromise the learned CT-DWI mapping and may influence the generalisability of the proposed approach to late-phase or post-intervention conditions. Additional validation of the proposed framework is warranted further in future papers with large multicenter datasets with larger timeframes and standardized imaging protocols to assess the robustness and clinical applicability of the proposed framework. Finally, while the current inference time allows generation of a synthesized 3D DWI volume within several minutes on a GPU, further optimization is required for seamless integration into time-critical acute stroke workflows.
One key limitation of the proposed method is the presence of hallucinated lesions in the synthesized diffusion MRI images shown in Figure 9. Such FP hyperintensities that do not correspond to true ischemic lesions on reference diffusion MRI might be caused by the model’s attempt to infer diffusion-like contrast from available structural and intensity data from non-contrast CT. From a clinical viewpoint, these hallucinations can be potentially serious, as images of spurious lesions can cause infarct extent or tissue status to be misread in assessing acute stroke. Similar artifacts were qualitatively present in other image synthesis techniques assessed in this work, and we did not conduct a systematic quantitative comparison of hallucination rates by image synthesis method. This is a notable limitation of the present work. Therefore, the constructed diffusion MRI produced in the proposed framework needs to be considered as a supportive visualization tool as opposed to a substitute for true diffusion imaging, and its results should always be considered in relation with actual CT images and clinical findings.
Another significant constraint lies in the architectural nature of our model. Although it was built within a 2D framework to facilitate training, this may limit its ability to capture the intricate subtleties and spatial relationships found in 3D brain anatomy. This could adversely affect the effectiveness of our model when applied to 3D medical imaging datasets commonly employed in clinical practice. In addition, the dependence on the quality and quantity of training data could be considered another major drawback. Like other deep learning models, the performance of ControlNet is highly sensitive to the diversity and representativeness of the training set. In this study, the model was trained on a limited set of NCCT and DWI images. Therefore, it may not be able to cope with rare stroke subtypes or cases where changes in imaging procedures result in an atypical anatomy. This may result in decreased accuracy in real-world clinical settings, where the modalities deployed and patients seen by doctors vary greatly. Therefore, future studies will need larger and more diverse datasets as well as further evaluation in multicentre trials to confirm stable behavior across varying clinical scenarios.
Conclusions
In conclusion, the advent of deep learning has engendered transformative potential for the expeditious diagnosis of AIS. The proposed approach demonstrates the potential for AIS NCCT to be converted into DWI, thereby emphasising the viability of modality transformation in the diagnosis of cerebral infarction. The potential exists for the translation of this work into clinical settings, with the objective of assisting physicians in the decision-making process for patients in emergency situations. This would be achieved by providing clinicians with high-quality MRI images.
Acknowledgments
None.
Footnote
Reporting Checklist: The authors have completed the TRIPOD+AI reporting checklist. Available at https://qims.amegroups.com/article/view/10.21037/qims-2025-2000/rc
Data Sharing Statement: Available at https://qims.amegroups.com/article/view/10.21037/qims-2025-2000/dss
Funding: This work was supported by
Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://qims.amegroups.com/article/view/10.21037/qims-2025-2000/coif). The authors have no conflicts of interest to declare.
Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study was approved by the Ethics Committee of Longgang Central Hospital of Shenzhen (No. 2023ECPJ077), and individual consent for this retrospective analysis was waived.
Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.
References
- Wang YJ, Li ZX, Gu HQ, Zhai Y, Zhou Q, Jiang Y, et al. China Stroke Statistics: an update on the 2019 report from the National Center for Healthcare Quality Management in Neurological Diseases, China National Clinical Research Center for Neurological Diseases, the Chinese Stroke Association, National Center for Chronic and Non-communicable Disease Control and Prevention, Chinese Center for Disease Control and Prevention and Institute for Global Neuroscience and Stroke Collaborations. Stroke Vasc Neurol 2022;7:415-50. [Crossref] [PubMed]
- Saceleanu VM, Toader C, Ples H, Covache-Busuioc RA, Costin HP, Bratu BG, Dumitrascu DI, Bordeianu A, Corlatescu AD, Ciurea AV. Integrative Approaches in Acute Ischemic Stroke: From Symptom Recognition to Future Innovations. Biomedicines 2023;11:2617. [Crossref] [PubMed]
- Huo X, Ma G, Tong X, Zhang X, Pan Y, Nguyen TN, et al. Trial of Endovascular Therapy for Acute Ischemic Stroke with Large Infarct. N Engl J Med 2023;388:1272-83. [Crossref] [PubMed]
- Ye Q, Zhai F, Chao B, Cao L, Xu Y, Zhang P, Han H, Wang L, Xu B, Chen W, Wen C, Wang S, Wang R, Zhang L, Jiao L, Liu S, Zhu YC, Wang LD. Rates of intravenous thrombolysis and endovascular therapy for acute ischaemic stroke in China between 2019 and 2020. Lancet Reg Health West Pac 2022;21:100406. [Crossref] [PubMed]
- Xiong Y, Wakhloo AK, Fisher M. Advances in Acute Ischemic Stroke Therapy. Circ Res 2022;130:1230-51. [Crossref] [PubMed]
- Darehed D, Blom M, Glader EL, Niklasson J, Norrving B, Eriksson M. In-Hospital Delays in Stroke Thrombolysis: Every Minute Counts. Stroke 2020;51:2536-9. [Crossref] [PubMed]
- Schellinger PD, Bryan RN, Caplan LR, Detre JA, Edelman RR, Jaigobin C, Kidwell CS, Mohr JP, Sloan M, Sorensen AG, Warach STherapeutics and Technology Assessment Subcommittee of the American Academy of Neurology. Evidence-based guideline: The role of diffusion and perfusion MRI for the diagnosis of acute ischemic stroke [RETIRED]: report of the Therapeutics and Technology Assessment Subcommittee of the American Academy of Neurology. Neurology 2010;75:177-85. Erratum in: Neurology 2010;75:938.
- Kaesmacher J, Cavalcante F, Kappelhof M, Treurniet KM, Rinkel L, Liu J, et al. Time to Treatment With Intravenous Thrombolysis Before Thrombectomy and Functional Outcomes in Acute Ischemic Stroke: A Meta-Analysis. JAMA 2024;331:764-77. [Crossref] [PubMed]
- Mosconi MG, Paciaroni M. Treatments in Ischemic Stroke: Current and Future. Eur Neurol 2022;85:349-66. [Crossref] [PubMed]
- Wardlaw JM, Murray V, Berge E, del Zoppo GJ. Thrombolysis for acute ischaemic stroke. Cochrane Database Syst Rev 2014;2014:CD000213. [Crossref] [PubMed]
- van Poppel LM, Majoie CBLM, Marquering HA, Emmer BJ. Associations between early ischemic signs on non-contrast CT and time since acute ischemic stroke onset: A scoping review. Eur J Radiol 2022;155:110455. [Crossref] [PubMed]
- Wardlaw JM, Mielke O. Early signs of brain infarction at CT: observer reliability and outcome after thrombolytic treatment--systematic review. Radiology 2005;235:444-53. [Crossref] [PubMed]
- Li Z, Huang H, Zhang Z, Shi G. Manifold-based multi-deep belief network for feature extraction of hyperspectral image. Remote Sens 2022;14:1484.
- Wintermark M, Sanelli PC, Albers GW, Bello J, Derdeyn C, Hetts SW, Johnson MH, Kidwell C, Lev MH, Liebeskind DS, Rowley H, Schaefer PW, Sunshine JL, Zaharchuk G, Meltzer CC. Imaging recommendations for acute stroke and transient ischemic attack patients: A joint statement by the American Society of Neuroradiology, the American College of Radiology, and the Society of NeuroInterventional Surgery. AJNR Am J Neuroradiol 2013;34:E117-27. [Crossref] [PubMed]
- Kamal H, Lopez V, Sheth SA. Machine Learning in Acute Ischemic Stroke Neuroimaging. Front Neurol 2018;9:945. [Crossref] [PubMed]
- Zhu JY, Park T, Isola P, Efros AA. editors. Unpaired image-to-image translation using cycle-consistent adversarial networks. Proceedings of the IEEE international conference on computer vision; 2017.
- Welander P, Karlsson S, Eklund A. Generative adversarial networks for image-to-image translation on multi-contrast MR images-a comparison of cyclegan and unit. arXiv preprint arXiv:180607777; 2018.
- Zhou L, Schaefferkoetter JD, Tham IWK, Huang G, Yan J. Supervised learning with cyclegan for low-dose FDG PET image denoising. Med Image Anal 2020;65:101770. [Crossref] [PubMed]
- Sun B, Jia S, Jiang X, Jia F. Double U-Net CycleGAN for 3D MR to CT image synthesis. Int J Comput Assist Radiol Surg 2023;18:149-56. [Crossref] [PubMed]
- Sun H, Xi Q, Fan R, Sun J, Xie K, Ni X, Yang J. Synthesis of pseudo-CT images from pelvic MRI images based on an MD-CycleGAN model for radiotherapy. Phys Med Biol 2022;
- Yang TH, Su YY, Tsai CL, Lin KH, Lin WY, Sung SF. Magnetic resonance imaging-based deep learning imaging biomarker for predicting functional outcomes after acute ischemic stroke. Eur J Radiol 2024;174:111405. [Crossref] [PubMed]
- Yahav-Dovrat A, Saban M, Merhav G, Lankri I, Abergel E, Eran A, Tanne D, Nogueira RG, Sivan-Hoffmann R. Evaluation of Artificial Intelligence-Powered Identification of Large-Vessel Occlusions in a Comprehensive Stroke Center. AJNR Am J Neuroradiol 2021;42:247-54. [Crossref] [PubMed]
- Sheth SA, Lopez-Rivera V, Barman A, Grotta JC, Yoo AJ, Lee S, Inam ME, Savitz SI, Giancardo L. Machine Learning-Enabled Automated Determination of Acute Ischemic Core From Computed Tomography Angiography. Stroke 2019;50:3093-100. [Crossref] [PubMed]
- Amukotuwa SA, Straka M, Smith H, Chandra RV, Dehkharghani S, Fischbein NJ, Bammer R. Automated Detection of Intracranial Large Vessel Occlusions on Computed Tomography Angiography: A Single Center Experience. Stroke 2019;50:2790-8. [Crossref] [PubMed]
- Chatterjee A, Somayaji NR, Kabakis IM. Abstract WMP16: artificial intelligence detection of cerebrovascular large vessel occlusion-nine month, 650 patient evaluation of the diagnostic accuracy and performance of the Viz. ai LVO algorithm. Stroke 2019;50:AWMP16.
- Jabal MS, Joly O, Kallmes D, Harston G, Rabinstein A, Huynh T, Brinjikji W. Interpretable Machine Learning Modeling for Ischemic Stroke Outcome Prediction. Front Neurol 2022;13:884693. [Crossref] [PubMed]
- Shaham U, Lederman RR. Learning by coincidence: Siamese networks and common variable learning. Pattern Recognition 2018;74:52-63.
- Tomita N, Jiang S, Maeder ME, Hassanpour S. Automatic post-stroke lesion segmentation on MR images using 3D residual convolutional neural network. Neuroimage Clin 2020;27:102276. [Crossref] [PubMed]
- Maier O, Schröder C, Forkert ND, Martinetz T, Handels H. Classifiers for Ischemic Stroke Lesion Segmentation: A Comparison Study. PLoS One 2015;10:e0145118. [Crossref] [PubMed]
- Kuang H, Wang Y, Liu J, Wang J, Cao Q, Hu B, Qiu W, Wang J. Hybrid CNN-Transformer Network With Circular Feature Interaction for Acute Ischemic Stroke Lesion Segmentation on Non-Contrast CT Scans. IEEE Trans Med Imaging 2024;43:2303-16. [Crossref] [PubMed]
- Gheibi Y, Shirini K, Razavi SN, Farhoudi M, Samad-Soltani T. CNN-Res: deep learning framework for segmentation of acute ischemic stroke lesions on multimodal MRI images. BMC Med Inform Decis Mak 2023;23:192. [Crossref] [PubMed]
- Kuang H, Tan X, Wang J, Qu Z, Cai Y, Chen Q, Kim BJ, Qiu W. Segmenting Ischemic Penumbra and Infarct Core Simultaneously on Non-Contrast CT of Patients with Acute Ischemic Stroke Using Novel Convolutional Neural Network. Biomedicines 2024;12:580. [Crossref] [PubMed]
- Yan S, Wang C, Chen W, Lyu J. Swin transformer-based GAN for multi-modal medical image translation. Front Oncol 2022;12:942511. [Crossref] [PubMed]
- Gutierrez A, Tuladhar A, Wilms M, Rajashekar D, Hill MD, Demchuk A, Goyal M, Fiehler J, Forkert ND. Lesion-preserving unpaired image-to-image translation between MRI and CT from ischemic stroke patients. Int J Comput Assist Radiol Surg 2023;18:827-36. [Crossref] [PubMed]
- Song T. Generative model-based ischemic stroke lesion segmentation. arXiv preprint arXiv:190602392; 2019.
- Perez E, Strub F, De Vries H, Dumoulin V, Courville A. editors. Film: Visual reasoning with a general conditioning layer. Proceedings of the AAAI Conference on Artificial Intelligence; 2018.


