Structural and Statistical Knowledge-Enhanced Attention Network for early Parkinson’s disease diagnosis
Original Article

Structural and Statistical Knowledge-Enhanced Attention Network for early Parkinson’s disease diagnosis

Yu Shen1,2 ORCID logo, Kai Qiao1, Jinjin Hai1, Xiangli Yang1, Tianqing Hu3, Qinglei Zhou3, Meiyun Wang2,4, Bin Yan1

1Henan Key Laboratory of Imaging and Intelligent Processing, Information Engineering University, Zhengzhou, China; 2Department of Radiology, Henan Provincial People’s Hospital, Zhengzhou, China; 3School of Computer and Artificial Intelligence, Zhengzhou University, Zhengzhou, China; 4Biomedical Research Institute, Henan Academy of Sciences, Zhengzhou, China

Contributions: (I) Conception and design: Y Shen, K Qiao, J Hai; (II) Administrative support: Y Shen, M Wang, Q Zhou, B Yan; (III) Provision of study materials or patients: Y Shen, M Wang; (IV) Collection and assembly of data: Y Shen, T Hu; (V) Data analysis and interpretation: Y Shen, X Yang; (VI) Manuscript writing: All authors; (VII) Final approval of manuscript: All authors.

Correspondence to: Qinglei Zhou, PhD. School of Computer and Artificial Intelligence, Zhengzhou University, No. 100 Kexue Avenue, Zhengzhou 450001, China. Email: ieqlzhou@zzu.edu.cn; Meiyun Wang, MD. Department of Radiology, Henan Provincial People’s Hospital, No. 7 Weiwu Road, Zhengzhou 450000, China; Biomedical Research Institute, Henan Academy of Sciences, No. 266-38 Mingli Road, Zhengzhou 450000, China. Email: mywang@zzu.edu.cn; Bin Yan, PhD. Henan Key Laboratory of Imaging and Intelligent Processing, Information Engineering University, No. 62 Kexue Avenue, Zhengzhou 450001, China. Email: ybspace@hotmail.com.

Background: Early diagnosis of Parkinson’s disease (PD) is crucial for prompt treatment and improved clinical outcomes, but accurate early diagnosis remains challenging. Deep learning (DL) methods have demonstrated significant potential for diagnosing neurological diseases using magnetic resonance imaging (MRI). However, most existing models use generic pattern recognition approaches without incorporating the distinctive characteristics of neuroimaging data and domain-specific knowledge. These limitations restrict both diagnostic accuracy (ACC) and clinical interpretability. This diagnostic ACC study aims to develop a specialized DL framework that effectively integrates neuroimaging domain knowledge to enhance both diagnostic performance and clinical interpretability for early PD detection.

Methods: In this study, we propose a Structural and Statistical Knowledge-Enhanced Attention Network (SSKEA-Net) for early PD diagnosis, consisting of two innovative cascaded modules: the Gray-White Interactive Modulation (GWIM) module, which utilizes a structurally specific gray-white matter separation mechanism to modulate channel-wise attention and further enhances tissue-specific features; and the Statistical Prior-Guided Attention (SPGA) module, which incorporates voxel-level statistically significant maps as spatial attention weights to guide feature extraction toward disease-related brain regions. By incorporating the domain knowledge-guided explicit feature enhancement strategy, SSKEA-Net effectively reduces feature interference between gray and white matter, significantly enhancing both model performance and interpretability. We evaluated the diagnostic performance of model using diffusion tensor imaging as input on a rigorously matched early-stage PD dataset with strictly controlled age and gender distributions, and employed activation heatmaps to visualize the interpretability of model.

Results: SSKEA-Net achieved superior diagnostic performance in five-fold cross-validation, with an ACC of 0.8798±0.0158, positive predictive value of 0.9185±0.0138, true positive rate of 0.8067±0.0267, specificity of 0.9406±0.0091, and an area under the curve of 0.9301±0.0146, outperforming both classic three-dimensional DL models and the current top-performing neuroimaging model, Simple Fully Convolutional Network. Compared to the baseline model, the cumulative activation heatmaps demonstrated that SSKEA-Net achieved precise anatomical localization, with activation specifically concentrated on the substantia nigra, putamen, midbrain, and corpus callosum, which correspond to clinically relevant brain structures associated with early PD, confirming the enhanced interpretability and clinical relevance.

Conclusions: The proposed SSKEA-Net effectively combines domain-specific structural and statistical knowledge with DL, achieving both high diagnostic ACC and clinical interpretability for early PD detection. By incorporating neuroimaging priors into the network architecture, this approach provides a valuable framework for developing more reliable and interpretable artificial intelligence systems in clinical neuroimaging applications.

Keywords: Parkinson’s disease (PD); neuroimaging; deep learning (DL); domain knowledge; model interpretability


Submitted Jun 30, 2025. Accepted for publication Dec 17, 2025. Published online Jan 23, 2026.

doi: 10.21037/qims-2025-1468


Introduction

Parkinson’s disease (PD) is a common neurodegenerative disorder, and over six million individuals worldwide are affected by PD. PD ultimately results in severe disability, as none of the available treatments are curative (1). Pharmacological interventions are effective in controlling symptoms, particularly in the early stage. Therefore, accurate and timely diagnosis of PD are paramount in clinical management (2). PD diagnosis traditionally relies on motor symptoms and patient history (3). By the time overt motor symptoms manifest, the disease has already progressed beyond the ideal period for therapeutic intervention. Therefore, early misdiagnosis rates remain high and diagnostic accuracy (ACC) based on clinical guidelines may only range from 73.8% to 83.9%, even influenced by the experience of a neurologist (4,5). For the reasons mentioned above, it is essential to identify reliable and consistent biomarkers for the early diagnosis of PD.

Magnetic resonance imaging (MRI) is widely used for early disease detection and diagnosis and it has confirmed microstructural changes in the brains of early-stage PD patients across various modalities (6). For example, structural MRI, such as T1-weighted (T1W) MRI could provide anatomical details of subcortical structures, which could reflect the macro-structural changes in nervous system disease (4,7). Additionally, diffusion MRI is highly sensitive to detecting early neurodegenerative changes by measuring the random motion of water molecules within tissue (6,8). Recent advancements in neuroscience and artificial intelligence (AI) have enabled the development of automated MRI-based methods for early diagnosis of PD. MRI scan for early PD diagnosis can offer standardized and radiation-free information, thereby significantly enhancing the ACC and clinical applicability of PD diagnostics to achieve early diagnosis and screening (9). Deep learning (DL) models have demonstrated to have high prediction ACC for diagnosing PD with neuroimaging MRI (10). Recent DL models for PD diagnosis using MRI have advanced with two-dimensional (2D) and three-dimensional (3D) architectures across various modalities. Currently, T1W and diffusion MRI are widely used for DL models in diagnosing neurodegenerative diseases. The PD diagnosis model using T1W MRI has achieved an ACC of over 70% and using diffusion MRI has achieved an ACC of over 77% (11).

Early research primarily focused on applying 2D network architectures to T1W MRI data, exploring various approaches to enhance diagnostic performance. An ensemble learning framework, incorporating award-winning convolutional neural network (CNN) models from the ImageNet, was employed to leverage the complementary strengths of each model and revealed that focusing on gray matter (GM) and white matter (WM) regions significantly improves PD classification performance (12). Using image enhancement techniques to extract the substantia nigra region as a preliminary attempt, using domain knowledge to guide model training and Residual Network-50 (ResNet-50), demonstrating the best performance. This enhancement approach represents a preliminary attempt at using domain knowledge to guide model training (13). Another study proposed a hybrid model that combines symptom data as domain knowledge and T1W MRI, enhancing model ACC of multi-stage PD classification and enriching feature representation. The hybrid model performed well across various stages. By incorporating clinical information beyond imaging as domain knowledge, this approach enriches feature representation (14). By preprocessing and optimizing T1W MRI, the study successfully developed a hybrid model that combines Grey Wolf Optimization with four DL models for diagnosing PD (15). Although 2D networks utilize different image views, they remain limited in capturing the spatial structural information of the brain compared to 3D networks, which constrains their further application in PD diagnosis. To overcome the limitations of 2D networks, a previous study developed a 35-layer 3D CNN model focusing on analyzing 3D T1W MRI for distinguishing PD patients from neurotypical subjects. Other studies proposed the 3D ResNet-18 architecture for PD diagnosis using 3D T1W images and employed Class Activation Maps (CAM) to identify clinically relevant brain regions and improve the interpretability (16,17). A study utilizing the largest T1W MRI dataset, with 2,041 subjects from 13 different research cohorts, applied the advanced Structured Feature Correlation Network model for PD classification. Saliency maps were used to emphasize critical regions, such as the orbitofrontal cortex, temporal lobes, and deep GM structures (18). These studies improved model interpretability by aligning activation maps with clinical diagnostic information, offering an alternative to domain knowledge-guided attention. Meanwhile, 3D networks leverage richer spatial and anatomical data, often improving diagnostic performance and interpretability.

Despite the greater availability of T1W MRI data, researchers have increasingly utilized diffusion MRI for its unique sensitivity to microstructural alterations that provide higher diagnostic specificity (SPE) for early PD detection. CNN using multi-modal diffusion MRI reached an ACC of 81%, integrated with a CAM module to validate whether the learned deep features conform to neurologic mechanisms of PD (17). A selective stacking approach, which selects fractional anisotropy (FA) and mean diffusivity image derived from diffusion tensor imaging (DTI) and brain regions of interest (ROI) was proposed to optimize 3D CNN network of PD diagnosis (19). Another study successfully applied autoencoders and variational autoencoders to train on DTI from healthy controls (HC), achieving anomaly detection in newly diagnosed PD patients, particularly revealing significant structural anomalies in the WM (20). In a large-scale study with 1,264 datasets from eight centers, the Simple Fully Convolutional Neural (SFCN) network optimized multimodal T1W and DTI combinations, achieving 80.8% ACC in early PD diagnosis. Although this ACC is lower than previous smaller-scale studies, the large dataset and structured organization enhance credibility and generalizability (6).

On the other hand, existing DL models typically utilize data-driven strategies without incorporating inherent neuroimaging properties and domain-specific knowledge, such as anatomical structures and tissue-specific characteristics. This limitation reduces the interpretability and clinical applicability of these models (21). Specifically, for neuroimaging-based diagnosis of neurological diseases, current models often ignore differences between distinct brain tissue types, potentially causing interference between tissue-specific features and reducing the sensitivity of the model to diagnostic biomarkers (21). In recent years, researchers have extensively explored model interpretability, which primarily encompasses attention-based approaches, domain-specific knowledge integration methods, and rule or example-based explanation techniques. Attention mechanisms are commonly used to increase model transparency by highlighting meaningful regions within medical images. For example, one study introduced a region-based integration-and-recalibration (RIR) attention module, which dynamically emphasized clinically important areas in optical coherence tomography (OCT) images, significantly improving interpretability in nuclear cataract classification (22). Another study proposed an Interpretable Dual-Attention Network (IDANet) that combined bidirectional spatial attention and parallel channel attention, along with the Local Interpretable Model-Agnostic Explanations (LIME) method, to visually demonstrate the decision-making process in diabetic retinopathy grading (23). Additionally, a Vision Transformer (ViT) based model integrating transformer and convolutional attention blocks was developed for laryngeal cancer grading. This method balanced interpretability and performance by generating visual explanations closely aligned with clinical areas of interest (24). Furthermore, focused attention methods within ViT architectures have been proposed, producing clear attribution maps that provide better interpretability and clinical acceptance compared to conventional CNN methods (25). Integration of domain-specific knowledge has also been explored to enhance clinical relevance and interpretability. One study presented a two-stage, expert-guided approach for diagnosing microvascular invasion in hepatocellular carcinoma, utilizing clinically defined imaging attributes as soft supervision and attention guidance, thereby improving both diagnostic ACC and interpretability (26). Another study developed a lightweight DL model for fetal ultrasound imaging, using knowledge-based scoring rules defined by clinical experts to select informative image frames, enhancing clinical interpretability and reliability (27). A hierarchical model using parameterized Gabor kernels was also introduced for sleep staging, explicitly incorporating expert electroencephalography (EEG) knowledge to establish a clear and interpretable decision-making process from detailed EEG features to broader temporal patterns (28). Several studies have developed transparent decision processes through rule-based or hypothesis-driven frameworks. For instance, a rule-based classifier called Find Rule Set (FIND-RS) was developed, which utilized Bayesian point selection methods to identify clear and representative decision rules, significantly enhancing the interpretability and clinical acceptance of the classification model (29). Moreover, a systematic review highlighted the potential of Neuro-Symbolic Learning (NSL), combining symbolic reasoning with DL, to improve model interpretability in brain tumor segmentation tasks (30). And hypothesis-driven explanatory frameworks have been integrated into Transformer models to improve robustness and interpretability (26,27). Other studies proposed sample-based interpretation strategies and general guidelines for interpretable AI model design. For example, one study introduced a multi-scale interpretable model, utilizing a feature pyramid structure to achieve multi-scale interpretable learning for breast mass margin classification, making the model’s decision process more consistent with clinical radiological practice (31). Another approach incorporated expert annotations and attention guidance maps into a dermatology diagnostic model, achieving expert-level interpretation without sacrificing predictive ACC (32). Additionally, a joint-learning framework named DL model of segmentation and classification for diagnosing coronavirus disease (DeepSC-COVID) was developed, integrating three-dimensional (3D) lesion segmentation and disease classification tasks for coronavirus disease 2019 (COVID-19) computed tomography (CT) imaging, providing lesion-based masks as visual evidence to enhance clinical interpretability (33).

Despite significant advances in model interpretability, existing approaches have not exploited the unique neuroanatomical properties and tissue-specific pathological patterns inherent in neuroimaging data, particularly the distinct characteristics of GM and WM alterations in PD progression. The pathophysiology of PD involves structural alterations in both GM and WM. Pathological alterations within GM nuclei appear during prodromal stage of PD (8), and WM microstructural abnormalities progressively diffuse with disease progression (34). Analyzing GM and WM separately provides the comprehensive characterization of pathological features across different stages. Considering that GM and WM exhibit different signal intensities in DTI, separate processing could reduce signal interference between tissues and improve diagnostic ACC and sensitivity. Additionally, traditional structural or intensity-based imaging features alone may be insufficient for accurate discrimination between PD patients and HC, due to the typically indistinct neuroimaging differences observed in PD (35). Statistical analysis and extraction of features may provide additional domain-specific prior knowledge, effectively guiding the network to selectively focus on disease-sensitive regions, thereby improving diagnostic ACC and model interpretability.

To enhance model performance and address the interpretability challenges in neuroimaging-based diagnosis, we follow the principles of tissue-specific analysis and statistical prior integration described above, proposing a Structural and Statistical Knowledge-Enhanced Attention Network (SSKEA-Net). SSKEA-Net employs FA images derived from DTI as input and incorporates two novel modules: Gray-White Interactive Modulation (GWIM) module and Statistical Prior-Guided Attention (SPGA) module. The GWIM module, designed based on distinct pathological and imaging characteristics of GM and WM in PD, structurally separates tissue-specific features and interactively modulates attention across them, effectively reducing feature interference between tissues. Meanwhile, the SPGA module utilizes voxel-level statistical priors derived from voxel-based analysis (VBA) to guide spatial attention, enabling the model to focus on anatomically and pathologically significant brain regions and improving the detection of subtle pathological changes. Through this integration of anatomical and statistical prior knowledge, SSKEA-Net leverages domain-specific clinical priors and diagnostic reasoning strategies, significantly improving practical diagnostic ACC and clinical interpretability. Experiments on age- and gender-matched DTI dataset demonstrate that SSKEA-Net achieved superior diagnostic performance, with attention map visualization further confirming the clinical interpretability. The proposed SSKEA-Net establishes a generalized and clinically effective analytical framework for early neurodegenerative disease diagnosis, providing a novel paradigm for designing interpretable and domain-guided DL architectures in neuroimaging. We hypothesized that incorporating tissue-specific (GM and WM) structural knowledge and statistical priors into DL architecture would significantly improve diagnostic ACC compared to conventional DL models for early PD detection. We present this article in accordance with the STARD reporting checklist (available at https://qims.amegroups.com/article/view/10.21037/qims-2025-1468/rc).


Methods

Data acquisition

This retrospective diagnostic ACC study analyzed existing DTI data from Parkinson’s Progression Markers Initiative (PPMI) (https://ppmi-info.org/), Alzheimer’s Disease Neuroimaging Initiative (ADNI) (https://adni.loni.usc.edu/) databases, and Henan Provincial People’s Hospital. We utilize multi-center data to ensure generalization capability of the model and a randomized allocation approach for training and validation sets.

The parameters of 3D T1W magnetization prepared rapid gradient echo (MPRAGE) and DTI imaging from PPMI and ADNI data are the following: 3D T1 MPRAGE, voxel size =1 mm × 1 mm × 1 mm, echo time (TE) =3.0 ms, repetition time (TR) =2,300 ms; DTI, b-values =0 and 1,000 s/mm2, nonzero b-value gradient directions =30 or 64, voxel size =2 mm × 2 mm × 2 mm. For the independently collected data at our institution, the protocols and methodologies were consistent with the PPMI data, ensuring uniformity in data acquisition and comparability of results. For independently collected data, the parameters were as follows: 3D T1 MPRAGE, voxel size =1 mm × 1 mm × 1 mm, TE =2.28 ms, TR =2,300 ms, field of view (FOV) =260 mm × 260 mm; DTI, b-values =0 and 1,000 s/mm2, nonzero b-value gradient directions =64, voxel size =2.2 mm × 2.2 mm × 2.2 mm.

Eligibility criteria for all participants included: (I) aged between 40 and 80 years; (II) no history of intracranial surgery or traumatic brain injury; and (III) no other psychiatric or central nervous system disorders. In addition to these general criteria, patients with PD were diagnosed according to the “Movement Disorder Society Clinical Diagnostic Criteria” for PD, with disease severity assessed using Movement Disorder Society Unified Parkinson’s Disease Rating Scale (MDS-UPDRS) and Hoehn & Yahr (H&Y) stage, and all patients with PD were in their early stages, with a H&Y stage ≤2.5. The MDS clinical diagnostic criteria were selected as the reference standard because they are the current international consensus gold standard for the clinical diagnosis of PD. Exclusion criteria included: (I) incomplete clinical diagnostic information; (II) excessive head motion during DTI acquisition (translation >3 mm or rotation >3°); (III) incomplete or corrupted DTI data.

Based on these criteria, participants were recruited from three sources: the PPMI database, the ADNI database, and our institution between February 2023 and February 2025, using a Siemens 3.0T Prisma MRI scanner with a 64-channel head-neck coil. And other sites also used Siemens 3.0T MRI scanners. Approvals for all cases were provided by relevant Institutional Review Boards and fully informed written consent was given by all participants from the affiliated hospital of the first author. Participants were recruited using a mixed sampling approach: random selection from PPMI and ADNI databases, and consecutive enrollment of eligible patients at our institution. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study was approved by the Ethical Committee of Henan Provincial People’s Hospital [No. (2024) Ethics Review-196] and all patients provided written informed consent.

MRI data pre-processing

DTI preprocessing

DTI preprocessing was performed using the Functional MRI of the Brain (FMRIB) Software Library (FSL 6.0; Oxford, UK; https://fsl.fmrib.ox.ac.uk/fsl/fslwiki/FSL). The preprocessing encompassed the following steps: (I) eddy correction using the FMRIB’s Diffusion Toolbox within FSL; (II) adjusting gradient directions based on the alterations induced by eddy correction; (III) skull stripping using the brain extraction tool within FSL; (IV) calculating the diffusion tensor metrics using the “dtifit” function within FSL for obtaining FA images. FA quantifies the degree of anisotropy of water diffusion, reflecting microstructural integrity of WM. In PD, alterations in FA values can indicate degeneration of WM tracts, which are associated with motor and cognitive impairments. We use these FA maps derived from DTI as the primary input for our DL model. The FA is computed as:

FA=3/2(λ1λ2)2+(λ2λ3)2+(λ3λ1)2λ12+λ22+λ32

where λ1, λ2and λ3 are the eigenvalues of the diffusion tensor.

Standardized preprocessing parameters were rigorously maintained across all subjects and acquisition centers to enhance model robustness and prevent the incorporation of technical variations as pathological features.

Image registration and segmentation

The registration approach involved linear and nonlinear registration of FA images using the FA brain template from the Montreal Neurological Institute (MNI). Deformation fields and transformation matrices were applied with affine transformations for linear registration and the symmetric normalization (SyN) algorithm for nonlinear registration, from Advanced Normalization Tools (https://stnava.github.io/ANTs/). FA images were resampled to 91×109×91 and resliced to 91×109×33 to reduce computational complexity. Additionally, image segmentation was conducted using Statistical Parametric Mapping 12 (SPM12) (https://www.fil.ion.ucl.ac.uk/spm/software/spm12/), resulting in probabilistic segmentation maps of GM and WM. These maps were subsequently threshold to generate binary tissue masks, explicitly delineating GM and WM regions for further analysis.

SSKEA-Net

This study proposes the SSKEA-Net, aiming to integrate structural priors and statistical significance from neuroimaging to enhance diagnostic performance and clinical interpretability in early-stage PD diagnosis. SSKEA-Net introduces a generalizable framework that can be integrated into various CNN architectures through systematic incorporation of two core modules: GWIM and SPGA modules. To demonstrate the implementation, we present an example using 3D ResNet-18 as the backbone architecture, as illustrated in Figure 1. The framework follows a hierarchical design principle where the SPGA module is embedded immediately after the initial convolutional layer to exploit statistical significance within shallow feature representations. In the shallow layers (Layer 1), we introduce Structural-and-Statistical Knowledge Residual Blocks (SSK-Res Blocks), which integrate both GWIM and SPGA modules after the convolutional operations for dual feature enhancement. For deeper layers (Layers 2–4), we employ Gray-White Interaction Residual Blocks (GWI-Res Blocks) containing only the GWIM module, as statistical priors are most effective in shallow representations while tissue-specific modeling remains beneficial throughout the network. This flexible design allows SSKEA-Net modules to be adapted to different network architectures while maintaining the core principle of domain knowledge integration.

Figure 1 The framework of SSKEA-Net for early diagnosis of PD, illustrated with 3D ResNet-18 as an example backbone architecture, using FA imaging derived from DTI as input. 3D, three-dimensional; Conv, convolution; DTI, diffusion tensor imaging; FA, fractional anisotropy; GM, gray matter; ResNet, Residual Network; SSKEA-Net, Structural and Statistical Knowledge-Enhanced Attention Network; WM, white matter.

Specifically, the GWIM module exploits tissue-specific features of GM and WM through domain knowledge-guided explicit feature enhancement strategy. Binary masks of GM and WM, derived from preprocessing steps, are applied to intermediate feature maps to achieve structure-specific feature separation. Channel-wise attention is computed separately for each tissue type, followed by interactive modulation between GM and WM channels to enhance discriminative features related to tissue-specific structural differences. The SPGA module utilizes voxel-level t-value map derived from VBA, generated by performing two-sample t-tests between PD and HC groups. The t-value map is then normalized to generate spatial attention weights, guiding the model to prioritize brain regions with significant statistical differences. This strategy enhances the representation of structurally abnormal regions related to PD pathology in shallow network layers.

To optimize feature extraction efficiency across network depths and ensure precise anatomical information alignment, we incorporated several crucial implementation details in SSKEA-Net. First, for Multi-scale Mask Adaptation (MsMA) strategy, the anatomical masks and statistical weighted maps must match the spatial dimensions of corresponding feature maps for element-wise multiplication. We addressed this through adaptive resizing using Global Average Pooling (GAP), which preserves spatial information while maintaining computational efficiency. Second, for module integration strategy, we implemented a serial connection within the SSK-Res blocks where GWIM processes features before SPGA. This design enables tissue-specific feature extraction followed by statistical attention refinement, effectively combining anatomical and statistical information. Third, for SPGA Module Embedding Position, we restricted SPGA to shallow layers as statistical priors provide stronger guidance in early feature representations, while GWIM is integrated throughout the entire network to ensure continuous tissue-specific modeling. The effectiveness of these configurations has been validated through systematic ablation experiments, with results consistent with our expectations. The technical implementation pathways of the two core modules (GWIM and SPGA) will be elaborated in detail in the following sections.

GWIM module

The GWIM module was designed based on tissue-specific pathological characteristics observed in neurodegenerative diseases, particularly targeting the distinct pathological patterns in GM and WM regions during PD progression. Previous studies have demonstrated that PD pathophysiology involves complex fiber connectivity and neural structural changes. Dopaminergic neuron degeneration in the nigrostriatal pathway initially affects GM nuclei, such as the substantia nigra and putamen (8,36), followed by subsequent involvement of WM tracts through axonal propagation (2). Traditional neuroimaging analyses and DL models typically process GM and WM jointly, ignoring these tissue-specific pathological differences, which reduces their sensitivity to subtle local pathological changes. This methodological limitation is especially pronounced in DTI analyses, where GM and WM exhibit fundamentally different signal characteristics. Specifically, GM regions exhibit lower FA and weaker signal intensity, whereas WM shows higher FA values and signal intensity. Without differentiating between these signals, computational models may excessively emphasize regions with high signal intensity (WM), while neglecting diagnostically important regions with lower intensity (GM nuclei). This introduces systematic bias and decreases sensitivity to early pathological changes. Inspired by the RIR Network proposed by Zhang et al. (22), we utilize structural GM and WM masks for tissue-specific feature separation and propose the GWIM module to separately extract and enhance structural features for each tissue type, thus improving the network’s ability to recognize and represent tissue-specific pathological characteristics. The GWIM module pipeline is shown in Figure 2. The implementation of the GWIM module comprises four core steps: (I) structure-specific feature splitting; (II) channel-wise attention mechanism; (III) channel-based interactive modulation; (IV) feature enhancement and integration.

Figure 2 The total GWIM module architecture. GM, gray matter; GWIM, Gray-White Interactive Modulation; WM, white matter.

Structure-specific feature splitting

Based on anatomical priors, the input feature map is decomposed into tissue-specific feature subsets. Specifically, pre-segmented GM and WM anatomical masks are applied to the input feature map F via element-wise multiplication to extract the corresponding tissue-specific features FGM and FWM:

FGM=FCMGM

FWM=FCMWM

Where FRH×W×D×C is the input feature map and FGM,FWMRH×W×D×C are the output tissue-specific feature. MGM,MWMRH×W×D×1 denote the GM and WM anatomical masks, respectively, and C is the Hadamard product with channel-dimension broadcasting. This anatomy-guided feature separation allows the network to process tissue through separate channel attention pathways, reducing the signal interference in standard feature extraction. To ensure anatomical masks maintain consistent resolution with feature maps across different network depths, we implement MsMA using convolution-like GAP. This explicit separation of tissue-specific representations creates distinct feature spaces, establishing the foundation for targeted feature enhancement in later network stages.

Channel-wise attention mechanism

For the separation of tissue-specific features, we combined the channel attention mechanism from Squeeze-and-Excitation Networks (SENet) (37) with the 3D ResNet-18 to construct a dual-pathway channel attention structure. Specifically, the two parallel pathways correspond to GM and WM-related features from the Structure-specific Feature Splitting module. Within each pathway, channel-wise attention weights are computed independently through GAP to enhance the representation of the most informative features for each tissue type, providing a foundation for subsequent cross-tissue interactive modulation. Specifically, we apply 3D GAP to aggregate spatial information across the entire volume for each channel.

gGM=1H×W×Dx=1Hy=1Wz=1DFGM(x,y,z)

gWM=1H×W×Dx=1Hy=1Wz=1DFWM(x,y,z)

Where gGM and gWMR1×1×1×C represent the channel-wise statistical vectors for WM and GM, respectively, with C denoting the number of feature channels.

Next, we establish channel relationships through a feature transformation module using two 1×1×1 convolutional layers with a bottleneck architecture:

gGM=W2GMReLu(W1GMgGM)

gWM=W2WMReLu(W1WMgWM)

Where W1GM and W1WM are weight matrices of the first 1×1×1 convolutions that perform channel dimensionality reduction from C to C/r channels, effectively compressing the channel information W2GM. And W2WM are weight matrices of the second 1×1×1 convolutions that restore the original dimension from C/r to C channels. The reduction ratio r (set to 16) controls the bottleneck width, balancing computational efficiency and representational capacity. This “reduce-activate-restore” design enables efficient modeling of channel-wise relationships with reduced parameters. gGM and gWMR1×1×1×C vectors encode transformed channel features after the non-linear mapping. While traditional SENet methods use fully connected layers for this transformation, requiring tensor reshaping operations (37), we adopt 1×1×1 convolutions following Zhang et al.’s RIR framework to effectively decrease computation.

Channel-based interactive modulate

After obtaining the channel-wise attention vectors gGM and gWM, we further implement a channel-based interactive modulation mechanism to explicitly model the interactive and competitive relationships between GM and WM features. This interactive modulation is based on a channel-level Softmax normalization function, explicitly representing the relative importance between GM and WM on each channel. The interactive attention weights are calculated as follows:

ωGM,C=exp(gGM,C)exp(gGM,C)+exp(gWM,C),C{1,2,...,C}

ωWM,C=exp(gWM,C)exp(gWM,C)+exp(gGM,C),C{1,2,...,C}

Where ωGM,C and ωWM,C represent the normalized interactive weights of GM and WM features on the c-th channel and ωGM,C+ωWM,C=1. gGM,C and gWM,C represent the c-th channel values of GM and WM features generated by the channel attention mechanism.

Feature enhancement and integration

After obtaining the interactive modulation weights, these weights are applied to the original tissue-specific features to enhance their channel-wise representations:

FGMenh=FGMSωGM

FWMenh=FWMSωWM

Where FGMenh and FWMenhRH×W×D×C, S represents the Hadamard product with spatial-dimension broadcasting. ωGM=[ωGM,1,ωGM,2,...,ωGM,C] and ωWM=[ωWM,1,ωWM,2,...,ωWM,C] are the channel interactive modulation weight vectors for GM and WM, respectively. This operation preserves the spatial information of the original features while interactive modulating them at the channel level according to the relative importance between tissues. Finally, the enhanced tissue-specific features are integrated through element-wise addition:

FGWIMenh=FGMenh+FWMenh

Where FenhRH×W×D×C is the structural enhanced feature. Combined with residual connections in the backbone network, the GWIM module separately models GM and WM features and introduces cross-tissue interactive channel modulation to effectively reduce inter-tissue feature interference. This explicit modeling of tissue-specific features not only enhances the model’s sensitivity to subtle pathological differences in early-stage PD but also improves clinical interpretability by directing the network’s focus toward diagnostically relevant, tissue-specific feature channels.

SPGA module

The SPGA module incorporates voxel-wise t-value map generated from voxel-based two-sample t-tests into the feature extraction stage, guiding the network to emphasize the brain regions with statistically significant differences between patient groups. Specifically, the t-value map quantifies statistical group differences at each voxel, highlighting the regions of greatest diagnostic relevance. These statistical priors are transformed into spatial attention weights, allowing the model to selectively enhance features from diagnostically informative regions. The SPGA module integrates statistical weighted maps into feature representations, directing attention toward brain regions with significant differences to improve diagnostic ACC and interpretability.

Weighted map calculation

At first, VBA was conducted to identify significant differences in brain structure between early PD patients and HCs. VBA is a technique that compares neuroimaging data voxel-by-voxel across different groups to detect localized differences. We employed the General Linear Model (GLM) to perform the analysis, which is widely used in neuroimaging studies. After spatial normalization to ensure anatomical correspondence across subjects, FA images were analyzed using the following GLM equation:

Y=Xβ+ε

Where Y is the vector of observed voxel intensities, X is the design matrix containing group membership indicators (PD vs. HC) and potential confounding variables (e.g., age, sex), β is the parameter vector to be estimated that reflects the magnitude of group differences, and ε is the error term.

To evaluate the statistical significance of these differences, t-statistics were calculated at each voxel using the following equation:

Ti=(Y¯PD,iY¯HC,i)(SPD,i2/nPD)+(SHC,i2/nHC)

Where Y¯PD,i and Y¯HC,i represent the mean FA values at voxel i for PD patients and HC respectively, SPD,i2 and SHC,i2 are the corresponding sample variances, and nPD and nHC are the sample size. The generated whole-brain t-value map, where Ti represents the t-statistic at voxel position i quantifying the statistical significance of FA differences between groups, highlights regions with statistically significant differences, with higher absolute t-values indicating stronger group differences.

The t-value map was normalized and converted into a continuous statistical attention weight map within the [0–1] range through a sigmoid transformation:

WSPGA=σ(Ti)=1/(1+exp(Ti))

This transformation ensures that voxels with higher absolute t-values, which indicate stronger statistical differences between PD patients and HC, receive higher attention weights, while areas with minimal differences receive weights closer to 0.5. The calculation pipeline of weighted map is shown in Figure 3.

Figure 3 The total SPGA Module architecture. The SPGA module consists of two main components: (left) weighted map calculation and (right) statistical enhanced feature. Incorporating VBA statistical priors as spatial attention weights, the module enables focused feature enhancement on discriminative brain regions for improved PD diagnosis. HC, healthy control; PD, Parkinson’s disease; SPGA, Statistical Prior-Guided Attention; VBA, voxel-based analysis.

Statistical enhanced feature

After obtaining the statistical attention weighted map WSPGA, we use convolution-like pooling to rescale WSPGA to align with feature map at different network depths. The rescaled weighted map W˜SPGA is applied to modulate feature maps extracted by the backbone network, directing computational focus toward brain regions with significant group differences. The basic feature modulation mechanism can be formulated as:

FSPGAenh=FCW˜SPGA

Where FRC×H×W×D is the input feature map, and FSPGAenhRC×H×W×D is the statistical enhanced feature. C is the Hadamard product with channel-dimension broadcasting. This operation selectively enhances features from regions showing significant statistical differences between patient and control groups, providing a data-driven attention mechanism based on neuroimaging statistics. By integrating t-value map into the feature processing pipeline, the network benefits from prior statistical knowledge about potential disease-related regions, complementing the standard feature learning process with clinically relevant spatial guidance.

Module integration strategy

To integrate the GWIM and SPGA modules, we adopt a direct sequential integration strategy that combines tissue-specific feature extraction with statistical prior knowledge. In this approach, features are first processed through the GWIM module, then the output features from GWIM are directly multiplied with the statistical weight map from SPGA:

Foutenh=FGWIMenhCW˜SPGA

Where FGWIMenhRH×W×D×C is the output feature map from GWIM module, WSPGARH×W×D is the weighted map from SPGA module and FGWIMenhRH×W×D×C is the finally integrated feature. C is the Hadamard product with channel-dimension broadcasting. This operation generates the final features by first enhancing structural attention through GWIM and then applying statistical attention weighting from SPGA.

Direct sequential integration employs a hierarchical feature processing approach that creates a coherent information flow from structural to statistical domain knowledge. This strategy enables the GWIM module to first extract tissue-specific features, followed by spatial enhancement from the SPGA module using statistical priors. Through ablation experiments, this design was validated against three alternative strategies: weighted sequential integration, direct parallel integration, and gated parallel integration. Ablation experiment results demonstrate that direct sequential integration achieved the best performance across all evaluation metrics.

In addition to the module integration strategy, we also investigated different MsMA strategies and SPGA module embedding position strategies. The detailed ablation experimental results are presented in the “Results” section.

Experimental details

The resolution original of FA images is 94×94×67. And the FA images are spatially normalized to MNI standard space, resampled to 91×109×91. Based on the salient imaging features of early PD, the resampled FA images were cropped to 91×109×33 to reduce computational complexity. And the resampled and cropped FA images were used as the model input. The model outputs probability scores [0–1], with a pre-specified threshold of 0.5 for binary classification as PD positive. The model was trained with diagnostic labels but evaluated on test sets in a blinded manner without additional clinical information. Clinical diagnoses were established prior to the development of the AI model, ensuring complete independence from the index test results.

The proposed SSKEA-Net was trained using an optimization protocol designed to ensure model convergence and generalization capabilities of neuroimaging classification. Training incorporates cosine annealing learning rate scheduling with warm up, where the learning rate starts at 1×10−4 and gradually decreases to 1×10−9 over 5,000 epochs. The gradual decrease in learning rate helps fine-tune the model and avoid overfitting in later stages. Network weights are initialized using a Gaussian distribution for a robust starting point. The stochastic gradient descent (SGD) optimizer, with momentum of 0.9 and weight decay of 0.01, accelerates convergence and helps regularize the model to prevent overfitting, especially with high-dimensional brain MRI data. Cross-entropy loss is used for binary classification (PD vs. HC). To ensure robustness and generalizability, a 5-fold cross-validation procedure is implemented, providing a reliable estimate of true performance. All the parameters were shown in Table 1. We used Python programming language and PyTorch DL framework implemented on an NVIDIA RTX A6000 GPU.

Table 1

Experimental parameters

Parameter Value
Training set, n 419
Validation set, n 105
Batch size, n 128
Epochs 5,000
Warmup epochs 500
Cosine decay epochs 5,000
Max learning rate 1×10−4
Min learning rate 1×10−9
Optimizer SGD
Momentum 0.9
Weight decay 0.01
Loss function Cross-entropy loss
Dropout rate 0.9
Original resolution of FA 91×109×91
Resliced resolution of FA 91×109×33

FA, fractional anisotropy; SGD, stochastic gradient descent.

Performance evaluation indicators

We employ five common classification metrics to evaluate the performance: ACC, precision [positive predictive value (PPV)], sensitivity [recall, true positive rate (TPR)], SPE, and the area under the receiver operating characteristic (ROC) curve (AUC). These metrics provide a comprehensive evaluation of the classification performance. The calculations of these metrics are based on the counts of true positives (TP), false negatives (FN), false positives (FP), and true negatives (TN). The definitions and corresponding equations for these indicators are as follows:

ACC=TP+TNTP+TN+FP+FN

PPV=TPTP+FP

TPR=TPTP+FN

SPE=TNTN+FP


Results

Participant characteristics

This study finally involved 524 participants (238 early PD patients and 286 HCs). Overall, 245 participants (115 early PD patients, 130 HCs) were from the PPMI dataset. And 210 participants (123 early PD patients, 87 HCs) were recruited from our institution. To match the age of HC and enrich the multicenter data, 69 HCs were selected from the ADNI database. Figure 4 illustrates the participant flow and selection process. Table 2 shows the clinical features and demographics. No significant differences were observed between PD and HC groups in age or sex distribution. All patients were in early-stage disease. In this study, imaging and clinical assessments were performed as part of routine clinical care or research protocols, with no additional interventions between tests.

Figure 4 Flow chart of participants inclusion. ADNI, Alzheimer’s Disease Neuroimaging Initiative; HC, healthy control; PD, Parkinson’s disease; PPMI, Parkinson’s Progression Markers Initiative.

Table 2

Demographic and clinical characteristics of all data

Characteristics PD (n=238) HC (n=286) P value
Age (years) 61.13±7.78 61.57±9.32 0.33
Sex (male/female) 143/95 150/136 0.08
H&Y stage 1.68±0.48

Data are presented as mean ± standard deviation or n. HC, healthy control; H&Y, Hoehn and Yahr; PD, Parkinson’s disease.

Results of proposed SSKEA-Net

The 5-fold cross-validation results of the proposed SSKEA-Net are showed in Table 3: ACC (0.9048, 0.8846, 0.8762, 0.8667, 0.8667) with an average of 0.8798±0.0158, PPV (0.9318, 0.9268, 0.9268, 0.9024, 0.9048) with an average of 0.9185±0.0138, TPR (0.8542, 0.8085, 0.7917, 0.7872, 0.7917) with an average of 0.8067±0.0267, SPE (0.9474, 0.9474, 0.9474, 0.931, 0.9298) with an average of 0.9406±0.0091, and AUC (0.9423, 0.9313, 0.9455, 0.9109, 0.9203) with an average of 0.9301±0.0146. Compared with the original framework, the proposed SSKEA-Net demonstrates significant improvements across all evaluation metrics. Specifically, ACC increased from 0.8454±0.0108 to 0.8798±0.0158 (an improvement of 3.44%), PPV improved from 0.9069±0.0116 to 0.9185±0.0123 (an improvement of 1.16%), TPR rose from 0.7354±0.0244 to 0.8067±0.0267 (a improvement of 7.13%), SPE enhanced from 0.9371±0.0091 to 0.9406±0.0091 (an improvement of 0.35%), and AUC increased from 0.9210±0.0118 to 0.9301±0.0146 (an improvement of 0.91%). These results validate the effectiveness of the proposed method in improving classification performance. Figures 5,6 present the ROC curve and confusion matrix of SSKEA-Net.

Table 3

Comparison of the proposed SSKEA-Net with other methods

Method ACC PPV TPR SPE AUC
Original framework 0.8454±0.0108 0.9069±0.0116 0.7354±0.0244 0.9371±0.0091 0.9210±0.0118
SSKEA-Net* 0.8798±0.0158* 0.9185±0.0138* 0.8067±0.0267* 0.9406±0.0091* 0.9301±0.0146*

Data are presented as mean ± standard deviation. *, the best performance. ACC, accuracy; AUC, area under the curve; PPV, positive predictive value; SPE, specificity; SSKEA-Net, Structural and Statistical Knowledge-Enhanced Attention Network; TPR, true positive rate.

Figure 5 The ROC curve of the proposed SSKEA-Net in the results of 1st to 5th fold cross-validation. AUC, area under the curve; ROC, receiver operating characteristic; SSKEA-Net, Structural and Statistical Knowledge-Enhanced Attention Network; std.dev, standard deviation.
Figure 6 The confusion matrixes of the proposed SSKEA-Net in the results of 1st to 5th fold cross-validation. HC, healthy control; PD, Parkinson’s disease; SSKEA-Net, Structural and Statistical Knowledge-Enhanced Attention Network.

Comparative analysis with state-of-the-art methods

To assess the effectiveness of our proposed method, we conducted a detailed comparison between SSKEA-Net and current representative DL methods for PD diagnosis. We compared SSKEA-Net with previously proposed 3D CNN for PD diagnosis, such as 3D ResNet-34, 3D ResNet-50, 3D MobileNet (14), 3D EfficientNet (38), Zhao et al. (19) model proposed, Zhang et al. (16) model proposed, Yang et al. (17) proposed model and particularly influential SFCN for neuroimaging analysis on our dataset. SFCN was initially developed for brain age prediction and sex classification tasks and achieved state-of-the-art performance on UK Biobank (39). SFCN has also demonstrated remarkable effectiveness in PD diagnosis, showing its versatility and robustness in various neuroimaging classification tasks (11,18).

As shown in Table 4, among the comparative methods, 3D ResNet, 3D EfficientNet, and PD-SFCN demonstrated relatively strong performance. As expected, the PD-SFCN model exhibited the best results among existing methods, achieving an ACC of 0.8742±0.0125, PPV of 0.8674±0.0147, TPR of 0.9018±0.0135, SPE of 0.8316±0.0240, and AUC of 0.9294±0.0196. Our proposed SSKEA-Net achieved the highest performance among all compared methods in most evaluation metrics, including ACC, PPV, SPE, and AUC, demonstrating the effectiveness of our approach in PD diagnosis.

Table 4

Comparison of the proposed SSKEA-Net with other methods

Method ACC PPV TPR SPE AUC
3D ResNet-34 0.8102±0.0143 0.7868±0.0271 0.7873±0.0408 0.8284±0.0327 0.9041±0.0185
3D ResNet-50 0.8430±0.0071 0.9000±0.0269 0.7395±0.0445 0.9295±0.0247 0.9155±0.0035
3D MobileNet (14) 0.7855±0.0205 0.8390±0.0198 0.6530±0.0693 0.8955±0.0233 0.7630±0.0283
3D EfficientNet (38) 0.8308±0.0108 0.8726±0.0496 0.7394±0.0457 0.9056±0.0443 0.8862±0.0264
Zhao et al. (19) 0.7800 0.8070 0.8070
Zhang et al., Yang et al. (16,17) 0.8454±0.0108 0.9069±0.0116 0.7354±0.0244 0.9371±0.0091 0.9210±0.0118
PD-SFCN (11,18) 0.8742±0.0125 0.8674±0.0147 0.9018±0.0135* 0.8316±0.0240 0.9294±0.0196
SSKEA-Net* 0.8798±0.0158* 0.9185±0.0138* 0.8067±0.0267 0.9406±0.0091* 0.9301±0.0146*

Data are presented as mean ± standard deviation or mean. , results from Zhao et al. are reported as mean values only, as stated in the original study (the code was unavailable). *, the best performance. 3D, three-dimensional; ACC, accuracy; AUC, area under the curve; PD, Parkinson’s disease; PPV, positive predictive value; ResNet, Residual Network; SFCN, Simple Fully Convolutional Neural; SPE, specificity; SSKEA-Net, Structural and Statistical Knowledge-Enhanced Attention Network; TPR, true positive rate.

Results of ablation experiment

To evaluate effectiveness of proposed SSKEA-Net architecture, we first conducted ablation studies to evaluate the effectiveness of the GWIM and SPGA modules individually and jointly, then we conducted ablation studies across three fundamental aspects: module embedding position, MsMA strategy, and module integration strategy, and all the ablation experiment options are described in Table 5. All experiments employed 5-fold cross-validation, with performance metrics reported as mean ± standard deviation.

Table 5

Summary of ablation experiment configurations for SSKEA-Net

Ablation experiment configuration Settings
Module integration strategy Direct sequential integration
Weighted sequential integration
Direct parallel integration
Gated parallel integration
MsMA strategy NNI
Convolution-like GAP
Convolution-like GMP
SPGA module embedding position Layer 1; Layers 1–2; Layers 1–3; Layers 1–4

GAP, Global Average Pooling; GMP, Global Max Pooling; MsMA, Multi-Scale Mask Adaptation; NNI, Nearest Neighbor Interpolation; SPGA, Statistical Prior-Guided Attention; SSKEA-Net, Structural and Statistical Knowledge-Enhanced Attention Network.

Module effectiveness analysis

The results in Table 6 demonstrate the individual and combined contributions of the proposed modules. When incorporating GWIM alone, the model achieved ACC of 0.8607±0.0142 and AUC of 0.9258±0.0125, representing improvements of 1.53% and 0.48% over the baseline, respectively. The most notable improvement was in TPR, which increased from 0.7354±0.0244 to 0.7858±0.0234, indicating enhanced sensitivity in detecting PD cases. When using SPGA alone, the model showed ACC of 0.8645±0.0168 and AUC of 0.9229±0.0198, with improvements of 1.91% and 0.19% over the baseline. SPGA particularly improved TPR to 0.7814±0.0300 while maintaining high PPV (0.9071±0.0116), suggesting effective spatial attention guidance.

Table 6

The experiment results of module effectiveness analysis

Model variant ACC PPV TPR SPE AUC
Baseline (without GWIM and SPGA) 0.8454±0.0108 0.9069±0.0116 0.7354±0.0244 0.9371±0.0091 0.9210±0.0118
Baseline + GWIM 0.8607±0.0142 0.8951±0.0243 0.7858±0.0234 0.9232±0.0178 0.9258±0.0125
Baseline + SPGA 0.8645±0.0168 0.9071±0.0116 0.7814±0.0300 0.9336±0.0079 0.9229±0.0198
Full model (GWIM + SPGA) 0.8798±0.0158* 0.9185±0.0138* 0.8067±0.0267* 0.9406±0.0091* 0.9301±0.0146*

Data are presented as mean ± standard deviation. *, the best performance. ACC, accuracy; AUC, area under the curve; GWIM, Gray-White Interactive Modulation; PPV, positive predictive value; SPGA, Statistical Prior-Guided Attention; SPE, specificity; TPR, true positive rate.

Combining both GWIM and SPGA modules achieved the highest performance across all metrics (ACC of 0.8798, TPR of 0.8067 and AUC of 0.9301), with improvements of 3.44%, 7.13%, and 0.91% over the baseline in ACC, TPR, and AUC, respectively. The synergistic effect is particularly evident in TPR, which showed greater improvement with combined modules than the sum of individual improvements, demonstrating that the integration of tissue-specific modeling and statistical prior guidance creates complementary enhancements for PD diagnosis.

The ablation experiment results of module integration strategy

Besides the direct sequential integration we used in SSKEA-Net, we evaluated three alternative integration strategies to comprehensively assess different approaches for combining GWIM and SPGA modules.

Weighted sequential integration

This method introduces learnable channel-specific weights to achieve fine-grained regulation of statistical priors:

Foutenh=FGWIMenhC[σ(vWSPGA)]

Where v=[v1,v2,...,vC]RC is a learnable channel modulation parameter vector and σ is Sigmoid function. This operation first extracts structural features through GWIM, then applies channel-specific statistical weighting to these features. Each channel is independently modulated by learnable parameters v, enabling adaptive emphasis of different feature types based on their statistical relevance, with weights normalized to [0, 1] by the sigmoid function.

Direct parallel integration

GWIM and SPGA modules process input features in parallel, then the output features from both are added:

Foutenh=FGWIMenh+FSPGAenh

Where FGWIMenh and FSPGAenhRH×W×D×C is the output feature map from GWIM and SPGA module, respectively. This operation directly sums the output features from GWIM and SPGA modules, enabling the parallel extraction of complementary information where structural SPE and statistical prior knowledge contribute equally to the integrated representation.

Gated parallel integration

This method uses an efficient dynamic gating mechanism that adaptively adjusts the contribution weights of the two modules by generating a shared spatial gating map:

Foutenh=GCFGWIMenh+(1G)CFSPGAenh

Where the shared gating function GRH×W×D×1 is a single-channel spatial gating map generated from gated function fgate:

G=fgate(Fconcate)=σ(Wg2ReLu(Wg1Fconcate))

Where Fconcate=concate(FGWIMenh,FSPGAenh), fgateis a small network composed of two 1×1×1 convolutional layers that maps the concatenated features to a single-channel spatial gating map. BN is the batch normalization. σ is the Sigmoid activation function that ensures the final gating values are within [0, 1]. This integration strategy balances parameter count and feature representation capability. The network can adjust the influence of structural and statistical information based on input features while using the same spatial weights across channels.

Table 7 shows that sequential strategies consistently outperformed parallel strategies across all evaluation metrics. Direct sequential integration achieved the best performance with ACC of 0.8798±0.0158, PPV of 0.9185±0.0138, and AUC of 0.9301±0.0146. Weighted sequential integration showed comparable results with slightly lower AUC (0.9276±0.0153). Both parallel strategies demonstrated significantly reduced performance across multiple metrics. Direct parallel achieved ACC of 0.8571±0.0078, TPR of 0.7639±0.0197, and AUC of 0.9123±0.0160, while gated parallel showed even lower performance with ACC of 0.8508±0.0055, TPR of 0.7500±0.0361, and AUC of 0.9270±0.0078, representing substantial decreases compared to sequential strategies.

Table 7

The ablation experiment results of module integration strategy

Module integration strategy ACC PPV TPR SPE AUC
Direct sequential 0.8798±0.0158* 0.9185±0.0138* 0.8067±0.0267* 0.9406±0.0091 0.9301±0.0146*
Weighted sequential 0.8779±0.0173 0.9180±0.0138 0.8024±0.0321 0.9461±0.0165* 0.9276±0.0153
Direct parallel 0.8571±0.0078 0.9093±0.0099 0.7639±0.0197 0.9357±0.0083 0.9123±0.0160
Gated parallel 0.8508±0.0055 0.9086±0.0217 0.7500±0.0361 0.9357±0.0202 0.9270±0.0078

Data are presented as mean ± standard deviation. *, the best performance. ACC, accuracy; AUC, area under the curve; PPV, positive predictive value; SPE, specificity; TPR, true positive rate.

The comparative results of SPGA module embedding position

As the GWIM module is embedded in all residual blocks by design, we focused our ablation study on SPGA module embedding position. Besides its fixed position after the 7×7×7 convolution, we evaluated four SPGA configurations: embedding in (I) Layer 1 only (adopted in our final method), (II) Layers 1–2, (III) Layers 1–3, and (IV) all layers. Table 8 shows that the Layer 1-only configuration achieved the best performance. Specifically, the Layers 1–2 configuration showed a decrease in ACC to 0.8730±0.0045, while embedding SPGA in Layers 1–3 and Layers 1–4 resulted in the lowest performance. As SPGA embedding depth increased, model performance declined, particularly in AUC values, suggesting that excessive statistical prior guidance in deeper layers may interfere with the ability to learn task-specific features of model.

Table 8

The ablation experiment results of SPGA module embedding position

SPGA position ACC PPV TPR SPE AUC
Layer 1 0.8798±0.0158* 0.9185±0.0138 0.8067±0.0267* 0.9406±0.0091 0.9301±0.0146*
Layers 1–2 0.8730±0.0045 0.9195±0.0104* 0.7917±0.0000 0.9415±0.0083* 0.9153±0.0072
Layers 1–3 0.8626±0.0206 0.9002±0.0366 0.7855±0.0266 0.9266±0.0300 0.8810±0.0417
Layers 1–4 0.8645±0.0165 0.9000±0.0248 0.7896±0.0317 0.9266±0.0203 0.8801±0.0409

Data are presented as mean ± standard deviation. *, the best performance. ACC, accuracy; AUC, area under the curve; PPV, positive predictive value; SPE, specificity; SPGA, Statistical Prior-Guided Attention; TPR, true positive rate.

The comparative results of MsMA strategy

Given the resolution mismatch between feature maps at different network depths and the original anatomical masks/statistical weight maps, we evaluated three mask adaptation strategies—(I) Nearest Neighbor Interpolation (NNI): maintains binary mask values through direct interpolation; (II) convolution-like GAP: computes local region averages using the same kernel parameters as convolutional layers (adopted in our final method); (III) convolution-like global maximum pooling (GMP): extracts maximum values from local regions using the same kernel parameters as convolutional layers. Table 9 shows that the convolution-like GAP strategy achieved the best overall performance with ACC of 0.8798±0.0158, PPV of 0.9185±0.0138, and AUC of 0.9301±0.0146. While convolution-like GMP showed a slight advantage in TPR (0.8125±0.0000 vs. 0.8067±0.0267), it underperformed in other metrics, particularly in AUC (0.9227±0.0171). The NNI strategy demonstrated the poorest performance across all evaluation metrics, with notably lower ACC (0.8603±0.0119) and AUC (0.9050±0.0049).

Table 9

The ablation experiment results of MsMA strategy

MsMA Strategy ACC PPV TPR SPE AUC
NNI 0.8603±0.0119 0.9102±0.0080 0.7708±0.0340 0.9357±0.0083 0.9050±0.0049
Convolution-like GAP 0.8798±0.0158* 0.9185±0.0138* 0.8067±0.0267* 0.9406±0.0091* 0.9301±0.0146*
Convolution-like GMP 0.8730±0.0119 0.9008±0.0256 0.8125±0.0000 0.9240±0.0219 0.9227±0.0171

Data are presented as mean ± standard deviation. *, the best performance. ACC, accuracy; AUC, area under the curve; GAP, Global Average Pooling; GMP, Global Max Pooling; MsMA, Multi-scale Mask Adaptation; NNI, Nearest Neighbor Interpolation; PPV, positive predictive value; SPE, specificity; TPR, true positive rate.

CAM and interpretability

To visually evaluate the decision-making process and clinical interpretability of SSKEA-Net in early PD diagnosis, CAM was generated using the Grad-CAM++ (40) method to illustrate the attention distribution over disease-related brain regions. Figure 7 shows the generated CAM of SSKEA-Net and its backbone network (3D ResNet-18).

Figure 7 Comparison of CAMs between the backbone network and SSKEA-Net. The first row represents the original 3D ResNet-18, and the second row represents SSKEA-Net. Each column displays the CAM results for the same subject across different models. 3D, three-dimensional; CAM, Class Activation Map; ResNet, Residual Network; SSKEA-Net, Structural and Statistical Knowledge-Enhanced Attention Network.

Visualization analysis in Figure 6 reveals fundamental differences in attention distribution between SSKEA-Net and the baseline 3D ResNet18 model. The baseline model exhibits diffuse activation patterns across extensive cortical regions, WM areas, and even cerebrospinal fluid spaces, indicating a lack of selectivity in identifying pathologically relevant structures. SSKEA-Net demonstrates precise anatomical localization with activation focused on core pathological structures, including the substantia nigra, putamen, midbrain, and corpus callosum.

In summary, SSKEA-Net achieves precise localization of pathologically significant anatomical structures, with its attention distribution accurately mapping the pathological progression from disease onset to clinical symptom emergence. This successful integration of DL models with neuropathological knowledge validates the model’s biological plausibility and provides a robust interpretability foundation for clinical applications.


Discussion

In this study, we developed SSKEA-Net by integrating tissue-specific modeling and statistical prior guidance into a DL architecture for early PD diagnosis. On a rigorously age- and gender-matched dataset of early-stage PD patients and HC, SSKEA-Net achieved 87.98% ACC and 0.9301 AUC, demonstrating substantial improvements over existing methods. Beyond performance gains, the CAMs revealed that SSKEA-Net focuses on clinically relevant brain regions known to be affected in early PD, including the substantia nigra, putamen, and motor cortex, indicating enhanced interpretability compared to conventional DL approaches. These results suggest that incorporating domain-specific neuroimaging knowledge can effectively guide the model to learn more discriminative and clinically meaningful features for PD diagnosis.

DTI preprocessing consistency significantly impacts model robustness. The initial preprocessing steps: eddy correction, gradient adjustment, brain extraction, and tensor fitting could determine FA map quality and anatomical boundary definition, directly affecting the model’s ability to distinguish genuine pathological variations from technical artifacts. Registration to MNI space enables precise spatial alignment essential for the SPGA module’s statistical priors and the GWIM module’s tissue masks to correspond to correct anatomical locations. Inconsistencies in either preprocessing or registration would compromise our knowledge-guided approach by introducing artificial variations or spatial misalignment. We therefore maintained strict standardization using identical FSL 6.0 parameters and ANTs registration settings across all subjects, ensuring that SSKEA-Net captures genuine disease patterns rather than technical variations, thereby preserving both diagnostic ACC and interpretability.

To further validate the effectiveness of knowledge-guided design, we compared SSKEA-Net with current state-of-the-art 3D DL methods. While 2D networks have obtained superior classification performance on individual MRI slices, they neglect essential 3D anatomical structures and connections between brain regions (41), limiting their ACC and reliability in diagnosing neurodegenerative diseases. Therefore, our comparison focused on 3D CNN architectures that use 3D convolutional kernels to extract voxel-level spatial features, preserving anatomical continuity between slices. Among the evaluated methods, SFCN represents impressive results in neuroimaging studies, as it uses a lightweight fully convolutional architecture inspired by Visual Geometry Group Network, organized into seven blocks that progressively extract deeper features. Previous studies have established SFCN as the current state-of-the-art for neurodegenerative disease diagnosis, making it a primary reference method for comparison (11,18). Our results confirm that SFCN achieved an ACC of 0.8742, significantly outperforming other classic models, including 3D ResNet variants, which validates its reported outstanding performance. However, SSKEA-Net achieved 87.98%, demonstrating measurable improvements over this strong baseline.

The performance advantage of SSKEA-Net over existing methods can be attributed to its domain knowledge-guided architecture. Unlike conventional 3D CNNs that process brain images as generic volumetric data, SSKEA-Net integrates two critical neuroimaging insights: tissue-specific pathological patterns and statistically significant regions. The GWIM module explicitly separates GM and WM features before applying interactive modulation, thereby reducing inter-tissue feature interference that often arises in standard convolution operations. This tissue-aware approach aligns with the distinct pathophysiology of PD across cortical and subcortical regions. Meanwhile, the SPGA module converts group-level statistical maps into spatial attention cues, allowing the model to focus on regions with established diagnostic relevance. By embedding anatomical and statistical priors directly into the feature learning pipeline, SSKEA-Net captures more clinically meaningful features while maintaining efficiency, which explains its consistent performance advantage over purely data-driven architectures. This comparison demonstrates that while previous methods have advanced automated PD diagnosis through general-purpose architectures, the integration of domain-specific neuroimaging knowledge at the architectural level provides a new pathway for improving diagnostic ACC.

To understand the contribution of each design element, we systematically evaluated four aspects of SSKEA-Net: module effectiveness, module integration strategy, SPGA embedding position, and MsMA approach. The module effectiveness analysis confirmed that both GWIM and SPGA contribute significantly to the model’s performance, with their combination yielding synergistic improvements, particularly in TPR for PD detection. The TPR improvement from 73.54% to 80.67% reflects enhanced sensitivity to both gray and WM pathology. The interactive modulation of GWIM module dynamically balances contributions from both tissue types at the channel level, leveraging the complementary diagnostic information from GM neuronal degeneration and WM microstructural changes. The SPGA module further enhances this dual-tissue sensitivity; examination of the t-value maps revealed significant group differences in both GM nuclei and WM tracts, confirming that both tissue types contribute to improved detection. This validates our design principle of integrating tissue-specific modeling with statistical prior guidance to complementary pathological features from both tissue types. For the optimal module integration strategy, the superior performance of sequential strategies may be attributed to their hierarchical feature processing approach. In sequential configurations, the GWIM module first extracts tissue-specific features, followed by spatial weighting enhancement from the SPGA module based on statistical priors, creating a coherent information flow from structural to statistical domains. In contrast, parallel strategies process different types of features simultaneously, potentially failing to fully leverage the complementary advantages of both modules. The slight advantage of direct sequential integration over weighted sequential integration may benefit from its simpler design, reducing overfitting risk, especially under conditions of relatively limited training samples. Regarding the optimal configuration of SPGA module embedding position, the superior performance of Layer 1-only SPGA embedding can be attributed to the nature of statistical prior information. Statistical priors are most effective at the earliest feature extraction stage, where spatial anatomical information is richest (41). As network depth increases, features become more abstract and task-specific, diminishing the effectiveness of statistical guidance (42). This finding confirms that introducing statistical prior guidance exclusively at Layer 1, while maintaining tissue-specific modeling throughout the network via GWIM, creates an optimal balance for enhancing PD-related feature learning. Concerning the selection of MsMA strategy, these results indicate that GAP provides superior performance by generating soft mask. Compared to GMP and NNI, which produce hard masks, GAP creates smoother attention weight distributions by averaging activation values within local regions. This smooth transition helps reduce noise effects and improve model robustness to anatomical variations. This finding emphasizes the importance of mask smoothness and boundary continuity during multi-scale feature propagation for enhancing PD-related feature extraction, especially when processing brain MRI images with complex anatomical structures.

The visualization analysis provides important insights into how SSKEA-Net achieves superior diagnostic ACC compared to conventional approaches. For the subjects that were misclassified by the baseline model but correctly identified by SSKEA-Net, the CAMs show clear differences in attention patterns. The activation maps of baseline model display diffuse and scattered patterns across non-specific brain regions, including cerebrospinal fluid spaces, peripheral WM, and extensive cortical areas that lack clear pathological relevance to PD. In contrast, SSKEA-Net precisely focuses on established PD-related anatomical structures, with concentrated activations in the substantia nigra, putamen, and midbrain regions. This focused attention pattern demonstrates how the integration of domain knowledge enables SSKEA-Net to learn diagnostically relevant features, while conventional networks may be distracted by spurious patterns in clinically irrelevant areas. This focused attention pattern of proposed SSKEA-Net precisely aligns with the neuropathological hallmarks of PD. Progressive loss of dopaminergic neurons in the substantia nigra pars compacta represents the cardinal pathological signature of PD, accompanied by α-synuclein aggregation and Lewy body formation that constitute the pathological foundation of the disease (36,42). The putamen, as a major component of the striatum, exhibits degenerative changes in dopaminergic innervation that result in basal ganglia circuit dysfunction, directly mediating motor symptom manifestation (8,36). Activation in the midbrain region indicates the model’s capacity to capture brainstem pathology, consistent with Braak staging theory that pathological changes progress from lower brainstem to higher brain regions (43). Furthermore, corpus callosum activation reflects early WM microstructural alterations in PD patients, particularly the decreased FA observed in DTI, widely confirmed as an early neuroimaging characteristic of PD (44). The ability of SSKEA-Net to automatically identify these disease-specific regions without explicit anatomical labeling demonstrates that the integration of tissue-specific modeling and statistical priors successfully guides the network toward learning clinically meaningful patterns. This enhanced interpretability not only explains the model’s improved diagnostic performance but also increases clinician confidence in the model’s predictions by providing anatomically grounded reasoning for its decisions.

Furthermore, the demonstrated ability to precisely identify disease-specific regions suggests promising extensions of the SSKEA-Net framework. This knowledge-guided approach could naturally adapt to other quantitative neuroimaging modalities beyond DTI. For instance, recent Quantitative Susceptibility Mapping (QSM) studies have shown its effectiveness in quantifying iron accumulation patterns in neurodegenerative diseases (45). Voxel-based QSM analyses have successfully diagnosed PD and PD-mild cognitive impairment (MCI) patients (46,47), demonstrating its capability to provide statistical priors required for SPGA module. Machine learning with ROI-based QSM has efficiently identified PD patients with MCI, confirming compatibility with GWIM (48). Therefore, adapting our framework to QSM imaging represents a valuable future direction. Beyond modality extensions, the modular architecture also shows potential for more complex clinical applications. The demonstrated ability to precisely identify disease-specific regions suggests promising extensions of the SSKEA-Net framework. The modular architecture could be adapted for differential diagnosis between PD and atypical Parkinsonian syndromes. The GWIM module would remain applicable as all syndromes involve both gray and WM pathology, while the SPGA module could incorporate multi-group statistical comparisons. When sufficient multi-syndrome datasets become available, this knowledge-guided framework, with its interpretability advantage, could address the challenging task of differential diagnosis, revealing which anatomical patterns distinguish each syndrome.

There are limitations in this study. The current implementation relies on preprocessed tissue segmentation and statistical significance maps, suggesting opportunities for future end-to-end learning approaches. While our validation utilized data from limited centers, broader multi-center validation would further strengthen the generalizability. Third, our study used only binary classification (PD vs. controls) and did not include differential diagnosis of atypical syndromes, since PPMI has few multiple system atrophy or progressive supranuclear palsy data and our institution has limited samples. Future work with comprehensive multi-syndrome datasets could test differential diagnosis performance. Additionally, the framework’s potential for longitudinal disease monitoring and application to other neurodegenerative conditions remains to be explored.


Conclusions

This study proposes SSKEA-Net for early PD diagnosis, which systematically integrates domain-specific neuroimaging knowledge into DL architecture to improve both diagnostic performance and model interpretability. By incorporating tissue-specific modeling through GWIM and statistical prior guidance through SPGA modules, our approach reduces tissue feature interference while precisely directing model attention to clinically relevant brain regions. This design demonstrates that explicit incorporation of anatomical and statistical prior knowledge can simultaneously enhance both diagnostic ACC and clinical interpretability.

Experiments on rigorously age- and gender-matched datasets demonstrate that SSKEA-Net achieved significant improvements compared to other state-of-the-art networks in early-stage PD diagnosis. Visualization analysis revealed that attention regions are highly consistent with current understanding of disease pathological progression. By achieving both high diagnostic ACC and clinically meaningful interpretability, SSKEA-Net demonstrates the potential of knowledge-guided DL in medical imaging applications. This knowledge-guided approach provides a new paradigm for neuroimaging-based AI systems, demonstrating how systematic integration of domain expertise with DL can effectively connect computational modeling with clinical reasoning, which is instrumental in developing trustworthy and interpretable AI systems for neuroimaging diagnostics.


Acknowledgments

None.


Footnote

Reporting Checklist: The authors have completed the STARD reporting checklist. Available at https://qims.amegroups.com/article/view/10.21037/qims-2025-1468/rc

Data Sharing Statement: Available at https://qims.amegroups.com/article/view/10.21037/qims-2025-1468/dss

Funding: This work was supported by the Medical Science and Technology Research Project of Henan Province (No. SBGJ202303007), the Science and Technology Project of Henan Province (No. 242102311029), and the Central Plains Science and Technology Innovation Leading Talent Project (No. 244200510015).

Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://qims.amegroups.com/article/view/10.21037/qims-2025-1468/coif). The authors have no conflicts of interest to declare.

Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study was approved by the Ethical Committee of Henan Provincial People’s Hospital [No. (2024) Ethics Review-196] and all patients provided written informed consent.

Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.


References

  1. Global, regional, and national burden of Parkinson's disease, 1990-2016: a systematic analysis for the Global Burden of Disease Study 2016. Lancet Neurol 2018;17:939-53. [Crossref] [PubMed]
  2. Tolosa E, Garrido A, Scholz SW, Poewe W. Challenges in the diagnosis of Parkinson's disease. Lancet Neurol 2021;20:385-97. [Crossref] [PubMed]
  3. Global, regional, and national burden of disorders affecting the nervous system, 1990-2021: a systematic analysis for the Global Burden of Disease Study 2021. Lancet Neurol 2024;23:344-81. [Crossref] [PubMed]
  4. Rizzo G, Copetti M, Arcuti S, Martino D, Fontana A, Logroscino G. Accuracy of clinical diagnosis of Parkinson disease: A systematic review and meta-analysis. Neurology 2016;86:566-76. [Crossref] [PubMed]
  5. Beach TG, Adler CH. Importance of low diagnostic Accuracy for early Parkinson's disease. Mov Disord 2018;33:1551-4. [Crossref] [PubMed]
  6. Andica C, Kamagata K, Saito Y, Uchida W, Fujita S, Hagiwara A, Akashi T, Wada A, Ogawa T, Hatano T, Hattori N, Aoki S. Fiber-specific white matter alterations in early-stage tremor-dominant Parkinson's disease. NPJ Parkinsons Dis 2021;7:51. [Crossref] [PubMed]
  7. Camacho M, Wilms M, Mouches P, Almgren H, Souza R, Camicioli R, Ismail Z, Monchi O, Forkert ND. Explainable classification of Parkinson's disease using deep learning trained on a large multi-center database of T1-weighted MRI datasets. Neuroimage Clin 2023;38:103405. [Crossref] [PubMed]
  8. Zhang D, Zhou L, Yao J, Shi Y, He H, Wei H, Tong Q, Liu J, Wu T. Increased Free Water in the Putamen in Idiopathic REM Sleep Behavior Disorder. Mov Disord 2023;38:1645-54. [Crossref] [PubMed]
  9. Shaban M. Deep Learning for Parkinson's Disease Diagnosis: A Short Survey. Computers 2023;12:58.
  10. Loh HW, Hong W, Ooi CP, Chakraborty S, Barua PD, Deo RC, Soar J, Palmer EE, Acharya UR. Application of Deep Learning Models for Automated Identification of Parkinson's Disease: A Review (2011-2021). Sensors (Basel) 2021;21:7034. [Crossref] [PubMed]
  11. Camacho M, Wilms M, Almgren H, Amador K, Camicioli R, Ismail Z, Monchi O, Forkert NDAlzheimer’s Disease Neuroimaging Initiative. Exploiting macro- and micro-structural brain changes for improved Parkinson's disease classification from MRI data. NPJ Parkinsons Dis 2024;10:43. [Crossref] [PubMed]
  12. Mostafa TA, Cheng I. Parkinson's Disease Detection Using Ensemble Architecture from MR Images. 2020 IEEE 20th International Conference on Bioinformatics and Bioengineering (BIBE); Cincinnati, OH, USA. IEEE; 2020:987-92.
  13. Bhan A, Kapoor S, Gulati M. Diagnosing Parkinsons disease in Early Stages using Image Enhancement, ROI Extraction and Deep Learning Algorithms. 2021 2nd International Conference on Intelligent Engineering and Management (ICIEM); London, United Kingdom. 2021:521-5.
  14. Zhu S. Early Diagnosis of Parkinson's Disease by Analyzing Magnetic Resonance Imaging Brain Scans and Patient Characteristic. 2022 10th International Conference on Bioinformatics and Computational Biology (ICBCB); Hangzhou, China. 2022:116-23.
  15. Majhi B, Kashyap A, Mohanty SS, Dash S, Mallik S, Li A, Zhao Z. An improved method for diagnosis of Parkinson's disease using deep learning models enhanced with metaheuristic algorithm. BMC Med Imaging 2024;24:156. [Crossref] [PubMed]
  16. Zhang Y, Lei H, Huang Z, Li Z, Liu CM, Lei B. Parkinson's Disease Classification with Self-supervised Learning and Attention Mechanism. 2022 26th International Conference on Pattern Recognition (ICPR); Montreal, QC, Canada. 2022:4601-7.
  17. Yang MJ, Huang XB, Huang LQ, Cai GE. Diagnosis of Parkinson's disease based on 3D ResNet: The frontal lobe is crucial. Biomed Signal Proces 2023;85:104904.
  18. Baagil H, Hohenfeld C, Habel U, Eickhoff SB, Gur RE, Reetz K, Dogan I. Neural correlates of impulse control behaviors in Parkinson's disease: Analysis of multimodal imaging data. Neuroimage Clin 2023;37:103315. [Crossref] [PubMed]
  19. Zhao H, Tsai CC, Zhou M, Liu Y, Chen YL, Huang F, Lin YC, Wang JJ. Deep learning based diagnosis of Parkinson's Disease using diffusion magnetic resonance imaging. Brain Imaging Behav 2022;16:1749-60. [Crossref] [PubMed]
  20. Muñoz-Ramírez V, Kmetzsch V, Forbes F, Meoni S, Moro E, Dojat M. Subtle anomaly detection: Application to brain MRI analysis of de novo Parkinsonian patients. Artif Intell Med 2022;125:102251. [Crossref] [PubMed]
  21. Yin C, Imms P, Cheng M, Amgalan A, Chowdhury NF, Massett RJ, Chaudhari NN, Chen X, Thompson PM, Bogdan P, Irimia AAlzheimer’s Disease Neuroimaging Initiative. Anatomically interpretable deep learning of brain age captures domain-specific cognitive impairment. Proc Natl Acad Sci U S A 2023;120:e2214634120. [Crossref] [PubMed]
  22. Zhang X, Xiao Z, Fu H, Hu Y, Yuan J, Xu Y, Higashita R, Liu J. Attention to region: Region-based integration-and-recalibration networks for nuclear cataract classification using AS-OCT images. Med Image Anal 2022;80:102499. [Crossref] [PubMed]
  23. Bhati A, Gour N, Khanna P, Ojha A, Werghi N. An interpretable dual attention network for diabetic retinopathy grading: IDANet. Artif Intell Med 2024;149:102782. [Crossref] [PubMed]
  24. Huang P, He P, Tian S, Ma M, Feng P, Xiao H, Mercaldo F, Santone A, Qin J. A ViT-AMC Network With Adaptive Model Fusion and Multiobjective Optimization for Interpretable Laryngeal Tumor Grading From Histopathological Images. IEEE Trans Med Imaging 2023;42:15-28. [Crossref] [PubMed]
  25. Playout C, Duval R, Boucher MC, Cheriet F. Focused Attention in Transformers for interpretable classification of retinal images. Med Image Anal 2022;82:102608. [Crossref] [PubMed]
  26. Zhou Y, Sun SW, Liu QP, Xu X, Zhang Y, Zhang YD. TED: Two-stage expert-guided interpretable diagnosis framework for microvascular invasion in hepatocellular carcinoma. Med Image Anal 2022;82:102575. [Crossref] [PubMed]
  27. Li JT, Gao Z, Wang CL, Pu B, Li KL. A rule-guided interpretable lightweight framework for fetal standard ultrasound plane capture and biometric measurement. Neurocomputing 2025;621:129290.
  28. Niknazar H, Mednick SC. A Multi-Level Interpretable Sleep Stage Scoring System by Infusing Experts' Knowledge Into a Deep Network Architecture. IEEE Trans Pattern Anal Mach Intell 2024;46:5044-61. [Crossref] [PubMed]
  29. Bergamin L, Polato M, Aiolli F. Improving rule-based classifiers by Bayes point aggregation. Neurocomputing 2025;613:128699.
  30. Hassan M, Fateh AA, Lin JQ, Zhuang YJ, Lin GS, Xiong HR, You Z, Qin PW, Zeng HW. Unfolding Explainable AI for Brain Tumor Segmentation. Neurocomputing 2024;599:128058.
  31. Yang JL, Barnett AJ, Donnelly J, Kishore S, Fang J, Schwartz FR, Chen CF, Lo JY, Rudin C. FPN-IAIA-BL: A Multi-Scale Interpretable Deep Learning Model for Classification of Mass Margins in Digital Mammography. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2024:5003-9.
  32. Jalaboi R, Faye F, Orbes-Arteaga M, Jørgensen D, Winther O, Galimzianova A, Derm X. An end-to-end framework for explainable automated dermatological diagnosis. Med Image Anal 2023;83:102647. [Crossref] [PubMed]
  33. Wang X, Jiang L, Li L, Xu M, Deng X, Dai L, Xu X, Li T, Guo Y, Wang Z, Dragotti PL. Joint Learning of 3D Lesion Segmentation and Classification for Explainable COVID-19 Diagnosis. IEEE Trans Med Imaging 2021;40:2463-76. [Crossref] [PubMed]
  34. Yang K, Wu Z, Long J, Li W, Wang X, Hu N, Zhao X, Sun T. White matter changes in Parkinson's disease. NPJ Parkinsons Dis 2023;9:150. [Crossref] [PubMed]
  35. Zhang J. Mining imaging and clinical data with machine learning approaches for the diagnosis and early detection of Parkinson's disease. NPJ Parkinsons Dis 2022;8:13. [Crossref] [PubMed]
  36. Zhou L, Li G, Zhang Y, Zhang M, Chen Z, Zhang L, Wang X, Zhang M, Ye G, Li Y, Chen S, Li B, Wei H, Liu J. Increased free water in the substantia nigra in idiopathic REM sleep behaviour disorder. Brain 2021;144:1488-97. [Crossref] [PubMed]
  37. Hu J, Shen L, Albanie S, Sun G, Wu E. Squeeze-and-Excitation Networks. IEEE Trans Pattern Anal Mach Intell 2020;42:2011-23. [Crossref] [PubMed]
  38. Hasan MM, Alfaz N, Alam MAM, Rahman A, Shakhawat HM, Rahman S. Detection of Parkinson's Disease from T2-Weighted Magnetic Resonance Imaging Scans Using EfficientNet-V2. 2023 26th International Conference on Computer and Information Technology (ICCIT); Cox's Bazar, Bangladesh. 2023:1-6.
  39. Peng H, Gong W, Beckmann CF, Vedaldi A, Smith SM. Accurate brain age prediction with lightweight deep neural networks. Med Image Anal 2021;68:101871. [Crossref] [PubMed]
  40. Chattopadhay A, Sarkar A, Howlader P, Balasubramanian VN. Grad-CAM++: Generalized Gradient-Based Visual Explanations for Deep Convolutional Networks. 2018 IEEE Winter Conference on Applications of Computer Vision (WACV); Lake Tahoe, NV, USA. 2018:839-47.
  41. Ramírez VM, Kmetzsch V, Forbes F, Dojat M, Ieee, editors. Deep learning models to study the early stages of parkinson's disease. 2020 IEEE 17th International Symposium on Biomedical Imaging (ISBI); Iowa City, IA, USA. 2020:1534-7.
  42. Pineda-Pardo JA, Sánchez-Ferro Á, Monje MHG, Pavese N, Obeso JA. Onset pattern of nigrostriatal denervation in early Parkinson's disease. Brain 2022;145:1018-28. [Crossref] [PubMed]
  43. Otero-Jimenez M, Wojewska MJ, Binding LP, Jogaudaite S, Gray-Rodriguez S, Young AL, Gentleman S, Alegre-Abarrategui J. Neuropathological stages of neuronal, astrocytic and oligodendrocytic alpha-synuclein pathology in Parkinson's disease. Acta Neuropathol Commun 2025;13:25. [Crossref] [PubMed]
  44. Amandola M, Sinha A, Amandola MJ, Leung HC. Longitudinal corpus callosum microstructural decline in early-stage Parkinson's disease in association with akinetic-rigid symptom severity. NPJ Parkinsons Dis 2022;8:108. [Crossref] [PubMed]
  45. Uchida Y, Kan H, Sakurai K, Oishi K, Matsukawa N. Quantitative susceptibility mapping as an imaging biomarker for Alzheimer's disease: The expectations and limitations. Front Neurosci 2022;16:938092. [Crossref] [PubMed]
  46. Uchida Y, Kan H, Sakurai K, Arai N, Kato D, Kawashima S, Ueki Y, Matsukawa N. Voxel-based quantitative susceptibility mapping in Parkinson's disease with mild cognitive impairment. Mov Disord 2019;34:1164-73. [Crossref] [PubMed]
  47. Uchida Y, Kan H, Sakurai K, Inui S, Kobayashi S, Akagawa Y, Shibuya K, Ueki Y, Matsukawa N. Magnetic Susceptibility Associates With Dopaminergic Deficits and Cognition in Parkinson's Disease. Mov Disord 2020;35:1396-405. [Crossref] [PubMed]
  48. Shibata H, Uchida Y, Inui S, Kan H, Sakurai K, Oishi N, Ueki Y, Oishi K, Matsukawa N. Machine learning trained with quantitative susceptibility mapping to detect mild cognitive impairment in Parkinson's disease. Parkinsonism Relat Disord 2022;94:104-10. [Crossref] [PubMed]
Cite this article as: Shen Y, Qiao K, Hai J, Yang X, Hu T, Zhou Q, Wang M, Yan B. Structural and Statistical Knowledge-Enhanced Attention Network for early Parkinson’s disease diagnosis. Quant Imaging Med Surg 2026;16(2):137. doi: 10.21037/qims-2025-1468

Download Citation