Alzheimer’s disease prediction algorithm based on hippocampal longitudinal hybrid morphological features
Original Article

Alzheimer’s disease prediction algorithm based on hippocampal longitudinal hybrid morphological features

Jiaojiao Feng1, Kok Pin Ng2,3, Hua Wang1, Tao Yao1, Qingtang Su1, Maowen Ba4,5, Gang Wang6

1School of Information and Electrical Engineering, Ludong University, Yantai, China; 2Department of Neurology, National Neuroscience Institute, Singapore, Singapore; 3Duke-NUS Medical School, Singapore, Singapore; 4Department of Neurology, Affiliated Yantai Yuhuangding Hospital of Qingdao University, Yantai, China; 5Shandong Provincial Key Laboratory of Neuroimmune Interaction and Regulation, Yantai, China; 6Ulsan Ship and Ocean College, Ludong University, Yantai, China

Contributions: (I) Conception and design: J Feng, G Wang; (II) Administrative support: G Wang, H Wang; (III) Provision of study materials or patients: T Yao, Q Su, H Wang; (IV) Collection and assembly of data: M Ba, KP Ng; (V) Data analysis and interpretation: J Feng; (VI) Manuscript writing: All authors; (VII) Final approval of manuscript: All authors.

Correspondence to: Gang Wang, PhD. Ulsan Ship and Ocean College, Ludong University, No. 186 Hongqi Middle Road, Zhifu District, Yantai 264025, China. Email: gangwang1970@ldu.edu.cn; Maowen Ba, PhD. Department of Neurology, Affiliated Yantai Yuhuangding Hospital of Qingdao University, No. 20 Yuhuangding East Road, Zhifu District, Yantai 264000, China; Shandong Provincial Key Laboratory of Neuroimmune Interaction and Regulation, Yantai 264000, China. Email: bamaowen@163.com.

Background: Alzheimer’s disease (AD) is a progressive neurodegenerative disorder characterized by cognitive decline; this decline is closely linked to hippocampal morphological changes observed in structural magnetic resonance imaging (MRI). However, the existing AD prediction models have not fully explored the spatiotemporal correlation of hippocampal morphological features. To address this limitation, this study aims to develop a longitudinal prediction framework that captures both the temporal evolution and spatial distribution of hippocampal morphological alterations.

Methods: In this paper, we propose a novel deep learning framework for predicting the clinical progression of AD, which consists of a multi-view feature fusion convolutional network (M-FCN) and a bidirectional gated recurrent unit (Bi-GRU). The proposed M-FCN is based on the three-dimensional (3D) topological structure features of the hippocampus that introduces thickness features and heat kernel signature (HKS) to encode hippocampal morphological atrophy features. We utilize these features to construct a deep 3D hippocampus features description system for capturing the micro and macro structural changes of hippocampus. Hence, the task driven attention mechanism for prediction can effectively identify significant morphological changes caused by AD. The Bi-GRU module identifies inter sequence patterns and studies the temporal correlation between longitudinal features of hippocampus.

Results: The proposed method was evaluated using longitudinal T1-weighted MRI data from the Alzheimer’s Disease Neuroimaging Initiative (ADNI) (n=221). Compared with the existing AD prediction models, the correspondence between AD-related structural changes and clinical neurodegeneration indicators can be more accurately captured by our proposed deep learning model. The predictive performance was evaluated using root mean square error (RMSE), correlation coefficient (CC), and 95% confidence interval (CI). For the prediction of Mini-Mental State Examination (MMSE) scores, the model achieved a RMSE of 2.34 (95% CI: 2.27–2.45, CC =0.72) at M18, 2.58 (95% CI: 2.52–2.64, CC =0.77) at M24, and 2.60 (95% CI: 2.54–2.66, CC =0.83) at M36.

Conclusions: These results highlight the effectiveness of the proposed model in leveraging the spatiotemporal correlation of hippocampal morphology to provide high accuracy and reliable predictions.

Keywords: Alzheimer’s disease (AD); hippocampal morphological atrophy features; multi-view feature; spatiotemporal correlation


Submitted Feb 16, 2025. Accepted for publication Dec 05, 2025. Published online Jan 23, 2026.

doi: 10.21037/qims-2025-377


Introduction

Alzheimer’s disease (AD) is a neurodegenerative disorder characterized by progressive cognitive decline that poses a significant threat to the physical and mental health of the elderly (1). Currently, more than 50 million people worldwide suffer from AD, a number that is projected to rise to 138 million by 2050 (2). This creates a substantial burden on individuals, families, and society at large. Although AD is irreversible and incurable, early intervention can slow its progression, especially in individuals who are asymptomatic or exhibit only mild symptoms. Mild cognitive impairment (MCI) is a prodromal stage between normal cognitive aging and the onset of AD pathology. It is estimated that 10–12% of MCI patients progress to AD annually, and 80% develop AD during a follow-up period of 6 years (3). Therefore, accurately predicting individuals in the MCI stage is crucial for AD prevention and potential therapeutic interventions.

Currently, the cognitive assessments such as the Mini-Mental State Examination (MMSE), the Alzheimer’s Disease Assessment Scale Cognitive Subscale (11-item) (ADAS11) and ADAS13 (4,5) have been validated to predict the progression from MCI to AD (6,7). However, these tests depend to some extent on participants’ responses and are influenced by subjective factors, which may limit the accuracy of the diagnostic process (8). Therefore, there is an urgent need to establish reliable biomarkers that can reflect the dysfunction of brain structural networks as predictor of progressive cognitive decline in elderly individuals with MCI. Structural magnetic resonance imaging (sMRI) is widely used in clinical diagnosis of AD due to its non-invasive nature and provides detailed anatomical information of brain structures (9-11). Changes in the whole and sub-regional brain have been demonstrated to be correlated with the progression of AD (12,13). Furthermore, brain morphological changes detected by sMRI is closer to clinical symptoms compared to amyloid-beta and tau pathology and occurs prior to the onset of clinical symptoms (14,15). Among the biomarkers of sMRI, hippocampal atrophy begins at the preclinical phase and progresses steadily throughout MCI and AD, which has also been shown to be associated with cognitive decline (16,17). Therefore, studying longitudinal changes in the hippocampus during disease progression can help us study the pathology of AD and serve as a biomarker for predicting neurodegeneration.

Deep learning has been widely applied in the field of medical image analysis and has achieved significant results in tasks such as segmentation and classification (18-20). In the field of AD research, deep learning has likewise shown great potential (21,22). For instance, a recent study proposed a multi-modal fusion framework named dual-3DM3-AD, which integrates MRI and positron emission tomography (PET) data for early Alzheimer’s diagnosis, achieving an accuracy of 98% and outperforming existing approaches in multi-class AD classification. In addition, many AD studies have utilized advanced technologies such as deep learning and machine learning, successfully transforming AD diagnosis and prognosis tasks into regression or classification problems (23-27). However, the prediction performance of these methods depends on the extraction of salient features at single time points. As the development of AD is a gradual process, existing models are often unable to fully explore the temporal correlation between morphological changes caused by AD, ultimately making it difficult to accurately predict the progression of AD.

To make more effective use of longitudinal data, several studies used deep recurrent neural networks (28) and their variants to mine the correlation of biological features in time series (29,30). For example, Jung et al. (31) used a deep recurrent network to capture potential temporal correlations in longitudinal data. Volumes of each anatomical region of interest (ROI) were extracted from the time series to predict the trajectory of clinical status at multiple future time points. However, a single volume feature is not conducive to comprehensively describe the topological and structural changes caused by AD, making it difficult to fully explore the correlation between temporal data. Moreover, another work (32) also focuses on predicting AD progression and has achieved promising results, but it relies on a framework that feature extraction is completely independent of model training. making it difficult to fully explore the correlation between temporal data, thereby affecting the accuracy of prediction.

In addition, it is worth noting that different neurodegenerative diseases affect distinct brain regions, which means that not every brain region is equally important when diagnosing a specific etiology of the neurodegenerative disease. For example, a complete topology and morphological description of the hippocampus may result in redundant feature information that is not closely related to the development of AD, making it difficult to identify significant biomarkers associated with AD pathology, and also leading to high computational complexity. Lian et al. (33) proposed an end-to-end deep architecture for a multi-task weakly-supervised attention network (MWAN), in which attention mechanisms are embedded to automatically identify subject specific discriminative positions, achieving a certain level of effectiveness. Therefore, it is necessary to incorporate feature weighted or filtering mechanisms, such as attention mechanisms and pooling layer, into the feature encoding and feature fusion modules in order to select important feature information.

In this study, we propose a deep learning framework that combines multi-view feature fusion convolutional network (M-FCN) feature extraction model and recurrent neural network for predicting clinical scores of AD. The main aims of this study are summarized as follows:

  • Firstly, we develop a multi-view feature extraction framework which integrates topological features and surface morphological features of hippocampal surface to comprehensively describe AD-induced morphological changes from local to global perspectives.
  • Next, a collaborative learning framework is developed to reduce cumulative errors and improve model performance. Our method enhances the algorithm’s generalizability by simultaneously performing feature extraction and sequence learning through forward propagation.
  • Lastly, we integrate attention into feature aggregation to enable the model to prioritize critical morphological features, enhancing predictive accuracy by focusing on disease-related morphological changes.

We present this article in accordance with the TRIPOD+AI reporting checklist (available at https://qims.amegroups.com/article/view/10.21037/qims-2025-377/rc).


Methods

This study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments.

Subject

We downloaded the training and testing data from the Alzheimer’s Disease Neuroimaging Initiative (ADNI) database (34), which can be publicly accessed on the website (https://adni.loni.usc.edu/). The main purpose of the ADNI database is to test biomarkers such as MRI and PET, as well as clinical and neuropsychological evaluations, combined with the information provided above to measure the progress of MCI and AD. ADNI is the result of efforts of many co-investigators from a broad range of academic institutions and private corporations. Subjects have been recruited from over 50 sites across the U.S. and Canada. Since the launch of data collection in 2004 (ADNI-1), the ADNI database has continuously included imaging, clinical, and genetic data from multiple patients in four studies: ADNI1, ADNI2, ADNIGO, and ADNI3. The inclusion criteria require participants to be at least 55 years old, have a neurological examination to rule out major illnesses, and be able to complete a full set of standardized neuropsychological assessments.

In this study, we used the first screening date of the subjects as the baseline and calculated the follow-up time from the baseline. M06 and M12 represent the 6th and 12th month after baseline, respectively. We used MRI data from six time points: baseline (BL), 6th month (M06), 12th month (M12), 18th month (M18), 24th month (M24), and 36th month (M36), and combined with four clinical indicators: MMSE, ADAS11, ADAS13, and Clinical Dementia Rating-Sum of Boxes (CDR-SB) for comprehensive analysis, missing values of continuous variables are filled with the training set mean (missing rate <5%), as shown in Table 1.

Table 1

Statistical information on clinical cognitive scores of participants during each time period

Longitudinal times CDR-SB ADAS11 ADAS13 MMSE
BL (n=221)
   NC (n=106) 0.02±0.09 5.85±2.82 8.97±4.02 29.18±0.94
   MCI (n=115) 1.63±0.84 12.08±4.52 19.73±6.50 26.85±1.77
M06 (n=221)
   NC (n=106) 0.05±0.23 6.06±3.01 9.24±4.48 29.06±0.99
   MCI (n=115) 2.01±1.22 12.98±5.50 21.11±7.33 26.07±2.65
M12 (n=221)
   NC (n=106) 0.09±0.32 5.56±2.94 8.75±4.58 29.19±1.13
   MCI (n=115) 2.41±1.42 13.90±6.16 22.07±7.95 25.64±2.83
M18 (n=115)
   MCI (n=115) 2.94±1.76 15.17±7.44 23.98±9.82 24.96±3.40
M24 (n=208)
   NC (n=106) 0.18±0.55 5.78±3.14 9.22±4.86 29.03±1.16
   MCI (n=102) 3.62±2.19 15.86±7.58 25.21±9.94 24.09±4.09
M36 (n=186)
   NC (n=106) 0.33±0.87 5.86±3.21 9.51±5.02 28.90±1.22
   MCI (n=80) 4.46±2.90 17.95±9.76 27.41±12.49 23.12±5.06

Data are presented as mean ± standard deviation. BL, baseline; M06, 6th month; M12, 12th month; M18, 18th month; M24, 24th month; M36, 36th month. ADAS11, Alzheimer’s Disease Assessment Scale Cognitive Subscale (11-item); ADAS13, Alzheimer’s Disease Assessment Scale Cognitive Subscale (13-item); CDR-SB, Clinical Dementia Rating-Sum of Boxes; MCI, mild cognitive impairment; MMSE, Mini-Mental State Examination; NC, negative control.

Prediction model framework

Our proposed deep learning framework consists of M-FCN and bidirectional gated recurrent unit (Bi-GRU) to predict clinical scores during the progression of AD. M-FCN is responsible for capturing the comprehensive features of the left and right hippocampus (RH) of the input individual at each time point. Bi-GRU takes a sequence of time-series data as input, capturing the temporal correlations of data from the same individual for predicting future clinical scores.

The proposed framework consists of three main stages (as shown in Figure 1). In stage 1, necessary preprocessing steps are performed for MRI data, including bias field correction, skull stripping, template space registration, segmentation, reconstruction of the hippocampus and surface registration. In stage 2, the M-FCN module extracts deep features from the input hippocampus. In order to capture longitudinal depth features, the triangular mesh representation of the same hippocampus at each time step is passed to M-FCN to obtain the depth features. In the final stage, the extracted features are input into the Bi-GRU model, where the deep features of each time point are used as inputs to learn the temporal correlation between longitudinal data.

Figure 1 Model architecture: M-FCN for feature extraction and dual-GRU for AD clinical score prediction. AD, Alzheimer’s disease; ADAS, Alzheimer’s Disease Assessment Scale Cognitive Subscale; ADNI, Alzheimer’s Disease Neuroimaging Initiative; BL, baseline; BN, batch normalization; CDR-SB, Clinical Dementia Rating-Sum of Boxes; CNN, convolutional neural network; GRU, gated recurrent unit; HKS, heat kernel signature; M06, 6th month; M-FCN, multi-view feature fusion convolutional network; MMSE, Mini-Mental State Examination.

Image preprocessing

The MRI image preprocessing used the FMRIB’s Integrated Registration and Segmentation Tool (FIRST) tool from FMRIB Software Library (FSL) (35). For each T1 weighted MRI image, automatic segmentation was first performed and then spatially normalized to the Montreal Neurological Institute (MNI) template space through linear transformation. During this process, head tilt and alignment correction were performed to ensure spatial consistency and accuracy of the image. Subsequently, based on the segmentation results, the surface of the hippocampus was reconstructed using an automatic reconstruction algorithm. In order to effectively deal with the noise introduced by image scanning, the reconstructed surface of the hippocampus was smoothed. Finally, the spherical harmonic registration method was used to establish precise one-to-one morphological correspondence between different hippocampal surfaces (36). Thus, all the registered hippocampal surfaces have exactly same number of vertices and triangles, which lays the foundation for subsequent analysis of hippocampal morphological features.

Feature learning via M-FCN

M-FCN is a novel deep learning-based feature extraction framework that adopts a dual-stream architecture that independently processes both topological features and morphological features from left hippocampus (LH) surface and RH surface. This architecture is designed to capture a comprehensive structure of hippocampal topological and morphological features from a multi-view perspective.

The topological features were derived from the triangular mesh representation of hippocampal surface. Specifically, two key topological properties were extracted: the center coordinates of the surface elements, denoted as cvecRN×3, where N denotes the number of triangular surface elements, and the corresponding normal vectors, denoted as nvecRN×3. These topological features encode the global spatial orientation and shape of hippocampal surface.

For the morphological features, we utilize two descriptors to quantify the degree of morphological deformation on hippocampal surface. The first descriptor is the HKS measurement (37), denoted as HRN×12, which captures the intrinsic shape properties, such as multi -scale local structural features through heat diffusion across hippocampal surface. The second descriptor is the local hippocampal thickness measurement (38), denoted as FRN×3, which represents the shortest distance from the mid-axis to each vertex on the hippocampal triangular mesh, and it can well reflect the expansion or atrophy of hippocampal surface. To facilitate features aggregation, we incorporate neighborhood indices that define the adjacency relationships between surface units, enabling localized interaction between features and enhancing capacity to learn spatial correlations across each hippocampal surface. These multi-features of LH or RH are processed through distinct convolutional layers in each stream of the M-FCN, allowing the network to learn hierarchical representations of hippocampal structural information. By jointly processing the LH and RH, the dual-stream architecture is able to obtain a more comprehensive structural changes of the bilateral hippocampi induced by AD.

Topological features encoder

For the center point of each face element (cvecRN×3) encoding, the multi-layer perceptron (MLP) (39) network is used for nonlinear mapping to extract relevant spatial points distribution characteristics. For the normal vector of each face element (nvecRN×3) encoding, we calculated the spatial correlation between a set of learnable kernel vectors and the normal vectors of face elements. Here, the learnable kernel vectors were defined as a set of task-oriented vector sets denote as kvec. In practical applications, since normal vectors were mostly unit vectors, we used a spherical coordinate system to model the learnable kernel vector set, effectively considering the spatial geometric and topological relationship between the learnable kernel vectors and the normal vectors. Next, we calculated the correlation matrix between the ith vector and the jth kernel vector are calculated by using the following formula:

KC(i,j)=exp(||nvecikvecj||22σ2)

· represents the distance between two vectors, hyperparameter σ strives for a balance between fitting ability and tolerance for complex data. By using M learnable kernel vectors, each face can obtain a multidimensional spatial correlation matrix of N×M, which can be used to describe local morphological differences compared to the M learnable discriminant vectors.

Morphological features encoder

Thickness calculation is a common measurement of surface morphology deformation that can reflect the expansion or atrophy of relevant brain regions. For the thickness measurement encoder, we first used our previous work (40) to obtain a series of longitudinal eigen loops for each subject via the eigen graph analysis. Next, we computed the center of each eigen loop to construct the geometric centerline of each hippocampal surface. In each eigen loop, the thickness of each vertex is defined as the radial distance between each vertex and the center point. Finally, the original thickness measurements of all the faces where the thickness of each face was composed of the thickness measurements of three vertices, were processed through two layers of one-dimensional convolution, normalization, and activation function to be encoded into a N×64 matrix.

HKS is a local morphological feature description method based on heat diffusion theory, whose core idea is to simulate the process of heat propagation on the surface of an object at different time scales, in order to quantify the influence of geometric structure on heat diffusion behavior. In the high curvature areas of the hippocampus (such as areas with more wrinkles), the complex geometric structure limits the propagation of heat, resulting in a slower rate of heat diffusion. On the contrary, in areas of smooth and low curvature on the surface of the hippocampus, the speed of heat diffusion is faster because simpler geometric structures allow for rapid heat propagation. Therefore, HKS can not only comprehensively capture the overall geometric features of the hippocampus, but also reflect in detail the local morphological changes in different regions of the hippocampus.

The encoding process for the HKS is elaborated as follows. Considering the amount of heat transferred from point x to point y at a given time t, we defined a heat kernel as the following:

kt(x,y)=i=0eλitϕi(x)ϕi(y)

Where λi and ϕi are the ith eigenvalue and ith eigenfunction of the Laplace-Beltrami operator (41) on Riemannian manifold M, the Laplace-Beltrami operator is a generalization of the Laplace operator to Riemannian manifolds, defined as the divergence of the gradient. It essentially characterizes how a function’s value at a given point differs from its neighborhood average, playing a fundamental role in geometric shape analysis.

The heat kernel can be interpreted as the transition density function of Brownian motion on the manifold, which can be used to stably capture all of information about the intrinsic morphology of the manifold. On this basis, Sun et al. (37) further proposed a heat kernel signature (HKS) (x) of a given point x on the M, which is defined as a function over the temporal domain:

HKS(x):+,HKS(x,t)=kt(x,x)=i=0eλitφi2(x)

Thus, the HKS{kt(x,x)}t>0 encodes the multi-scale morphological information about the neighborhoods of a point x through setting different time scales. After we calculated the HKS features of a triangular mesh and obtained an N×T matrix, where N is the number of triangular mesh elements and T is the time scale of heat diffusion, an encoder step was applied which involves a two-layer one-dimensional convolution process, coupled with normalization and activation functions. Finally, we obtained N×64 encoded HKS representation.

Feature combination and feature aggregation

To enhance the representation ability of the encoded features, and capture more discriminative and comprehensive patterns from the hippocampus, we adopted a method of common feature fusion (42), amalgamating different forms of the encoded features to generate a richer feature representation. That is, we executed a two-layer feature-level fusion to combine the multiple features other than HKS, yielding a novel feature representation. Following this, an MLP was applied to transform the initial features of varying facets into an N×512 feature matrix. Additionally, a skip connection (43) mechanism was employed to merge features twice during the feature fusion process.

To broaden the perceptual scope of each face region, we performed feature aggregation based on previously encoded features and neighborhood indices. By concatenating rich features of each face with those of its neighboring faces along a specific dimension, we constructed novel aggregated features (44) to ensure robustness of the feature representation. Due to the inherent multi-scale feature representation in HKS, we chose not to include HKS features in the feature aggregation process to mitigate potential overfitting of the model. During the feature aggregation process, we aggregated features of local face elements, thus excluding spatial positional information (center).

Specifically, we define a connectivity aggregation that concatenates the self-features of each face element with its neighbor features, followed by processing the aggregated features through an MLP network. This process involved dimension transformation, activation, linear transformation operations on the aggregated features to extract output features for each face. Subsequently, we extracted the maximum value of each output feature through pooling operations to obtain advanced features connected with the fused features. Finally, the fused features were concatenated with the encoded HKS features to generate a comprehensive discriminative feature representation of the hippocampus. The process of feature combination and feature aggregation is shown in Figure 2.

Figure 2 Feature combination and feature aggregation process. HKS, heat kernel signature; MLP, multi-layer perceptron.

Attention mechanism

We introduced an attention mechanism to refine the feature representation of hippocampal structures by focusing on relevant regions. Specifically, let the thickness measurements of the hippocampal surface be denoted as FN×3, where N is the number of faces on the surface, and each element FiF represents the thickness value of the ith face. For each face i, let N(i) denote the set of indices of its adjacent faces (including itself), with N(i)=K (where K is the neighborhood size, set to 3 in our implementation based on mesh connectivity). Then, the contextual thickness features FconcatN×2 is obtained by maximum pooling and average pooling within each face. The formula is expressed as follows:

Favg=avgjN(i)Fj,Fmax=maxjN(i)FjFconcat=[Favg,Fmax]

Next, a shared 2×1 convolution is applied to Fconcat, followed by a sigmoid activation function to generate the attention map Wc of thickness measurements across all faces. Finally, the original thickness feature F is reweighted by performing element-wise multiplication with the attention map Wc:

WC=σ[Conv(Fconcat)]F=WCF

This attention map represents the weights of thickness measures on all faces driven by the prediction task. These weights are applied to the input thickness measurements, allowing the network to focus on the local morphological measurements with significant predictive effects while suppressing those with less significant predictive effects. The attention module is shown in Figure 3.

Figure 3 Spatial attention mechanism diagram.

Sequential feature learning based on Bi-GRU

GRU is an advanced variant of RNN that utilizes gating mechanisms to learn long-term dependencies of sequential data (45). Compared to long short-term memory (LSTM), GRU reduces the number of learning parameters by simplifying gated control unit structure, achieving comparable performance and improving computational efficiency. In order to better capture contextual information in sequential data, we used a Bi-GRU to simultaneously focus on the past and future time information of the input data. The input of the Bi-GRU includes the features of LH and RH generated by the M-FCN at different time points and the output of the Bi-GRU from the connection of the forward and backward GRU layers is used to create the prediction vectors for clinical indicators.

Assuming there are N longitudinal hybrid features generated by the M-FCN module {Xn}n=1N, where Xn={x1i,xti,,xTi}RI×T, here n represents nth subject, T and I represent the dimensions of time and features, respectively. Then we use Bi-GRU network shown in Figure 1 to capture the temporal dependencies within these hybrid features.

To model the temporal dependencies within this multivariate data, we employed the GRU network, which exceled in capturing sequential patterns and relationships over time. The specific internal operations of our encoding module are as follows:

uti=σ(Wuxti+Uuht1i+bz)rti=σ(Wrxti+Urht1i+br)h˜ti=tanh[Whxti+Uh(rtht1i)+bh]hti=uth˜ti+(1ut)ht1i

Where represents Hadamard product, when AD induces significant short-term structural changes in the hippocampus, the reset gate rti in the GRU can quickly respond and capture these changes, accurately reflecting the current state of the hippocampus. In contrast, the update gate uti captures the long-term morphological trends exhibited by the hippocampus during the progression of AD by controlling the proportion of historical state information that is retained. By combining these two gating mechanisms, the GRU model can capture both short-term and long-term structural changes in the hippocampus. Bi-GRU consists of two parallel GRU layers that handle time-series inputs from both forward and backward directions, respectively. The calculation formula is as follows:

hti=GRUf(xti,ht1i)hti=GRUb(xti,ht1i)

In the Bi-GRU model, GRUf and GRUb refer to the input layers in the forward and backward directions, respectively. The spatiotemporal features learned by Bi-GRU are obtained through forward and backward output vectors hf and hb to represent. We trained the model at BL, M06, and M12 time points. During the training process, we fully considered the spatial correlation between hippocampal morphology at different time points and the temporal correlation between different time points. To verify the effectiveness and generalization ability of the trained model, we used the data from the M18, M24, and M36 as the validation test set.


Results

Experimental settings

This study was conducted on NVIDIA Titan XP servers based on the PyTorch (46) framework. For practical deployment, the minimum hardware requirements include an NVIDIA GPU (≥8 GB VRAM) and 16 GB system memory. To ensure clinical practicality, the proposed framework was systematically evaluated in terms of computational efficiency and scalability. The complete pipeline consists of two stages: hippocampal segmentation and feature extraction, followed by the predictive modeling stage. The initial hippocampal processing requires only standard neuroimaging preprocessing time (a few minutes per scan), and this step can be seamlessly integrated into existing hospital workflows.

The core prediction model contains approximately 7.52 million parameters with a computational cost of 1.17 GFLOPs, representing a lightweight architecture that is substantially more efficient than conventional three-dimensional (3D)-convolutional neural network (CNN)-based models while maintaining high predictive performance. Memory usage analysis across different sequence lengths revealed a consistently low footprint of about 28.7 MB during inference, confirming the scalability of the model for longitudinal data processing without memory bottlenecks. This decoupled design allows the computationally intensive feature extraction to align with standard clinical pipelines, enabling near real-time prediction once features are prepared.

During the model prediction process, due to the significant impact of hyperparameter selection on experimental results, we set appropriate parameters during the training process to ensure the optimal performance of the model. Specifically, the learning rate, the batch size, the number of training epochs and the L1 regularization coefficient are set as 0.00008, 8, 100, and 0.1, respectively. These parameter settings have been validated through multiple experiments, aiming to balance the training efficiency and prediction accuracy of the model. In order to optimize the performance of the prediction task, we chose the mean squared error (MSE) as the loss function. MSE can effectively measure the difference between predicted values and true values, helping the model continuously adjust its parameters during training to minimize errors.

In our experiment, we used the proposed method to predict four clinical scores including MMSE, ADAS11, ADAS13 and CDR-SB for participants at M18, M24, and M36, all clinical scores are normalized. We calculated the root mean square error (RMSE) (47) and Pearson correlation coefficient (CC) between the predicted value and the actual clinical score as evaluation criteria, using the following formula:

RMSE=1Ni=1N(yiy^i)2,CC=cov(Y,Y^)σ(Y)σ(Y^)

Where Y and Y^ represent the true value vector and the predicted score vector for the four clinical scores, respectively. yi is the factor of Y, y^i is the factor of Y^. RMSE reflects the degree of deviation between predicted clinical scores and actual clinical scores, The smaller the RMSE value, the stronger the predictive ability of the model, if the RMSE is less than the Minimum Clinically Important Difference of a certain scale, it is generally considered that the model’s prediction results have certain reference value at the individual level. CC reflects the correlation between actual clinical scores and predicted clinical scores. The absolute value of the CC closer to 1 indicate a strong linear relationship, while the absolute values closer to 0 suggest weak or no linear association.

Experimental result

Longitudinal data predictions

To evaluate the effectiveness of our proposed M-FCN combined with Bi-GRU prediction model, we used BL, M06, and M12 as training sets and validated our prediction model on the clinical scores of M18, M24, and M36 validation sets. The following is the clinical score predicted using Bi-GRU after feature extraction using pretrained M-FCN. As shown in Figures 4,5, the RMSE and CC gradually increased with increasing prediction time. Specifically, the RMSE of MMSE forecasts was 2.34 [95% confidence interval (CI): 2.29–2.39] at M18, 2.58 (95% CI: 2.52–2.64) at M24, and 2.60 (95% CI: 0.54–2.66) at M36. This trend suggests that the accumulation of prediction errors over longer time frames may lead to an increase in RMSE. However, despite the slightly higher RMSE, the increase in CC values (0.72, 0.77, and 0.83 for M18, M24, and M36, respectively) suggests that the model is better able to capture long-term patterns of cognitive decline, making the predicted scores more consistent with actual clinical trajectories.

Figure 4 RMSE of MMSE, ADAS13, ADAS11, and CDR-SB at three time points. M18, 18th month; M24, 24th month; M36, 36th month. ADAS11, Alzheimer’s Disease Assessment Scale Cognitive Subscale (11-item); ADAS13, Alzheimer’s Disease Assessment Scale Cognitive Subscale (13-item); CI, confidence interval; CDR-SB, Clinical Dementia Rating-Sum of Boxes; CI, confidence interval; MMSE, Mini-Mental State Examination; RMSE, root mean square error.
Figure 5 Scatter plots of true and predicted values for the four indicators at three time points: (A) MMSE; (B) ADAS13; (C) ADAS11; and (D) CDR-SB (with time points M18, M24, and M36 presented for each indicator). M18, 18th month; M24, 24th month; M36, 36th month. ADAS11, Alzheimer’s Disease Assessment Scale Cognitive Subscale (11-item); ADAS13, Alzheimer’s Disease Assessment Scale Cognitive Subscale (13-item); CC, correlation coefficients; CDR-SB, Clinical Dementia Rating-Sum of Boxes; MMSE, Mini-Mental State Examination.

Comparison of different recurrent neural network

We compared the performance of Bi-GRU and Transformer architectures using MMSE as the prediction target. As shown in Figure 6, the predictive score coefficients of Transformer for M18, M24, and M36 are 0.64, 0.67, and 0.75, respectively. Compared with the performance of Bi-GRU mentioned above, the predictive performance of Transformer has slightly decreased. The performance difference is mainly attributed to the data characteristics rather than the model capacity. The longitudinal AD dataset in this study has limited sample size and short temporal sequences (3–5 time points per subject). Under such conditions, the Transformer’s global attention mechanism may not fully exploit its advantages and tends to be sensitive to local noise. In contrast, the Bi-GRU’s gated structure provides a more stable learning process for short, gradual temporal dependencies commonly observed in AD progression.

Figure 6 Predicted versus actual MMSE trajectories generated by the Transformer model at three time points: (A) M18; (B) M24; and (C) M36. M18, 18th month; M24, 24th month; M36, 36th month. CC, correlation coefficient; MMSE, Mini-Mental State Examination.

Furthermore, to demonstrate the effectiveness of Bi-GRU in our experiments, we compared Bi-GRU with commonly used recurrent neural networks (RNNs). As shown in the Table 2, clearly indicate that our proposed M-FCN combined with Bi-GRU achieves the best predictive performance among the commonly used RNNs. Compared to bidirectional LSTM (Bi-LSTM), the RMSE of ADAS11 score predictions at M18, M24, and M36 decreased by 0.67, 0.73, and 1.16 respectively. Compared to bidirectional RNN (Bi-RNN), the RMSE of ADAS11 score predictions at M18, M24, and M36 decreased by 0.93, 1.2, and 1.54 respectively. These results demonstrate the advantages of Bi-GRU, further validating its effectiveness in neurodegenerative disease prediction and providing strong support for related research.

Table 2

Comparison of RMSE and CC for different recurrent neural networks

Method Times RMSE CC
MMSE CDR-SB ADAS11 ADAS13 MMSE CDR-SB ADAS11 ADAS13
Bi-GRU M18 2.34 1.22 4.47 6.43 0.72 0.77 0.82 0.85
M24 2.58 1.96 4.76 6.74 0.77 0.72 0.82 0.88
M36 2.60 2.11 5.50 6.80 0.83 0.84 0.87 0.90
Bi-RNN M18 2.58 1.24 5.40 8.10 0.65 0.73 0.67 0.74
M24 2.89 1.92 5.96 8.76 0.70 0.71 0.64 0.83
M36 2.84 2.28 7.04 7.79 0.78 0.71 0.66 0.81
Bi-LSTM M18 2.43 1.43 5.14 6.85 0.72 0.75 0.85 0.86
M24 2.69 2.13 5.62 8.38 0.77 0.81 0.83 0.89
M36 2.64 2.44 6.66 9.31 0.86 0.89 0.86 0.89

M18, 18th month; M24, 24th month; M36, 36th month. ADAS11, Alzheimer’s Disease Assessment Scale Cognitive Subscale (11-item); ADAS13, Alzheimer’s Disease Assessment Scale Cognitive Subscale (13-item); Bi-GRU, bidirectional gated recurrent unit; Bi-LSTM, bidirectional long short-term memory; Bi-RNN, bidirectional recurrent neural network; CC, correlation coefficient; CDR-SB, Clinical Dementia Rating-Sum of Boxes; MMSE, Mini-Mental State Examination; RMSE, root mean square error.

Performance comparison using different time series data as input

To verify whether using multiple time series data as input is beneficial for improving the predictive performance of the model, we designed two different longitudinal training sets to predict the clinical scores at future time points. Firstly, we use training data from BL and M06 time points to predict future clinical scores. In this experiment, the model was trained by capturing the patient’s changing trends in the initial stage. However, due to the slow disease progression of the subjects during early follow-up, this method may not provide accurate predictions. We further combined data from three time points (BL, M06, and M12) in our model. As illustrated in Figure 7, by adding data from this time point of 12 months, the model can better predict the patient’s cognitive changing patterns over time and capture more temporal dynamic features. This result is consistent with the findings in the study (31). Our findings indicate that compared to less time series data, using multiple time series data as inputs is more conducive to improving the predictive performance of the model.

Figure 7 Comparison of experimental performance between two training sets. (A) CC of MMSE; (B) the CC of ADAS11; (C) the CC of ADAS13; (D) the CC of CDR-SB; (E) RMSE value of MMSE; (F) the RMSE value of ADAS11; (G) the RMSE value of ADAS13; (H) the RMSE value of CDR-SB. BL, baseline; M06, 6th month; M12, 12th month; M18, 18th month; M24, 24th month; M36, 36th month. ADAS11, Alzheimer’s Disease Assessment Scale Cognitive Subscale (11-item); ADAS13, Alzheimer’s Disease Assessment Scale Cognitive Subscale (13-item); CC, correlation coefficient; CDR-SB, Clinical Dementia Rating-Sum of Boxes; MMSE, Mini-Mental State Examination; RMSE, root mean square error.

Comparison of topological and morphological features with volumetric features

In this work, we used topological and morphological features as comprehensive structural characteristics of hippocampal surface, which can capture the subtle changes affected by AD. It is worth noting that the majority of current research (48,49) uses the volume of brain regions as the features to perform AD detection tasks. These methods attempt to identify the progression of AD by analyzing the volume changes in different regions of brain. To verify the predictive performance of multiple morphological features based on the M-FCN and volume measures, we conducted the comparative experiment using the same cohort in Table 1. The volume measurements in the experiment were set as the sum of the volumes of LH and RH. As shown in Table 3, the comparative results revealed that multiple morphological features extracted based on M-FCN had higher accuracy and robustness in terms of predicting individual clinical scores. This finding indicated that volume measurement may not fully capture subtle changes in hippocampus structure, especially in the early stages of AD.

Table 3

Comparison of RMSE and CC when using hippocampal volume and M-FCN to extract features for prediction

Method Times RMSE CC
MMSE CDR-SB ADAS11 ADAS13 MMSE CDR-SB ADAS11 ADAS13
vol M18 2.76 1.46 6.23 7.65 0.70 0.79 0.85 0.88
M24 3.71 2.17 6.94 8.32 0.69 0.76 0.84 0.89
M36 4.10 2.56 7.87 9.61 0.76 0.82 0.85 0.89
M-FCN M18 2.34 1.22 4.47 6.43 0.72 0.77 0.82 0.85
M24 2.58 1.96 4.89 6.74 0.77 0.72 0.82 0.88
M36 2.60 2.11 5.50 6.80 0.83 0.84 0.87 0.90

M18, 18th month; M24, 24th month; M36, 36th month. ADAS11, Alzheimer’s Disease Assessment Scale Cognitive Subscale (11-item); ADAS13, Alzheimer’s Disease Assessment Scale Cognitive Subscale (13-item); CC, correlation coefficient; CDR-SB, Clinical Dementia Rating-Sum of Boxes; M-FCN, multi-view feature fusion convolutional network; MMSE, Mini-Mental State Examination; RMSE, root mean square error; vol, volume.

Ablation experiment

To verify the importance of heat kernel features and thickness features in clinical score prediction, an ablation experiment was conducted. Specifically, we sequentially removed thickness features and heat kernel features from the model, and evaluated the predictive performance of four clinical scores (MMSE, ADAS11, ADAS13, and CDR-SB) at three follow-up time points (M18, M24, M36). The complete M-FCN model, which integrates both HKS and thickness features, was used as the baseline for all ablation comparisons.

The experimental results after removing thickness features are shown in Figure 8. Taking MMSE prediction as an example: compared with the baseline, the CC decreased from 0.72 to 0.69 at M18 (a 4.17% reduction), with a corresponding increase in RMSE from 2.34 to 2.37 (1.28% increase); at M24, CC dropped from 0.77 to 0.74 (3.90% reduction) and RMSE slightly rose from 2.58 to 2.59 (0.39% increase); at M36, CC decreased by 2.41% (from 0.83 to 0.81) and RMSE increased by 1.54% (from 2.60 to 2.74). Notably, the performance degradation became more pronounced as the prediction horizon extended, indicating that thickness features contribute more to long-term clinical score prediction. For other clinical scores (ADAS11, ADAS13, CDR-SB), consistent trends were observed, which confirming the indispensable role of thickness features.

Figure 8 CC of four clinical indicators at M18, M24, and M36 after removing thickness features. (A) Predicted vs. actual scatter plot of MMSE; (B) predicted vs. actual scatter plot of ADAS11; (C) predicted vs. actual scatter plot of ADAS13; (D) predicted vs. actual scatter plot of CDR-SB. M18, 18th month; M24, 24th month; M36, 36th month. ADAS11, Alzheimer’s Disease Assessment Scale Cognitive Subscale (11-item); ADAS13, Alzheimer’s Disease Assessment Scale Cognitive Subscale (13-item); CC, correlation coefficient; CDR-SB, Clinical Dementia Rating-Sum of Boxes; MMSE, Mini-Mental State Examination.

The experimental results obtained by removing heat kernel features are presented in Figure 9. Consistent with the above observations, the model also exhibited significant performance degradation relative to the baseline. For MMSE prediction: at M18, the CC decreased from 0.72 to 0.68 and RMSE increased from 2.34 to 2.40. For ADAS11, ADAS13, and CDR-SB, removing heat kernel features also led to reduced CCs and elevated RMSE values. These results demonstrate that heat kernel features, like thickness features, play a critical role in supporting the model’s clinical score prediction performance.

Figure 9 CC of four clinical indicators at M18, M24, and M36 after removing HKS features. (A) Predicted vs. actual scatter plot of MMSE; (B) predicted vs. actual scatter plot of ADAS11; (C) predicted vs. actual scatter plot of ADAS13; (D) predicted vs. actual scatter plot of CDR-SB. M18, 18th month; M24, 24th month; M36, 36th month. ADAS11, Alzheimer’s Disease Assessment Scale Cognitive Subscale (11-item); ADAS13, Alzheimer’s Disease Assessment Scale Cognitive Subscale (13-item); CC, correlation coefficient; CDR-SB, Clinical Dementia Rating-Sum of Boxes; HKS, heat kernel signature; MMSE, Mini-Mental State Examination.

External validation

To evaluate the generalization capability of the proposed model, an external validation was conducted on the Minimum Information Representation In Image Analysis and Data (MIRIAD) database, which is independent of the training and validation datasets. Longitudinal data at multiple time points, including BL, 6 weeks, 26 weeks, 52 weeks, and 24 months, were utilized for evaluation. For 46 subjects from the MIRIAD cohort, Figure 10A-10C respectively shows the scatter plots of the model’s predictions for MMSE scores at 26 weeks, 52 weeks, and 24 months. The model achieved RMSE values of 2.45, 2.6, and 2.8 for the three time points, respectively. In comparison, although the slightly higher error in the external test can be attributed to the limited sample size and potential inter-dataset variability, the predicted trends remained consistent with the ground truth. These results demonstrate that the proposed model maintains stable predictive performance across different datasets, indicating good generalization capability.

Figure 10 Prediction results of MMSE scores at three time points on the MIRIAD dataset: (A) 26 weeks; (B) 52 weeks; (C) 24 months. CC, correlation coefficient; MIRIAD, Minimal Interval Resonance Imaging in Alzheimer’s Disease; MMSE, Mini-Mental State Examination.

Discussion

Apolipoprotein E (APOE) genotype-specific predictive performance across follow-up

To evaluate whether the framework’s predictive ability varies by APOE genotype over longitudinal follow-up, we analyzed the CCs between predicted and actual outcomes across three time points (M18, M24, and M36) in APOE ε4 carriers (n=40) and non-carriers (n=40). As shown in Figures 11,12, the framework maintained consistent predictive performance across genotypes at early follow-up, with a slight divergence emerging in later stages: at M18, both subgroups exhibited comparable correlations: 0.69 in APOE ε4 carriers (95% CI: 0.50–0.82) and 0.68 in non-carriers (95% CI: 0.49–0.81), with no statistical divergence (z=0.07, P=0.944). At M24, the correlation remained similar between groups but showed a modest increase from M18: 0.75 in carriers (95% CI: 0.59–0.85) and 0.71 in non-carriers (95% CI: 0.54–0.83), with no statistical divergence (z=0.27, P=0.787). At M36, a more pronounced gap emerged: carriers exhibited a substantially higher correlation (0.88, 95% CI: 0.80–0.93) compared to non-carriers (0.77, 95% CI: 0.64–0.86). Fisher’s test indicated this difference was statistically significant (z=1.99, P=0.047), with the correlation in carriers exceeding that in non-carriers by 0.11.

Figure 11 Comparison of predicted MMSE scores between APOE carriers and non-carriers: (A) M18; (B) M24; (C) M36. M18, 18th month; M24, 24th month; M36, 36th month. APOE, apolipoprotein E; CC, correlation coefficient; MMSE, Mini-Mental State Examination.
Figure 12 Predictive correlation of MMSE across follow-up time by APOE genotype. M18, 18th month; M24, 24th month; M36, 36th month. APOE, apolipoprotein E; MMSE, Mini-Mental State Examination.

This time-dependent pattern may reflect the biological interplay between APOE genotype and disease progression. APOE ε4 carriers are known to exhibit accelerated pathological changes over time, which becomes more pronounced by M36, which aligns with the framework’s enhanced ability to capture progression trajectories in this subgroup.

Notably, even at the earliest time point (M18), both subgroups maintained clinically meaningful correlations (>0.65), supporting the framework’s utility for early-stage prediction across genotypes. The significant divergence at M36 highlights its particular value for tracking APOE ε4 carriers, a high-risk subgroup where precise longitudinal prediction is critical for timely intervention.

Novelty of using HKS features with different time scales

In our feature extraction module, we introduced the HKS as a multi-scale morphological features extraction method to describe the hippocampal surface structural changes induced by AD. By relying on the rate of heat propagation at different time scales, the HKS is able to capture the morphological differences of different regions on the hippocampal surface at different spatial scales. Therefore, this study demonstrated that HKS using the appropriate time scale has strong structural discrimination sensitivity.

Existing studies have shown that AD affects local areas of the hippocampus in MCI stage (50), leading to neuronal loss and volume reduction in these areas. In order to accurately capture these local morphology changes, we focused on the HKS at shorter time scales, as it can better reflect the local surface morphological characteristics induced by AD.

Next, we used the ADAS11 score at three time points (M18, M24, M36) as the target for prediction, focusing on the discrimination power of the HKS at different time scales. As shown in Figure 13, the results indicated that when time scales =12, the deviations between the predicted scores and actual scores were lower, and the correlations were higher. Therefore, the HKS value at time scales =12 is more suitable for our task.

Figure 13 A comparison of experimental results based on HKS values at different time points. M18, 18th month; M24, 24th month; M36, 36th month. CC, correlation coefficient; HKS, heat kernel signature; RMSE, root mean square error.

Uniform manifold approximation and projection (UMAP) visualization for the features based on M-FCN feature extraction: comparison of MCI converters and MCI nonconverters

We divided the MCI dataset into MCI converters (individuals who eventually converted to AD) and MCI nonconverters (individuals who did not convert to AD) based on disease progression (as shown in Table 4). In the prediction process of clinical tasks, we performed UMAP (51) dimensionality reduction on the features of BL and M18 (18 months later) to visualize the distribution of high-dimensional data in low dimensional space. The analysis results, as shown in Figure 14, indicated that in the BL stage, the distribution of the reduced dimensional features of the MCI converters and MCI nonconverters were difficult to be distinguished in the BL stage. However, in the M18 stage, the two sets of features gradually showed a separation trend. This reflected the potential of M-FCN feature extraction in capturing disease progression.

Table 4

Demographic information of converters and non-converters

Variables Converter (n=47) Non-converter (n=68) χ2/F P
Gender (male/female) 25/22 40/28 0.37 0.54
Age 74.90±6.96 72.90±7.30 1.59 0.21
MMSE 22.91±3.01 26.40±2.89 12.88 <0.001

Data are presented as number or mean ± standard deviation. MMSE, Mini-Mental State Examination.

Figure 14 2D UMAP embedding visualization: BL and M18 are based on the dimensionality reduction results of M-FCN extracted features during the prediction process. BL, baseline; M18, 18th month. 2D, two-dimensional; M-FCN, multi-view feature fusion convolutional network; MCI, mild cognitive impairment; UMAP, uniform manifold approximation and projection.

Ordinary differential equations (ODE) predictive performance

Neural manifold ODE (52) were applied for continuous-time modeling and to estimate the hidden states associated with missing data. These manifold ODEs operate through distinct forward and backward processes defined in separate spaces. For the forward pass, they rely on the Riemannian exponential map, which allows mapping points from the tangent space onto the manifold, thus capturing the non-Euclidean structure. This method uses the Euler solver to advance the solution along time steps:

dh(t)dt=f[h(t),t,θ]

f represents a Neural Network, and t{0T}. Starting from the GRU initial state h(0), we model continuous dynamic changes. By solving the ODE, we can obtain the state h(T) at any time point.

To handle real-world data challenges—including missing timepoints, irregular sampling intervals, and varying follow-up durations—our ODE model learns continuous progression trajectories in a latent manifold. By integrating the system dynamics using actual elapsed time, it reconstructs missing states and generalizes predictions without relying on biased imputation or fixed time bins. This inherent robustness to temporal variability is reflected in the stable performance across different time horizons.

In this study, we evaluated the predictive performance of incorporating ODEs using MMSE and ADAS11 as assessment measures. As shown in Table 5, the results presented a mix of outcomes across different time points (M18, M24, and M36), with minor fluctuations in RMSE and CC values. For MMSE, the RMSE values with ODE integration showed slight reductions at M18 and M24 (from 2.34 to 2.33 and from 2.58 to 2.51, respectively), suggesting a marginal improvement. However, at M36, RMSE for the ODE model shows a slight increase (2.64 vs. 2.60). The CC values similarly reflected minor fluctuations but indicated a relatively stable performance across the time points, with slight gains in some instances. Previous studies have demonstrated the effectiveness of ODEs. This inconsistent impact may be due to the limited observation points in our study, as ODEs generally perform better with longer time series data.

Table 5

Comparison of predictive performance using ODE

Clinical indicators Evaluation M18 M24 M36
ODE w/o ODE ODE w/o ODE ODE w/o ODE
MMSE RMSE 2.33 2.34 2.51 2.58 2.64 2.60
CC 0.72 0.72 0.79 0.77 0.84 0.83
ADAS11 RMSE 4.32 4.47 4.9 4.9 5.5 5.5
CC 0.84 0.82 0.87 0.82 0.85 0.9

M18, 18th month; M24, 24th month; M36, 36th month. ADAS11, Alzheimer’s Disease Assessment Scale Cognitive Subscale (11-item); CC, correlation coefficient; ODE, ordinary differential equations; MMSE, Mini-Mental State Examination; RMSE, root mean square error; w/o, without.

Attention mechanism results

We visualized the results of the embedded attention mechanism, as shown in Figure 15, which illustrates the affected regions of the hippocampus on the left (LH) and right (RH) sides at different time points (M18, M24, and M36). The color coding indicates that red regions represent areas with more significant structural changes, while blue regions indicate relatively less affected areas. The analysis reveals that the CA1 region is the most commonly affected area, followed by the subiculum region. Over time, from M18 to M36, the affected regions progressively expand, with more pronounced changes observed at M36. Additionally, the extent of changes in the LH is slightly greater than in the right, showing a certain degree of lateralization.

Figure 15 Visualization results of affected areas on the left (L) and right (R) sides of the hippocampus at different time points (M18, M24, and M36). M18, 18th month; M24, 24th month; M36, 36th month.

These findings are consistent with the pathological characteristics of AD, where early involvement of the CA1 region is considered a critical marker of memory decline, while further involvement of the subiculum region may impact overall cognitive function.

Limitations

The proposed deep learning framework for predicting clinical scores of AD progression may provide a new way to monitor the progression of individuals with AD based on the morphological features of their MRI brain images. Nonetheless, there are several limitations of this study. First, a relatively small number of subjects are included as the research objects, e.g., there are 221 subjects in training stage, which are relatively small in sample size to fully capture the correlation of morphological features in spatial and temporal series. Second, in this study, the diagnostic classification of AD and MCI was based on clinical criteria provided by the ADNI dataset, without requiring biomarker confirmation of amyloid (Aβ) or tau positivity. According to the current NIA-AA and A/T/N research frameworks, AD is biologically defined, and clinical diagnosis alone may not fully distinguish AD from other neurodegenerative conditions. Consequently, the cohort used in this work may include non-AD pathologies, which could reduce the specificity of the proposed predictive model. Future work will incorporate subjects with biomarker-confirmed A+ status to enhance diagnostic accuracy and biological interpretability of the predictive framework.

Moreover, this study primarily models hippocampal-driven structural changes and does not account for the full anatomical heterogeneity of AD. As recent evidence indicates the presence of multiple AD subtypes with distinct atrophy trajectories (53), our current framework captures only hippocampal-related progression. Future work will integrate multi-regional morphological features to explore subtype-specific degeneration patterns.

Despite this limitation, our results show that this prediction model has strong sensitivity to the morphological changes induced by AD and can fully explore the long- and short-term dependencies on sequential morphological features.


Conclusions

In this article, we developed an AD progression prediction model based on hippocampal morphology that combines multi-view feature extraction method M-FCN and sequential feature learning method Bi-GRU to predict future clinical scores. By utilizing task driven feature learning modules to generate comprehensive, reliable, and closely related structural features to the progression of AD, a solid foundation is laid for the temporal correlation mining of time-series structural features. Furthermore, through training Bi-GRU module to capture both short-term and long-term structural changes of hippocampus, our model can provide a comprehensive perspective on the morphological evolution of the hippocampus. The empirical results demonstrated the potential to accurately grasp the correspondence between structural changes and clinical indicators of neurodegeneration from the complex morphological changes caused by AD.

Currently, the amyloid-tau-neurodegeneration (A/T/N) framework has become the mainstream paradigm in AD research (54) providing a unified biological model for understanding the sequential pathological cascade underlying AD progression.

In the context of our work, the hippocampal morphological features derived from MRI primarily capture the downstream neurodegeneration (N) component of this pathological cascade. The observed structural alterations can thus be interpreted as macroscopic manifestations of cumulative Aβ and tau pathology. To further enhance the biological interpretability and clinical utility of the proposed framework, future extensions will incorporate multimodal biomarkers, including Aβ and tau PET signals (55) to represent molecular-level pathology and FDG-PET measures of neuronal metabolism. By fusing these modalities with MRI-based morphological descriptors through shared latent representation learning or cross-modal attention mechanisms, it becomes possible to model both upstream molecular alterations and downstream morphological consequences in a unified predictive space. Such integration is expected to yield more accurate and interpretable predictions of cognitive decline, improve early diagnostic sensitivity, and facilitate the translation of the model into clinical decision-support systems.


Acknowledgments

None.


Footnote

Reporting Checklist: The authors have completed the TRIPOD+AI reporting checklist. Available at https://qims.amegroups.com/article/view/10.21037/qims-2025-377/rc

Funding: This work was supported by the National Natural Science Foundation of China (No. 62171209); The Joint Fund of the National Natural Science Foundation of China (No. U24A20328); The Youth Innovation Technology Project of Higher School in Shandong Province (No. 2023KJ212).

Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://qims.amegroups.com/article/view/10.21037/qims-2025-377/coif). H.W. reports that this work was supported by the Joint Fund of the National Natural Science Foundation of China (No. U24A20328) and the Youth Innovation Technology Project of Higher School in Shandong Province (No. 2023KJ212). G.W. reports that this work was supported by the National Natural Science Foundation of China (No. 62171209). The other authors have no conflicts of interest to declare.

Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. This study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments.

Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.


References

  1. Knopman DS, Amieva H, Petersen RC, Chételat G, Holtzman DM, Hyman BT, Nixon RA, Jones DT. Alzheimer disease. Nat Rev Dis Primers 2021;7:33. [Crossref] [PubMed]
  2. Alzheimer’s Association. 2019 Alzheimer’s disease facts and figures. Alzheimers Dement 2019;15:321-87.
  3. Langa KM, Levine DA. The diagnosis and management of mild cognitive impairment: a clinical review. JAMA 2014;312:2551-61. [Crossref] [PubMed]
  4. Pisani S, Mueller C, Huntley J, Aarsland D, Kempton MJ. A meta-analysis of randomised controlled trials of physical activity in people with Alzheimer's disease and mild cognitive impairment with a comparison to donepezil. Int J Geriatr Psychiatry 2021;36:1471-87. [Crossref] [PubMed]
  5. Kueper JK, Speechley M, Montero-Odasso M. The Alzheimer's Disease Assessment Scale-Cognitive Subscale (ADAS-Cog): Modifications and Responsiveness in Pre-Dementia Populations. A Narrative Review. J Alzheimers Dis 2018;63:423-44. [Crossref] [PubMed]
  6. Wu Y, Zhang X, He Y, Cui J, Ge X, Han H, Luo Y, Liu L, Wang X, Yu HAlzheimer's Disease Neuroimaging Initiative. Predicting Alzheimer's disease based on survival data and longitudinally measured performance on cognitive and functional scales. Psychiatry Res 2020;291:113201. [Crossref] [PubMed]
  7. Arevalo-Rodriguez I, Smailagic N, Roqué I, Figuls M, Ciapponi A, Sanchez-Perez E, Giannakou A, Pedraza OL. Bonfill Cosp X, Cullum S. Mini-Mental State Examination (MMSE) for the detection of Alzheimer's disease and other dementias in people with mild cognitive impairment (MCI). Cochrane Database Syst Rev 2015;2015:CD010783. [Crossref] [PubMed]
  8. Wessels AM, Dowsett SA, Sims JR. Detecting Treatment Group Differences in Alzheimer's Disease Clinical Trials: A Comparison of Alzheimer's Disease Assessment Scale - Cognitive Subscale (ADAS-Cog) and the Clinical Dementia Rating - Sum of Boxes (CDR-SB). J Prev Alzheimers Dis 2018;5:15-20. [Crossref] [PubMed]
  9. Karikari TK, Benedet AL, Ashton NJ, Lantero Rodriguez J, Snellman A, Suárez-Calvet M, Saha-Chaudhuri P, Lussier F, Kvartsberg H, Rial AM, Pascoal TA, Andreasson U, Schöll M, Weiner MW, Rosa-Neto P, Trojanowski JQ, Shaw LM, Blennow K, Zetterberg HAlzheimer’s Disease Neuroimaging Initiative. Diagnostic performance and prediction of clinical progression of plasma phospho-tau181 in the Alzheimer's Disease Neuroimaging Initiative. Mol Psychiatry 2021;26:429-42. [Crossref] [PubMed]
  10. Frisoni GB, Fox NC, Jack CR Jr, Scheltens P, Thompson PM. The clinical use of structural MRI in Alzheimer disease. Nat Rev Neurol 2010;6:67-77. [Crossref] [PubMed]
  11. Benoit JS, Chan W, Piller L, Doody R. Longitudinal Sensitivity of Alzheimer's Disease Severity Staging. Am J Alzheimers Dis Other Demen 2020;35:1533317520918719. [Crossref] [PubMed]
  12. Carlson NE, Moore MM, Dame A, Howieson D, Silbert LC, Quinn JF, Kaye JA. Trajectories of brain loss in aging and the development of cognitive impairment. Neurology 2008;70:828-33. [Crossref] [PubMed]
  13. Zhang J, Xie L, Cheng C, Liu Y, Zhang X, Wang H, Hu J, Yu H, Xu J. Hippocampal subfield volumes in mild cognitive impairment and alzheimer's disease: a systematic review and meta-analysis. Brain Imaging Behav 2023;17:778-93. [Crossref] [PubMed]
  14. Liu M, Li F, Yan H, Wang K, Ma Y, Shen L, Xu M. A multi-model deep convolutional neural network for automatic hippocampus segmentation and classification in Alzheimer's disease. Neuroimage 2020;208:116459. [Crossref] [PubMed]
  15. Walhovd KB, Fjell AM, Sørensen Ø, Mowinckel AM, Reinbold CS, Idland AV, Watne LO, Franke A, Dobricic V, Kilpert F, Bertram L, Wang Y. Genetic risk for Alzheimer disease predicts hippocampal volume through the human lifespan. Neurol Genet 2020;6:e506. [Crossref] [PubMed]
  16. Gao N, Liu Z, Deng Y, Chen H, Ye C, Yang Q, Ma T. MR-based spatiotemporal anisotropic atrophy evaluation of hippocampus in Alzheimer's disease progression by multiscale skeletal representation. Hum Brain Mapp 2023;44:5180-97. [Crossref] [PubMed]
  17. Kantarci K, Weigand SD, Przybelski SA, Preboske GM, Pankratz VS, Vemuri P, Senjem ML, Murphy MC, Gunter JL, Machulda MM, Ivnik RJ, Roberts RO, Boeve BF, Rocca WA, Knopman DS, Petersen RC, Jack CR Jr. MRI and MRS predictors of mild cognitive impairment in a population-based sample. Neurology 2013;81:126-33. [Crossref] [PubMed]
  18. Mushtaq N, Khan AA, Khan FA, Ali MJ, Ali Shahid MM, Wechtaisong C, Uthansakul P. Brain tumor segmentation using multi-view attention based ensemble network. Computers, Materials & Continua 2022;72:5793-806.
  19. Rastogi D, Johri P, Donelli M, Kadry S, Khan AA, Espa G, Feraco P, Kim J. Deep learning-integrated MRI brain tumor analysis: feature extraction, segmentation, and Survival Prediction using Replicator and volumetric networks. Sci Rep 2025;15:1437. [Crossref] [PubMed]
  20. Thapar P, Rakhra M, Prashar D, Mrsic L, Khan AA, Kadry S. Skin cancer segmentation and classification by implementing a hybrid FrCN-(U-NeT) technique with machine learning. PLoS One 2025;20:e0322659. [Crossref] [PubMed]
  21. Khan AA, Mahendran RK, Perumal K, Faheem M. Dual-3DM(3)AD: Mixed Transformer Based Semantic Segmentation and Triplet Pre-Processing for Early Multi-Class Alzheimer's Diagnosis. IEEE Trans Neural Syst Rehabil Eng 2024;32:696-707. [Crossref] [PubMed]
  22. Kujur A, Raza Z, Khan AA, Wechtaisong C. Data Complexity Based Evaluation of the Model Dependence of Brain MRI Images for Classification of Brain Tumor and Alzheimer’s Disease. IEEE Access 2022;10:112117-33.
  23. Liu F, Yuan S, Li W, Xu Q, Wu X, Han K, Wang J, Miao S. Multi-task joint learning network based on adaptive patch pruning for Alzheimer’s disease diagnosis and clinical score prediction. Biomedical Signal Processing and Control 2024;95:106398.
  24. QiuSJoshiPSMillerMIXueCZhouXKarjadiC,
  25. et al. Development and validation of an interpretable deep learning framework for Alzheimer's disease classification. Brain 2020;143:1920-33.
  26. Katabathula S, Wang Q, Xu R. Predict Alzheimer's disease using hippocampus MRI data: a lightweight 3D deep convolutional network model with visual and global shape representations. Alzheimers Res Ther 2021;13:104. [Crossref] [PubMed]
  27. Liu M, Zhang J, Lian C, Shen D. Weakly Supervised Deep Learning for Brain Disease Prognosis Using MRI and Incomplete Clinical Scores. IEEE Trans Cybern 2020;50:3381-92. [Crossref] [PubMed]
  28. Cui R, Liu MAlzheimer's Disease Neuroimaging Initiative. RNN-based longitudinal analysis for diagnosis of Alzheimer's disease. Comput Med Imaging Graph 2019;73:1-10. [Crossref] [PubMed]
  29. Mienye ID, Swart TG, Obaido G. Recurrent Neural Networks: A Comprehensive Review of Architectures, Variants, and Applications. Information 2024;15:517.
  30. Rahim N, El-Sappagh S, Ali S. Muhammad Khan, Del Ser J, Abuhmed T. Prediction of Alzheimer's progression based on multimodal Deep-Learning-based fusion and visual Explainability of time-series data. Information Fusion 2023;92:363-88.
  31. Lei B, Liang E, Yang M, Yang P, Zhou F, Tan EL, Lei Y, Liu CM, Wang T, Xiao X, Wang S. Predicting clinical scores for Alzheimer’s disease based on joint and deep learning. Expert Systems with Applications 2022;187:115966.
  32. Jung W, Jun E, Suk HIAlzheimer’s Disease Neuroimaging Initiative. Deep recurrent model for individualized prediction of Alzheimer's disease progression. Neuroimage 2021;237:118143. [Crossref] [PubMed]
  33. Dong Q, Li Z, Liu W, Chen K, Su Y, Wu J, Caselli RJ, Reiman EM, Wang Y, Shen J. Correlation studies of Hippocampal Morphometry and Plasma NFL Levels in Cognitively Unimpaired Subjects. IEEE Trans Comput Soc Syst 2023;10:3602-8. [Crossref] [PubMed]
  34. Lian C, Liu M, Wang L, Shen D. Multi-Task Weakly-Supervised Attention Network for Dementia Status Estimation With Structural MRI. IEEE Trans Neural Netw Learn Syst 2022;33:4056-68. [Crossref] [PubMed]
  35. Jack CR Jr, Bernstein MA, Fox NC, Thompson P, Alexander G, Harvey D, et al. The Alzheimer's Disease Neuroimaging Initiative (ADNI): MRI methods. J Magn Reson Imaging 2008;27:685-91. [Crossref] [PubMed]
  36. Patenaude B, Smith SM, Kennedy DN, Jenkinson M. A Bayesian model of shape and appearance for subcortical brain segmentation. Neuroimage 2011;56:907-22. [Crossref] [PubMed]
  37. Shen L, Huang H, Makedon F, Saykin AJ. Efficient registration of 3D SPHARM surfaces. Fourth Canadian Conference on Computer and Robot Vision (CRV '07); May 2007; Los Alamitos, CA, USA. IEEE; 2007: 81-8.
  38. Sun J, Ovsjanikov M, Guibas L. A concise and provably informative multi-scale signature based on heat diffusion. Computer Graphics Forum 2009;28:1383-92.
  39. Styner M, Oguz I, Xu S, Brechbühler C, Pantazis D, Levitt JJ, Shenton ME, Gerig G. Framework for the Statistical Shape Analysis of Brain Structures using SPHARM-PDM. Insight J 2006;242-50.
  40. Taud H, Mas JF. Multilayer perceptron (MLP). In: Camacho Olmedo MT, Paegelow M, Mas JF, Escobar F, editors. Geomatic approaches for modeling land change scenarios. Cham: Springer; 2018:451-5.
  41. Li N, Su Q, Yao T, Ba M, Wang G. Landmark-based spherical quasi-conformal mapping for hippocampal surface registration. Quant Imaging Med Surg 2024;14:3997-4014. [Crossref] [PubMed]
  42. Boscain U, Laurent C. The Laplace-Beltrami operator in almost-Riemannian geometry. An Inst Fourier Grenoble 2013;63:1739-70.
  43. Apostolova LG, Mosconi L, Thompson PM, Green AE, Hwang KS, Ramirez A, Mistur R, Tsui WH, de Leon MJ. Subregional hippocampal atrophy predicts Alzheimer's dementia in the cognitively normal. Neurobiol Aging 2010;31:1077-88. [Crossref] [PubMed]
  44. Adaloglou N. Intuitive explanation of skip connections in deep learning. AI Summer 2020. Available online: https://theaisummer.com/skip-connections/
  45. Muresan MP, Nedevschi S. Multi-object tracking of 3D cuboids using aggregated features. 2019 IEEE 15th International Conference on Intelligent Computer Communication and Processing (ICCP); 05-07 September 2019; Cluj-Napoca, Romania; IEEE, 2019:11-18.
  46. Dey R, Salem FM. Gate-variants of gated recurrent unit (GRU) neural networks. 2017 IEEE 60th international midwest symposium on circuits and systems (MWSCAS); 06-09 August 2017; Boston, MA, USA; IEEE, 2017: 1597-1600.
  47. Imambi S, Prakash KB, Kanagachidambaresan GR. PyTorch. In: Prakash KB, Kanagachidambaresan GR, editors. Programming with TensorFlow: Solution for Edge Computing Applications. Cham: Springer; 2021:87-104.
  48. Hodson TO. Root mean square error (RMSE) or mean absolute error (MAE): When to use them or not. Geosci Model Dev 2022;15:5481-7.
  49. Fujita S, Mori S, Onda K, Hanaoka S, Nomura Y, Nakao T, Yoshikawa T, Takao H, Hayashi N, Abe O. Characterization of Brain Volume Changes in Aging Individuals With Normal Cognition Using Serial Magnetic Resonance Imaging. JAMA Netw Open 2023;6:e2318153. [Crossref] [PubMed]
  50. Alves F, Kalinowski P, Ayton S. Accelerated Brain Volume Loss Caused by Anti-β-Amyloid Drugs: A Systematic Review and Meta-analysis. Neurology 2023;100:e2114-24. [Crossref] [PubMed]
  51. Dong Q, Zhang W, Wu J, Li B, Schron EH, McMahon T, Shi J, Gutman BA, Chen K, Baxter LC, Thompson PM, Reiman EM, Caselli RJ, Wang Y. Applying surface-based hippocampal morphometry to study APOE-E4 allele dose effects in cognitively unimpaired subjects. Neuroimage Clin 2019;22:101744. [Crossref] [PubMed]
  52. Allaoui M., Kherfi ML, Cheriet A. Considerably Improving Clustering Algorithms Using UMAP Dimensionality Reduction Technique: A Comparative Study. In: El Moataz A, Mammass D, Mansouri A, Nouboud F, editors. Image and Signal Processing. Cham: Springer; 2020:317-325.
  53. Jeong S, Jung W, Sohn J, Suk HI. Deep Geometric Learning With Monotonicity Constraints for Alzheimer's Disease Progression. IEEE Trans Neural Netw Learn Syst 2025;36:7090-102. [Crossref] [PubMed]
  54. Zhang B, Lin L, Wu S, Al-Masqari ZHMA. Multiple Subtypes of Alzheimer's Disease Base on Brain Atrophy Pattern. Brain Sci 2021;11:278. [Crossref] [PubMed]
  55. Xiong X, He H, Ye Q, Qian S, Zhou S, Feng F, Fang EF, Xie C. Alzheimer's disease diagnostic accuracy by fluid and neuroimaging ATN framework. CNS Neurosci Ther 2024;30:e14357. [Crossref] [PubMed]
  56. Jack CR Jr, Bennett DA, Blennow K, Carrillo MC, Feldman HH, Frisoni GB, Hampel H, Jagust WJ, Johnson KA, Knopman DS, Petersen RC, Scheltens P, Sperling RA, Dubois B. A/T/N: An unbiased descriptive classification scheme for Alzheimer disease biomarkers. Neurology 2016;87:539-47. [Crossref] [PubMed]
Cite this article as: Feng J, Ng KP, Wang H, Yao T, Su Q, Ba M, Wang G. Alzheimer’s disease prediction algorithm based on hippocampal longitudinal hybrid morphological features. Quant Imaging Med Surg 2026;16(2):170. doi: 10.21037/qims-2025-377

Download Citation