Effect of reconstruction settings on radiomics feature robustness in positron emission tomography images: a clinical study
Introduction
Positron emission tomography-computed tomography (PET/CT) has changed the field of oncology, allowing for comprehensive cancer detection, diagnosis, staging, and treatment monitoring. PET imaging can provide functional imaging data that are critical for evaluating tumor characteristics, and efforts to optimize PET imaging have partly focused on optimizing quantification and image quality. For clinical 18F-fluorodeoxyglucose (18F-FDG) PET/CT oncology imaging, standardized uptake value (SUV) metrics including maximum SUV (SUVmax), mean SUV (SUVmean), and peak SUV (SUVpeak) are now routinely utilized in single-time-point images for quantifying glucose metabolic activity (1). However, these parameters may not comprehensively capture the underlying spatial distribution of 18F-FDG and the intricacies of spatial tumor complexity (2). In recent years, radiomics has emerged as a medical image analysis technology. The concept of radiomics was first proposed by Lambin et al. in 2012 (3) and involves the high-throughput extraction of quantitative features from different modalities images (e.g., PET and CT) and the conversion of the image data of the region of interest into analyzable high-dimensional data (4,5). The high-throughput features, including intensity, shape, volume and texture, can effectively address the limitations of traditional quantitative methods (4,5).
However, although an attractive direction for research, radiomics has yet to be translated into routine clinical or diagnostic applications. Concerns regarding the quality and robustness of features in radiomics analysis have been raised, and highlighting the necessity for standardized validation of PET radiomics prior to its implementation in clinical practice (6). It has been indicated that radiomic features are sensitive to image acquisition settings, reconstruction algorithms, and image processing (5). An earlier study demonstrated that different image features have varying sensitivities to reconstruction settings in 18F-FDG PET (2). Another study also indicated that the robustness of 18F-FDG PET/CT image radiomics features varied in advanced reconstruction settings, with different settings exerting a variety effect depending on the feature (7). These findings emphasize the importance of understanding the impact of advanced reconstruction methods on these features and of determining whether variations in feature values are due to a true physiological reaction or to image generation parameters of the methodology (8).
Currently, the ordered subset expectation maximization (OSEM) algorithm is the prevalent statistical reconstruction algorithm used in clinical routines. For the purpose of overcoming the drawbacks of increased noise and local convergence in the OSEM algorithm, a new Bayesian penalized likelihood algorithm (Hyper Iterative), was introduced by United Imaging Healthcare, which incorporates a penalty term to suppress the image noise during the iterative reconstruction process (9-11). Deep learning for PET reconstruction is capable of achieving superior performance in denoising PET images (12,13), the deep progressive reconstruction (DPR) algorithm employs two neural networks to respectively suppress the noise in the input image and generate high-contrast images, with its training images being derived from uEXPLORER, the world’s first clinical total-body PET scanner (12). These advanced reconstruction methods have been reported to provide superior image quality characteristics as compared to OSEM. Since radiomic features are highly sensitive to image quality and noise characteristics, these improvements may also affect the stability and reliability of quantitative features extracted from PET images. Shiiba et al. conducted a phantom study to compare the stability of radiomic features reconstructed with the OSEM, Hyper Iterative, and DPR algorithms and concluded that advanced reconstruction methods provided enhanced stability of radiomic features as compared with OSEM (14). Interestingly, they also noted that the number of features with a coefficient of variation (COV) less than 10% remained comparable across all reconstruction methods, indicating a baseline level of feature stability irrespective of the algorithm used (14). However, these findings were derived exclusively from phantom data, and it remains uncertain whether similar patterns hold in clinical settings. Moreover, the robustness of radiomic features across these reconstruction algorithms has not been systematically assessed with clinical PET/CT images. Therefore, in this study, we sought to evaluate the stability of radiomic features reconstructed with OSEM, Hyper Iterative, and DPR using real clinical data.
Our study, based on clinical PET/CT images, aimed to evaluate the impact of reconstruction algorithms, their parameter settings, and acquisition duration on the robustness of radiomic features. As emphasized in a recent review (15), reporting all relevant parameters throughout the radiomics workflow is essential to ensuring a study’s reproducibility and transparency. In accordance with these recommendations, we documented all parameters in detail at each step of the workflow, including image acquisition, reconstruction, segmentation, preprocessing, and feature extraction. To assess the effects of these algorithms from multiple perspectives, this study compared the features extracted from images reconstructed using three different algorithms—OSEM, Hyper Iterative, and DPR—across various acquisition times. Furthermore, the impact of different β parameter values on feature extraction was also assessed during the Hyper Iterative reconstruction process.
Methods
Patients
This study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments and was approved by the Ethics Committee of The First Affiliated Hospital, Zhejiang University School of Medicine ([2025B] IIT Ethics approval No. 0921). The requirement for informed consent was waived due to the retrospective nature of the analysis.
Anonymized and de-identified data were selected via random sampling from the institutional database of the Department of Nuclear Medicine, The First Affiliated Hospital, Zhejiang University School of Medicine. We initially enrolled 62 oncological patients who underwent clinical 18F-FDG PET/CT examinations at The First Affiliated Hospital, Zhejiang University School of Medicine, from June 2024 to November 2024. Patients with histologically confirmed lung cancer were included. Patients were instructed to fast and to avoid liquids other than water for at least 6 hours prior to scanning. Blood glucose levels were measured immediately before the 18F-FDG injection and were confirmed to be ≤11 mmol/mL. The injection dose of 18F-FDG was based on the patient’s weight (3.54±0.21 MBq/kg). The exclusion criteria included no uptake of 18F-FDG in lung lesions, claustrophobia or an inability to maintain stable for 10 min, and any abnormal conditions such as artifacts that prevented the adequate evaluation of the lung and liver background regions. For the revised analysis, 58 patients were retained after 4 cases whose target lesions could not be reliably delineated on CT images were excluded. In these patients, the anatomical boundaries of the lesions were not clearly distinguishable due to low contrast with surrounding tissues, respiratory motion artifacts, or lesion proximity to complex anatomical structures, which precluded accurate CT-based segmentation.
PET/CT acquisition and reconstruction
All images were acquired with a uMI 780 PET/CT scanner (United Imaging Healthcare, Shanghai, China). Patients were first scanned via CT with a tube voltage of 120 kV and a reference tube current of 150 mAs to provide anatomical information and attenuation correction for PET images. Subsequently, PET scans were acquired in list mode for 2 min per bed position and 4–6 bed positions in step-and-shoot mode (the number of bed positions was determined based on patient height), with an overlap of 50%. PET data were reconstructed with three algorithms: OSEM, Hyper Iterative reconstruction, and DPR). Hyper Iterative is a Bayesian penalized-likelihood reconstruction algorithm, and DPR is a deep learning-based method. Both Hyper Iterative and DPR have been developed and commercialized by United Imaging Healthcare and have been increasingly adopted in clinical practice (9-13).
OSEM was performed with 2 iterations and 20 subsets and followed by a Gaussian post-smoothing filter (full width at half maximum =2 mm). Both Hyper Iterative and DPR reconstructions were performed without post-filtering. All reconstructions were performed with an identical matrix size (192×192), field of view (FOV; 600 mm), and slice thickness (2.68 mm) and included time-of-flight (TOF) and point-spread-function (PSF) modeling.
To evaluate the effect of acquisition duration, the 2-min raw data were rebinned into 30-s, 60-s, 90-s, and 120-s durations and reconstructed with both OSEM and DPR. Additionally, to assess the reproducibility of radiomic features across different time intervals with identical acquisition duration, the 2-min list-mode data were divided into four nonoverlapping 30-s segments (0–30, 31–60, 61–90, and 91–120 s), and each segment was independently reconstructed with the same OSEM parameters.
Within the Hyper Iterative reconstruction, two types of analyses were designed to separately evaluate the effects of reconstruction parameters and acquisition duration. Firstly, to assess the effect of the β parameter, 120-s full-duration data were reconstructed with β values of 0.2, 0.3, 0.4, 0.5, and 0.6. This β range was selected based on a previous 18F-FDG-PET study, in which β values of 0.3 and 0.5 demonstrated good image quality and lesion detectability (11). Second, based on our clinical experience and prior finding (11), β=0.5 was selected as the fixed standard setting for subsequent temporal reconstruction (30, 60, 90, and 120 s) via Hyper Iterative. This separation of β variation and acquisition time allowed for independent evaluation of parameter effects on radiomic features.
In total, 20 reconstruction settings were generated per patient. Among these, 16 were included in the five predefined analysis groups to evaluate the effects of the reconstruction algorithm, parameter settings, and acquisition time (as detailed in Table 1), which included four acquisition durations for OSEM, four for DPR, and eight for Hyper Iterative (four durations at β=0.5 and five β values at 120 s; note that β=0.5 at 120 s is counted in both categories). In addition, four nonoverlapping 30-s segments (0–30, 31–60, 61–90, and 91–120 s) were reconstructed with OSEM to independently assess the impact of counting statistics on the image features.
Table 1
| Reconstruction method | Compared reconstructions | Assigned analysis group | Tested variable |
|---|---|---|---|
| OSEM | 30 s, 60 s, 90 s, 120 s† | OSEM-Time | Acquisition time |
| Hyper Iterative | 30 s, 60 s, 90 s, 120 s† (β =0.5) | HI-Time | Acquisition time |
| β =0.2, 0.3, 0.4, 0.5†, 0.6 (120 s) | HI-Para | β parameter | |
| DPR | 30 s, 60 s, 90 s, 120 s† | DPR-Time | Acquisition time |
| All algorithms (120 s) | OSEM: 120 s†, hyper iterative: 120 s† (β =0.5), DPR: 120 s† | ReconGroup | Algorithm type |
†, used in more than one analysis group. DPR, deep progressive reconstruction; HI, Hyper Iterative; OSEM, ordered subset expectation maximization.
Radiomics feature extraction
In line with previous studies (7,16), only lesions with a volume greater than 1.7 cm3 were included in the subsequent analysis in order to successfully calculate partial texture parameters and reduce partial volume effect interference. Lesion segmentations were performed with an advanced workstation (uWS-MI, United Imaging Healthcare). Lesion volumes of interest (VOIs) were manually delineated by two board-certified nuclear medicine physicians based on the anatomical boundaries observed on the corresponding CT images. Manual contouring was performed on the CT scans in consensus with a visual assessment approach (8,17). A total of 93 lesions from the revised cohort of 58 patients were delineated. These CT-based VOIs were then applied to all reconstructed PET images across different acquisition durations and algorithms. This strategy was chosen to eliminate the potential bias introduced by the use of SUV-based PET contours, which are known to be affected by reconstruction parameters, and to ensure the comparability of radiomic features across datasets.
LIFEx 7.6.0 software (http://www.lifexsoft.org; French Atomic Energy and Alternative Energies Commission) was used to extract radiomics features from the VOIs of PET images (18). All images were analyzed at their original voxel size of 3.125×3.125×2.68 mm3, consistent with the PET/CT reconstruction protocol. The intensity was discretized with a fixed bin width of 64 gray levels between 0 and 20 SUV, which were automatically adjusted via LIFEx software. A total of 129 radiomic features were initially extracted from each PET image according to a previous study (6) and based on the standardized definitions of the Image Biomarker Standardisation Initiative (IBSI) (19). To ensure analytical relevance and comparability across reconstruction conditions, features were excluded if they (I) contained missing values [not available (NA)] in any reconstruction condition or (II) were constant across all settings (19,20). After these criteria were applied, a final set of 106 features was retained for further analysis. The remaining features were categorized into three main classes: intensity-based (i.e., SUV-based; n=24), intensity histogram-based (n=30), and texture-based (n=52) including gray-level co-occurrence matrix (GLCM; n=20), gray-level run-length matrix (GLRLM; n=11), neighborhood gray-tone difference matrix (NGTDM; n=5), and gray-level size zone matrix (GLSZM; n=16) (Table S1).
Given that radiomic features are deterministic functions of voxel-level intensity values, their sensitivity to reconstruction-induced variation is expected to depend on their mathematical formulation and category (19,21). In particular, first-order features primarily capture global intensity distributions, whereas texture features (e.g., GLCM and GLRLM) encode spatial relationships and gray-level patterns, which may amplify the effects of noise, resolution, and quantization (19-21). Accordingly, in our evaluation of feature robustness, we considered the inherent sensitivity of different feature categories to image perturbations—as inferred from their mathematical definitions—to guide the interpretation of the observed variability.
Data analysis
To assess the variability of radiomic features across different reconstruction settings, we analyzed five predefined comparison groups using 16 reconstructed image sets. These sets were selected from the total of 20 available reconstructions per lesion, with some sets assigned to multiple analysis groups, as detailed in Table 1. These groups included: (I) ReconGroup, with the three algorithms (OSEM, Hyper Iterative, and DPR) being compared using the 120-s reconstructed images; (II) HI-Para, with different β values (0.2–0.6) in Hyper Iterative being compared using 120-s data; (III) OSEM-Time (OSEM with 30–120 s); (IV) HI-Time (Hyper Iterative at β=0.5 with 30–120 s); and (V) DPR-Time (DPR with 30–120 s).
The COV was calculated as follows:
where SD and mean refer to the standard deviation and mean values of each radiomics feature over each analysis group. The variability of each feature was characterized by the mean value of the COV of all lesions for the feature. Subsequently, the features were categorized into four groups based on COV, including very small variations (COV ≤5%), small variations (5%< COV ≤10%), intermediate variations (10%< COV ≤20%), and large variations (COV >20%). To further examine the relationship between the influence of different factors on each feature, pairwise Pearson correlation analysis was conducted to compare the mean COV across all features between each pair of analysis groups.
Lin’s concordance correlation coefficient (CCC) was used to assess the agreement of radiomic features across different acquisition duration groups (22). For each algorithm, 120-s reconstruction images were used as the reference images, and the CCC values of 90 vs. 120 s, 60 vs. 120 s, 30 vs. 120 s were calculated, respectively, to evaluate the reproducibility of features extracted from short-time reconstruction images. Additionally, OSEM-reconstructed 2-min images were used as the reference to assess the reproducibility of short-duration image features from the DPR and Hyper Iterative algorithms to the full-duration features of OSEM. The CCC was calculated as follows:
where ρ is the Pearson correlation coefficient between features of the selected two groups; and 𝜎x and 𝜎y are the standard deviations of the two groups, respectively; and 𝜇x and 𝜇y are the means of the two groups, respectively.
In addition, we calculated the mean percentage difference (%Diff) between the selected pair of duration groups (23) as follows:
We followed the convention proposed in a previous (24), in which features with a CCC value below 0.80 are classified as non-repeatable, values of 0.80 or higher as repeatable, and values of 0.95 or above as highly repeatable. Furthermore, features with CCC >0.8 and %Diff between –10% and 10% were considered to have high reproducibility, as defined in other work (23,25).
In addition to the five comparison groups, we further assessed the reproducibility of radiomic features across different temporal segments under identical acquisition duration. Feature-wise COV, along with pairwise CCC and %Diff, were computed to quantify variability across segments. The four nonoverlapping 30-s segments (0–30, 31–60, 61–90, 91–120 s) were reconstructed from the same 120-s list-mode PET acquisition via the OSEM algorithm. This segment-based analysis was performed independently, as it was specifically designed to evaluate the impact of counting statistics on radiomic features.
The Kruskal-Wallis rank-sum test, followed by the Dunn post hoc test for multiple comparisons, was employed to compare the mean COV (averaged across lesions) between ReconGroup, HI-Para, and OSEM-Time to assess the effect of various factors on the features. Similarly, to evaluate the relative robustness of different reconstruction algorithms across varying acquisition durations, both the mean COV and CCC values were compared across the OSEM-Time, HI-Time, and DPR-Time groups via the Friedman test. Post hoc pairwise comparisons were performed with the Wilcoxon signed-rank test and Bonferroni correction. A P value less than 0.05 was considered statistically significant.
Results
Ultimately, 58 patients (33 males and 25 females; mean age 65.74±10.86 years) with suspected and diagnosed lung cancer were enrolled, and the presence of lesions was confirmed by CT, enhanced CT, 18F-FDG PET, and biochemical findings. The detailed characteristics of the study population are summarized in Table 2. PET images of a representative patient with lung cancer reconstructed by the different reconstruction settings are shown in Figure 1. Meanwhile, Figures 2,3 provide an overview of radiomic features’ COV results across five analysis groups, including ReconGroup, HI-Para, OSEM-Time, HI-Time, and DPR-Time. Figure 2 displays the variability values, while Figure 3 presents the heat maps for the categorical feature stability. These figures visualize the detailed group-wise analyses presented in the following sections.
Table 2
| Characteristic | Value |
|---|---|
| Age (year) | 65.74±10.86 [27, 86] |
| Sex (male/female) | 33/25 |
| Height (m) | 1.629±0.088 [1.49, 1.85] |
| Weight (kg) | 58.66±10.96 [39, 85] |
| Body mass index (kg/m2) | 22.05±3.46 [16.37, 32.39] |
| Uptake time (min) | 68.98 ±11.04 [50, 87] |
| Injected dose (MBq/kg) | 3.54±0.21 [2.88, 3.78] |
| Blood glucose (mmol/L) | 6.04±1.15 [4.9, 10.70] |
The values are presented as the mean ± SD [range]. SD, standard deviation.
Effect of reconstruction algorithms on the stability of image features
For ReconGroup, across different reconstruction algorithms images, 32% (34/106), 37% (39/106), 19% (20/106), and 12% (13/106) of the 106 radiomics features exhibited very small, small, medium, and large COV ranges, respectively (Table S2 and Figures 2,3). The features that were insensitive to algorithm changes (COV ≤5%) included 22 intensity and intensity histogram-based features, including SUVmean, the median and area under the curve (AUC) of the cumulative intensity volume histogram (CIVH), total lesion glycolysis (TLG), and 7 GLCM-derived, 4 GLRLM-derived and 1 GLSZM-derived features. Additionally, 20 intensity and corresponding histogram features and 19 texture features measuring heterogeneity showed small variations (5%< COV ≤10%). Moderate variability (10%< COV ≤20%) was observed in 5 first-order/histogram features and 15 texture features. The features that demonstrated a large variation (COV >20%) included seven intensity and corresponding histogram-based features (e.g., skewness, kurtosis, and histogram gradient metrics) and six texture features including joint maximum, cluster shade, and cluster prominence of GLCM; large-zone emphasis; and zone-size variance of GLSZM.
Effect of parameters on the stability of image features
Regarding the effect of β value changes in Hyper Iterative reconstruction, 68% (72/106) of features exhibited very small variability (COV ≤5%), 18% (19/106) showed small variability (5%< COV ≤10%), 9% (10/106) showed intermediate variability (10%< COV ≤20%), and only 5% (5/106) exhibited large variability (COV >20%) (Table S3 and Figures 2,3). Among the 54 intensity-based features, 43 demonstrated strong robustness, while out of the 52 texture features, 26 exhibited significant robustness. These robust features largely overlapped with those in the ReconGroup, with the COV being in the very small (100%), small (87%), and medium (40%) ranges. Only five features had a COV >20% under the β variation, including two intensity/histogram features (kurtosis of intensity and histogram) and three texture features—cluster shade (GLCM), large-zone low-gray-level emphasis (GLSZM), and zone-size variance (GLSZM).
Effect of duration on the stability of image features
Across different acquisition durations in the OSEM reconstruction algorithm, 34% (36/106), 39% (41/106), 17% (18/106), and 10% (11/106) of the 106 radiomics features in this study exhibited very small, small, medium, and large COV ranges, respectively (Table S4). Across different acquisition durations in the DPR reconstruction algorithm, 35% (37/106), 39% (41/106), 16% (17/106), and 10% (11/106) exhibited very small, small, medium, and large COV ranges, respectively (Table S5). Across different acquisition durations in the Hyper Iterative reconstruction algorithm, 37% (39/106), 36% (38/106), 19% (20/106), and 8% (9/106) demonstrated very small, small, medium, and large COV range, respectively (Table S6). A subset of 33 features consistently demonstrated high stability (COV ≤5%) across all three reconstruction algorithms and acquisition durations. These included 19 intensity- and histogram-based features (e.g., SUVmean, percentiles, AUC of CIVH, entropy) and 14 texture features from the GLCM, GLRLM, and GLSZM families (e.g., joint average, sum average, and zone-size entropy). Although a small number of features exhibited large variability (COV >20%), their identities were largely consistent across algorithms.
Relationship of COV among analysis groups
The feature-level correlation analysis results showed that all five analysis groups were significantly positively correlated with one another (Figure 4), indicating consistent patterns of variability across different reconstruction and parameter settings. The features showing very low variability (COV <5%) in all analysis groups are summarized in Table 3. The multiple comparison results demonstrated that the HI-Para group had a significantly lower average COV of features as compared to both the ReconGroup (P<0.001) and OSEM-Time (P<0.001); meanwhile, there were no statistically significant differences between ReconGroup and OSEM-Time, indicating the alteration of algorithmic parameters (HI-Para) exerted the least impact on the features.
Table 3
| Feature group | Feature (n) | Radiomics feature names |
|---|---|---|
| Intensity | 9 | Mean, medial, 25th intensity percentile, 50th intensity percentile, 75th intensity percentile, 90th intensity percentile, area under the curve CIVH, root mean square intensity, total lesion glycolysis |
| Intensity histogram | 10 | Mean, median, 25th percentile, 50th percentile, 75th percentile, 90th percentile, entropy log10, entropy log2, area under the curve CIVH, root mean square |
| GLCM | 7 | Joint average, joint entropy log2, joint entropy log10, sum average, inverse difference, normalized inverse difference, normalized inverse difference moment |
| GLRLM | 4 | Short-run emphasis, long-run emphasis, run-length nonuniformity, run percentage |
| NGTDM | – | – |
| GLSZM | 1 | Zone-size entropy |
CIVH, cumulative intensity volume histogram; COV, coefficient of variation; GLCM, gray-level co-occurrence matrix; GLRLM, gray-level run-length matrix; GLSZM, gray-level size-zone matrix; NGTDM, neighborhood gray-tone difference matrix.
In a separate analysis designed to evaluate the relative robustness of different reconstruction algorithms across varying acquisition durations, the mean COV values of OSEM-Time, HI-Time, and DPR-Time were compared. No statistically significant differences were observed among the three groups, indicating that the variability of radiomic features with respect to acquisition time was similar across all reconstruction methods.
Concordance of radiomics features across the PET image reconstruction methods
Figure 5 presents the heatmap of CCC matrices comparing short (30, 60, and 90 s) and full (120 s) acquisition durations across the three reconstruction algorithms. The majority of radiomic features demonstrated high temporal agreement with CCC >0.8 in most comparisons. Specifically, for OSEM reconstructions, 105 (99%), 98 (92%), and 104 (98%) of the 106 features achieved CCC >0.8 in the 90 vs. 120 s, 60 vs. 120 s, and 30 vs. 120 s comparisons, respectively; the corresponding values for Hyper Iterative were 98%, 96%, and 90%, while those for DPR were 95%, 94%, and 80%.
To identify robust features, CCC values were combined with %Diff thresholds (CCC >0.8 and –10%< %Diff <10%). Under this composite criterion, 59 features showed consistent reproducibility. These included 35 intensity- and histogram-based metrics [e.g., SUVmean, SUVmax, TLG, and root mean square (RMS)] and 24 texture features, including GLCM joint average and joint entropy and multiple GLRLM- and GLSZM-derived metrics (Table 4).
Table 4
| Feature group | Features (n) | Radiomics feature names |
|---|---|---|
| Intensity | 17 | Mean, median, maximum, 10th intensity percentile, 25th intensity percentile, 50th intensity percentile, 75th intensity percentile, 90th intensity percentile, standard deviation, intensity interquartile range, mean absolute deviation, robust mean absolute deviation, median absolute deviation, area under the curve CIVH, energy, root mean square intensity, total lesion glycolysis |
| Intensity histogram | 18 | Mean, median, 10th percentile, 25th percentile, 50th percentile, 75th percentile, 90th percentile, standard deviation, maximum gray-level, mode, interquartile range, mean absolute deviation, robust mean absolute deviation, entropy log10, entropy log2, area under the curve CIVH, root mean square, maximum histogram gradient gray level |
| GLCM | 10 | Joint average, joint entropy log10, joint entropy log2, difference average, sum average, inverse difference, correlation, autocorrelation, normalized inverse difference, normalized inverse difference moment |
| GLRLM | 5 | High gray-level run emphasis, short-run high-gray-level emphasis, long-run high-gray-level emphasis, gray-level nonuniformity, run-length nonuniformity |
| NGTDM | 0 | – |
| GLSZM | 9 | Small-zone emphasis, high-gray-level zone emphasis, small-zone high-gray-level emphasis, gray-level nonuniformity, zone-size nonuniformity, normalized zone-size nonuniformity, zone percentage, gray-level variance, zone-size entropy |
CCC, concordance correlation coefficient; CIVH, cumulative intensity volume histogram; %Diff, percentage difference; GLCM, gray-level co-occurrence matrix; GLRLM, gray-level run-length matrix; GLSZM, gray-level size zone matrix; NGTDM, neighborhood gray-tone difference matrix.
Among them, 27 features also exhibited low variability (COV ≤5%) across all five analysis groups, indicating high stability across both reconstruction algorithms and acquisition durations. These key intensity- and histogram-based metrics included SUVmean, SUV percentiles, TLG, entropy measures, and texture metrics such as GLCM joint average and GLSZM zone-size entropy (Table S7).
The Friedman test was used to assess the differences in the reproducibility (based on CCC values) between the three algorithms at each acquisition duration. Under 90-s reconstruction, the CCC values obtained by OSEM were significantly higher than those obtained by DPR (P<0.001) and Hyper Iterative (P<0.001). At 60 s, Hyper Iterative outperformed OSEM (P<0.001) and DPR (P<0.001). At 30 s, the CCC values obtained by OSEM were also significantly higher than those obtained by DPR (P=0.006) and Hyper Iterative (P<0.001).
Segment-based temporal reproducibility analysis
In addition to the five predefined comparison groups, we further calculated the feature-wise COV and pairwise CCC and%Diff of radiomic features across the four nonoverlapping 30-s segments (0–30, 31–60, 61–90, and 91–120 s) reconstructed via OSEM. As shown in Table S8, only 10 features exhibited very small variability (COV ≤5%) across all four temporal segments. These included key intensity histogram features (entropy log10 and entropy log2), several GLCM-based texture features, GLRLM features such as run-length nonuniformity, and one GLSZM feature (zone-size entropy).
In terms of agreement across temporal segments, the pairwise CCC values across the four 30-s temporal segments remained generally high, as shown in Figure 6, indicating the stable reproducibility of radiomic features within different subsets of the same acquisition window. The highest reproducibility was observed between 0–30 s and 31–60 s. Based on the combined criteria of CCC >0.8 and –10< %Diff <10, a total of 72 features—42 intensity/histogram-based and 30 texture-based—demonstrated high reproducibility (Table S9). Among them, six features also exhibited very small variability across segments (COV ≤5%), suggesting particularly strong robustness.
Discussion
Radiomics analysis is an emerging field in medical imaging and has been acknowledged as a promising classification tool with the inherent potential to revolutionize disease diagnosis specially cancer, provided that the features are accurate, robust, and reproducible. Reconstruction is one of the key factors influencing radiomics analysis in 18F-FDG PET/CT imaging, especially as next-generation algorithms continue to emerge. Therefore, selecting reproducible and robust features constitutes meaningful work, enabling the full utilization of imaging modalities for radiomics analysis. In this study, we aimed to investigate the robustness and reproducibility of radiomics features in 18F-FDG lung tumor PET images under different reconstruction algorithms and parameters.
Our COV correlation analysis revealed a significant positive relationship among the COVs of various features across different algorithms, parameter settings, and acquisition times. This suggests that the variability patterns of individual features are consistent across these factors, indicating that features with high or low stability tend to maintain their relative robustness regardless of the specific reconstruction condition. This observation is consistent with previous studies that reported the radiomic features are influenced by a combination of factors related to image acquisition, reconstruction, and processing, with these factors affecting features in a correlated manner, impacting their robustness (19,20). This consistency is critical for the identification of robust radiomic biomarkers, as it suggests that certain features may be less susceptible to technical variations and more reliable for clinical applications (26). For example, features with low COV and high CCC values across multiple conditions may serve as candidates for reproducible tumor quantification in multicenter studies, as they demonstrate insensitivity to changes in imaging protocols (6,7).
By integrating reproducibility (CCC) and variability (COV) criteria, we identified a subset of radiomic features that consistently demonstrated stability across different reconstruction algorithms and acquisition durations, underscoring their potential as reliable biomarkers for clinical PET imaging. The stability of SUV-based metrics such as SUVmean, intensity percentiles (25th–90th), and TLG may be related to their inherent dependence on SUV normalization protocols in which SUV-based metrics are explicitly calibrated to patient weight and injected dose (27), which systematically reduces technical variability across algorithms. Similarly, histogram-driven parameters (e.g., AUC of CIVH and RMS values) exhibited strong reproducibility, supporting their suitability for streamlining clinical workflows without compromising reproducibility, which is in line with efforts to optimize PET protocols for radiomics (28,29).
For texture features, the stability of GLCM-derived parameters and GLRLM features attests to their relative insensitivity to algorithm-specific implementations. This finding supports Larue et al.’s hypothesis (30), which holds that features relying on spatial averaging (e.g., sum average) are less affected by computational pipelines than are higher-order statistics. Together, these findings suggest that a carefully selected group of first- and second-order radiomics features, especially those capturing global intensity distribution and texture, remain robust against reconstruction algorithms. This stability enhances their potential utility for accurately characterizing tumor metabolism and heterogeneity in clinical applications.
The identification of radiomic features that are robust to variations in reconstruction algorithms and acquisition parameters is relevant for clinical and translational applications. In multicenter clinical trials, differences in scanner hardware and reconstruction settings often introduce technical variability that undermines the reproducibility of radiomics biomarkers (19,20,31). Robust features, as identified in this study, offer the potential for standardized image-based quantification across heterogeneous imaging protocols. This is especially valuable in longitudinal disease monitoring or treatment response assessment, where imaging conditions may change over time. Recent consensus recommendations have emphasized the need for reproducibility as a prerequisite for radiomic feature selection in clinical workflows (32,33). Moving forward, robust features should be prioritized as candidate biomarkers, but only those that also demonstrate prognostic or predictive utility in outcome-based models should be translated into clinical endpoints (34).
Although all features in this study were derived from SUV images, our analysis revealed that feature robustness varies not only across feature categories (e.g., first-order vs. texture) but also within them. For example, certain GLCM-based features such as joint average and sum average showed strong stability across different reconstruction conditions, suggesting robustness to voxel-level variation. In contrast, some first-order features (e.g., kurtosis and entropy) were relatively more sensitive to changes in algorithm settings and acquisition duration. In line with this, Fooladi et al. reported that among the different reconstruction algorithms and parameter settings they examined, the textural-based feature group showed the highest robustness, with 11 texture features with stable values being identified across prostate-specific membrane antigen PET image reconstructions (6). Additionally, a previous phantom-based study using the same Hyper Iterative and DPR algorithms as those applied in the present study revealed that certain radiomic features maintain a COV below 10% across reconstruction techniques, indicating a baseline level of stability inherent to the original features regardless of the reconstruction technique employed (14).
Higher-order texture methods, based on both grayscale values and spatial orientation, generate metrics that quantify gray-level differences between neighboring pixels or voxels and encode higher-order information such as spatial intensity distributions (e.g., GLCM and GLRLM), gradient patterns, and regional texture complexity (21). It has also been reported that these features tend to be more sensitive to image resolution, noise, and quantization effects than to global SUV stability (19,20,35). These findings suggest that the robustness of certain features we observed may not solely arise from SUV-level consistency, potentially indicating the inherent stability of specific feature formulations.
To address the influence of count statistics on radiomic features, we analyzed four nonoverlapping 30-s segments within the 2-min acquisition window. Notably, only 10 features exhibited very small variability (COV ≤5%) across all segments (Table S8). Interestingly, these included just two intensity histogram features (entropy log10 and entropy log2), while other first-order features commonly identified as highly robust in the main analysis—such as SUVmean, percentiles, and TLG—did not meet the COV threshold in this setting. In contrast, 8 of the 10 stable features were texture-based, consistent with those identified as robust across reconstruction algorithms, acquisition times, and parameters. This discrepancy suggests that short-term count variability may disproportionately affect the absolute stability of voxel-level intensity features, while texture metrics—by virtue of spatial averaging and relative voxel contrast—may be more resilient to the noise introduced by temporal undersampling. Notably, despite the limited COV-stable features, a large proportion of features still showed high reproducibility based on CCC and %Diff, indicating that count variability primarily influences value fluctuation rather than structural consistency. Importantly, these findings underscore that certain texture features remain reliable even under reduced or temporally segmented acquisitions, reaffirming their robust reproducibility and potential clinical value.
In the algorithms comparison, no significant difference was found in the average COV of features between the three algorithms, indicating that the stability performance of the two new algorithms, Hyper Iterative and DPR, under various acquisition durations, appeared to be comparable to that of the OSEM algorithm. This observation differed from the conclusions of previous phantom study, which suggested superior stability in advanced algorithms as compared to OSEM (14). The discrepancy in stability outcomes between patient and phantom studies might arise from fundamental differences in study conditions. Phantom studies, conducted under controlled environments with known ground-truth activity distributions and minimal noise variability, inherently favor advanced algorithms by allowing full exploitation of their noise-suppression capabilities (36). In contrast, patient studies introduce biological variability (e.g., respiratory motion and heterogeneous tracer uptake) and suboptimal count statistics due to clinical time constraints, which may diminish the theoretical advantages of advanced algorithms (37). This highlights the critical importance of validating phantom-based algorithm performance claims in clinical settings with patient-derived confounding factors.
Differences in feature reproducibility across reconstruction algorithms may not directly reflect algorithm superiority but rather stem from the interaction between reconstruction characteristics, segmentation strategy, and feature-specific sensitivity. In our study, VOIs were delineated via CT-based anatomical segmentation, which inherently includes both high-uptake tumor regions and adjacent peripheral zones with lower activity. Advanced reconstruction algorithms such as Hyper Iterative and DPR enhance image contrast and better define subtle gray-level transitions, particularly in boundary and low-intensity regions (9-13). This enhancement, while beneficial for lesion detection, may inadvertently introduce variability in features that are sensitive to low gray levels or large homogeneous zones. Specifically, features such as GLRLM long-run low-gray-level emphasis and GLSZM large-zone emphasis and low-gray-level zone emphasis are known to be affected by subtle intensity fluctuations and spatial texture regularity (19,20,35). As shown in Figure 5, these features exhibited lower CCC values in reconstructions using advanced algorithms, suggesting that contrast enhancement may amplify minor voxel-level variations in heterogeneous VOIs. Importantly, prior studies have emphasized that the robustness of radiomic features depends not only on noise characteristics but also on segmentation methods and reconstruction-feature compatibility (6,8,21,30). These findings indicate that there is a need to align reconstruction methods, segmentation strategies, and feature selection to optimize reproducibility in radiomics workflows.
This study involved certain limitations that should be addressed. First, we did not consider the effect of quantification and segmentation methods on the results, and future work should examine the impact of these methods on the radiomics characteristics of patients with lung cancer. Second, the initial experience gained in this study was based on a single center and a single disease type (lung cancer), and multicenter data collection is needed to validate the results of this study. Finally, while we focused on identifying radiomic features that are robust across different reconstruction settings, we acknowledge that robustness alone does not guarantee clinical relevance or discriminative power. Therefore, future studies should combine robustness evaluation with outcome-driven analyses, such as predictive modeling and survival analysis, to identify features with both stability and clinical utility.
Conclusions
In this study, we systematically evaluated the effect of different reconstruction algorithms, parameter settings, and acquisition durations on PET-based radiomic features. Our findings demonstrate that a subset of radiomic features—specifically those reflecting global intensity distribution and spatial texture patterns—consistently exhibit high robustness across reconstruction algorithms, parameter settings, acquisition durations, and even under variable count statistics, supporting their potential as reliable and reproducible imaging biomarkers for clinical PET applications. Count variability primarily affects absolute feature values rather than reproducibility, especially for intensity-based features. Reconstruction methods, segmentation strategies, and feature selection should be jointly considered to optimize reproducibility in radiomics workflows. The robust features identified in this study are promising candidates for consistent and robust tumor quantification, particularly in multicenter studies, longitudinal monitoring, and prospective trials, where imaging protocols may vary.
Acknowledgments
We would like to thank Ying Wang and Shengle Wang for their helpful comments and assistance in manuscript editing.
Footnote
Data Sharing Statement: Available at https://qims.amegroups.com/article/view/10.21037/qims-2025-958/dss
Funding: This study was supported by
Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://qims.amegroups.com/article/view/10.21037/qims-2025-958/coif). S.Z. is an employee of United Imaging Healthcare Group Co., Ltd. The other authors have no conflicts of interest to declare.
Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study was approved by the Ethics Committee of The First Affiliated Hospital, Zhejiang University School of Medicine ([2025B] IIT Ethics approval No. 0921) and individual consent for this retrospective analysis was waived.
Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.
References
- Wahl RL, Jacene H, Kasamon Y, Lodge MA. From RECIST to PERCIST: Evolving Considerations for PET response criteria in solid tumors. J Nucl Med 2009;50:122S-50S. [Crossref] [PubMed]
- Yan J, Chu-Shern JL, Loi HY, Khor LK, Sinha AK, Quek ST, Tham IW, Townsend D. Impact of Image Reconstruction Settings on Texture Features in 18F-FDG PET. J Nucl Med 2015;56:1667-73. [Crossref] [PubMed]
- Lambin P, Rios-Velazquez E, Leijenaar R, Carvalho S, van Stiphout RG, Granton P, Zegers CM, Gillies R, Boellard R, Dekker A, Aerts HJ. Radiomics: extracting more information from medical images using advanced feature analysis. Eur J Cancer 2012;48:441-6. [Crossref] [PubMed]
- Sun Y, Ge X, Niu R, Gao J, Shi Y, Shao X, Wang Y, Shao X. PET/CT radiomics and deep learning in the diagnosis of benign and malignant pulmonary nodules: progress and challenges. Front Oncol 2024;14:1491762. [Crossref] [PubMed]
- Mayerhoefer ME, Materka A, Langs G, Häggström I, Szczypiński P, Gibbs P, Cook G. Introduction to Radiomics. J Nucl Med 2020;61:488-95. [Crossref] [PubMed]
- Fooladi M, Soleymani Y, Rahmim A, Farzanefar S, Aghahosseini F, Seyyedi N, Sh, Zadeh P. Impact of different reconstruction algorithms and setting parameters on radiomics features of PSMA PET images: A preliminary study. Eur J Radiol 2024;172:111349. [Crossref] [PubMed]
- Shiri I, Rahmim A, Ghaffarian P, Geramifar P, Abdollahi H, Bitarafan-Rajabi A. The impact of image reconstruction settings on 18F-FDG PET radiomic features: multi-scanner phantom and patient studies. Eur Radiol 2017;27:4498-509. [Crossref] [PubMed]
- van Velden FH, Kramer GM, Frings V, Nissen IA, Mulder ER, de Langen AJ, Hoekstra OS, Smit EF, Boellaard R. Repeatability of Radiomic Features in Non-Small-Cell Lung Cancer [(18)F]FDG-PET/CT Studies: Impact of Reconstruction and Delineation. Mol Imaging Biol 2016;18:788-95.
- Xu L, Li RS, Wu RZ, Yang R, You QQ, Yao XC, Xie HF, Lv Y, Dong Y, Wang F, Meng QL. Small lesion depiction and quantification accuracy of oncological (18)F-FDG PET/CT with small voxel and Bayesian penalized likelihood reconstruction. EJNMMI Phys 2022;9:23. [Crossref] [PubMed]
- Xu L, Cui C, Li R, Yang R, Liu R, Meng Q, Wang F. Phantom and clinical evaluation of the effect of a new Bayesian penalized likelihood reconstruction algorithm (HYPER Iterative) on (68)Ga-DOTA-NOC PET/CT image quality. EJNMMI Res 2022;12:73. [Crossref] [PubMed]
- Sui X, Tan H, Yu H, Xiao J, Qi C, Cao Y, Chen S, Zhang Y, Hu P, Shi H. Exploration of the total-body PET/CT reconstruction protocol with ultra-low 18F-FDG activity over a wide range of patient body mass indices. EJNMMI Phys 2022;9:17. [Crossref] [PubMed]
- Lv Y, Xi C. PET image reconstruction with deep progressive learning. Phys Med Biol 2021; [Crossref]
- Wang T, Qiao W, Wang Y, Wang J, Lv Y, Dong Y, Qian Z, Xing Y, Zhao J. Deep progressive learning achieves whole-body low-dose (18)F-FDG PET imaging. EJNMMI Phys 2022;9:82. [Crossref] [PubMed]
- Shiiba T, Watanabe M. Stability of radiomic features from positron emission tomography images: a phantom study comparing advanced reconstruction algorithms and ordered subset expectation maximization. Phys Eng Sci Med 2024;47:929-37. [Crossref] [PubMed]
- Stefano A. Challenges and limitations in applying radiomics to PET imaging: Possible opportunities and avenues for research. Comput Biol Med 2024;179:108827. [Crossref] [PubMed]
- Orlhac F, Soussan M, Maisonobe JA, Garcia CA, Vanderlinden B, Buvat I. Tumor texture analysis in 18F-FDG PET: relationships between texture parameters, histogram indices, standardized uptake values, metabolic volumes, and total lesion glycolysis. J Nucl Med 2014;55:414-22. [Crossref] [PubMed]
- Pfaehler E, Beukinga RJ, de Jong JR, Slart RHJA, Slump CH, Dierckx RAJO, Boellaard R. Repeatability of (18) F-FDG PET radiomic features: A phantom study to explore sensitivity to image reconstruction settings, noise, and delineation method. Med Phys 2019;46:665-78. [Crossref] [PubMed]
- Nioche C, Orlhac F, Boughdad S, Reuzé S, Goya-Outi J, Robert C, Pellot-Barakat C, Soussan M, Frouin F, Buvat I. LIFEx: A Freeware for Radiomic Feature Calculation in Multimodality Imaging to Accelerate Advances in the Characterization of Tumor Heterogeneity. Cancer Res 2018;78:4786-9. [Crossref] [PubMed]
- Zwanenburg A, Vallières M, Abdalah MA, Aerts HJWL, Andrearczyk V, Apte A, et al. The Image Biomarker Standardization Initiative: Standardized Quantitative Radiomics for High-Throughput Image-based Phenotyping. Radiology 2020;295:328-38. [Crossref] [PubMed]
- Traverso A, Wee L, Dekker A, Gillies R. Repeatability and Reproducibility of Radiomic Features: A Systematic Review. Int J Radiat Oncol Biol Phys 2018;102:1143-58. [Crossref] [PubMed]
- Varghese BA, Fields BKK, Hwang DH, Duddalwar VA, Matcuk GR Jr, Cen SY. Spatial assessments in texture analysis: what the radiologist needs to know. Front Radiol 2023;3:1240544. [Crossref] [PubMed]
- Lin LI. A concordance correlation coefficient to evaluate reproducibility. Biometrics 1989;45:255-68.
- Fukai S, Daisaki H, Ishiyama M, Shimada N, Umeda T, Motegi K, Ito R, Terauchi T. Reproducibility of the principal component analysis (PCA)-based data-driven respiratory gating on texture features in non-small cell lung cancer patients with (18) F-FDG PET/CT. J Appl Clin Med Phys 2023;24:e13967. [Crossref] [PubMed]
- Crandall JP, Fraum TJ, Lee M, Jiang L, Grigsby P, Wahl RL. Repeatability of (18)F-FDG PET Radiomic Features in Cervical Cancer. J Nucl Med 2021;62:707-15. [Crossref] [PubMed]
- Faist D, Jreige M, Oreiller V, Nicod Lalonde M, Schaefer N, Depeursinge A, Prior JO. Reproducibility of lung cancer radiomics features extracted from data-driven respiratory gating and free-breathing flow imaging in [18F]-FDG PET/CT. Eur J Hybrid Imaging 2022;6:33.
- Leijenaar RT, Carvalho S, Velazquez ER, van Elmpt WJ, Parmar C, Hoekstra OS, Hoekstra CJ, Boellaard R, Dekker AL, Gillies RJ, Aerts HJ, Lambin P. Stability of FDG-PET Radiomics features: an integrated analysis of test-retest and inter-observer variability. Acta Oncol 2013;52:1391-7. [Crossref] [PubMed]
- Boellaard R, Delgado-Bolton R, Oyen WJ, Giammarile F, Tatsch K, Eschner W, et al. FDG PET/CT: EANM procedure guidelines for tumour imaging: version 2.0. Eur J Nucl Med Mol Imaging 2015;42:328-54. [Crossref] [PubMed]
- Hatt M, Tixier F, Pierce L, Kinahan PE, Le Rest CC, Visvikis D. Characterization of PET/CT images using texture analysis: the past, the present… any future? Eur J Nucl Med Mol Imaging 2017;44:151-65. [Crossref] [PubMed]
- van Timmeren JE, Leijenaar RTH, van Elmpt W, Wang J, Zhang Z, Dekker A, Lambin P. Test-Retest Data for Radiomics Feature Stability Analysis: Generalizable or Study-Specific? Tomography 2016;2:361-5. [Crossref] [PubMed]
- Larue RTHM, van Timmeren JE, de Jong EEC, Feliciani G, Leijenaar RTH, Schreurs WMJ, Sosef MN, Raat FHPJ, van der Zande FHR, Das M, van Elmpt W, Lambin P. Influence of gray level discretization on radiomic feature stability for different CT scanners, tube currents and slice thicknesses: a comprehensive phantom study. Acta Oncol 2017;56:1544-53. [Crossref] [PubMed]
- Yip SS, Aerts HJ. Applications and limitations of radiomics. Phys Med Biol 2016;61:R150-66. [Crossref] [PubMed]
- Lambin P, Leijenaar RTH, Deist TM, Peerlings J, de Jong EEC, van Timmeren J, Sanduleanu S, Larue RTHM, Even AJG, Jochems A, van Wijk Y, Woodruff H, van Soest J, Lustberg T, Roelofs E, van Elmpt W, Dekker A, Mottaghy FM, Wildberger JE, Walsh S. Radiomics: the bridge between medical imaging and personalized medicine. Nat Rev Clin Oncol 2017;14:749-62. [Crossref] [PubMed]
- Ibrahim A, Primakov S, Beuque M, Woodruff HC, Halilaj I, Wu G, Refaee T, Granzier R, Widaatalla Y, Hustinx R, Mottaghy FM, Lambin P. Radiomics for precision medicine: Current challenges, future prospects, and the proposal of a new framework. Methods 2021;188:20-9. [Crossref] [PubMed]
- Kickingereder P, Burth S, Wick A, Götz M, Eidel O, Schlemmer HP, Maier-Hein KH, Wick W, Bendszus M, Radbruch A, Bonekamp D. Radiomic Profiling of Glioblastoma: Identifying an Imaging Predictor of Patient Survival with Improved Performance over Established Clinical and Radiologic Risk Models. Radiology 2016;280:880-9. [Crossref] [PubMed]
- Orlhac F, Boughdad S, Philippe C, Stalla-Bourdillon H, Nioche C, Champion L, Soussan M, Frouin F, Frouin V, Buvat I. A Postreconstruction Harmonization Method for Multicenter Radiomic Studies in PET. J Nucl Med 2018;59:1321-8. [Crossref] [PubMed]
- Erlandsson K, Buvat I, Pretorius PH, Thomas BA, Hutton BF. A review of partial volume correction techniques for emission tomography and their applications in neurology, cardiology and oncology. Phys Med Biol 2012;57:R119-59. [Crossref] [PubMed]
- Nehmeh SA, Erdi YE. Respiratory motion in positron emission tomography/computed tomography: a review. Semin Nucl Med 2008;38:167-76. [Crossref] [PubMed]

