Large language models for analyzing contrast-enhanced ultrasound reports of pancreatic cystic lesions
Original Article

Large language models for analyzing contrast-enhanced ultrasound reports of pancreatic cystic lesions

Yuming Shao, Yang Gui, Xiaoyi Yan, Tianjiao Chen, Xueqi Chen, Li Tan, Jing Zhang, Hua Liang, Wanying Jia, Huijia Zhao, Baoquan Chen, Yuxin Jiang, Ke Lv

Department of Ultrasound, Peking Union Medical College Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College, Beijing, China

Contributions: (I) Conception and design: Y Shao, Y Gui, K Lv; (II) Administrative support: Y Jiang, K Lv; (III) Provision of study materials or patients: Y Gui, T Chen, X Chen, L Tan, J Zhang, H Liang, W Jia, H Zhao, B Chen; (IV) Collection and assembly of data: Y Shao, X Yan; (V) Data analysis and interpretation: Y Shao, X Yan; (VI) Manuscript writing: All authors; (VII) Final approval of manuscript: All authors.

Correspondence to: Ke Lv, MD. Department of Ultrasound, Peking Union Medical College Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College, No. 1 Shuaifuyuan, Dongcheng District, Beijing 100730, China. Email: lvke@163.com.

Background: Pancreatic cystic lesions (PCLs) require precise imaging characterization to guide clinical management. Contrast-enhanced ultrasound (CEUS) reports contain operator-dependent narratives that challenge clinicians. Large language models (LLMs) show potential in medical text analysis but lack validation for pancreatic CEUS interpretation. This study primarily aimed to evaluate the diagnostic accuracy of LLMs in interpreting Chinese CEUS reports of PCLs.

Methods: We retrospectively analyzed 80 pathologically confirmed PCLs (21 benign, 25 borderline malignant, 34 malignant). Four LLMs (GPT-4o, Claude 3.7 Sonnet, Gemini 2.0, DeepSeek-R1) and eight radiologists (senior/junior =4:4) independently interpreted reports under three input modalities: grayscale-only (IM1), grayscale + CEUS (IM2), and demographics + grayscale + CEUS (IM3). A weighted scoring system (0–100 points per case, yielding a maximum total of 8,000 points) quantified alignment with pathology-defined categories (benign/borderline/malignant). LLM errors were categorized into four reasons. Junior radiologists reinterpreted cases with LLM assistance.

Results: Under CEUS input conditions (IM2/IM3), LLMs showed no statistically significant difference from senior radiologists and significantly outperformed junior radiologists. The median score per case (out of 100), after averaging across the four LLMs, increased with input complexity (IM1: 47.50; IM2: 53.75; IM3: 73.75). Diagnostic accuracy varied by pathology: malignant lesions scored highest, while benign serous cystic neoplasms scored lowest due to suboptimal CEUS visualization and LLMs’ knowledge gaps. LLM guidance elevated junior radiologists’ accuracy to senior levels.

Conclusions: LLMs show promising capability in interpreting CEUS reports of PCLs, with no statistically significant difference from senior radiologists in this dataset. Their integration significantly improves junior clinicians’ interpretation of CEUS reports. These findings support further investigation of LLMs as potential auxiliary tools for ultrasound text analysis, though external validation is needed.

Keywords: Large language model (LLM); contrast-enhanced ultrasound (CEUS); pancreatic cystic lesions (PCLs); ultrasound


Submitted Feb 04, 2026. Accepted for publication Jul 06, 2026. Published online Aug 05, 2026.

doi: 10.21037/qims-2026-1-0307


Introduction

Pancreatic cystic lesions (PCLs) encompass a heterogeneous spectrum of lesions, including cystic neoplasms [e.g., serous cystic neoplasm (SCN), mucinous cystic neoplasm (MCN), intraductal papillary mucinous neoplasm (IPMN), solid pseudopapillary tumor (SPT)], and cystic degeneration of solid tumors [e.g., pancreatic ductal adenocarcinoma (PDAC)], and non-neoplastic entities (e.g., pseudocysts) (1,2). These lesions share overlapping imaging features such as cystic components, septations, or solid nodules, complicating radiological differentiation. Crucially, their management diverges significantly: SCNs are typically benign and require surveillance only, while MCNs/IPMNs harbor malignant potential necessitating resection or rigorous follow-up, and PDACs are malignant tumors (3,4). This diagnostic ambiguity underscores the imperative for precise lesion characterization to guide clinical decision-making.

While contrast-enhanced computed tomography (CT) and magnetic resonance imaging (MRI) remain cornerstone modalities for evaluating PCLs, contrast-enhanced ultrasound (CEUS) has gained recognition for its real-time assessment of microvascular perfusion, lack of ionizing radiation, and cost-effectiveness (5). Studies report comparable diagnostic efficacy between CEUS and MRI in diagnosing PCLs (6,7). Several guidelines also endorse CEUS for identifying vascularized mural nodules, septa, and differentiating avascular mucin plugs from enhancing solid components (8,9). However, a significant challenge limits its impact: CEUS findings are communicated through written reports that are highly operator-dependent. For clinicians who are not ultrasound specialists, these narrative reports can be difficult to interpret fully, as they use specialized terminology. In practice, management decisions often rely more on interpreting this text than on viewing the images directly. This creates a dependency on individual radiologist expertise, leading to potential variability in how these critical lesions are managed.

Large language models (LLMs) are language models trained on massive amounts of text data, most commonly from the internet (10). They demonstrate strong capabilities in natural language processing, including the ability to extract relevant information from unstructured text and identify patterns that support diagnostic classification. In medical imaging, LLMs have been leveraged for generating structured reports (11), severity grading (12), and report explanation (13). Nevertheless, existing research predominantly focuses on CT or MRI. A growing body of work has explored the application of LLMs across various ultrasound disciplines, including thyroid (14,15), breast (16,17), liver (18), ovarian (19), and musculoskeletal imaging (20,21), where they have been used for tasks such as report generation, risk stratification, and diagnostic classification. However, research on LLMs in the context of CEUS remains sparse, with only a handful of studies evaluating their feasibility for specific tasks such as liver Liver Imaging Reporting and Data System (LI-RADS) categorization (22). In the specific domain of PCLs, existing LLMs research has been almost exclusively focused on CT and MRI, addressing tasks such as automated feature extraction from radiology reports (23), risk classification of IPMN (24), and provision of management recommendations (25). Nevertheless, a pronounced gap remains in applying LLMs to ultrasound texts, particularly for complex CEUS narratives for PCLs. Pancreatic CEUS reports pose unique linguistic challenges as they integrate morphological descriptions from gray-scale ultrasound with dynamic contrast enhancement features.

Therefore, this study primarily aimed to evaluate the diagnostic accuracy of LLMs in interpreting Chinese CEUS reports of PCLs. Building on this primary accuracy analysis, we further sought to characterize the sources of diagnostic errors and to assess whether LLM-derived guidance could improve the diagnostic performance of junior radiologists. We present this article in accordance with the STARD reporting checklist (available at https://qims.amegroups.com/article/view/10.21037/qims-2026-1-0307/rc).


Methods

Participants

This study retrospectively and consecutively included all inpatients with surgically confirmed PCLs who underwent preoperative CEUS at Peking Union Medical College Hospital (PUMCH) between March 2017 and June 2024. No eligible patients were excluded; a total of 80 patients were enrolled. All CEUS examinations were performed within one month before the surgical pathology was obtained. Demographic information including age and sex alongside CEUS reports were collected. According to World Health Organization classification and clinical management principles, patients were stratified into three categories: Category 1 (benign) comprising SCN and other cystic lesions (hematomas, pseudocysts, lymphoepithelial cysts); Category 2 (borderline malignant) including IPMN, MCN, and SPT; and Category 3 (malignant) consisting of PDAC, malignant IPMN, malignant MCN, and malignant SPT. This grouping, while not a formal pathological category, reflects the clinical risk stratification that guides patient management. This study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. This study was approved by the Ethics Committee of Peking Union Medical College Hospital (approval No. I-23PJ898). Informed consent was taken from all individual participants.

CEUS report characteristics

All reports originated from nine radiologists using a standardized template, which individual radiologists could adapt to describe lesion-specific findings; the reports are therefore best characterized as semi-structured. Reports systematically described grayscale features (location, size, echogenicity, morphology, boundaries, septations, solid components, pancreatic duct status, calcifications), color Doppler characteristics (intralesional vascularity, relationship with adjacent vessels), and CEUS features (enhancement pattern in arterial phase, washout velocity in venous phase) in sequential order. While adhering to this structure, radiologists may supplement descriptions for atypical findings such as honeycomb septations, distribution of anechoic areas, or centripetal enhancement based on lesion-specific manifestations.

LLM evaluation

We assessed four LLMs (GPT-4o, Claude 3.7 Sonnet, Gemini 2.0, DeepSeek-R1) using three standardized input modalities: IM1 (grayscale report only), IM2 (grayscale + CEUS reports), and IM3 (demographics (age/sex) + grayscale + CEUS reports). All four LLMs were accessed via their respective web interfaces between April 14 and April 21, 2025, with default platform settings and no manual adjustment of temperature, top-p, or maximum output tokens. No system prompt was employed; the standardized user prompt was entered directly. Each case was submitted as an independent single-turn query, without any follow-up conversation or reference to prior model responses. Although conversation history was not explicitly cleared between cases, the single-turn nature of each query (with no model output being incorporated into subsequent prompts) ensured that each case was evaluated independently. Each case was evaluated only once per model per input modality. All model inputs and outputs were in Chinese. The English translation of the prompt (shown below) is provided for the reader’s convenience only and was not used for model inference. We are analyzing a new ultrasound report describing a lesion predominantly composed of cystic components, potentially with some solid elements. Known pancreatic cystic neoplasms include serous cystadenoma, mucinous cystadenoma, IPMN, and SPT. Other conditions broadly classified as PCLs include pseudocysts, hematomas, and lymphoepithelial cysts. Some predominantly solid lesions may undergo cystic degeneration, including PDAC, mucinous cystadenocarcinoma, and intraductal papillary mucinous carcinoma. Based on the above pathological types and the following ultrasound report, what is your assessment? Original Chinese prompts are available in Appendices 1,2. For IM3, demographic information preceded the imaging description. It is worth noting that the content input to the models consisted solely of the de-identified lesion descriptions, without any impression/diagnostic opinion section or other statements that might reveal a human diagnostic conclusion. Laboratory values such as CA19-9 were excluded to prevent the models from overweighting biomarkers relative to imaging descriptors. Ultrasound images were also excluded, as our preliminary validation confirmed that all four evaluated LLMs could identify ultrasound images generically but were unable to localize or characterize PCLs on these images.

Output scoring system

A weighted scoring system allocated a maximum of 100 points per case based on diagnostic quantity and ranking. Points were distributed according to predefined rules where a single diagnosis received 100 points, two diagnoses followed either equal (50/50) or weighted (70/30) partitioning, three diagnoses used distributions of 60/30/10, 33/33/33, 60/20/20, or 40/40/20, and four diagnoses employed schemes of 35/35/15/15 or 60/30/5/5. Crucially, points were awarded exclusively when the predicted diagnostic groups—benign, borderline malignant, or malignant—matched the pathological gold standard. The scoring system was designed to prioritize category-level agreement over exact pathological matching, reflecting the clinical reality that risk stratification—rather than precise pathological subtyping—often drives management decisions for PCLs. For instance, an IPMN case ranked as “IPMN (highest probability) > PDAC (middle probability) > MCN (lowest probability)” received 70 points, comprising 60 points for IPMN (highest probability) and 10 points for MCN (lowest probability), since both IPMN and MCN belong to the borderline malignant group. Ranked diagnoses were directly extracted from the free-text outputs based on the explicit ordering language used in each response (e.g., “most likely…, followed by…”), and their subsequent assignment to the three pathological categories was rule-based. This entire procedure was performed by a single investigator and independently verified by a second; no discrepancies were found. To supplement the weighted scoring system, 3×3 confusion matrices and per-category sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) were also calculated for each reader group under IM3.

Human reader assessment

Four senior radiologists (>25, >20, >15, and >10 years’ experience, and authors of 7, 13, 29, and 4 reports respectively) and four junior radiologists (<5 years’ experience) independently interpreted de-identified reports using identical prompts. Responses were scored per the same protocol. To prevent carryover effects, readers sequentially answered prompts (IM1 → IM2 → IM3) without modifying prior responses after accessing additional information. Four weeks later, junior radiologists reinterpreted the same cases. For each case, they were provided with the diagnoses suggested by the four LLMs, but remained blinded to their own initial responses. The same scoring criteria were used.

Error classification

Low-scoring cases were classified into four reasons. Reason 1 encompassed errors attributed to suboptimal image visualization, resulting in incomplete or misleading depiction of lesion features in the report. The downstream misinterpretation by both human readers and LLMs was therefore a shared consequence of this upstream imaging limitation. Reason 2 represented cases with correct human and report interpretation where lesions exhibited atypical imaging presentations confounding diagnosis. Reason 3 denoted failures in LLM comprehension despite accurate human and report interpretation. Reason 4 captured instances where both the report and LLM interpretation were correct, but human readers committed comprehension errors. Error categories were assigned through consensus discussion between two investigators who reviewed all relevant imaging, pathological, and model output data together. Given the retrospective nature of this qualitative analysis, formal blinding to the reference standard was not feasible. Notably, some cases were assigned more than one reason.

Statistical analysis

Continuous variables are expressed as median (interquartile range). For paired comparisons, the Friedman test was used, followed by Wilcoxon signed-rank tests with Bonferroni correction for pairwise comparisons. For independent group comparisons, the Kruskal-Wallis test was used, with Bonferroni-corrected pairwise comparisons (Mann-Whitney U tests). Median differences and 95% confidence intervals (CIs) were estimated using the Hodges-Lehmann method. For comparisons across reader groups (LLMs, senior radiologists, junior radiologists), the unit of analysis was the per-case mean score averaged across the four LLMs, four senior radiologists, and four junior radiologists, respectively. These per-case scores were treated as paired observations across groups (n=80 per group) and analyzed using the Friedman test. All analyses used SPSS 26.0 (IBM Corp.) with statistical significance defined as P<0.05.


Results

The study included 80 patients with PCLs, comprising 45 males and 35 females. Pathological stratification revealed 21 benign cases, with SCN predominating (n=14). Borderline malignant cases totaled 25, predominantly IPMN (n=15). The malignant group contained 34 cases, primarily PDAC and malignant IPMN (each n=14). Complete demographic and pathological distributions are presented in Table 1.

Table 1

Clinical and pathological characteristics of 80 patients with PCLs

Pathological type Cases Age, years Gender, male/female
Benign 21 54.00 (25.00) 10/11
   SCN 14 57.00 (28.75) 6/8
   Other benign cysts 7 54.00 (20.00) 4/3
Borderline 25 57.00 (30.00) 17/8
   IPMN 15 65.00 (18.00) 12/3
   MCN 2 68.50 (NA) 0/2
   SPT 8 31.00 (10.75) 5/3
Malignant 34 60.50 (14.50) 18/16
   PDAC 14 65.00 (13.00) 9/5
   Malignant MCN/SPT 6 50.50 (19.75) 3/3
   Malignant IPMN 14 65.00 (18.00) 6/8
Total 80 58.00 (22.75) 45/35

Data are presented as median (IQR) or number. IPMN, intraductal papillary mucinous neoplasms; IQR, interquartile range; MCN, mucinous cystic neoplasm; NA, not applicable; PCL, pancreatic cystic lesion; PDAC, pancreatic ductal adenocarcinoma; SCN, serous cystic neoplasm; SPT, solid pseudopapillary tumor.

For the three input modalities, the median scores per case (out of 100), after averaging across the four LLMs, for the four LLMs were 47.50 (IM1), 53.75 (IM2), and 73.75 (IM3). Among all four LLMs, GPT-4o achieved the highest total score in the IM2 (4,520/8,000, or 56.50 per case), while Deepseek-R1 performed best with IM3 scoring (4,800/8,000, or 60.00 per case). Notably, four LLMs maintained acceptable internal consistency and no statistically significant differences emerged in LLMs scores within each input modality (P=0.298 for IM1, 0.513 for IM2, and 0.962 for IM3).

Compared with radiologists, significant differences emerged in mean scores between groups establishing the consistent performance hierarchy of senior radiologists > LLMs > junior radiologists (P=0.015 for IM1, P=0.002 for IM2, P<0.001 for IM3). Specifically, no statistically significant difference was observed between senior radiologists and LLMs under CEUS input modalities in IM2 (P=0.193, median difference 3.750 points, 95% CI: 0.000 to 10.000) or IM3 (P=0.387, median difference 0.000 points, 95% CI: −3.750 to 8.750), though they differed significantly in IM1 (P=0.014, median difference 8.750 points, 95% CI: 1.250 to 17.500). Conversely, LLMs consistently outperformed junior radiologists in CEUS input modalities with significant advantages in IM2 (P=0.034, median difference −6.250 points, 95% CI: −12.500 to 0.000) and IM3 (P<0.001, median difference −12.500 points, 95% CI: −16.250 to −3.750), while their IM1 performance showed no statistical difference (P=0.412, median difference −2.500 points, 95% CI: −8.750 to 2.500) (Figure 1).

Figure 1 Scores of LLM, senior, and junior radiologists under each input modality. Scores represent the mean score per case (out of 100) from the weighted scoring system. ns, not significant; *, P<0.05; **, P<0.01. IM, input modality; LLM, large language model.

Standard 3×3 confusion matrices and derived per-category sensitivity and specificity for each reader group under IM3 are provided in Tables S1-S4. The overall three-category accuracy was 60.6% for LLMs, 61.3% for senior radiologists, and 45.6% for junior radiologists. These results are consistent with the performance hierarchy and the pattern of category-level scores from the weighted scoring system.

Progressive score improvement occurred in LLMs as information increased from IM1 to IM3. Human radiologists exhibited similar overall trends though 3 junior radiologists showed paradoxical score reduction when demographic data was added (Table S5). For LLMs, the largest score gain from added demographic data was seen in SPT, while malignant lesions showed a more modest increase. In contrast, modest score decreases were observed for other benign cysts and IPMN, suggesting that the added demographic data may have introduced noise rather than useful priors for these particular subtypes (Table S6).

Given that LLMs achieved their highest scores under IM3, we focused on this input modality to analyze diagnostic performance across pathological subtypes. Mean scores for each pathological category by LLMs, senior radiologists, and junior radiologists are presented in Table 2. LLMs demonstrated significantly varying performance across pathological groups with the lowest scores for benign lesions, intermediate for borderline malignant lesions, and highest for malignant lesions (P<0.001), mirroring the pattern observed in senior radiologists. Benign lesions scored significantly lower than borderline malignant (P=0.001, median difference –47.500, 95% CI: –75.000 to –17.500) and malignant lesions (P<0.001, median difference –52.500, 95% CI: –75.000 to –25.000); borderline malignant and malignant lesions did not differ (P=1.000, median difference 0.000, 95% CI: –15.000 to 7.500). This trend however differed among junior radiologists where borderline malignant lesions received higher scores than malignant lesions. Within the benign group, SCN scored lowest, particularly for LLMs and junior radiologists. On the other hand, SCNs represented the lowest-scoring pathological subtype for LLMs and junior radiologists. Among borderline malignant lesions, SPT consistently achieved the highest scores across all reader groups. IPMN showed intermediate scores ranging between 50–67.5 points. The small sample size (n=2) precludes meaningful comparison for MCN. Malignant lesions generally received high scores, with PDAC reaching up to 100 points in some cases for LLM assessments.

Table 2

Score for each pathological type under IM3

Pathological type LLM Senior radiologists Junior radiologists
Benign, n=21 5.00 (47.50) 42.50 (50.00) 0.00 (32.50)
   SCN, n=14 0.00 (28.75) 37.50 (50.00) 0.00 (25.00)
   Other benign cysts, n=7 42.50 (67.50) 50.00 (50.00) 57.50 (50.00)
Borderline, n=25 85.00 (63.50) 67.50 (63.75) 67.50 (55.00)
   IPMN, n=15 50.00 (82.50) 67.50 (50.00) 67.50 (75.00)
   MCN, n=2 50.00 (NA) 42.50 (NA) 41.25 (NA)
   SPT, n=8 96.25 (20.63) 80.00 (73.75) 71.25 (43.75)
Malignant, n=34 87.50 (50.63) 92.50 (50.00) 51.25 (56.25)
   PDAC, n=14 100.00 (11.88) 100 (18.13) 75.00 (52.50)
   Malignant MCN/SPT, n=6 96.25 (41.88) 96.25 (63.50) 75.00 (81.25)
   Malignant IPMN, n=14 57.50 (71.25) 50.00 (81.25) 25.00 (69.38)
Total, n=80 73.75 (86.25) 67.50 (71.88) 50.00 (50.00)

Values are presented as median (IQR) within each pathological category. , scores represent per-case averages across the four LLMs, four senior radiologists, and four junior radiologists, respectively. , subgroups with small sample sizes should be interpreted with caution. IM3, input modality 3; IPMN, intraductal papillary mucinous neoplasm; IQR, interquartile range; LLM, LLM, large language model; MCN, mucinous cystic neoplasm; NA, not applicable; PDAC, pancreatic ductal adenocarcinoma; SCN, serous cystic neoplasm; SPT, solid pseudopapillary tumor.

Given junior radiologists’ limited comprehension of PCLs, we analyzed cases where the average score of LLMs and senior radiologists fell below 50 points (n=30). Benign lesions predominated (15 cases), with SCN constituting the majority. Nine SCN cases were attributed to Reason 1 (suboptimal CEUS visualization to recognize microcystic architecture, leading to erroneous characterization as solid masses, Figure 2A). Two SCN cases demonstrated Reason 3 (one was also attributed to Reason 1), where LLMs failed to interpret a key imaging feature. Radiologists had documented peripheral microcystic distribution—a characteristic SCN feature—which LLMs lacked prior knowledge to contextualize (Figure 2B). These cases showed stark performance contrasts (senior radiologists had 75 and 67.5 points, while LLMs had 0 and 0 points). Five additional benign cases exhibited Reason 2 due to ambiguous imaging features (simple cysts or cysts with thin septum indistinguishable from MCN).

Figure 2 Representative cases highlighting diagnostic challenges. (A) SCN (white arrows) demonstrating hyper-enhanced pattern like a solid tumor on CEUS, corresponding to honeycomb microcystic architecture. (B) Peripheral microcysts (red arrows) of a SCN (white arrows), recognized by radiologists as characteristic of SCN, were not comprehended by LLMs. (C) Malignant IPMN (white arrows) showing thick septations (red arrow) without solid components on CEUS. (D) Malignant IPMN with mucin lake formation (white arrow) on grayscale ultrasound. CEUS confirmed non-enhancing content, but LLMs failed to correlate this finding with adjacent malignancy. CEUS, contrast-enhanced ultrasound; IPMN, intraductal papillary mucinous neoplasms; LLMs, large language models; SCN, serous cystic neoplasms.

Among borderline malignant lesions, two IPMN showed Reason 1 (undocumented ductal communication). Three cases reflected Reason 2 (solid components complicating differentiation from malignant transformation). One SPT case revealed discordance (25, 67.5, and 100, for senior, LLMs, and junior, respectively), potentially representing stochastic variation.

Among the malignant lesions, seven were malignant IPMNs. Three of these were attributed to both Reason 1 (poorly visualized communication to the pancreatic duct) and 2 (poorly visualized solid components, Figure 2C). One case was attributed to Reason 1 alone, one to Reason 2 alone, and two to Reason 3 (LLMs misunderstanding the descriptions of mucin lake, Figure 2D). One PDAC was assigned to Reason 2 (few solid components) and one malignant MCN was misclassified due to unclear imaging (solid components mischaracterized as dense septations) (Table 3).

Table 3

Classifications of lower scores for LLMs and senior radiologists under IM3

Pathological type Cases Reason 1 Reason 2 Reason 3 Reason 4
Benign 15 9 5 2 0
   SCN 11 9 1 2 0
   Other benign cysts 4 0 4 0 0
Borderline 6 2 3 0 1
   IPMN 5 2 3 0 0
   MCN 0 0 0 0 0
   SPT 1 0 0 0 1
Malignant 9 5 5 2 0
   PDAC 1 0 1 0 0
   Malignant MCN/SPT 1 1 0 0 0
   Malignant IPMN 7§ 4 4 2 0
Total 30 16 14 4 1

Reason 1: Suboptimal image visualization; Reason 2: Atypical imaging presentation; Reason 3: Failure in LLM interpretation; Reason 4: Failure in human interpretation. , some cases may have more than one reason. , one case was assigned both Reason 1 and 3. §, three cases were assigned both Reason 1 and 2. IM3, input modality 3; IPMN, intraductal papillary mucinous neoplasm; LLM, large language model; MCN, mucinous cystic neoplasm; PDAC, pancreatic ductal adenocarcinoma; SCN, serous cystic neoplasm; SPT, solid pseudopapillary tumor.

We ultimately evaluated the supplementary value of LLMs for junior radiologists’ diagnoses. Given the limited diagnostic utility of LLMs under grayscale-only reports (IM1), analyses focused on IM2 and IM3. LLM-assisted interpretation significantly enhanced junior radiologists’ scores compared to their initial performance (P=0.003, median difference 7.500, 95% CI: 0.000 to 12.500 for IM2 and P=0.001, median difference 12.500, 95% CI: 0.000 to 12.500 for IM3). Following LLM guidance, junior radiologists’ diagnostic accuracy approached that of senior radiologists, with no statistically significant differences observed (P=0.489, median difference 0.000, 95% CI: −8.750 to 3.750 for IM2 and P=0.338, median difference −1.250, 95% CI: −10.000 to 0.000 for IM3) (Figure 3).

Figure 3 Supplementary value of LLM for junior radiologists. Scores represent the mean score per case (out of 100). Mean scores of 4 senior radiologists and 4 LLMs are shown in dotted lines under IM2 (A) and IM3 (B), respectively. **, P<0.01. IM, input modality; LLM, large language model.

Discussion

Artificial intelligence has shown promise in assisting pancreatic imaging interpretation, such as quantifying necrosis volume on CT (26). However, the narrative and operator-dependent nature of CEUS reports poses a distinct challenge that requires linguistic comprehension beyond image analysis. This study therefore supports the feasibility of LLMs in interpreting CEUS reports for PCLs. Diagnostically, LLMs showed no statistically significant difference from senior radiologists when provided with CEUS features, demonstrating consistent performance across four distinct models. Besides, our error analysis revealed critical limitations in both CEUS visualization and LLM comprehension. Crucially, LLM-assisted interpretation raised junior radiologists’ diagnostic performance such that it no longer differed significantly from that of senior radiologists, suggesting potential utility for clinician training and decision support. However, larger studies are still needed to confirm these findings and to rule out clinically meaningful differences that may not have been detected owing to limited statistical power.

It is noteworthy that the performance of LLMs under CEUS input modalities underscores the richness of diagnostic information contained in CEUS reports. This capability is paramount for identifying key discriminators such as the presence and enhancement pattern of mural nodules, differentiating true enhancing solid components from avascular debris or mucin plugs, and characterizing septal vascularity—features often ambiguous on conventional grayscale ultrasound or even CT/MRI at times. Our findings, where the addition of CEUS features consistently and significantly boosted LLM diagnostic scores across most lesion types, provide evidence reinforcing the high diagnostic utility of CEUS as endorsed by existing guidelines (8,9).

Several aspects of our prompt design are worth noting. First, listing specific pathological entities may have limited the models’ outputs to the listed categories. This was considered acceptable because it reflects the clinical differential diagnosis framework. Second, output flexibility was preserved, as forcing categorical classifications proved unnecessary. LLMs consistently gave specific diagnoses before assigning categories even when instructed otherwise. Third, the weighted scoring system prioritized category-level agreement over exact pathological matching, consistent with how clinicians use risk stratification to guide management. While not designed to assess fine-grained diagnostic accuracy, this approach ensures consistency across different assessors.

Regarding linguistic considerations in input/output design, we note that most prior studies utilized English inputs, whereas our study employed original Chinese reports. Existing literature comparing LLM performance across languages (e.g., Chinese, Arabic, Korean) reports minor—though statistically significant—accuracy variations attributable to linguistic differences (27-29). We preserved Chinese reports without translation both to prevent information distortion during translation and to maintain clinical applicability in Chinese-speaking healthcare settings. Fundamentally, language serves as a medium for information transfer where logical relationships between concepts remain invariant across linguistic systems. To mitigate potential language bias, we incorporated four LLMs with diverse linguistic optimizations, including GPT-4o, Claude 3.7 Sonnet, and Gemini 2.0—primarily English-optimized yet multilingual—alongside DeepSeek-R1, which may feature Chinese-specific semantic enhancements. Although some studies report statistically significant inter-model diagnostic variations for certain diseases, the absolute performance differences are typically marginal (30,31). In our study, the absence of statistically significant differences among all four models (P=0.962 for IM3) suggests that language may not be a major source of variability in this specific task. However, given the modest sample size and the limited number of models evaluated, these findings should be regarded as preliminary; further multilingual validation is needed.

Under IM3 (demographic-augmented input), obvious diagnostic score variations persisted across pathological subtypes. SCN, which rarely undergo malignant transformation (32), exhibited the lowest scores, mainly because of Reason 1. Specifically, this reflects suboptimal CEUS visualization of characteristic microcystic architecture, aligning with prior reports on CEUS’s inherent challenges in SCN characterization (33). Two SCN cases belong to Reason 3, where LLMs failed to interpret the diagnostic significance of peripherally distributed microcysts—a descriptor rooted in single-center expertise rather than published consensus. These cases illustrate the need for explicit feature explanations during LLM fine-tuning to address such institution-specific lexicons. This is clinically relevant because SCN is a common benign lesion that generally requires only surveillance; poor diagnostic performance in this subgroup could lead to unnecessary anxiety and intervention, thereby limiting the clinical utility of LLM-assisted interpretation for PCLs. IPMN showed intermediate performance, though junior radiologists marginally outperformed seniors and LLMs, likely due to juniors overemphasizing ductal communication described in reports. SPT achieved the highest scores. While the characteristic rim enhancement on CEUS contributed to strong performance under IM2, the substantial score increase from IM2 to IM3 suggests that the well-known association of SPT with younger age also strongly influenced the models’ diagnostic outputs when demographic data were provided (34). PDAC demonstrated diagnostic excellence among seniors and LLMs, owing to hypo-enhancement patterns and vascular invasions that generated highly concordant textual descriptors (35,36). However, the occasional false-negative classification of PDAC as benign or borderline may lead to underestimation of disease severity and should be guarded against in clinical application.

This study has some limitations. First, it was a retrospective single-center study with a modest sample size; several pathological subgroups were small, and rare PCL subtypes were not represented, limiting generalizability. Second, the analysis was limited to report text and basic demographics—clinical information, laboratory biomarkers, and ultrasound images were not provided, although these contribute to multimodal diagnosis in practice. Third, the weighted scoring system was not externally validated, the error classification relied on consensus review rather than independent blinded assessment. Fourth, findings may be specific to the model versions, languages, and semi-structured report format used, and the semi-structured nature of the reports means it is possible that LLMs partly relied on statistical term–diagnosis associations captured in repeated descriptors. Averaging scores across readers may have suppressed within-group heterogeneity. Some reports may have been interpreted by their original authors, potentially introducing recall bias, though this was mitigated by the time interval between reporting and study interpretation. A test-retest learning effect may exist for junior radiologists in addition to the contribution of LLM guidance.


Conclusions

In conclusion, this study demonstrates that LLM shows promising diagnostic capability in interpreting CEUS reports for PCLs. No statistically significant difference was observed between LLMs and senior radiologists under CEUS input conditions, underscoring the informational value of CEUS reports. Crucially, LLM-assisted interpretation raised junior radiologists’ performance such that it no longer differed significantly from that of senior radiologists, highlighting their potential to help narrow expertise gaps in medical imaging.


Acknowledgments

None.


Footnote

Reporting Checklist: The authors have completed the STARD reporting checklist. Available at https://qims.amegroups.com/article/view/10.21037/qims-2026-1-0307/rc

Data Sharing Statement: Available at https://qims.amegroups.com/article/view/10.21037/qims-2026-1-0307/dss

Funding: This work was supported by Noncommunicable Chronic Diseases-National Science and Technology Major Project (Nos. 2025ZD0552400 and 2025ZD0552407, to K.L.), the Chinese Academy of Medical Sciences Innovation Fund for Medical Sciences (CIFMS) (No. 2023-I2M-C&T-A-005, to K.L.) and the Beijing Natural Science Foundation (No. 4262001, to K.L.).

Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://qims.amegroups.com/article/view/10.21037/qims-2026-1-0307/coif). The authors have no conflicts of interest to declare.

Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. This study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. This study was approved by the Ethics Committee of Peking Union Medical College Hospital (Approval No. I-23PJ898). Informed consent was taken from all individual participants.

Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.


References

  1. Gupta A, Chennatt JJ, Mandal C, Gupta J, Krishnasamy S, Bose B, Solanki P. H S, Singh SK, Gupta S. Approach to Cystic Lesions of the Pancreas: Review of Literature. Cureus 2023;15:e36827. [Crossref] [PubMed]
  2. Klöppel G, Kosmahl M. Cystic lesions and neoplasms of the pancreas. The features are becoming clearer. Pancreatology 2001;1:648-55.
  3. Aziz H, Acher AW, Krishna SG, Cloyd JM, Pawlik TM. Comparison of Society Guidelines for the Management and Surveillance of Pancreatic Cysts: A Review. JAMA Surg 2022;157:723-30. [Crossref] [PubMed]
  4. Kruse DE, Paulson EK. The Incidental Pancreatic Cyst: When to Worry About Cancer. Korean J Radiol 2024;25:559-64. [Crossref] [PubMed]
  5. Faccioli N, Santi E, Foti G, D'Onofrio M. Cost-effectiveness analysis of including contrast-enhanced ultrasound in management of pancreatic cystic neoplasms. Radiol Med 2022;127:349-59. [Crossref] [PubMed]
  6. Wang Y, Wang Y, Fan Z, Shan J, Yan K. The Value of Contrast-Enhanced Ultrasound Classification in Diagnosis of Pancreatic Cystic Lesions. Biomed Res Int 2019;2019:5698140. [Crossref] [PubMed]
  7. D'Onofrio M, Megibow AJ, Faccioli N, Malagò R, Capelli P, Falconi M, Mucelli RP. Comparison of contrast-enhanced sonography and MRI in displaying anatomic features of cystic pancreatic masses. AJR Am J Roentgenol 2007;189:1435-42. [Crossref] [PubMed]
  8. Sidhu PS, Cantisani V, Dietrich CF, Gilja OH, Saftoiu A, Bartels E, et al. The EFSUMB Guidelines and Recommendations for the Clinical Practice of Contrast-Enhanced Ultrasound (CEUS) in Non-Hepatic Applications: Update 2017 (Short Version). Ultraschall Med 2018;39:154-80. [Crossref] [PubMed]
  9. D'Onofrio M, de Sio I, Mirk P, Vidili G, Bertolotto M, Cantisani V, Schiavone C. SIUMB recommendations for focal pancreatic lesions. J Ultrasound 2020;23:599-606. [Crossref] [PubMed]
  10. Bhayana R. Chatbots and Large Language Models in Radiology: A Practical Primer for Clinical and Research Applications. Radiology 2024;310:e232756. [Crossref] [PubMed]
  11. Adams LC, Truhn D, Busch F, Kader A, Niehues SM, Makowski MR, Bressem KK. Leveraging GPT-4 for Post Hoc Transformation of Free-text Radiology Reports into Structured Reporting: A Multilingual Feasibility Study. Radiology 2023;307:e230725. [Crossref] [PubMed]
  12. Bhayana R, Nanda B, Dehkharghanian T, Deng Y, Bhambra N, Elias G, Datta D, Kambadakone A, Shwaartz CG, Moulton CA, Henault D, Gallinger S, Krishna S. Large Language Models for Automated Synoptic Reports and Resectability Categorization in Pancreatic Cancer. Radiology 2024;311:e233117. [Crossref] [PubMed]
  13. Li H, Moon JT, Iyer D, Balthazar P, Krupinski EA, Bercu ZL, Newsome JM, Banerjee I, Gichoya JW, Trivedi HM. Decoding radiology reports: Potential application of OpenAI ChatGPT to enhance patient understanding of diagnostic reports. Clin Imaging 2023;101:137-41. [Crossref] [PubMed]
  14. Dong A, Zhang L, Ma J, Li W, Hong G. Evaluation of large language models for TI-RADS category extraction from free-text thyroid US reports: a comparative study with human readers. Endocrine 2026;91:156. [Crossref] [PubMed]
  15. Yang Z, Huang T, Huang L, Yao J, Jiang L, Wu W, Xie X, Xu M, Zhang X. Evaluating Consistency and Accuracy of GPT-4 Omni to Analyze Thyroid Ultrasound Features and ACR TR Categories to Aid Report Generation. Curr Med Imaging 2026;22:e15734056437835. [Crossref] [PubMed]
  16. Abbasian Ardakani A, Mohammadi A, Kuzan TY, Kuzan BN, Khorshidi H, Ghorbani A, Mohebbi A, Faeghi F, Hatamikia S, Acharya UR. Ultrasound-based Detection and Malignancy Prediction of Breast Lesions Eligible for Biopsy: A Multi-center Clinical-scenario Study Using Nomograms, Large Language Models, and Radiologist Evaluation. Acad Radiol 2026;33:2869-86. [Crossref] [PubMed]
  17. Azhar K, Lee BD, Byon SS, Lee S, Cho KR, Song SE. Semiautomated breast ultrasound report generation using multimodal large language models and deep learning. Front Med (Lausanne) 2026;13:1679203. [Crossref] [PubMed]
  18. Lin S, Li Y, Mao R, Zou X, Hu Y, Ye H, Wu X, Yang L, He J, Lu S, Li L, Zhou J. Large Language Models for the Differentiation of Benign and Malignant Liver Nodules based on Multimodal Prompts in Liver US Cases. Ultrasound Med Biol 2026;52:1374-81. [Crossref] [PubMed]
  19. Guo Y, Gong J, Jiang R, Agarwal A, Goel R, Selingreund R, Liu Y, Ren M. Automated O-RADS Risk Stratification Using a Large Language Model Analysis of Narrative Ultrasound Reports. Ultrasound Med Biol 2026;52:1363-73. [Crossref] [PubMed]
  20. Akkus I, Deniz G. Diagnostic Performance of Large Language Models in Musculoskeletal Ultrasound: A Comparative Evaluation of ChatGPT-5.1 and Gemini for Plantar Fasciitis. J Imaging Inform Med 2026; Epub ahead of print. [Crossref]
  21. Tang M, Xu L, Zhu L, Maimaitiabula A, Zhang X, Pei C, Hu L. Large language models for structuring knee ultrasound reports-A comparative evaluation of DeepSeek R1, gemini 2.5 flash, and GPT-4o. BMC Med Inform Decis Mak 2026;26:115. [Crossref] [PubMed]
  22. Huang J, Yang R, Huang X, Zeng K, Liu Y, Luo J, Lyshchik A, Lu Q. Feasibility of large language models for CEUS LI-RADS categorization of small liver nodules in patients at risk for hepatocellular carcinoma. Front Oncol 2024;14:1513608. [Crossref] [PubMed]
  23. Choubey AP, Eguia E, Hollingsworth A, Chatterjee S, D'Angelica MI, Jarnagin WR, Wei AC, Schattner MA, Do RK, Soares KC. Data Extraction and Curation from Radiology Reports for Pancreatic Cyst Surveillance Using Large Language Models. J Am Coll Surg 2025;241:766-72. [Crossref] [PubMed]
  24. Sato M, Yasaka K, Abe S, Kurashima J, Asari Y, Kiryu S, Abe O. Efficacy of a large language model in classifying branch-duct intraductal papillary mucinous neoplasms. Abdom Radiol (NY) 2026;51:417-23. [Crossref] [PubMed]
  25. Sengul I, Sengul D. Exegesis on using a customized GPT to provide guideline-based recommendations for the management of pancreatic cystic lesions. Endosc Int Open 2025;13:a26051215. [Crossref] [PubMed]
  26. Lu CX, Zhou J, Feng YC, Meng SJ, Guo XL, Su WS, Ngo T, Hsu TH, Lin P, Huang J, Liu ST, Palacio MLB, Change WL, Qin G, Hu YQ, Zhan LH. Artificial intelligence models assisting physicians in quantifying pancreatic necrosis in acute pancreatitis. Quant Imaging Med Surg 2025;15:135-48. [Crossref] [PubMed]
  27. Sallam M, Alasfoor IM, Khalid SW, Al-Mulla RI, Al-Farajat A, Mijwil MM, Zahrawi R, Sallam M, Egger J, Al-Adwan AS. Chinese generative AI models (DeepSeek and Qwen) rival ChatGPT-4 in ophthalmology queries with excellent performance in Arabic and English. Narra J 2025;5:e2371. [Crossref] [PubMed]
  28. Zhong W, Liu Y, Liu Y, Yang K, Gao H, Yan H, Hao W, Yan Y, Yin C. Performance of ChatGPT-4o and Four Open-Source Large Language Models in Generating Diagnoses Based on China’s Rare Disease Catalog: Comparative Study. J Med Internet Res 2025;27:e69929. [Crossref] [PubMed]
  29. Oh N, Kim J, Park S, An S, Lee E, Do H, et al. Large Language Model-Assisted Surgical Consent Forms in Non-English Language: Content Analysis and Readability Evaluation. J Med Internet Res 2025;27:e73222. [Crossref] [PubMed]
  30. Wu J, Wang Z, Qin Y. Performance of DeepSeek-R1 and ChatGPT-4o on the Chinese National Medical Licensing Examination: A Comparative Study. J Med Syst 2025;49:74. [Crossref] [PubMed]
  31. Zhou H, Wang Z, Wang R, Jiang L, Zhu C, Guo H, Song T, Yin N. DeepSeek Versus GPT: Evaluation of Large Language Model Chatbots' Responses on Orofacial Clefts. J Craniofac Surg 2025;36:2197-201. [Crossref] [PubMed]
  32. Sakorafas GH, Smyrniotis V, Reid-Lombardo KM, Sarr MG. Primary pancreatic cystic neoplasms revisited. Part I: serous cystic neoplasms. Surg Oncol 2011;20:e84-92.
  33. Hashimoto S, Hirooka Y, Kawabe N, Nakaoka K, Yoshioka K. Role of transabdominal ultrasonography in the diagnosis of pancreatic cystic lesions. J Med Ultrason (2001) 2020;47:389-99. [Crossref] [PubMed]
  34. Fan Z, Li Y, Yan K, Wu W, Yin S, Yang W, Xing B, Li X, Zhang X. Application of contrast-enhanced ultrasound in the diagnosis of solid pancreatic lesions--a comparison of conventional ultrasound and contrast-enhanced CT. Eur J Radiol 2013;82:1385-90. [Crossref] [PubMed]
  35. Jia WY, Gui Y, Chen XQ, Tan L, Zhang J, Xiao MS, Chang XY, Dai MH, Guo JC, Cheng YJ, Wang X, Zhang JH, Zhang XQ, Lv K. Efficacy of color Doppler ultrasound and contrast-enhanced ultrasound in identifying vascular invasion in pancreatic ductal adenocarcinoma. Insights Imaging 2024;15:181. [Crossref] [PubMed]
  36. Jia WY, Gui Y, Chen XQ, Zhang XQ, Zhang JH, Dai MH, Guo JC, Chang XY, Tan L, Bai CM, Cheng YJ, Li JC, Lv K, Jiang YX. Evaluation of the diagnostic performance of the EFSUMB CEUS Pancreatic Applications guidelines (2017 version): a retrospective single-center analysis of 455 solid pancreatic masses. Eur Radiol 2022;32:8485-96. [Crossref] [PubMed]
Cite this article as: Shao Y, Gui Y, Yan X, Chen T, Chen X, Tan L, Zhang J, Liang H, Jia W, Zhao H, Chen B, Jiang Y, Lv K. Large language models for analyzing contrast-enhanced ultrasound reports of pancreatic cystic lesions. Quant Imaging Med Surg 2026;16(9):726. doi: 10.21037/qims-2026-1-0307

Download Citation