Artificial intelligence in assisted localization of fetal ultrasound standard planes: a narrative review
Review Article

Artificial intelligence in assisted localization of fetal ultrasound standard planes: a narrative review

Kezhen Wang1,2,3,4#, Lingling Lei1,2#, Zichao Liu1,2,4,5, Yiyang Huang1,2,5,6, Qiao Wei1,2,4,7, Wen He6,8, Linyuan Jin6, Meng Du1,2,4,6, Zhiyi Chen1,2,4,6

1Key Laboratory of Medical Imaging Precision Theranostics and Radiation Protection, College of Hunan Province, The Changsha Central Hospital, Hengyang Medical School, University of South China, Changsha, China; 2Institute of Medical Imaging, Hengyang Medical School, University of South China, Hengyang, China; 3College of Mechanical Engineering, University of South China, Hengyang, China; 4Institute for Future Sciences, University of South China, Changsha, China; 5The Seventh Affiliated Hospital, Hunan Veterans Administration Hospital, Hengyang Medical School, University of South China, Changsha, China; 6Department of Ultrasound, Department of Medical Imaging, the Affiliated Changsha Central Hospital, Hengyang Medical School, University of South China, Changsha, China; 7School of Computer and Software, University of South China, Hengyang, China; 8Department of Ultrasound, Beijing Tiantan Hospital, Capital Medical University, Beijing, China

Contributions: (I) Conception and design: All authors; (II) Administrative support: M Du, Z Chen; (III) Provision of study materials or patients: M Du, Z Chen; (IV) Collection and assembly of data: K Wang, L Lei, Y Huang, Z Liu; (V) Data analysis and interpretation: K Wang, L Lei, Z Liu, Q Wei; (VI) Manuscript writing: All authors; (VII) Final approval of manuscript: All authors.

#These authors contributed equally to this work and shared the first authorship.

Correspondence to: Linyuan Jin, MB. Department of Ultrasound, Department of Medical Imaging, the Affiliated Changsha Central Hospital, Hengyang Medical School, University of South China, No. 161, South Shaoshan Road, Changsha, China. Email: 371938634@qq.com; Meng Du, PhD; Zhiyi Chen, PhD. Key Laboratory of Medical Imaging Precision Theranostics and Radiation Protection, College of Hunan Province, The Changsha Central Hospital, Hengyang Medical School, University of South China, No. 161, South Shaoshan Road, Changsha, China; Institute of Medical Imaging, Hengyang Medical School, University of South China, Hengyang, China; Institute for Future Sciences, University of South China, Changsha, China; Department of Ultrasound, Department of Medical Imaging, the Affiliated Changsha Central Hospital, Hengyang Medical School, University of South China, Changsha, China. Email: dumeng_work@126.com; zhiyi_chen@usc.edu.cn.

Background and Objective: Ultrasound imaging has become a vital clinical diagnostic tool due to its cost-effectiveness and absence of ionizing radiation. In fetal ultrasound, the precise localization of standard ultrasound planes is a prerequisite for subsequent quantitative analysis and diagnosis. However, manual identification is highly operator-dependent and subjective, and the multi-planar data acquired by three-dimensional (3D) fetal ultrasound present challenges for efficient localization within large datasets. The objective of this narrative review is to systematically summarize the research progress and methodologies of artificial intelligence (AI) in the automated and precise localization of fetal ultrasound standard planes, while discussing the current challenges and future directions for its clinical application.

Methods: We conducted structured literature search in the PubMed and Web of Science databases to identify studies applying AI to ultrasound imaging, with a primary focus on fetal standard plane recognition. The search focused on all types of publications involving AI-based standard plane recognition in ultrasound between 1998 and 2025. Only peer-reviewed articles and review papers published in English were included. Titles and abstracts were screened to determine eligibility.

Key Content and Findings: This review details the transformative impact of AI, particularly deep learning, on automating standard plane localization in ultrasound, with a primary emphasis on fetal applications. Key findings reveal a progression from handcrafted features to advanced architectures like convolutional neural network and reinforcement learning, achieving expert-level accuracy across fetal standard planes, including abdominal, facial, brain, and cardiac views, while representative adult applications (e.g., hepatobiliary and thyroid imaging) are included as supplementary examples. The review identifies technical challenges, including data variability and computational costs, which are being addressed via transfer learning, attention mechanisms, and lightweight network design. The emergence of multi-organ models highlights a trend towards comprehensive AI systems for holistic screening.

Conclusions: AI significantly advances ultrasound by automating standard plane localization, particularly in fetal ultrasound, enhancing diagnostic consistency and workflow efficiency. Integration into clinical practice promises to reduce operator dependency and improve screening accessibility. Persistent challenges include model generalizability, data scarcity, and real-world clinical integration. Future work should prioritize multi-center validation, explainable AI, and end-to-end system development to fully realize this technology’s potential.

Keywords: Artificial intelligence (AI); ultrasound imaging; standard plane localization; prenatal screening


Submitted Oct 31, 2025. Accepted for publication Jun 05, 2026. Published online Aug 05, 2026.

doi: 10.21037/qims-2025-aw-2271


Introduction

Since its initial clinical application in the 1950s, ultrasound imaging has gradually evolved into a vital diagnostic tool in modern medicine (1), valued for its cost-effectiveness, portability, absence of ionizing radiation, and high soft-tissue resolution (2). Among its diverse clinical applications, fetal ultrasound represents one of the most standardized and widely adopted fields, where accurate identification of anatomical planes is essential for routine screening and prenatal diagnosis. In the diagnostic ultrasound workflow, precise positioning of standardized anatomical sections (clinically defined as ultrasound standard planes) is essential for subsequent biometry, quantitative analysis and diagnosis, and provides an anatomical reference frame of reference for subsequent evaluations (3,4). Operators typically use these planes as a foundation for dynamic scanning from multiple angles and sections, thereby achieving a comprehensive diagnosis. Furthermore, distinctive anatomical landmarks within the standard planes hold significant diagnostic value. For instance, in cardiac ultrasonography, the apical four-chamber view enables the assessment of chamber dimensions, ventricular septal integrity, and valvular structures, allowing for the quantitative evaluation of cardiac function. Similarly, in obstetric ultrasound, the biparietal diameter plane is essential for accurate measurement of the fetal head, aiding in the assessment of fetal development (5-7). Therefore, the accurate positioning of the standard plane not only affects the accuracy and reproducibility of quantitative analysis, but also ultimately determines the reliability and clinical reference value of the diagnostic results.

In clinical practice, the identification of two-dimensional (2D) ultrasound standard planes is performed manually, which is highly dependent on the operator’s anatomical knowledge and experience, making it subjective, as even small probe movements can cause significant deviations (8). Unlike static 2D images, 2D ultrasound videos provide more contextual information for interpretation. In contrast, three-dimensional (3D) ultrasound with matrix array probes requires the physician to maintain stable probe contact while an electronic system automates scanning in sector or pyramidal patterns (9). This captures a volumetric data block encompassing the entire anatomical region in a short time, containing multiple standard planes (10). This method reduces operator dependency, improves scanning efficiency, and offers richer spatial and diagnostic data not available in 2D imaging (11). However, both 2D videos and 3D datasets pose challenges for standard plane localization due to vast search spaces and anatomical variability, making manual identification within large image sets or videos difficult (10). These technical challenges exacerbate the growing global shortage of trained sonographers, a crisis that is particularly acute in rural and remote regions where access to expert diagnostic services is severely limited. Additionally, the prolonged and repetitive physical strain required to manually locate optimal acoustic windows constitutes a significant occupational hazard, leading to work-related musculoskeletal disorders that further constrain the active workforce.

Artificial intelligence (AI), as a transformative force in modern healthcare, has demonstrated remarkable potential across various medical imaging domains (12,13), including intracranial aneurysms detection in computed tomography (CT) (14), and image quality and visualization of metabolic activity enhancement in nuclear medicine (15,16). Within ultrasound imaging, a particularly promising application is the automatic localization of standard planes, which is expected to streamline clinical workflows and reduce dependence on sonographers (Figure 1). Crucially, AI-guided systems address the workforce gap by enabling less experienced operators to acquire diagnostic-quality images, thereby extending critical imaging services to underserved areas. Furthermore, by automating the recognition of standard planes, these systems significantly reduce the time and physical effort required for scanning, thereby mitigating the risk of scanning-related injuries and preserving the health of sonographers. The evolution of AI, from traditional machine learning to deep learning models, and further to methodologies such as supervised and reinforcement learning (RL), has enabled remarkable advancements in ultrasound standard plane localization (17-19). These AI-based techniques now achieve accuracy that matches or surpasses human experts, while significantly exceeding the speed of conventional methods. The integration of AI into ultrasound acquisition and analysis enhances standardization in prenatal, hepatobiliary, echocardiography, and thyroid evaluations, thereby boosting operational efficiency and reducing diagnostic errors.

Figure 1 Challenges of manual ultrasound standard plane localization and advantages of AI-assisted localization. Challenges in manual ultrasound standard plane localization and the advantages of AI-assisted localization. Manual localization in 2D and 3D ultrasound is constrained by operator dependency, subjective assessment, image misalignment, anatomical variability, and labor-intensive workflows. AI-based methods improve standardization, efficiency, and diagnostic accuracy. AI, artificial intelligence; 2D, two-dimensional; 3D, three-dimensional.

Persistent challenges in traditional ultrasound underscore the need for such innovation. Accordingly, this review primarily focuses on recent advances in AI-based localization of fetal ultrasound standard planes, while also summarizing representative adult applications as supplementary references. On the basis, this review examines the application of AI in ultrasound standard plane localization, focusing on recent advancements, persisting challenges, and future directions. Given the heterogeneous characteristics of datasets across anatomical categories, the discussion is structured according to dataset classification. We present this article in accordance with the Narrative Review reporting checklist (available at https://qims.amegroups.com/article/view/10.21037/qims-2025-aw-2271/rc).


Methods

We conducted a structured literature search in the PubMed and Web of Science databases to identify studies applying AI to ultrasound imaging (Table 1). The search covered publications from 1998 to 2025 and focused on peer-reviewed articles published in English. The retrieved literature was screened for relevance to the scope of this review, with emphasis on studies that informed the narrative synthesis of recent advances, major challenges, and future directions in the field. To improve literature coverage, we expanded the search strategy by incorporating broader technical and clinical keywords, including specific algorithm-related terms [e.g., convolutional neural network (CNN), transformer, RL] and organ-specific terms (e.g., brain, cardiac, heart). The specific PubMed search terms are presented in Table 2.

Table 1

The search strategy summary

Items Specification
Date of search 28 July 2025
Databases and searched PubMed, Web of Science
Search terms used Artificial intelligence-related terms, ultrasound-related terms, standard plane/view/section terms, algorithm-specific terms, and organ-specific terms
Timeframe 1998 to 2025
Eligibility considerations Peer-reviewed English-language articles relevant to artificial intelligence-based ultrasound standard plane localization were considered
Literature screening approach The retrieved literature was reviewed by the authors for relevance to the scope and narrative synthesis of this review

Table 2

The search terms used in PubMed

Search terms
   (“Artificial Intelligence” or “AI” or “Deep Learning” or “Machine Learning”) AND “Medical Imaging” AND (“Ultrasonic Standard Plane” or “Ultrasonic Standard View” or “Ultrasonic Standard Section”) AND “Image Recognition”
   (“Artificial Intelligence” or “AI” or “Deep Learning” or “Machine Learning”) AND “Medical Imaging” AND “Image Recognition”
   (“Artificial Intelligence” or “AI” or “Deep Learning” or “Machine Learning”) AND (“Ultrasonic Standard Plane” or “Ultrasonic Standard View” or “Ultrasonic Standard Section”) AND “Image Recognition”
   (“Artificial Intelligence” or “AI” or “Deep Learning” or “Machine Learning”) AND “Medical Imaging”
   (“Artificial Intelligence” or “AI” or “Deep Learning” or “Machine Learning”) AND “Ultrasound Imaging”
   (“Artificial Intelligence” or “AI” or “Deep Learning” or “Machine Learning”) AND (“Ultrasonic Standard Plane” or “Ultrasonic Standard View” or “Ultrasonic Standard Section”)
   (“Artificial Intelligence” OR AI OR “Deep Learning” OR “Machine Learning” OR CNN OR Transformer OR “Reinforcement Learning”)
   (“Brain OR Cardiac OR Heart OR Abdominal OR Facial”)
   (“Fetal OR Prenatal”)

Fundamentals of AI in standard plane localization

AI was initially applied to standard plane localization in 2D ultrasound videos using traditional machine learning algorithms like AdaBoost and random forests (20). These methods localized planes by detecting predefined morphological features, such as the stomach bubble (SB), spine (SP) curvature, and umbilical vein (UV) course (21). However, this approach heavily relied on expert-designed features. With the rise of deep learning, CNN became popular for localizing standard planes. By learning from large annotated datasets, CNN can automatically extract relevant features, improving localization accuracy (22). Unlike traditional methods, CNN can achieve a deeper semantic understanding. While CNNs are highly efficient for local feature extraction, their restricted receptive fields often limit their ability to capture the overarching global anatomical context—a critical requirement for interpreting complex ultrasound views where spatial relationships between distant structures are key. To address this, Transformer-based architectures have recently been introduced. Unlike CNNs, Transformers leverage self-attention mechanisms to model long-range global dependencies, significantly improving the contextual understanding of spatially distant anatomical landmarks. However, Transformers typically require larger datasets to overcome their lack of inductive bias and entail higher computational overhead, which can challenge real-time deployment on mobile ultrasound devices. Further progress introduced recurrent neural networks (RNNs), which analyze temporal context in ultrasound videos, leveraging information from previous frames to better mimic the sonographer’s cognitive process (23).

While 2D ultrasound allows frame-by-frame analysis, 3D ultrasound involves large volumetric data, making sequential frame processing inefficient (10). However, 3D ultrasound provides critical spatial information, such as the coronal plane of the uterus, essential for diagnosing uterine anomalies and intrauterine devices (11). Therefore, automating standard plane localization in 3D ultrasound is crucial. Researchers have developed new approaches for 3D ultrasound, including supervised learning and RL (24,25). Supervised learning enables models to map 3D data to target planes, but requires large annotated datasets and results in complex, nonlinear tasks. In contrast, RL mimics the sonographer’s process by allowing an agent to learn optimal search policies through trial and error, simplifying learning (26). RL offers a unique advantage by shifting from static image classification to a dynamic, sequential decision-making paradigm, which is particularly effective for navigating through 3D volumes. Yet, the technical application of RL remains bottlenecked by training instability, low sample efficiency, and the significant challenge of defining robust reward functions that can generalize across the high anatomical variability of different patients. Early RL methods localized only one plane at a time, but multi-agent reinforcement learning (MARL) frameworks have since been developed to localize multiple planes simultaneously, ensuring accuracy and spatial coherence (27) (Figure 2). AI-assisted standard plane localization has proven valuable in clinical applications such as prenatal screening, liver lesion differentiation, and thyroid nodule risk stratification.

Figure 2 Schematic diagram of AI-assisted ultrasound standard plane localization. Schematic illustration of AI-assisted ultrasound standard plane localization. In 2D ultrasound, standard planes are identified from videos or individual frames using traditional machine learning and deep learning methods. In 3D ultrasound, supervised learning and reinforcement learning-based methods enable navigation within volumetric data for more efficient and accurate target plane localization. AI, artificial intelligence; CNN, convolutional neural network; 2D, two-dimensional; 3D, three-dimensional; RNN, recurrent neural network.

Application of AI in fetal ultrasound standard planes localization

Suboptimal growth assessment is a global health issue, with millions of infants born with low birth weight due to intrauterine growth restriction (IUGR) or preterm birth, leading to significant morbidity and mortality. Ultrasound is the gold standard for prenatal screening, enabling accurate biometry, growth assessment, and early detection of structural anomalies by capturing standardized fetal planes, including the head, abdomen, face, brain, and heart. These planes clearly show key anatomical features, such as cranial structures, spine, and cardiac elements, essential for biometric measurements, anomaly detection, and growth evaluation. Reliable identification of standard ultrasound planes is critical for accurate fetal biometry and early intervention. However, studies show considerable variability in prenatal anomaly detection sensitivity across institutions (27.5% to 96%), with inconsistent plane acquisition contributing to this disparity. AI has the potential to enhance sensitivity by enabling precise localization of standard planes.

Fetal abdominal ultrasound standard plane localization

Accurate localization of the fetal abdominal standard plane (FASP) is crucial for assessing fetal growth and diagnosing conditions like IUGR. This plane includes the SB, UV, and SP, and its precise acquisition is vital for biometric measurements like abdominal circumference (28). However, factors such as acoustic artifacts, image distortion, fetal position, and probe orientation often cause intra-class variation, as structures like the abdominal aorta and gallbladder can resemble the key features (29).

In recent years, AI technologies have offered promising solutions to address these challenges (Table 3). As early as 2011, researchers began applying AI to automate the selection of standard planes from 3D ultrasound volumes (30). Using a combination of the 1CTC joint classifier and the 2STC independent classifier—both based on Haar wavelet features and the AdaBoost algorithm—they extracted region features of the SB and UV and applied a scoring system to identify the optimal plane. The 2STC model achieved a plane selection accuracy of 91.29% on a test set of 30 volumes, significantly reducing the manual effort required by clinicians, particularly in cases with suboptimal fetal positioning. In 2013, Ni et al. proposed an approach that mimicked clinical reasoning by modeling spatial relationships among anatomical structures in ultrasound images (28). Their selective search and sequential detection method employed a two-stage framework: an AdaBoost classifier first identified the region of interest containing the fetal abdomen to narrow the search area, and a vascular probability map was used to enhance the vascular features of the UV and SP. Combined with grayscale segmentation of the SB, a geometric constraint model was constructed, assuming the SB is positioned to the left of the UV and the SP is located approximately 1–2 cm anterior to the UV. Experiments on 100 fetal abdominal videos demonstrated an 81% FASP localization accuracy—significantly outperforming the traditional AdaBoost method (31%)—validating the importance of spatial relationships in excluding anatomically similar distractors.

Table 3

Summary of selected representative studies on AI-assisted localization of fetal abdominal ultrasound standard planes

Author Ultrasound data Ultrasound system Year AI model Results
Ni D et al. (28) 1,664 images and 100 videos Siemens 2013 AdaBoost Acc: NR, Pre: NR, Rec: NR, F1: NR, mAP: NR, accuracy over conventional AdaBoost methods
Chen Het al. (29) 11,942 images and 219 videos Siemens 2015 CNN Acc: 90.4%, Pre: 90.8%, Rec: 99.5%, F1: 95.0%, mAP: NR
Rahmatullah B et al. (30) 45 scan volumes Philips HD9 2011 AdaBoost Acc: NR, Pre: 76.29%, Rec: 91.29%, F1: NR, mAP: NR
Ahmed M et al. (31) 1,000 images and 230 3D volumes 2016 AdaBoost Acc: 92.5%, Pre: NR, Rec: NR, F1: NR, mAP: NR
Cai Y et al. (32) 33 videos containing 1,616 frames Philips 2018 CNN Acc: NR, Pre: 96.5%, Rec: 99.0%, F1: 97.8%, mAP: NR
Cai Y et al. (33) 33 videos containing 1,616 frames Philips 2018 CNN + GAN Acc: NR, Pre: 96.8%, Rec: 96.2%, F1: 96.5%, mAP: NR

Recall: a metric quantifying the proportion of actual positives correctly identified by a model. F1-score: the harmonic mean of precision and recall that provides a balanced measure of a model’s performance. Acc, accuracy; AdaBoost, adaptive boosting; AI, artificial intelligence; CNN, convolutional neural network; 3D, three-dimensional; GAN, generative adversarial network; mAP, mean average precision; NR, not reported; Pre, precision; Rec, recall.

With advances in deep learning, Chen et al. introduced a domain-transferred deep CNN, which leverages statistical similarities in edge and texture between natural and medical images. They first pre-trained a base network on ImageNet to extract general visual features, then transferred the convolutional layers to the FASP localization task and fine-tuned the fully connected layers to adapt to the ultrasound domain (29). This method achieved an image-level localization accuracy of 89.6% and a video-level F1 score of 95.0% across 11,942 training and 8,718 testing images. The deep features from the nF7 layer effectively distinguished FASP from non-FASP regions, exhibiting strong robustness to noise and variability.

Researchers further refined algorithms by incorporating clinicians’ visual search strategies. Eye-tracking analysis of ten sonographers revealed that 32% of fixations focused on the spine, while 13% and 11% fell within the bounding boxes of the SB and UV, respectively. Moreover, browsing speed decreased by 40% when abdominal structures became visible (31). Inspired by this behavior, a standardized plane detection framework was designed, consisting of four stages: abdominal region detection, SB and UV candidate identification, SP candidate detection, and graphical structural modeling. This approach achieved 92.5% plane detection accuracy on 80 3D volumes, with an average processing time of 10.1 seconds per volume. Building on this, subsequent studies transformed eye-tracking data into visual heatmaps and fused them with ultrasound images via a pre-trained SonoNet-16 branch to construct the SonoEyeNet model family. Among them, the SEN-Late FT model performed best, attaining 96.5% precision, 99.0% recall, and a 97.8% F1 score (32). Expanding on this, Cai et al. proposed a multi-task generative adversarial network framework, multi-task SonoEyeNet (M-SEN), which generates physician-like attention maps to assist in detecting fetal standard planes. This approach overcomes the dependency on eye-tracking inputs by inferring visual focus from image data alone. The M-SEN BCE + GAN model achieved 96.8% precision, 96.2% recall, and a 96.5% F1 score—outperforming conventional image-only models (33). These deep learning approaches, which emulate clinical reasoning and integrate diagnostic heuristics, bring AI-assisted ultrasound plane localization closer to real-world clinical workflows.

Fetal facial ultrasound standard plane (FFUSP) localization

Accurate localization of FFUSP is critical for subsequent biometric assessments and the early detection of congenital anomalies (34). In early pregnancy, particularly at 11–12 weeks of gestation, increased nuchal translucency (NT) has been closely linked to chromosomal abnormalities (34). Moreover, standard facial planes allow clinicians to observe the contour and structural integrity of the fetal face, including indicators such as nasal bone hypoplasia or absence, frontal bossing, interruptions in the echogenicity of the hard palate, and discontinuities along the lip contour. These features serve as key diagnostic markers for craniofacial malformations and genetic syndromes. Early identification of such anomalies can provide essential guidance for personalized prenatal care strategies, benefiting both expectant mothers and healthcare providers. Cleft lip and palate, as the most common facial anomalies in fetuses, can be detected early through FFUSP, enabling timely clinical consultation, psychological preparation, and reducing emotional burden on families. Standard FFUSP typically includes the mid-sagittal plane (MSP), nasal-lip coronal plane (NCP), and orbital axial plane, all of which offer critical anatomical information for assessing facial development and identifying abnormalities (35).

Initial applications of AI in FFUSP identification were grounded in traditional image processing and handcrafted feature extraction techniques. Feature selection was considered pivotal, with single descriptors such as Haar features, histogram of oriented gradients (HOG), and scale-invariant feature transform (SIFT), as well as multi-dimensional features combining motion, intensity, shape, and edge cues, being widely employed. These features were typically fed into classifiers like support vector machines (SVM) for FFUSP detection. For instance, a 2015 study introduced an approach based on dense RootSIFT features encoded via Fisher vectors (FV), with spatial information captured through a multi-layer Fisher network and classified using an SVM trained with the stochastic dual coordinate ascent algorithm (35). This method achieved a mean precision of 99.19% and an average accuracy of 93.27%, outperforming prior handcrafted methods (34). Similarly, the LH-SVM method integrated local binary patterns (LBP) and HOG features to leverage their complementary strengths in texture and edge detection. Tested on 943 ultrasound samples, it achieved 94.67% classification accuracy and significantly improved recall—by 12%—in detecting the complex NCP view (35).

In 2016, a deep convolutional neural network (DCNN) was introduced for FFUSP localization, utilizing compact 3×3 convolution filters to enhance feature learning while reducing computational cost. Compared to handcrafted features like HOG and LBP, DCNN improved localization accuracy on sagittal plane images from 85% to 92%, demonstrating enhanced robustness against speckle noise and fetal positional variation (36,37). However, conventional DCNN often suffered from redundant computations, limiting real-time applicability. Xue et al.’s YOLO framework addressed this by reframing object detection as a regression task, enabling end-to-end optimization without proposal generation. Building on YOLO, lightweight architectures were incorporated into FFUSP tasks. Replacing YOLOv4’s backbone with GhostNet reduced the model size by 85% (to 42.7 MB), while preserving high detection accuracy [mean average precision (mAP) of 98.06%] and achieving real-time inference speeds of 0.07 seconds per frame. This approach autonomously localized standard planes with key anatomical landmarks such as the eye axis and nasolabial groove, overcoming the limitations of conventional single-plane classification (38).

To meet the specific clinical needs of early NT screening, Xue et al. proposed a GhostNet-enhanced YOLOv4 model incorporating attention mechanisms. The system achieved 97.20% and 99.07% accuracy in detecting the MSP and retronasal triangle views, respectively (39). Notably, this study introduced a “clinical control scheme” that evaluated the completeness of anatomical structures—such as the nasal bone and palate line—using automated scoring, thereby extending AI’s role from detection to clinical decision support. For example, when the nasal bone appeared underdeveloped or unclear, the system flagged the image as suboptimal and suggested probe adjustment. More recently, real-time YOLOv5s has been employed for multi-slice FFUSP structure detection and classification, garnering localization from ultrasound experts for its capacity to identify standard anatomical landmarks across varying views (40). While further validation and refinement are necessary, these advancements illustrate the promising clinical potential of AI-driven FFUSP localization systems (Table 4).

Table 4

Summary of selected representative studies on AI-assisted localization of fetal facial ultrasound standard planes

Author Ultrasound data Ultrasound system Year AI model Results
Wang X et al. (35) 1,293 images Philips, General Electric 2021 SVM Acc: 94.67%, Pre: NR, Rec: 93.88%, F1: 94.08%, mAP: 94.27%
Zhen Y et al. (36) 1,870 images Siemens 2016 CNN Acc: 96.99%, Pre: NR, Rec: NR, F1: NR, mAP: NR
Yu Z et al. (37) 7,267 images Siemens 2018 CNN Acc: 96.53%, Pre: 96.98%, Rec: 97. 00%, F1: 96.99%, mAP: NR
Xue H et al. (38) 3,462 images Philips 2022 CNN Acc: NR, Pre: NR, Rec: NR, F1: NR, mAP: 98.06%
Xue H et al. (39) 1,635 images Philips, General Electric 2023 CNN Acc: 94.16%, Pre: NR, Rec: NR, F1: NR, mAP: NR
Liu Z et al. (40) 3,190 images Philips, General Electric 2024 CNN Acc: NR, Pre: 98.1%, Rec: 98.2%, F1: NR, mAP: 98.1%

Acc, accuracy; AI, artificial intelligence; CNN, convolutional neural network; mAP, mean average precision; NR, not reported; Pre, precision; Rec, recall; SVM, support vector machine.

Fetal brain ultrasound standard plane localization

Fetal brain ultrasound is a cornerstone of prenatal diagnosis, offering essential insights into fetal development and enabling the detection of intracranial anomalies such as hydrocephalus, microcephaly, and schizencephaly (41-43). Accurate prenatal diagnosis of fetal brain abnormalities is critical for clinical decision-making, especially when pregnancy termination is under consideration, as diagnostic precision directly affects subsequent management. For example, a study in 2000 reported only 43% concordance between prenatal ultrasound and autopsy findings for Dandy-Walker malformation—a rare congenital neurological defect—highlighting the limitations of traditional approaches (41). With the advancement of AI, its application in recognizing standard planes in fetal brain ultrasound has transitioned from early image processing algorithms to deep learning-based automated systems, significantly improving both diagnostic accuracy and efficiency (Table 5).

Table 5

Summary of selected representative studies on AI-assisted localization of fetal brain ultrasound standard planes

Author Ultrasound data Ultrasound system Year AI model Results
Liu X et al. (44) 257 images N/S (retrospective) 2012 AAM Acc: NR, Pre: NR, Rec: NR, F1: NR, mAP: NR. EER =28%
Qu R et al. (45) Dataset1: 30,000 images; Dataset2: 1,200 images ALOKA 2020 CNN Dataset 1: Acc: 91.0%, Pre: 85.5%, Rec: 90. 1%, F1: 90.0%, mAP: NR; Dataset 2: Acc: 89.1%, Pre: 85.3%, Rec: 86. 4%, F1: 90.1%, mAP: NR
Li Y et al. (46) 72 three-dimensional volumes N/S 2018 ITN Acc: NR, Pre: NR, Rec: NR, F1: NR, mAP: NR. Transventricular plane: average positional error =3.83 mm and rotational error =12.7°; transcerebellar plane: average positional error =3.80 mm and rotational error =12.6°
Li Y et al. (47) 72 three-dimensional volumes N/S 2018 CNN regression model Acc: NR, Pre: NR, Rec: NR, F1: NR, mAP: NR. Positional error =5.85 mm, rotational error =12.9°
Qu R et al. (48) 30,000 ultrasound images ALOKA 2020 CNN Acc: 92.93%, Pre: 92.62%, Rec: 92.39%, F1: 93.53%, mAP: NR
Zhao L et al. (49) 5,709 images Samsung, Siemens, SonoScape 2022 USPD Acc: NR, Pre: NR, Rec: NR, F1: NR, mAP: 85.37%
Lin M et al. (50) 43,890 images and 169 videos General Electric, etc. 2022 PAICS Acc: 96.3%, Pre: NR, Rec: NR, F1: NR, mAP: NR

AAM, active appearance model; Acc, accuracy; AI, artificial intelligence; CNN, convolutional neural network; EER, equal error rate; ITN, iterative transformation network; mAP, mean average precision; NR, not reported; N/S, not stated; PAICS, prenatal ultrasound diagnosis artificial intelligence conduct system; Pre, precision; Rec, recall; USPD, ultrasound standard plane detection.

In 2012, researchers introduced an automated scan plane localization algorithm based on active appearance models. The algorithm identified key anatomical landmarks—such as the “butterfly-shaped” paired thalami and the cavum septum pellucidum—to evaluate scan plane quality. This approach assisted novice clinicians in locating standard planes (44). Nevertheless, due to the high visual similarity in the shape, texture, and grayscale characteristics of fetal brain images, traditional features—such as gray-level co-occurrence matrices and HOG descriptors—struggled to differentiate among subtle anatomical variations. Consequently, these methods were limited in their ability to handle small inter-class differences and large intra-class variations.

The introduction of CNN marked a turning point in fetal brain standard plane (FBSP) localization. In 2017, the first CNN model capable of identifying six standard fetal brain planes—including the thalamic and lateral ventricle axial views—was proposed. This model employed data augmentation and domain adaptation to address limited training data, achieving high localization accuracy across a dataset of 30,000 ultrasound images (45). CNN demonstrated the ability to autonomously learn complex textural and morphological features—such as midline brain structures and lateral ventricles—surpassing the performance of conventional handcrafted methods. The proliferation of 3D ultrasound further enhanced AI capabilities. In 2018, iterative transformation networks were developed to predict spatial plane parameters (position and orientation) in 3D volumes using a joint image-geometric loss function, achieving automatic localization of standard planes such as transventricular and transcerebellar planes with errors limited to 3.8 mm/12.7° and a detection speed of 0.5 seconds per plane (46). Subsequent improvements reduced localization errors further to 3.45 mm and 12.4° (47). Despite the success of CNN in medical imaging, their deployment in clinical practice is often hampered by computational costs and limited generalizability, especially when trained on small ultrasound datasets. Furthermore, CNN pre-trained on natural images are not inherently suited for FBSP analysis. To address the challenges posed by immature fetal brain structures and fine-grained anatomical differences, a Differential-CNN was introduced in 2020. This model enhanced directional pattern extraction in pixel neighborhoods using differential operators, boosting fine-detail localization without increasing computational load. It achieved 92.93% accuracy, particularly excelling in differentiating visually similar non-standard planes, such as the anterior horn coronal view and the MSP (48). Concurrently, AI research began to emphasize multi-task learning and model interpretability. For instance, Zhao et al. proposed an ultrasound standard plane detection (USPD) model for fetal head imaging with KASR and UPC modules. KASR improves recognition of key structures via hybrid knowledge graphs, while UPC classifies planes as standard or not. Together, they boost performance and offer interpretable results. These modules can be seamlessly integrated into broader frameworks, enhancing both accuracy and clinical usability (49).

Parallel advancements have emerged in related segmentation tasks. The Residual U-Net, for example, has demonstrated high performance in fetal head circumference segmentation by leveraging residual and skip connections to preserve detailed spatial information, achieving accurate delineation on a dataset of 1,000 training images. Building on this progress, the prenatal ultrasound diagnostic AI system (PAICS) was developed to support central nervous system anomaly screening in clinical settings. Using data from two tertiary hospitals in China, the system was trained and validated on 43,890 ultrasound images and 169 video clips from 16,297 pregnancies, with data split into training, fine-tuning, and validation sets in an 8:1:1 ratio (50). PAICS is powered by the real-time CNN-based detection algorithm YOLOv3. It achieved expert-level performance in macro- and micro-averaged accuracy, sensitivity, and area under the curve (AUC), while dramatically reducing diagnostic time—processing each image in just 0.025 seconds compared to 4.4 seconds for human experts (P<0.001). This efficiency makes PAICS particularly valuable in primary care settings, where it can effectively address the shortage of experienced practitioners.

Fetal heart ultrasound standard plane localization

According to global case investigations (51), congenital anomalies represent the fourth leading cause of neonatal mortality, affecting an estimated 7.9 million newborns worldwide each year. Among these, congenital heart disease (CHD) is the most prevalent, accounting for 30–40% of all congenital anomalies and impacting approximately 1% of neonates globally (52). Fetal echocardiography is the cornerstone of prenatal CHD screening, with diagnostic accuracy hinging on the precise acquisition of standard views, such as the four-chamber view and the left/right ventricular outflow tracts. However, due to the complexity of fetal cardiac anatomy, variability in fetal positioning, and the susceptibility of ultrasound images to artifacts, the diagnostic sensitivity of the four-chamber view alone varies widely between 43% and 96%. In this context, AI has proven instrumental in enhancing the consistency of structural detection by automatically extracting textural and morphological features—particularly in delineating fine structures such as the endocardial borders and atrioventricular septum (53) (Table 6).

Table 6

Summary of selected representative studies on AI-assisted localization of fetal heart ultrasound standard planes

Author Ultrasound data Ultrasound system Year AI model Results
Dong J et al. (54) 1,991 images N/S 2019 ARVBNet Acc: NR, Pre: NR, Rec: NR, F1: NR, mAP: 93.52%
Li F et al. (55) 3,360 images N/S 2024 FHUSP-NET Acc: 96.4%, Pre: NR, Rec: NR, F1: NR, mAP: NR
He J et al. (56) 29,218 images Siemens, Philips, Samsung, Sonoscape 2024 FCUM Acc: NR, Pre: 98.20%, Rec: NR, F1: NR, mAP: NR
Pu B et al. (57) 115,706 frames (from 643 videos) N/S 2024 HFSCCD (YOLOv5-based hybrid architecture) Acc: NR, Pre: NR, Rec: NR, F1: NR, mAP: NR. Localization error (ES/ED): 2.53/1.30 frames. Standard plane F1-score: 90.92%. Clinical accuracy (4-frame): 98.25%

Acc, accuracy; AI, artificial intelligence; ARVBNet, aggregated residual visual block network; ED, end-diastole; ES, end-systole; FCUM, fetal cardiac ultrasound standard plane detection model; FHUSP-NET, fetal heart ultrasound standard plane localization network; HFSCCD, hybrid neural network for fetal standard cardiac cycle detection in ultrasound videos; mAP, mean average precision; NR, not reported; N/S, not stated; Pre, precision; Rec, recall.

Initial AI research in fetal cardiac ultrasound primarily focused on identifying anatomical structures within the four-chamber view. A notable example is the ARVBNet model (aggregated residual visual block network) proposed in 2019, which emulated the receptive field mechanism of human vision to enable real-time detection of central cardiac structures, such as atria and ventricles, achieving a mAP of 93.52% and a detection speed of 101 frames per second (54). Although effective in enhancing feature representation, this model did not extend to multi-plane localization. As deep learning methodologies matured, researchers began integrating multi-task learning strategies to improve the localization of standard cardiac planes. In 2024, FHUSP-NET (fetal heart ultrasound standard plane network) was introduced, combining localization of five standard planes—including the four-chamber and outflow tract views—with the detection of ten critical anatomical features. Employing spatial pyramid pooling and squeeze-and-excitation networks, the model significantly improved the identification of small structures such as moderator bands and pulmonary veins. It achieved a localization accuracy of 96.4% with an average processing time of only 13.6 milliseconds per image, marking the first demonstration of real-time multi-plane assessment suitable for clinical use (55). In the same year, another model, FCUM (fetal cardiac ultrasound model), incorporated multi-task learning and hybrid attention mechanisms to detect 11 anatomical structures and classify three standard planes. By sequentially integrating channel and spatial attention, the model outperformed baseline architectures like YOLOv8 in classification accuracy, achieving inference speeds of 0.03 seconds per image, although its parameter count and computational complexity remain high (56).

A pivotal advancement in AI-assisted fetal echocardiography is the incorporation of dynamic ultrasound video analysis. The hybrid fetal standard cardiac cycle detection (HFSCCD) network represents a significant step forward, achieving end-to-end detection of standard cardiac cycles from ultrasound video data for the first time (57). The model integrates three key modules: anatomical structure detection, cardiac cycle localization, and standard plane localization. Utilizing a joint probabilistic framework, it accurately identifies end-systolic and end-diastolic frames, while the XGBoost algorithm captures structural correlations to effectively filter out background and non-standard views. The system demonstrated a marked improvement in detecting key frames within clinical video datasets, establishing a foundation for the automated measurement of cardiac functional parameters, such as ejection fraction (EF).

Overall, AI has significantly advanced the automation and accuracy of fetal cardiac ultrasound standard plane localization, evolving from static image analysis to dynamic video interpretation. Core technological enablers include multi-task learning, attention mechanisms, and real-time optimization. However, challenges remain regarding data diversity, model lightweighting, and clinical integration—areas that will likely shape the future trajectory of AI-driven fetal echocardiography.

Multiple fetal organ types ultrasound standard plane localization

While early AI approaches have demonstrated remarkable performance in recognizing standard planes of single fetal organs, their limitations in cross-organ generalizability have become increasingly apparent. Notably, Ultrasound images of different fetal organs, such as the brain, heart, and abdomen, exhibit significant modality heterogeneity. For instance, midline structures in cranial views differ considerably from gastrointestinal outlines in abdominal views, posing challenges for a single model to generalize across various organ types. In recent years, with the standardization of multicenter datasets, the development of universal neural networks capable of multi-organ localization has gained growing attention (Table 7).

Table 7

Summary of selected representative studies on AI-assisted localization of ultrasound standard planes across multiple fetal organ types

Author Target organ Ultrasound data Ultrasound system Year AI model Results
Baumgartner C et al. (58) Brain, abdomen, spine, extremities, heart, etc. 34,587 ultrasound images, 2,638 ultrasound videos General Electric 2017 SonoNet Acc: NR, Pre: NR, Rec: NR, F1: 79.8%, mAP: NR
Pu B et al. (59) Abdomen, brain, lumbosacral spine 1,199 ultrasound videos N/S 2021 FUSPR Acc: 87.38%, Pre: NR, Rec: NR, F1: 88.96%, mAP: NR
Guo J et al. (60) Femur, thalamus, abdomen 14,954 ultrasound images Samsung, Siemens 2023 Coarse-to-Fine Multi-Task Learning Framework Acc: 95.2%, Pre: NR, Rec: NR, F1: 95.2%, mAP: 93.5%
Yu T et al. (61) Abdomen, brain, femur, thorax, maternal cervix, etc. Burgos-Artizzu data set: 12,400 ultrasound images N/S 2024 LPC-SonoNet Six-category classification—Acc: 97.0%, Pre: NR, Rec: NR, F1: NR, mAP: NR. Nine-category classification—Acc: 91.9%, Pre: NR, Rec: NR, F1: NR, mAP: NR
Krishna et al. (62) Abdomen, brain, femur, thorax, maternal cervix, etc. Burgos-Artizzu data set: 12,707 ultrasound images N/S 2024 ACW-CNN Acc: NR, Pre: 95.80%, Rec: 95.06%, F1: 95.43%, mAP: NR
Migliorelli G et al. (63) Brain, femur, thorax, maternal cervix Burgos-Artizzu data set: 10,425 ultrasound images General Electric, ALOKA 2024 SimClr-based Contrastive Learning Model 10% labeling data—Acc: NR, Pre: NR, Rec: NR, F1: 73.06%, mAP: NR. 99% labeling data—Acc: NR, Pre: NR, Rec: NR, F1: 83.08%, mAP: NR

Acc, accuracy; ACW-CNN, adaptive channel weighting convolutional neural network; AI, artificial intelligence; FUSPR, fetal ultrasound standard plane localization; LPC-SonoNet, light pyramid convolution sonography network; mAP, mean average precision; NR, not reported; N/S, not stated; Pre, precision; Rec, recall; SonoNet, sonography network.

A pioneering effort in this direction was SonoNet, introduced in 2017, which enabled real-time detection of 13 fetal standard views encompassing cranial, cardiac, abdominal, and spinal planes. The core innovation of this approach lies in its use of weakly supervised localization. It relies solely on image-level labels to train the model, while employing class activation maps (CAM) to indirectly highlight key anatomical structures (e.g., SB, UV). This approach eliminated the need for labor-intensive pixel-wise annotations. The model achieved a mean F1 score of 0.798 in simulated real-time detection, and a retrospective frame retrieval accuracy of 90.09%, demonstrating the feasibility of multi-organ plane localization (58). Building upon this, the 2021 FUSPR model integrated industrial internet of things (IIoT) technologies to create a distributed ultrasound processing platform. By employing a dual-stream CNN-RNN network, it concurrently learned spatial features (such as thalamic morphology in cranial views) and temporal patterns (such as transducer movement in abdominal views), enabling dynamic localization of four types of planes: abdomen, thalamus, cerebellum, and lumbosacral spine. The model achieved an accuracy of 87.38%, representing the first engineering framework for real-time, synchronized multi-organ data processing via IIoT (59).

To address the challenge of inter-class similarity—such as the high grayscale overlap between non-standard abdominal and cranial planes—a coarse-to-fine multi-task learning framework was proposed in 2023. This framework introduced a hierarchical classification strategy: initially assigning broad plane categories via coarse classification, then performing fine-grained discrimination between standard and non-standard planes. Applied to six subcategories across three anatomical planes (femur, thalamus, abdomen), the model achieved an overall accuracy of 95.2%, outperforming baseline models such as ResNet18 (84.03%) and EfficientNet (92.81%). Notably, the detection accuracy for non-standard planes (e.g., fetal femur non-standard plane, FFNSP) improved from 65.7% to 72.6% (60). In 2024, LPC-SonoNet aimed to enhance both model efficiency and classification scope. By integrating lightweight pyramid convolution (LPC) blocks into the original SonoNet framework, the parameter count was reduced from 14.9 million to 4.3 million. Simultaneously, the model supported classification of nine standard views—including refined subdivisions of cranial views into transventricular, transthalamic, and transcerebellar types. It achieved 97.0% accuracy for six categories and 91.9% for nine categories on the Burgos-Artizzu dataset, establishing a strong foundation for deployment on portable devices (61).

Krishna et al. proposed a deep learning-based approach combining CNNs with adaptive channel weighting (ACW) to tackle low inter-class variability and mitigate class imbalance. This is particularly effective for distinguishing structurally similar cardiac views and learning from class-imbalanced datasets with rare malformations. Their model embedded a multi-scale channel attention module (MSCAM) into a visual geometry group-19 (VGG-19) backbone, dynamically adjusting channel weights to amplify key features—such as the prominent echogenic signal of the spine in abdominal views. This approach achieved an overall accuracy of 98.20% across six diagnostic categories, including abdomen, femur, and thorax, marking a 5.7% improvement compared to conventional CNN (62). Addressing the high similarity between transventricular and transthalamic cranial planes, Migliorelli et al. employed contrastive learning, specifically SimCLR, to minimize intra-class feature distances and maximize inter-class separation. Even with a 50% reduction in labeled data, the model maintained an F1 score exceeding 95%, demonstrating the efficacy of unsupervised pretraining for small-sample, multi-organ classification tasks (63).

These multi-organ AI models have enabled comprehensive evaluations using a single device, significantly enhancing screening efficiency. However, most current models are trained on retrospective datasets, highlighting the need for prospective, multicenter clinical trials to validate their real-world diagnostic performance, particularly with respect to missed diagnosis rates. Looking forward, the integration of multimodal data—such as ultrasound, magnetic resonance imaging (MRI), and fetal electrocardiography—into cross-modal feature fusion frameworks holds promise for improving the detection of complex fetal anomalies.


Application of AI in other ultrasound standard planes localization

Ultrasound imaging has become a cornerstone in the screening and diagnosis of adult organ diseases—particularly of the heart, thyroid and liver —due to its non-invasiveness, real-time feedback, and cost-effectiveness. In the adult cardiovascular field, recognizing standard views is the essential first step for any comprehensive interpretation, as these views provide complementary information regarding cardiac anatomy and function. For instance, the apical four-chamber view enables the assessment of chamber dimensions and valvular structures, while other views like the parasternal long axis are critical for measuring ventricular wall thickness. In clinical thyroid examinations, ultrasound provides a clear depiction of nodules, enlargement, and other abnormalities, making it the preferred modality for evaluating both functional and structural aspects of the gland. Similarly, liver ultrasound enables direct visualization of intrahepatic vascular patterns and space-occupying lesions, playing an essential role in the early detection of conditions such as cirrhosis and hepatocellular carcinoma. However, these organs face similar imaging challenges, particularly the need for standardized scanning planes—such as the apical views for the heart, transverse isthmus view for the thyroid and the first porta hepatis plane for the liver—such as the transverse isthmus view for the thyroid and the first porta hepatis plane for the liver—which are critical for accurate parameter measurements and lesion localization. The precision of recognizing these standard planes directly influences subsequent clinical decision-making. Deep learning techniques in AI have substantially advanced the automation of standard plane localization by overcoming the time-consuming and subjective limitations of manual localization (Table 8).

Table 8

Summary of selected publications on other ultrasound standard plane localization

Author Target organ Ultrasound data Ultrasound system Year AI model Results
Madani A et al. (64) Heart 834,267 images General Electric, Philips, Siemens 2018 VGG-16 Video—Acc: NR, Pre: NR, Rec: NR, F1: NR, mAP: NR. Image—Acc: NR, Pre: NR, Rec: NR, F1: NR, mAP: NR
Kusunose K et al. (65) Heart 17,000 images EPIQ, Vivid E9/E95, iE33, Preirus, SSA-770A 2020 CNN Acc: 98.1%, Pre: NR, Rec: NR, F1: NR, mAP: NR
Li X et al. (66) Heart 170,311 images Philips, GE, Siemens, Mindray 2024 VGG16 + FPN + Transformer Acc: 97.8%, Pre: 94.8%, Rec: 95.3%, F1: 95.1%, mAP: NR
Steffner KR et al. (67) Heart 3,432 TEE videos N/S 2024 EchoNet Acc: NR, Pre: NR, Rec: NR, F1: NR, mAP: NR. Micro-averaged AUC: 91.9%, Macro-averaged AUC: 90.1%
Wu J et al. (68) Liver 4,682 LUSP images SuperSonic Imagine, General Electric 2021 CNN Acc: 92.31%, Pre: NR, Rec: NR, F1: NR, mAP: NR
Zhang J et al. (69) Liver 14,900 LUSP images N/S 2023 Ultra-Attention Acc: 93.2%, Pre: NR, Rec: NR, F1: NR, mAP: NR
Guo M et al. (70) Thyroid 4,509 TUSP images N/S 2019 ResNet-18 Acc: 83.88%, Pre: NR, Rec: NR, F1: NR, mAP: NR
Guo M et al. (71) Thyroid 5,500 TUSP images N/S 2021 ResNet-18 Acc: 91.07%, Pre: NR, Rec: NR, F1: NR, mAP: NR
Zhu Y et al. (72) Thyroid N/S 2023 ResNet-cbam Acc: 89%, Pre: 91%, Rec: 91%, F1: 89%, mAP: NR
Zeng P et al. (73) Thyroid 9,778 TUSP images Myriad, General Electric, Hitachi, Siemens 2023 TUSPM-NET Acc: 96.57%, Pre: 97.79%, Rec: 96.8%, F1: NR, mAP: NR, specificity: 99.73%

Acc, accuracy; AI, artificial intelligence; AUC, area under the curve; CNN, convolutional neural network; FPN, feature pyramid network; LUSP, liver ultrasound standard plane; mAP, mean average precision; NR, not reported; N/S, not stated; Pre, precision; Rec, recall; ResNet-18, residual network-18; ResNet-cbam, resnet-convolutional block attention module; TEE, transesophageal echocardiography; TUSP, thyroid ultrasound standard plane; TUSPM-NET, thyroid ultrasound standard plane multi-task network; VGG-16, visual geometry group-16.

In the realm of adult echocardiography, deep learning has demonstrated expert-level performance in multi-view classification. A landmark study in 2018 utilized a CNN to simultaneously classify 15 standard views (including 12 video and 3 still-image views), achieving an overall test accuracy of 97.8% on videos (64). Notably, the model achieved 91.7% accuracy on single low-resolution images, significantly outperforming board-certified echocardiographers whose accuracy ranged from 70.2% to 84.0%. Beyond pure classification, research has focused on the clinical feasibility of these models in real-world settings. A 2020 study demonstrated that a CNN-based classifier could accurately recognize 5 standard views with 98.1% accuracy, even across different equipment vendors and a wide range of left ventricular ejection fractions (LVEF) (65). This study also established that a mislabeling rate of up to 1.9% in the training database did not significantly impact the accuracy of subsequent clinical prediction models for EF, highlighting the robustness of AI for high-throughput clinical analysis.

Addressing the clinical need for image quality control, a 2024 study proposed a multi-task deep learning approach that integrates view recognition with real-time quality assessment (66). This model, trained on over 170,000 images, achieved a classification accuracy of 97.8% and an inference time of only 2.8 ms per frame, meeting the requirements for immediate feedback during image acquisition. By utilizing a hierarchical neck network to fuse multi-scale features, the model effectively simulated the human visual system’s processing, ensuring that only high-quality images are selected for downstream AI-assisted diagnosis. Furthermore, AI applications have extended to transesophageal echocardiography (TEE), which is vital for intraoperative management. A 2024 multi-institutional study developed a spatiotemporal CNN model to classify 8 standardized TEE views, achieving an overall micro-averaged AUC of 0.919. This work represents the first application of deep learning to unstructured, clinically acquired TEE video data, providing a foundation for automated evaluation in complex, dynamic surgical environments (67).

In liver ultrasound, a 2021 study addressed the automatic classification of liver ultrasound standard planes (LUSP) by introducing a method based on a pre-trained CNN. This model achieved the first automated classification of eight LUSP types, including views of hepatic and inferior vena cava veins. Employing a 13-layer CNN architecture combined with transfer learning (pre-trained convolutional layers and fine-tuned fully connected layers), the model effectively mitigated overfitting caused by limited medical datasets. It achieved a classification accuracy of 92.31% on a real-world clinical dataset, significantly outperforming traditional approaches that used handcrafted features (e.g., HOG, LBP) with SVM classifiers, which reached only about 85% accuracy. This study demonstrated the feasibility of deep learning for liver ultrasound plane classification (72). As clinical demands have become more refined, a 2023 study expanded the classification task from 8 to 13 LUSP types and proposed the ultra-attention structured perception strategy. Mimicking the visual attention mechanism of sonographers, this strategy emphasized key anatomical regions (e.g., hepatic hilum, hepatobiliary interface). The model, based on ResNet, incorporated a visual attention module to enhance feature extraction in low-contrast, high-noise images, raising classification accuracy to 93.2%. Additionally, a weakly supervised cropping method—requiring only 150 annotated samples—was employed to guide data augmentation, significantly reducing annotation costs and offering a scalable solution for large-scale clinical deployment (69). Similarly, in thyroid ultrasound examinations, early efforts utilized the ResNet-18 architecture to classify thyroid ultrasound standard plane (TUSP) images. The inclusion of residual units helped alleviate the vanishing gradient problem during deep network training and facilitated effective feature extraction (70,71). Building on this foundation, a 2023 study addressed challenges associated with the low contrast and high structural similarity of thyroid images by proposing a ResNet-CBAM model. This model integrated channel and spatial attention mechanisms to emphasize critical features, resulting in a classification accuracy of 89%—a 5% improvement over the baseline ResNet (68). In the same year, the TUSPM-NET multitask model marked a breakthrough by simultaneously performing classification of eight TUSP types and detection of key anatomical structures such as the isthmus and lobe boundaries. By introducing a plane-specific loss function and location filtering module, the model achieved a mAP (mAP@0.5) of 95% for anatomical detection, while the localization precision and recall for standard planes reached 97.79% and 96.80%, respectively. The model processed each image in just 19.9 milliseconds, meeting the requirements for real-time clinical scanning (73).

Overall, to address the scarcity of medical imaging data, mainstream approaches have relied on pre-trained models for network initialization, combined with data augmentation techniques [e.g., rotation, cropping, and hue, saturation, value (HSV) transformation] to improve data efficiency. Innovations in attention mechanisms—from global to structured local attention—have significantly enhanced feature discrimination in low-contrast ultrasound images. Nevertheless, while commercial systems currently enable the near-automatic fusion of ultrasound with MRI or CT for image-guided interventions (such as biopsy and ablation), current automated detection algorithms remain largely confined to single-modality ultrasound. This disjunction limits the comprehensive assessment of complex pathologies, such as spatially associated hepatic lesions. Consequently, extending multi-modal imaging fusion from procedural guidance to automated detection systems represents a critical trajectory for future research.


Future prospects

Currently, the accuracy of automatic localization models remains a major challenge in advancing AI-assisted standard plane localization, particularly in fetal ultrasound where precise identification of multiple standardized planes is essential for reliable prenatal assessment. Enhancing the accuracy of these models relies critically on the integration of clinical anatomical prior knowledge with deep learning frameworks. By incorporating prior information such as spatial relationships of key anatomical structures and landmark positions, the model can better interpret tissue distribution in images, thereby improving the precision of standard plane localization (74). This is particularly important in fetal ultrasound, as the distinction between “acceptable” and “unacceptable” views often depends on whether subtle anatomical landmarks are fully visible and whether the scan meets standardized requirements (75). Moreover, the introduction of prior knowledge enhances model interpretability, helping clinicians understand the diagnostic basis of the model and increasing trust in AI-assisted decision-making.

Beyond model optimization, high-quality datasets are also fundamental to technological progress. A critical bottleneck for real-world deployment is model generalizability across different institutions and imaging devices. Ultrasound images are highly sensitive to hardware specifications and operator settings, leading to significant “domain shift” when a model trained on data from one hospital is applied to another. Future research must prioritize the development of domain-invariant architectures and robust data augmentation strategies to mitigate inter-vendor and inter-institutional variability (76,77). However, most current studies still rely on private datasets for training, especially in fetal ultrasound where expert annotation of standard planes requires substantial clinical experience, which not only raises the barrier to entry but also hinders fair development and reproducibility in the field. Medical data sharing faces multiple challenges, including patient privacy protection, high costs of multi-institutional coordination, and scarcity of expert annotations. Although publicly available medical datasets have gradually increased in recent years, their scale and quality still lag far behind those in the natural image domain (78).

Furthermore, a significant challenge in cross-study comparison is the lack of standardized evaluation metrics across the current literature. As summarized in Tables 3-8, while some studies rely heavily on comprehensive metrics like precision, recall, and F1-score, others report only overall accuracy or specific spatial error distances. This heterogeneity impedes a fair, quantitative benchmarking of different state-of-the-art models. Future research should strive to establish a consensus-based, standardized evaluation framework specifically tailored for ultrasound standard plane localization to facilitate transparent and reproducible comparisons.

In resource-constrained embedded devices and mobile ultrasound systems, lightweight algorithm design is essential for achieving efficient and real-time diagnosis (79). Lightweighting is generally pursued under the premise of “no significant degradation in localization accuracy”, with a focus on reducing computational complexity, memory usage, and improving processing speed. Commonly used strategies include model pruning (80), quantization (81), knowledge distillation (82), and lightweight network architecture design (83). These methods optimize convolution operations and network structures, significantly reducing computational resource consumption and meeting the real-time requirements of ultrasound image processing. When it comes to fetal ultrasound, future lightweight systems should not be limited to recognizing standard cross-sections in single frames; they should also be capable of handling continuous video streams and 3D volumetric data, as a real prenatal examination is a dynamic scanning process (84).

Standard plane localization is only the foundational step for accurate diagnosis; a complete ultrasound diagnostic process also includes subsequent stages such as lesion detection, interpretation, and diagnosis. Successful clinical translation requires seamless integration into existing clinical workflows rather than serving as an isolated tool. For example, AI modules should ideally operate in the background of the ultrasound machine, providing real-time quality feedback without interrupting the sonographer’s scanning rhythm. Therefore, integrating the localization model with other functional modules can provide more comprehensive intelligent assistance. For example, combining it with lesion detection models (85) enables simultaneous standard plane localization and abnormal region identification, along with preliminary report generation. Integration with robotic systems (86) allows automatic standard plane acquisition, optimizing the allocation of medical resources. The integration of the localization model with a conversational AI system will drive the automation of community screening. This automated workflow covers key steps from medical history collection and standard plane localization to preliminary diagnosis, ultimately alleviating the heavy reliance on specialized human resources (87). Furthermore, by incorporating probe pose sensing technology (88), real-time spatial localization and guidance during scanning can be achieved, improving the consistency and accuracy of image acquisition, while also offering visual feedback for clinical teaching and training.

In the future, ultrasound standard plane localization technology will shift from “isolated algorithm optimization” to “systematic solutions addressing clinical needs”. The future development of fetal ultrasound standard plane localization should move from single-plane identification to multi-plane collaboration, from static image classification to dynamic video understanding, and from retrospective algorithm evaluation to prospective validation in real clinical workflows. In this process, building an end-to-end intelligent system that integrates image acquisition, plane localization, quality control, abnormality detection, and decision support is more likely to truly improve the consistency, accessibility, and clinical effectiveness of prenatal screening.

Currently, ultrasound standard plane localization technology is shifting from “isolated algorithm optimization” to “systematic solutions addressing clinical needs” (89). The growing focus is on building end-to-end intelligent ultrasound systems that integrate image acquisition, plane localization, lesion analysis, and decision support (90,91), thereby promoting the automation of clinical workflows and standardization of diagnoses.

The clinical translation of AI in ultrasonography has transitioned from theoretical modeling to the deployment of regulatory-approved commercial systems. Currently, several platforms, including Caption AI (GE HealthCare), UltraSight AI Guidance, the ScanNav suite (Anatomy PNB & Fetal), and the LVivo series (DiA Imaging Analysis), have successfully secured Food and Drug Administration (FDA) clearance and Conformité Européenne (CE) marking. The regulatory approval of these systems was primarily underpinned by pivotal multicenter clinical trials—most notably for Caption AI and UltraSight—which demonstrated that AI-guided image acquisition by non-experts is non-inferior to that of expert sonographers in obtaining diagnostic-quality views. However, while these systems show promise, there remains a pressing need for more large-scale prospective validation in diverse clinical settings. Most existing literature focuses on retrospective performance on curated datasets, which may not reflect real-world “noise”, such as difficult-to-image patients [e.g., those with high body mass index (BMI)] or emergency scenarios. Moving forward, prospective trials measuring clinical outcomes—such as diagnostic turnaround time and referral rates—will be essential to prove the true value of AI in the clinical pipeline.


Conclusions

The integration of AI, particularly deep learning, for the automatic localization of standard ultrasound planes represents a paradigm shift in medical imaging. In clinical practice, AI-assisted localization extends beyond merely improving operational efficiency by reducing sonographers’ repetitive tasks. More critically, it serves as a clinical decision-support tool that mitigates the heavy reliance on operator expertise, thereby enhancing diagnostic consistency and accuracy. This is particularly vital in fetal ultrasound and prenatal screening, where the acquisition of standardized planes forms the basis of reliable biometric measurements and reproducible diagnoses, while additional applications in abdominal sonography and superficial organ assessment are considered supportive extensions of this technology. The clinical integration of such systems can empower a broader range of practitioners to perform high-quality examinations, potentially expanding access to expert-level ultrasound services.

Future research in ultrasound plane localization should transition from isolated algorithmic advances to integrated clinical solutions. Key priorities include developing interpretable AI through anatomical knowledge integration, creating scalable datasets for robust generalization, and implementing lightweight systems that combine localization with lesion detection and robotic assistance. In the field of fetal ultrasound, future research should place greater emphasis on multicenter validation across different ultrasound devices, operators, and gestational ages, while also promoting the transition of related models from static image recognition toward real-time scanning assistance, image quality control, and integration into the complete prenatal screening workflow. This holistic approach will enable end-to-end diagnostic workflows by seamlessly integrating precise localization with objective quantitative analysis, ultimately standardizing ultrasound practice and expanding clinical accessibility.


Acknowledgments

None.


Footnote

Reporting Checklist: The authors have completed the Narrative Review reporting checklist. Available at https://qims.amegroups.com/article/view/10.21037/qims-2025-aw-2271/rc

Funding: The work was supported by the National Natural Science Foundation of China (Nos. 82572257, 82272028); Health Research Project of Hunan Provincial Health Commission (No. W20241010); Clinical Research 4310 Program of the Affiliated Changsha Central Hospital of the University of South China (No. 20214310NHYCG06); Hunan Provincial Health High-Level Talent Scientific Research Project (No. R2023010) and Postgraduate Scientific Research Innovation Project of Hunan Province (Nos. CX20251494, CX20251496).

Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://qims.amegroups.com/article/view/10.21037/qims-2025-aw-2271/coif). The authors have no conflicts of interest to declare.

Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved.

Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.


References

  1. Newman PG, Rozycki GS. The history of ultrasound. Surg Clin North Am 1998;78:179-95. [Crossref] [PubMed]
  2. Carson PL. Ultrasound: Imaging, development, application. Med Phys 2023;50:35-9. [Crossref] [PubMed]
  3. Alansary A, Folgoc LL, Vaillant G, Oktay O, Li Y, Bai W, Passerat-Palmbach J, Guerrero R, Kamnitsas K, Hou B, McDonagh S, Glocker B, Kainz B, Rueckert D. Automatic View Planning with Multi-scale Deep Reinforcement Learning Agents. Cham: Springer International Publishing, 2018.
  4. Wang B, Wang Y, Chu Y, Zhang K, Liu L, Zhang K, Zhu B, Wang D, Jiang T. AbVLM-Q: intelligent quality assessment for abdominal ultrasound standard planes via vision-language modeling. BMC Med Imaging 2025;25:344. [Crossref] [PubMed]
  5. Yu Y, Chen Z, Zhuang Y, Yi H, Han L, Chen K, Lin J. A guiding approach of Ultrasound scan for accurately obtaining standard diagnostic planes of fetal brain malformation. J Xray Sci Technol 2022;30:1243-60. [Crossref] [PubMed]
  6. Gaccioli F, Aye ILMH, Sovio U, Charnock-Jones DS, Smith GCS. Screening for fetal growth restriction using fetal biometry combined with maternal biomarkers. Am J Obstet Gynecol 2018;218:S725-37. [Crossref] [PubMed]
  7. Faletra FF, Ho SY, Leo LA, Paiocchi VL, Mankad S, Vannan M, Moccetti T. Which Cardiac Structure Lies Nearby? Revisiting Two-Dimensional Cross-Sectional Anatomy. J Am Soc Echocardiogr 2018;31:967-75.
  8. Ni D, Yang X, Chen X, Chin CT, Chen S, Heng PA, Li S, Qin J, Wang T. Standard plane localization in ultrasound by radial component model and selective search. Ultrasound Med Biol 2014;40:2728-42. [Crossref] [PubMed]
  9. Fenster A, Downey DB. Three-dimensional ultrasound imaging. Annu Rev Biomed Eng 2000;2:457-75. [Crossref] [PubMed]
  10. Yang X, Huang Y, Huang R, Dou H, Li R, Qian J, Huang X, Shi W, Chen C, Zhang Y, Wang H, Xiong Y, Ni D. Searching collaborative agents for multi-plane localization in 3D ultrasound. Med Image Anal 2021;72:102119. [Crossref] [PubMed]
  11. Wong L, White N, Ramkrishna J, Araujo Júnior E, Meagher S, Costa Fda S. Three-dimensional imaging of the uterus: The value of the coronal plane. World J Radiol 2015;7:484-93. [Crossref] [PubMed]
  12. Amisha Malik P, Pathania M, Rathaur VK. Overview of artificial intelligence in medicine. J Family Med Prim Care 2019;8:2328-31. [Crossref] [PubMed]
  13. Litjens G, Kooi T, Bejnordi BE, Setio AAA, Ciompi F, Ghafoorian M, van der Laak JAWM, van Ginneken B, Sánchez CI. A survey on deep learning in medical image analysis. Med Image Anal 2017;42:60-88. [Crossref] [PubMed]
  14. Park A, Chute C, Rajpurkar P, Lou J, Ball RL, Shpanskaya K, et al. Deep Learning-Assisted Diagnosis of Cerebral Aneurysms Using the HeadXNet Model. JAMA Netw Open 2019;2:e195600. [Crossref] [PubMed]
  15. Ullah MN, Levin CS. Application of Artificial Intelligence in PET Instrumentation. PET Clin 2022;17:175-82. [Crossref] [PubMed]
  16. Nensa F, Demircioglu A, Rischpler C. Artificial Intelligence in Nuclear Medicine. J Nucl Med 2019;60:29S-37S. [Crossref] [PubMed]
  17. Burgos-Artizzu XP, Coronado-Gutiérrez D, Valenzuela-Alcaraz B, Bonet-Carne E, Eixarch E, Crispi F, Gratacós E. Evaluation of deep convolutional neural networks for automatic classification of common maternal fetal ultrasound planes. Sci Rep 2020;10:10200. [Crossref] [PubMed]
  18. Shen YT, Chen L, Yue WW, Xu HX. Artificial intelligence in ultrasound. Eur J Radiol 2021;139:109717. [Crossref] [PubMed]
  19. Shen J, Zhang CJP, Jiang B, Chen J, Song J, Liu Z, He Z, Wong SY, Fang PH, Ming WK. Artificial Intelligence Versus Clinicians in Disease Diagnosis: Systematic Review. JMIR Med Inform 2019;7:e10010. [Crossref] [PubMed]
  20. Zhang L, Chen S, Chin CT, Wang T, Li S. Intelligent scanning: automated standard plane selection and biometric measurement of early gestational sac in routine ultrasound examination. Med Phys 2012;39:5015-27. [Crossref] [PubMed]
  21. Yang X, Ni D, Qin J, Li S, Wang T, Chen S, Heng PA. Standard plane localization in ultrasound by radial component. 2014 IEEE 11th International Symposium on Biomedical Imaging (ISBI). 2014 29 April-2 May 2014.
  22. Kim B, Kim KC, Park Y, Kwon JY, Jang J, Seo JK. Machine-learning-based automatic identification of fetal abdominal circumference from ultrasound images. Physiol Meas 2018;39:105007. [Crossref] [PubMed]
  23. Azizi S, Bayat S, Yan P, Tahmasebi A, Kwak JT, Xu S, Turkbey B, Choyke P, Pinto P, Wood B, Mousavi P, Abolmaesumi P. Deep Recurrent Neural Networks for Prostate Cancer Detection: Analysis of Temporal Enhanced Ultrasound. IEEE Trans Med Imaging 2018;37:2695-703. [Crossref] [PubMed]
  24. Dadoun H, Rousseau AL, de Kerviler E, Correas JM, Tissier AM, Joujou F, Bodard S, Khezzane K, de Margerie-Mellon C, Delingette H, Ayache N. Deep Learning for the Detection, Localization, and Characterization of Focal Liver Lesions on Abdominal US Images. Radiol Artif Intell 2022;4:e210110. [Crossref] [PubMed]
  25. Fan Z, Gong P, Tang S, Lee CU, Zhang X, Song P, Chen S, Li H. Joint localization and classification of breast masses on ultrasound images using an auxiliary attention-based framework. Med Image Anal 2023;90:102960. [Crossref] [PubMed]
  26. Barata C, Rotemberg V, Codella NCF, Tschandl P, Rinner C, Akay BN, Apalla Z, Argenziano G, Halpern A, Lallas A, Longo C, Malvehy J, Puig S, Rosendahl C, Soyer HP, Zalaudek I, Kittler H. A reinforcement learning model for AI-based decision support in skin cancer. Nat Med 2023;29:1941-6. [Crossref] [PubMed]
  27. Vlontzos A, Alansary A, Kamnitsas K, Rueckert D, Kainz B. Multiple Landmark Detection Using Multi-agent Reinforcement Learning. Cham: Springer International Publishing, 2019.
  28. Ni D, Li T, Yang X, Qin J, Li S, Chin C-T, Ouyang S, Wang T, Chen S. Selective Search and Sequential Detection for Standard Plane Localization in Ultrasound. Berlin, Heidelberg: Springer Berlin Heidelberg, 2013.
  29. Chen H, Ni D, Qin J, Li S, Yang X, Wang T, Heng PA. Standard Plane Localization in Fetal Ultrasound via Domain Transferred Deep Neural Networks. IEEE J Biomed Health Inform 2015;19:1627-36. [Crossref] [PubMed]
  30. Rahmatullah B, Papageorghiou A, Noble JA. Automated Selection of Standardized Planes from Ultrasound Volume. Berlin, Heidelberg: Springer Berlin Heidelberg, 2011.
  31. Ahmed M, Noble JA. An eye-tracking inspired method for standardised plane extraction from fetal abdominal ultrasound volumes. 2016 IEEE 13th International Symposium on Biomedical Imaging (ISBI). 2016 13-16 April 2016.
  32. Cai Y, Sharma H, Chatelain P, Noble JA. SonoEyeNet: Standardized Fetal Ultrasound Plane Detection Informed by Eye Tracking. Proc IEEE Int Symp Biomed Imaging 2018;2018:1475-8. [Crossref] [PubMed]
  33. Cai Y, Sharma H, Chatelain P, Noble JA. Multi-task SonoEyeNet: Detection of Fetal Standardized Planes Assisted by Generated Sonographer Attention Maps. Med Image Comput Comput Assist Interv 2018;11070:871-9. [Crossref] [PubMed]
  34. Lei B, Tan EL, Chen S, Zhuo L, Li S, Ni D, Wang T. Automatic Recognition of Fetal Facial Standard Plane in Ultrasound Image via Fisher Vector. PLoS One 2015;10:e0121838. [Crossref] [PubMed]
  35. Wang X, Liu Z, Du Y, Diao Y, Liu P, Lv G, Zhang H. Recognition of Fetal Facial Ultrasound Standard Plane Based on Texture Feature Fusion. Comput Math Methods Med 2021;2021:6656942. [Crossref] [PubMed]
  36. Yu Zhen, Ni Dong, Chen Siping, Li Shengli, Wang Tianfu, Lei Baiying. Fetal facial standard plane recognition via very deep convolutional networks. Annu Int Conf IEEE Eng Med Biol Soc 2016;2016:627-30. [Crossref] [PubMed]
  37. Yu Z, Tan EL, Ni D, Qin J, Chen S, Li S, Lei B, Wang T. A Deep Convolutional Neural Network-Based Framework for Automatic Fetal Facial Standard Plane Recognition. IEEE J Biomed Health Inform 2018;22:874-85. [Crossref] [PubMed]
  38. Xue H, Liu Z, Yu W, Liu PJ. Automatic recognition of fetal facial ultrasound standard planes based on improved YOLOv4. 2022 IEEE 16th International Conference on Anti-counterfeiting, Security, and Identification (ASID), Xiamen, China. 2022:110-3.
  39. Xue H, Yu W, Liu Z, Liu P. Early Pregnancy Fetal Facial Ultrasound Standard Plane-Assisted Recognition Algorithm. J Ultrasound Med 2023;42:1859-80. [Crossref] [PubMed]
  40. Liu Z, Yu W, Wu X, Yang T, Lyu G, Liu P, Xue H. Detection of fetal facial anatomy in standard ultrasonographic sections based on real-time target detection network. Int J Gynaecol Obstet 2024;165:916-28. [Crossref] [PubMed]
  41. Carroll SG, Porter H, Abdel-Fattah S, Kyle PM, Soothill PW. Correlation of prenatal ultrasound diagnosis and pathologic findings in fetal brain abnormalities. Ultrasound Obstet Gynecol 2000;16:149-53. [Crossref] [PubMed]
  42. Cavalheiro S, Moron AF, Almodin CG, Suriano IC, Hisaba V, Dastoli P, Barbosa MM. Fetal hydrocephalus. Childs Nerv Syst 2011;27:1575-83. [Crossref] [PubMed]
  43. Monteiro de Castro Doin Trigo LA, Benini-Junior JR, Brito LGO, Marba STM, Amaral E. Ultrasound diagnosis of microcephaly: a comparison of three reference curves and postnatal diagnosis. Arch Gynecol Obstet 2019;300:1211-9. [Crossref] [PubMed]
  44. Liu X, Annangi P, Gupta M, Yu B, Padfield D, Banerjee J, Krishnan K, Doyley MM. Learning-based scan plane identification from fetal head ultrasound images. Medical Imaging 2012: Ultrasonic Imaging, Tomography, and Therapy. 2012.
  45. Qu R, Xu G, Ding C, Jia W, Sun M. Deep Learning-Based Methodology for Recognition of Fetal Brain Standard Scan Planes in 2D Ultrasound Images. IEEE Access 2020;8:44443-51.
  46. Li Y, Khanal B, Hou B, Alansary A, Cerrolaza JJ, Sinclair M, Matthew J, Gupta C, Knight C, Kainz B, Rueckert D. Standard Plane Detection in 3D Fetal Ultrasound Using an Iterative Transformation Network. Cham: Springer International Publishing, 2018.
  47. Li Y, Cerrolaza JJ, Sinclair M, Hou B, Alansary A, Khanal B, Matthew J, Kainz B, Rueckert D. Standard Plane Localisation in 3D Fetal Ultrasound Using Network with Geometric and Image Loss. Medical Imaging with Deep Learning (MIDL 2018) 2018.
  48. Qu R, Xu G, Ding C, Jia W, Sun M. Standard Plane Identification in Fetal Brain Ultrasound Scans Using a Differential Convolutional Neural Network. IEEE Access 2020;8:83821-30.
  49. Zhao L, Li K, Pu B, Chen J, Li S, Liao X. An ultrasound standard plane detection model of fetal head based on multi-task learning and hybrid knowledge graph. Future Generation Computer Systems 2022;135:234-43.
  50. Lin M, He X, Guo H, He M, Zhang L, Xian J, Lei T, Xu Q, Zheng J, Feng J, Hao C, Yang Y, Wang N, Xie H. Use of real-time artificial intelligence in detection of abnormal image patterns in standard sonographic reference planes in screening for fetal intracranial malformations. Ultrasound Obstet Gynecol 2022;59:304-16. [Crossref] [PubMed]
  51. Li ZY, Chen YM, Qiu LQ, Chen DQ, Hu CG, Xu JY, Zhang XH. Prevalence, types, and malformations in congenital anomalies of the kidney and urinary tract in newborns: a retrospective hospital-based study. Ital J Pediatr 2019;45:50. [Crossref] [PubMed]
  52. Ewer AK, Middleton LJ, Furmston AT, Bhoyar A, Daniels JP, Thangaratinam S, Deeks JJ, Khan KS. Pulse oximetry screening for congenital heart defects in newborn infants (PulseOx): a test accuracy study. Lancet 2011;378:785-94. [Crossref] [PubMed]
  53. Gudigar A. Role of Four-Chamber Heart Ultrasound Images in Automatic Assessment of Fetal Heart: A Systematic Understanding. Informatics 2022;9:34.
  54. Dong J, Liu S, Wang T. ARVBNet: Real-Time Detection of Anatomical Structures in Fetal Ultrasound Cardiac Four-Chamber Planes. Machine Learning and Medical Engineering for Cardiovascular Health and Intravascular Imaging and Computer Assisted Stenting. Lecture Notes in Computer Science 2019:130-7.
  55. Li F, Li P, Wu X, Zeng P, Lyu G, Fan Y, Liu P, Song H, Liu Z. FHUSP-NET: A Multi-task model for fetal heart ultrasound standard plane recognition and key anatomical structures detection. Comput Biol Med 2024;168:107741. [Crossref] [PubMed]
  56. He J, Yang L, Liang B, Li S, Xu C. Fetal cardiac ultrasound standard section detection model based on multitask learning and mixed attention mechanism. Neurocomputing 2024;579.
  57. Pu B, Li K, Chen J, Lu Y, Zeng Q, Yang J, Li S. HFSCCD: A Hybrid Neural Network for Fetal Standard Cardiac Cycle Detection in Ultrasound Videos. IEEE J Biomed Health Inform 2024;28:2943-54. [Crossref] [PubMed]
  58. Baumgartner CF, Kamnitsas K, Matthew J, Fletcher TP, Smith S, Koch LM, Kainz B, Rueckert D. SonoNet: Real-Time Detection and Localisation of Fetal Standard Scan Planes in Freehand Ultrasound. IEEE Trans Med Imaging 2017;36:2204-15. [Crossref] [PubMed]
  59. Pu B, Li K, Li S, Zhu N. Automatic Fetal Ultrasound Standard Plane Recognition Based on Deep Learning and IIoT. IEEE Transactions on Industrial Informatics 2021;17:7771-80.
  60. Guo J, Tan G, Wu F, Wen H, Li K. Fetal Ultrasound Standard Plane Detection With Coarse-to-Fine Multi-Task Learning. IEEE J Biomed Health Inform 2023;27:5023-31. [Crossref] [PubMed]
  61. Yu T, Tsui PH, Leonov D, Wu S, Bin G, Zhou Z. LPC-SonoNet: A Lightweight Network Based on SonoNet and Light Pyramid Convolution for Fetal Ultrasound Standard Plane Detection. Sensors (Basel) 2024;24:7510. [Crossref] [PubMed]
  62. Krishna TB, Kokil P. Integration of a Deep Convolutional Neural Network With Adaptive Channel Weight Technique for Automated Identification of Standard Fetal Biometry Planes. IEEE Transactions on Instrumentation and Measurement 2024;73:1-11.
  63. Migliorelli G, Fiorentino MC, Di Cosmo M, Villani FP, Mancini A, Moccia S. On the use of contrastive learning for standard-plane classification in fetal ultrasound imaging. Comput Biol Med 2024;174:108430. [Crossref] [PubMed]
  64. Madani A, Arnaout R, Mofrad M, Arnaout R. Fast and accurate view classification of echocardiograms using deep learning. NPJ Digit Med 2018;1:6. [Crossref] [PubMed]
  65. Kusunose K, Haga A, Inoue M, Fukuda D, Yamada H, Sata M. Clinically Feasible and Accurate View Classification of Echocardiographic Images Using Deep Learning. Biomolecules 2020;10:665. [Crossref] [PubMed]
  66. Li X, Zhang H, Yue J, Yin L, Li W, Ding G, Peng B, Xie S. A multi-task deep learning approach for real-time view classification and quality assessment of echocardiographic images. Sci Rep 2024;14:20484. [Crossref] [PubMed]
  67. Steffner KR, Christensen M, Gill G, Bowdish M, Rhee J, Kumaresan A, He B, Zou J, Ouyang D. Deep learning for transesophageal echocardiography view classification. Sci Rep 2024;14:11. [Crossref] [PubMed]
  68. Wu J, Zeng P, Liu P, Lv G. Automatic classification method of liver ultrasound standard plane images using pre-trained convolutional neural network. Connection Science 2021;34:975-89.
  69. Zhang J, Chen Y, Zeng P, Liu Y, Diao Y, Liu P. Ultra-Attention: Automatic Recognition of Liver Ultrasound Standard Sections Based on Visual Attention Perception Structures. Ultrasound Med Biol 2023;49:1007-17. [Crossref] [PubMed]
  70. Guo M, Du Y. Classification of Thyroid Ultrasound Standard Plane Images using ResNet-18 Networks. 2019 IEEE 13th International Conference on Anti-counterfeiting, Security, and Identification (ASID), Xiamen, China. 2019:324-8.
  71. Guo M, Wang K, Liu S, Du Y, Liu P, Su Q, Lv G. Recognition of Thyroid Ultrasound Standard Plane Images Based on Residual Network. Comput Intell Neurosci 2021;2021:5598001. [Crossref] [PubMed]
  72. Zhu Y, Yao L, Zhou X, An X, Yue Y, Xu S. Classification of thyroid ultrasound standard section based on ResNet-cbam. Third International Conference on Digital Signal and Computer Communications 2023 (DSCC 2023).
  73. Zeng P, Liu S, He S, Zheng Q, Wu J, Liu Y, Lyu G, Liu P. TUSPM-NET: A multi-task model for thyroid ultrasound standard plane recognition and detection of key anatomical structures of the thyroid. Comput Biol Med 2023;163:107069. [Crossref] [PubMed]
  74. Hao M, Guo J, Liu C, Chen C, Wang S. Development and preliminary testing of a prior knowledge-based visual navigation system for cardiac ultrasound scanning. Biomed Eng Lett 2024;14:307-16. [Crossref] [PubMed]
  75. Salomon LJ, Alfirevic Z, Berghella V, Bilardo CM, Chalouhi GE, Da Silva Costa F, Hernandez-Andrade E, Malinger G, Munoz H, Paladini D, Prefumo F, Sotiriadis A, Toi A, Lee W. ISUOG Practice Guidelines (updated): performance of the routine mid-trimester fetal ultrasound scan. Ultrasound Obstet Gynecol 2022;59:840-56. [Crossref] [PubMed]
  76. Wang SY, Pershing S, Lee AY. Big data requirements for artificial intelligence. Curr Opin Ophthalmol 2020;31:318-23. [Crossref] [PubMed]
  77. Cima G, Console M, Lenzerini M, Poggi A. A review of data abstraction. Front Artif Intell 2023;6:1085754. [Crossref] [PubMed]
  78. Li J, Zhu G, Hua C, Feng M, Bennamoun B, Li P, Lu X, Song J, Shen P, Xu X, Mei L, Zhang L, Shah SAA, Bennamoun M. A Systematic Collection of Medical Image Datasets for Deep Learning. ACM Computing Surveys 2023;56:116.
  79. Protserov S, Hunter J, Zhang H, Mashouri P, Masino C, Brudno M, Madani A. Development, deployment and scaling of operating room-ready artificial intelligence for real-time surgical decision support. NPJ Digit Med 2024;7:231. [Crossref] [PubMed]
  80. Zhang M, Yu X, Rong J, Ou L. Graph pruning for model compression. Applied Intelligence 2022;52:11244-56.
  81. Gong C, Chen Y, Lu Y, Li T, Hao C, Chen D, Vec Q. Minimal Loss DNN Model Compression With Vectorized Weight Quantization. IEEE Transactions on Computers 2021;70:696-710.
  82. Chen G, Fan Z, Zhu Y, Zhang T. Multi-Time Knowledge Distillation. Neurocomputing 2025;623:9.
  83. Zhu M, Min W, Han Q, Zhan G, Fu Q, Li J. ShuffleNeMt: modern lightweight convolutional neural network architecture. Pattern Analysis and Applications 2024;27:123.
  84. Men Q, Zhao H, Drukker L, Papageorghiou AT, Noble JA. ScanAhead: Simplifying standard plane acquisition of fetal head ultrasound. Med Image Anal 2025;104:103614. [Crossref] [PubMed]
  85. Han D, Ibrahim N, Lu F, Zhu Y, Du H, AlZoubi A. Automatic Detection of Thyroid Nodule Characteristics From 2D Ultrasound Images. Ultrason Imaging 2024;46:41-55. [Crossref] [PubMed]
  86. Hidalgo EM, Wright L, Isaksson M, Lambert G, Marwick TH. Current Applications of Robot-Assisted Ultrasound Examination. JACC Cardiovasc Imaging 2023;16:239-47. [Crossref] [PubMed]
  87. Lin Z, Du Y, Yang W, Qian Y, Li J. Conversational AI in Medicine: Redefining Diagnostics with AMIE. BIO Integration 2025;6.
  88. Khaw KL, Sridharan A, Poznick L, Kilbaugh TJ, Hwang M. Probe position sensor to track image location in 2D ultrasound. Ultrasonics 2020;103:106084. [Crossref] [PubMed]
  89. Baloescu C, Bailitz J, Cheema B, Agarwala R, Jankowski M, Eke O, Liu R, Nomura J, Stolz L, Gargani L, Alkan E, Wellman T, Parajuli N, Marra A, Thomas Y, Patel D, Schraft E, O’Brien J, Moore CL, Gottlieb M. Artificial Intelligence-Guided Lung Ultrasound by Nonexperts. JAMA Cardiol 2025;10:245-53. [Crossref] [PubMed]
  90. Su K, Liu J, Ren X, Huo Y, Du G, Zhao W, Wang X, Liang B, Li D, Liu PX. A fully autonomous robotic ultrasound system for thyroid scanning. Nat Commun 2024;15:4004. [Crossref] [PubMed]
  91. Jiang Z, Salcudean SE, Navab N. Robotic ultrasound imaging: State-of-the-art and future perspectives. Med Image Anal 2023;89:102878. [Crossref] [PubMed]
Cite this article as: Wang K, Lei L, Liu Z, Huang Y, Wei Q, He W, Jin L, Du M, Chen Z. Artificial intelligence in assisted localization of fetal ultrasound standard planes: a narrative review. Quant Imaging Med Surg 2026;16(9):745. doi: 10.21037/qims-2025-aw-2271

Download Citation