Artificial intelligence-based prenatal ultrasound for diagnosing fetal lip and palate during the second trimester
Original Article

Artificial intelligence-based prenatal ultrasound for diagnosing fetal lip and palate during the second trimester

Wei Jiang, Huixian Pang, Shihua Deng, Chaoting Chen

Department of Ultrasound, Huazhong University of Science and Technology Union Shenzhen Hospital (Nanshan Hospital), Shenzhen, China

Contributions: (I) Conception and design: W Jiang, H Pang; (II) Administrative support: W Jiang, H Pang; (III) Provision of study materials or patients: W Jiang, H Pang; (IV) Collection and assembly of data: W Jiang, S Deng, C Chen; (V) Data analysis and interpretation: W Jiang, S Deng, C Chen; (VI) Manuscript writing: All authors; (VII) Final approval of manuscript: All authors.

Correspondence to: Wei Jiang, MD. Department of Ultrasound, Huazhong University of Science and Technology Union Shenzhen Hospital (Nanshan Hospital), No. 89 Taoyuan Road, Nanshan District, Shenzhen 518052, China. Email: 2646748175@qq.com.

Background: Fetal facial anomalies are among the most common congenital conditions, with cleft lip and palate being the most prevalent. This study aimed to evaluate the clinical value of convolutional neural networks in automatically identifying standard ultrasonic cross-sectional images of the fetal lip and palate during the second trimester.

Methods: From September 2021 to December 2022, prenatal sonographers collected dynamic videos of the lip and palate of 700 fetuses at 20–24 weeks of gestation, including 5 standard cross-sectional images, along with nonstandard cross-sectional images and background images. A YOLOv5-based artificial intelligence (AI) model was used as the object detection network. The sonographers manually marked the images of 500 fetuses (450 in the training set and 50 in the validation set). The AI model, a midcareer sonographer, and a junior sonographer were involved in the classification and identification of standard fetal lip and palate ultrasound images in the test set (200 fetuses), and the results were compared with the standard results obtained by the senior prenatal sonographer. The receiver operating characteristic curve was plotted, and the sensitivity, specificity, and accuracy were calculated for the AI model, midcareer sonographer, and junior sonographer.

Results: For the standard coronal section of the nasal lip, the area under the curve (AUC) of the AI model, the midcareer sonographer, and the junior sonographer was 0.971, 0.935, and 0.880, respectively. For the standard midsagittal section of the face, the AUC of the AI model, the midcareer sonographer, and the junior sonographer was 0.988, 0.939, and 0.904, respectively. For the standard upper alveolar ridge section, the AUC of the AI model, the midcareer sonographer, and the junior sonographer was 0.977, 0.840, and 0.824, respectively. For all standard sections, the AI model demonstrated significantly higher AUC values as compared to both the midcareer and junior sonographers (P<0.05).

Conclusions: The AI model demonstrated higher classification efficacy than did the midcareer and junior sonographers and performed more quickly and efficiently.

Keywords: Fetal; lips and palate; standard ultrasonic section; artificial intelligence (AI); quality control


Submitted Jul 17, 2024. Accepted for publication May 21, 2025. Published online Jul 30, 2025.

doi: 10.21037/qims-24-1454


Introduction

Fetal facial anomalies are the most common congenital anomalies, while cleft lip and palate are the most common facial anomalies (1). The presence of lip and palate anomalies in fetuses aggravates the difficulty of infant feeding, increases family economic burden, and affects the healthy development of the child’s psychology. Midterm prenatal ultrasound examination has the advantages of safety, noninvasiveness, and economy and is recognized as the best imaging method for examining fetal growth and development (2). Therefore, prenatal detection of cleft lip and palate is crucial as it allows for better pregnancy counseling, enables early intervention planning, facilitates parental psychological preparation, and ensures appropriate delivery planning at tertiary care centers with multidisciplinary expertise.

However, due to the interobserver variability among sonographers, the detection rate for fetal cleft lip and palate varies significantly, resulting in both false-positive and false-negative diagnoses (3). In a recent systematic review and meta-analysis of high-risk pregnancies, the pooled sensitivity of the ultrasound for the fetal cleft lip and palate was 87% [95% confidence interval (CI): 71–95%], with a false-negative rate (FNR) of 13%, while the specificity was 98% (95% CI: 90–100%), with a false-positive rate (FPR) of 2% (4). The ultrasound examination of the lip and palate is affected by fetal position, amniotic fluid volume, peripheral bone tissue echoes, and maternal factors (including maternal body mass index, subcutaneous tissue thickness, and abdominal wall characteristics). The examination requires careful attention to multiple anatomical points and processes, often necessitating repeated scanning positions by sonographers, which can lead to occupational musculoskeletal strain from prolonged and repetitive movements. In daily prenatal screening clinical work, different standard sections contain a variety of important anatomical structures, such as the nose, upper and lower lips, upper alveolar ridge, hard palate, and soft palate. Prenatal senior sonographers judge whether fetal lip and palate anomalies exist based on whether key structures are complete and the morphology is normal, among many other factors.

Deep learning consists of multiple layers of artificial neural networks, which can automatically extract image features without human intervention, and is now commonly used in the recognition and classification of medical images (5-9). Artificial intelligence (AI) has been applied to various aspects of prenatal ultrasound, including standard view classification, fetal biometry, and anomaly detection (10). However, no studies have addressed the AI-based recognition of standard fetal facial ultrasound sections (including the upper alveolar, under-lip oblique coronal, and under-jaw coronal views) and their variations. Therefore, this study evaluated the feasibility and effectiveness of AI in classifying both standard and nonstandard fetal facial ultrasound images at 20–24 weeks of gestation, comparing its performance with that of midcareer and junior sonographers.

This study used senior sonographers’ image classifications as the reference standard and compared the performance of AI with that of midcareer and junior sonographers in recognizing standard fetal lip and palate ultrasound images. Our aim was to evaluate the clinical utility of AI models in automatically identifying standard ultrasound views. This technology could assist in clinical examinations of the fetal lip and palate at 20–24 weeks of gestation, support junior sonographers in learning standard ultrasound planes, and help reduce the workload of image quality control. We present this article in accordance with the TRIPOD+AI reporting checklist (available at https://qims.amegroups.com/article/view/10.21037/qims-24-1454/rc).


Methods

Ethical consideration

The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments and was approved by the Institutional Ethics Committee of Huazhong University of Science and Technology Union Shenzhen Hospital (No. KY-2022-12-12). Informed consent was obtained from all individual participants.

Patient selection

This study retrospectively collected prenatal ultrasound videos from women who underwent routine ultrasound examination at 20–24 weeks of gestation at Huazhong University of Science and Technology Union Shenzhen Hospital (Nanshan Hospital) between September 2021 and December 2022. The inclusion criteria for images were as follows: clear images, target structures located in the center of the images, pure backgrounds without artifacts, and no overlapping color Doppler images, measurement data, text labels, or annotations included in the images. The exclusion criteria for images were as follows: images that were blurred or trailing due to maternal obesity, fetal position changes, or image jitter and fetal limbs or accessories occupying more than two-thirds of the overall image.

Ultrasound detection

The Voluson E10 ultrasound device (GE HealthCare, Chicago, IL, USA) was used with two types of probes, a two-dimensional (2D) C2-9 probe and an RM6C volume probe. The pregnant woman was in a supine position and instructed to breathe calmly with her hands placed on both sides of the body, fully exposing the abdomen. The ultrasound examinations were performed by a team of six prenatal ultrasound specialists, all certified in obstetric ultrasound imaging. The team consisted of two senior sonographers (>10 years of experience), two midcareer sonographers (5 years of experience), and two junior sonographers (1–2 years of experience). All team members were trained in standard fetal anatomical scanning protocols and regularly performed prenatal ultrasound examinations in daily clinical practice. Each sonographer obtained dynamic video sequences by maintaining a stable probe position at the standard plane while making subtle angular adjustments to capture both standard and nonstandard views. Each recorded video sequence contained 75–150 frames and was stored in the ultrasound device for subsequent analysis.

Standard ultrasound section of fetal lip and palate judgement criteria

The following criteria were used to classify standard and nonstandard sections of fetal lip and palate images.

  • Nasolabial coronal section: (I) for standard section, both nostrils and nasal wings were displayed completely and symmetrically, with the upper and lower lips, as well as the chin, being clearly visible. (II) For nonstandard section, both nostrils were not fully displayed, the upper and lower lips could be visible, and the chin was unclear. (III) The key anatomical structures included the nose, upper and lower lips, and chin.
  • Midsagittal section of the face: (I) for the standard section, the nasal area was visible as a nasal pillar rather than nostrils, with the upper and lower lips clearly displayed, and the chin appearing as a round dot. (II) For nonstandard section, the nasal area appeared as nostrils, or the chin appeared as a long strip of the jawbone, with poor visualization of the upper and lower lips. (III) The key anatomical structures included the nose, upper and lower lips, and chin.
  • Transverse section of the upper alveolar ridge: (I) for standard section, the sound beam was as perpendicular as possible to the upper alveolar ridge (deviation angle ≤30°), with the upper alveolar ridge and primary incisors symmetrically displayed and clearly visible. At least two primary incisors were required to be displayed. (II) For nonstandard section, the upper alveolar ridge was asymmetrically displayed, or the primary incisors were unclear. (III) The key anatomical structures included the upper alveolar ridge and primary incisors.
  • Oblique coronal section through the lower lip or mandible (showing the soft palate): (I) for the standard section, the sound beam was as perpendicular as possible to the soft palate (deviation angle ≤30°), and the soft palate was required to appear as an equals sign, with the nasopharynx and tongue clearly displayed. (II) For nonstandard section, the soft palate was not clearly visible. (III) The key anatomical structures included the soft palate, nasopharynx, and tongue.
  • Background images: images of the fetal lip and palate that did not belong to any of the abovementioned sections were classified as background images.
  • Handling of suboptimal images: suboptimal images, such as those affected by maternal obesity, fetal position, or image artifacts, were excluded during the dataset preparation to ensure the quality of the training and validation datasets. However, we acknowledge that such suboptimal conditions are common in clinical practice and pose challenges for both human sonographers and AI models. (I) In real-time scanning, human sonographers can adjust the transducer’s position and angle, optimize imaging parameters, and use their clinical judgment to overcome challenges posed by maternal obesity or unfavorable fetal positions. For example, they may reposition the mother or wait for the fetus to move into a more favorable position. (II) Unlike human sonographers, the AI model does not have the capability to adjust scanning parameters or reposition the patient. Instead, it processes the input data as it is. To address this limitation, data augmentation techniques (e.g., scaling, flipping, and rotation) were applied during training to improve the model’s robustness against image variations.

Diagrams of the standard and nonstandard sections are shown in Figure 1.

Figure 1 Standard and nonstandard ultrasonic sections of fetal facial regions. (A) Standard coronal section of the nasal and labial regions. (B) Nonstandard coronal section of the nasal and labial regions. (C) Standard midsagittal section of the facial region. (D) Nonstandard midsagittal section of the facial region. (E) Standard transverse section through the upper dental arch. (F) Nonstandard transverse section through the upper dental arch. (G) Standard oblique coronal section through the lower lip or jaw (showing the soft palate). (H) Nonstandard oblique coronal section through the lower lip or jaw (showing the soft palate). (I) Background image.

Data annotation process

The computer hardware and software environment used for the study were as follows: (I) CPU, Intel® Xeon® CPU E5-2650v4 at 2.20 GHz (Intel Corporation, USA); (II) GPU, TITAN V (NVIDIA Corporation, USA); (III) VRAM, 12 G; (IV) RAM, 32 G; (V) computer operating system, 64-bit Ubuntu 18.04 (Canonical Ltd., UK); and (VI) programming code, Python 3.8 (Python Software Foundation, USA). The research assistants imported dynamic video data into the SonoKit mapping software (SonoKit, Sonoscape Medical Corp, China).

The research assistants imported dynamic video data into the SonoKit mapping software. A senior sonographer then extracted static images based on predefined inclusion and exclusion criteria. To ensure consistency and accuracy, the extracted images were annotated with specialized software according to standardized criteria for each lip and palate ultrasound section.

All image annotations were subjected to a rigorous review process. Specifically, after the initial labeling by the senior sonographer, all annotations were independently verified by a second senior sonographer. This double-review process was implemented to minimize subjectivity, reduce errors, and ensure the reliability of the dataset. Any discrepancies between the two reviewers were resolved through discussion until a consensus was reached.

Image preprocess

The image was scaled proportionally to 640×640, and the empty edge parts of the image were filled with pixels of 0. Horizontal flip and proportional scaling (0.7–1.3 times) were used for data augmentation. The batch size was set to 32, and a total of 70 epochs were trained. The decreasing curve of the loss and the precision, recall, and mean average precision of the validation set are shown in Figure 2.

Figure 2 The plot of the descending trend of loss and the changing trend of the validation set performance. The detection box loss, foreground loss, and classification loss of the training set all show a decreasing trend, but the loss in the validation set generally becomes flat or no longer decreases after about 50 epochs. Meanwhile, the accuracy, recall rate, and mAP of the validation set also remained flat or no longer increased after about 50 epochs. mAP, mean average precision; Val, validation.

Algorithm model

The YOLOv5 model was used as the object detection network, and the basic modules included the Focus, Conv (convolution), BottleneckCSP (cross stage partial Bottleneck), SPP (spatial pyramid pooling), Concat (concatenation), and Detect modules were used in this study. The Conv module is a common Conv + BatchNorm (batch normalization) + Act [activation function, such as ReLU (rectified linear unit) or HardSwish] module, as shown in Figure 3. The object detection network was used to extract features and recognize images and output the sectional area and structural area of the image and its corresponding category. The category of the sectional area was used to determine which section the current image belonged to and to further determine whether it was a standard image by determining whether all the structures to be recognized were included.

Figure 3 Network architecture diagram. Conv, convolution; CSP, cross stage partial.

Statistical analysis

Statistical analysis was performed with SPSS 26.0 software (IBM Corp., Armonk, NY, USA). With the expert’s image classification results serving as the standard, the recall, specificity, precision, positive predictive value (PPV), negative predictive value (NPV), FPR, FNR, F1 score, and accuracy of the AI model, midcareer sonographer, and junior sonographer in classifying and identifying ultrasound sections of fetal lip and palate were calculated. The count data were analyzed via the chi-squared test. The receiver operating characteristic (ROC) curves were plotted, and the area under the curve (AUC) was calculated to evaluate the classification performance of the three groups. A Z-test was used to compare the differences between the AUC values, and P<0.05 was considered statistically significant.


Results

Patients’ basic information

A total of 700 fetuses at 20–24 weeks of gestation were included in this study. Among them, there were 500 fetuses in the training and validation sets, and 1,679 related ultrasound dynamic videos of the lip and palate were collected. There were 445 videos of the nasal-labial coronal section, 392 videos of the midsagittal section of the face, 324 videos of the transverse section of the maxilla, 262 videos of the oblique coronal section through the lower lip or jaw (showing the hard palate), and 256 videos of the oblique coronal section through the lower lip or jaw (showing the soft palate). The test set included 200 fetal samples with 510 related ultrasound dynamic videos of the lip and palate being collected. Among them, there were 107 videos of the nasal-labial coronal section, 93 videos of the midsagittal section of the face, 103 videos of the transverse section of the maxilla, 112 videos of the oblique coronal section through the lower lip or jaw (showing the hard palate), and 95 videos of the oblique coronal section through the lower lip or jaw (showing the soft palate).

The senior sonographer categorized a total of 8,160 ultrasound images of fetal lip and palate in the training and validation sets, along with an additional 4,416 ultrasound images in the test set (Table 1).

Table 1

The images of each section classified by the senior sonographer

Ultrasound section Training and validation set (n=500, images =8,160) Test set (n=200, images =4,416)
Standard section
   Nasal-labial coronal section 1,095 610
   Midsagittal section of the face 1,023 540
   Transverse section of the maxilla 558 456
   Oblique coronal section through the lower lip or jaw (showing the hard palate) 812 547
   Oblique coronal section through the lower lip or jaw (showing the soft palate) 640 553
Nonstandard section
   Nasal-labial coronal section 1,106 346
   Midsagittal section of the face 909 274
   Transverse section of the maxilla 623 281
   Oblique coronal section through the lower lip or jaw (showing the hard palate) 718 360
   Oblique coronal section through the lower lip or jaw (showing the soft palate) 540 302
Background image 136 147

The classification results of the AI model, midcareer sonographer, and junior sonographer

In the test set, confusion matrices for the AI system, midcareer sonographer, and junior sonographer are shown in Tables 2-4, respectively. The sensitivity, specificity, PPV, FNR, and F1 scores for standard sections are presented in Table 5. The ROC curves for all three groups are displayed in Figure 4. For the classification and recognition of standard fetal lip and palate ultrasound sections, the AUC of the AI model was superior to that of the midcareer sonographer, whose AUC was superior to that of the junior sonographer. Pairwise comparisons demonstrated statistically significant differences between all three groups (the AI system, midcareer sonographer, and junior sonographer) (Table 6).

Table 2

Confusion matrix for the artificial intelligence model’s classification results

True label Prediction label
0 1 2 3 4 5
0 584 1 0 0 0 25
1 0 531 0 0 0 9
2 0 0 448 0 0 8
3 0 0 0 532 0 15
4 9 6 10 8 355 165
5 57 21 105 42 106 1,379

0: standard nasal-labial coronal section. 1: standard midsagittal section of the face. 2: standard transverse section of the maxilla. 3: standard oblique coronal section through the lower lip or jaw (showing the hard palate). 4: standard oblique coronal section through the lower lip or jaw (showing the soft palate). 5: background image.

Table 3

Confusion matrix of the midcareer sonographer’s results

True label Prediction label
0 1 2 3 4 5
0 553 0 0 0 0 57
1 1 486 0 0 0 53
2 1 0 326 1 0 128
3 0 0 0 426 1 120
4 0 0 0 3 314 236
5 138 84 136 84 108 1,160

0: standard nasal-labial coronal section. 1: standard midsagittal section of the face. 2: standard transverse section of the maxilla. 3: standard oblique coronal section through the lower lip or jaw (showing the hard palate). 4: standard oblique coronal section through the lower lip or jaw (showing the soft palate). 5: background image.

Table 4

Confusion matrix of the junior sonographer’s results

True label Prediction label
0 1 2 3 4 5
0 484 0 0 0 0 126
1 0 452 0 0 0 88
2 0 0 320 17 16 103
3 0 0 1 373 9 164
4 18 9 8 10 314 194
5 111 100 133 134 258 974

0: standard nasal-labial coronal section. 1: standard midsagittal section of the face. 2: standard transverse section of the maxilla. 3: standard oblique coronal section through the lower lip or jaw (showing the hard palate). 4: standard oblique coronal section through the lower lip or jaw (showing the soft palate). 5: background image.

Table 5

The classification results of the AI model, midcareer sonographer, and junior sonographer

Variables Sensitivity (%) Specificity (%) PPR (%) FNR (%) F1 score
AI
   Nasal-labial coronal section 95.73 98.27 89.85 99.30 0.927
   Midsagittal section of the face 98.33 99.28 94.99 99.77 0.966
   Transverse section of the maxilla 98.24 97.10 79.57 99.79 0.879
   Oblique coronal section through the lower lip or jaw (showing the hard palate) 97.26 98.71 91.41 99.61 0.942
   Oblique coronal section through the lower lip or jaw (showing the soft palate) 64.20 97.26 77.00 94.99 0.700
Midcareer sonographer
   Nasal-labial coronal section 90.65 96.32 79.79 98.46 0.849
   Midsagittal section of the face 90.00 97.83 85.26 98.59 0.876
   Transverse section of the maxilla 71.49 96.56 70.56 95.98 0.710
   Oblique coronal section through the lower lip or jaw (showing the hard palate) 77.87 97.72 82.87 96.89 0.803
   Oblique coronal section through the lower lip or jaw (showing the soft palate) 56.78 97.17 79.49 94.01 0.643
Junior sonographer
   Nasal-labial coronal section 79.34 96.61 78.95 96.68 0.791
   Midsagittal section of the face 83.70 97.18 80.57 97.71 0.821
   Transverse section of the maxilla 70.17 96.41 69.26 96.56 0.697
   Oblique coronal section through the lower lip or jaw (showing the hard palate) 68.19 95.83 69.85 95.51 0.717
   Oblique coronal section through the lower lip or jaw (showing the soft palate) 56.78 92.67 52.59 93.74 0.546

AI, artificial intelligence; F1 score, a measure of a test’s accuracy, calculated as the harmonic mean of precision and recall; FNR, false negative rate; PPR, positive predictive value.

Figure 4 ROC of the AI model, midcareer sonographer, and junior sonographer. (A) Nasal-labial coronal section. (B) Midsagittal section of the face. (C) Transverse section of the maxilla. (D) Oblique coronal section through the lower lip or jaw (showing the hard palate). (E) Oblique coronal section through the lower lip or jaw (showing the soft palate). The blue line represents the result from the AI model, the yellow line represents the result from the midcareer sonographer, and the green line represents the result from the junior sonographer. AI, artificial intelligence; ROC, receiver operating characteristic.

Table 6

The comparison of AUC between the AI model, midcareer sonographer, and junior sonographer

Variables Z value P
Nasal-labial coronal section
   AI vs. midcareer sonographer 7.273 <0.05
   AI vs. junior sonographer 11.688 <0.05
   Midcareer sonographer vs. junior sonographer 8.568 <0.05
Midsagittal section of the face
   AI vs. midcareer sonographer 8.109 <0.05
   AI vs. junior sonographer 10.861 <0.05
   Midcareer sonographer vs. junior sonographer 6.585 <0.05
Transverse section of the maxilla
   AI vs. midcareer sonographer 13.127 <0.05
   AI vs. junior sonographer 13.623 <0.05
   Midcareer sonographer vs. junior sonographer 2.728 <0.05
Oblique coronal section through the lower lip or jaw (showing hard palate)
   AI vs. mid-career sonographer 11.984 <0.05
   AI vs. junior sonographer 16.280 <0.05
   Midcareer sonographer vs. junior sonographer 9.011 <0.05
Oblique coronal section through the lower lip or jaw (showing the soft palate)
   AI vs. midcareer sonographer 6.713 <0.05
   AI vs. junior sonographer 10.299 <0.05
   Midcareer sonographer vs. junior sonographer 13.497 <0.05

AI, artificial intelligence; AUC, area under the curve.

The classification efficiency of the AI, midcareer sonographer, and junior sonographer

The total classification time of the AI mode was 48 minutes and 35 seconds, with an average classification time of 0.7 seconds per image. The total time for the midcareer physician classification process was 5 hours 12 minutes and 32 seconds, with an average time of about 4.2 seconds per image. The total time for the junior physician classification process was 5 hours 21 minutes and 20 seconds, with an average time of about 4.4 seconds per image. The AI model’s classification image time was much shorter than that of the midcareer and junior physicians. The average classification time per image for both the midcareer and junior physicians was highly similar.


Discussion

In this study, classification results from a senior sonographer were taken as the standard. The differences in the classification of standard and nonstandard ultrasound sections of the lip and palate in 20- to 24-week-old fetuses were compared between the AI model, midcareer sonographer, and junior sonographer. The results showed that the AI model performed better than the midcareer and junior sonographers. In addition, the AI model’s classification time per image was much shorter than that of the midcareer and junior sonographers.

Cleft lip and palate is the most common facial congenital anomaly and the second most common birth defect (11). Whether evaluating fetal growth status or anatomical abnormalities, ultrasound technicians need to obtain corresponding standard ultrasound sections from the complex fetal structures and perform identification and analysis. For example, in the standard coronal section of the nose and lip, the fetal nose, upper and lower lips, and philtrum can be used to screen for an absent nose, proboscis, oral tumors, cleft lip, etc.; the standard midline sagittal section of the face displays the forehead, nose, nasal bone, upper and lower lips, and mandible can be used for screening microcephaly, mandibular hypoplasia, and nasal bone absence, among other conditions (12); and the standard oblique coronal section through the lower lip or mandible (showing the soft palate), displaying the soft palate and tongue, can be used to screen for soft palate cleft (13).

However, this identification and screening process is not only time-consuming but also extremely dependent on the sonographer’s work experience, which introduces various problems in clinical practice. First, this cumbersome examination process often relies on the sonographer’s identification and manual acquisition based on past work experience. This not only requires multiple repeated image acquisition steps in real-time work but may also affect the quality of fetal ultrasound examination or cause missed diagnoses if the display of these sections is not standard or if corresponding nonstandard sections are incorrectly identified as standard sections. Second, the ultrasound image quality control process is labor-intensive and time-consuming for prenatal ultrasound experts. Therefore, a method that can automatically complete the identification of standard ultrasound sections and has high accuracy is necessary to alleviate these problems and may also serve as a teaching tool to improve the ability of junior sonographers to recognize fetal standard sections.

AI has shown great potential in the classification, detection, and segmentation of prenatal ultrasound images (14-21). In this study, the AI model performed better than did both the midcareer and junior sonographers in classifying various sections of the fetal lip and palate, but this difference was more pronounced in the standard alveolar section and the standard oblique coronal section through the lower lip or mandible (showing the hard palate). Considering that these two sections are more commonly used in the III-level prenatal ultrasound examination—and not routinely retained sections for I- and II-level prenatal ultrasound examination—and that they are difficult to scan and display due to positional relationships, we speculate that they may be poorly understood by midcareer and junior sonographers. Compared with the classification of the other four sections, that of the standard oblique coronal section through the lower lip or mandible (showing the soft palate) by the AI model, midcareer sonographer, and junior sonographer was poorer according to various classification efficiency indicators. This may be related to the key anatomical structure of the soft palate, which is the critical structure for recognizing the standard oblique coronal section through the lower lip or mandible (showing the soft palate). The ultrasonic image of the soft palate is a slightly echogenic band formed by the mucosa of the soft palate, sandwiched between two low-echogenic bands, forming an equals sign (22). This anatomical structure is small and compared with the strong echo of the forehead, nasal bone, and hard palate plate in the standard midline sagittal section or the strong echo of the hard palate plate in the standard oblique coronal section through the lower lip or mandible (showing the hard palate), the image features of the soft palate are difficult to recognize. This may explain the performance of the AI model, midcareer sonographer, and junior sonographer in recognizing this section.

In terms of overall recognition, the AI demonstrated a classification efficiency higher than that of both the midcareer and junior physicians, and its classification time was also shorter than that of the midcareer and junior physicians, indicating that AI performs more quickly and efficiently in the classification of fetal lip and palate standard ultrasound sections than do midcareer and junior physicians. Thus, this model is worth promoting in the clinical practice, as it can assist in the recognition and acquisition of standard sections in prenatal examinations and alleviate the workload of examining physicians.

In this study, we investigated the clinical implications of implementing AI in prenatal ultrasound practice. The AI model demonstrated significant potential in improving workflow efficiency by automating the identification of standard ultrasound sections, thus reducing the workload of sonographers and saving time in busy clinical settings. Furthermore, AI can serve as a valuable training tool for junior sonographers, providing consistent and accurate guidance to enhance their recognition skills, especially for challenging sections. Despite these advantages, barriers such as implementation costs, resistance from practitioners, and the need for extensive validation across diverse clinical settings remain challenges to adoption. Addressing these issues in future research will be critical to ensuring the seamless integration of AI into routine prenatal care, ultimately enhancing efficiency and diagnostic accuracy.

Identifying standard ultrasound sections in fetal lip and palate remains a challenging task for both sonographers and AI models due to the changes in standard sections caused by different fetal positions and scanning orientations, as well as the presence of artifacts in ultrasound images (23,24). This study not only made it more difficult for AI to identify standard lip and palate sections by adding nonstandard sections of each lip and palate and background images but also made the model’s design better serve clinical practice, and thus our findings may be particularly valuable and innovative. The focus of prenatal ultrasound examinations assisted by AI technology includes both the detection of normal fetal growth and development and the screening of fetal structural abnormalities (25,26). This finding suggests that after successfully recognizing normal fetal lip and palate sections, the AI models developed in this study should be further trained and optimized to detect and classify fetal lip and palate abnormalities. With the advancement in technology, we believe that the intelligent diagnosis and classification of fetal lip and palate abnormalities via AI will be realized.

This study involved several notable limitations that should be considered in the interpretation of our findings. First, the use of a limited number of human interpreters (one per experience level) may reduce the generalizability of our comparisons, as performance could vary across a broader group of sonographers. Second, despite the substantial dataset, the full breadth of the anatomical variations or challenging cases might not have been captured, which could affect the AI’s performance in a diversity of clinical scenarios. Third, as a single-center study, the findings may not directly apply to institutions with different protocols or patient populations, highlighting the need for further multicenter validation. Fourth, our study only evaluated 2D ultrasound images, potentially limiting the application of findings to advanced three-dimensional imaging technologies. Finally, the focus of this study was on standard plane identification rather than the diagnosis of anomalies, which suggests that further research is needed to evaluate and enhance the AI’s capability to detect cleft lip and palate abnormalities. These limitations should be considered when interpreting the study’s conclusions and designing future research.


Conclusions

In this study, a YOLOv5-based AI model was used as an object detection network for the classification of fetal lip and palate standard sections. The results indicated that the classification efficiency of the AI model was greater than that of the junior and midcareer sonographers, and the classification process was fast and efficient. This AI model can be used as an auxiliary tool for obtaining standard sections of the fetal lip and palate and has certain clinical value in the quality evaluation of ultrasound images.


Acknowledgments

None.


Footnote

Reporting Checklist: The authors have completed the TRIPOD+AI reporting checklist. Available at https://qims.amegroups.com/article/view/10.21037/qims-24-1454/rc

Data Sharing Statement: Available at https://qims.amegroups.com/article/view/10.21037/qims-24-1454/dss

Funding: This study was supported by Commission of Science and Technology of Shenzhen (No. JCYJ20220530142002005).

Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://qims.amegroups.com/article/view/10.21037/qims-24-1454/coif). The authors have no conflicts of interest to declare.

Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. This study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments and was approved by the Institutional Ethics Committee of Huazhong University of Science and Technology Union Shenzhen Hospital (approval No. KY-2022-12-12). Informed consent was obtained from all individual participants.

Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.


References

  1. Wong FW, King NM. A review of the rate of occurrence of cleft lip and palate in Chinese people. Hong Kong Med J 1997;3:96-100.
  2. Divya K, Iyapparaja P, Raghavan A, Diwakar MP. Accuracy of Prenatal Ultrasound Scans for Screening Cleft Lip and Palate: A Systematic Review. J Med Ultrasound 2022;30:169-75. [Crossref] [PubMed]
  3. Maarse W, Bergé SJ, Pistorius L, van Barneveld T, Kon M, Breugem C, Mink van der Molen AB. Diagnostic accuracy of transabdominal ultrasound in detecting prenatal cleft lip and palate: a systematic review. Ultrasound Obstet Gynecol 2010;35:495-502. [Crossref] [PubMed]
  4. Lai GP, Weng XJ, Wang M, Tao ZF, Liao FH. Diagnostic Accuracy of Prenatal Fetal Ultrasound to Detect Cleft Palate in High-Risk Fetuses: A Systematic Review and Meta-Analysis. J Ultrasound Med 2022;41:605-14. [Crossref] [PubMed]
  5. LeCun Y, Bengio Y, Hinton G. Deep learning. Nature 2015;521:436-44. [Crossref] [PubMed]
  6. Litjens G, Kooi T, Bejnordi BE, Setio AAA, Ciompi F, Ghafoorian M, van der Laak JAWM, van Ginneken B, Sánchez CI. A survey on deep learning in medical image analysis. Med Image Anal 2017;42:60-88. [Crossref] [PubMed]
  7. Cheng JZ, Ni D, Chou YH, Qin J, Tiu CM, Chang YC, Huang CS, Shen D, Chen CM. Computer-Aided Diagnosis with Deep Learning Architecture: Applications to Breast Lesions in US Images and Pulmonary Nodules in CT Scans. Sci Rep 2016;6:24454. [Crossref] [PubMed]
  8. Han S, Kang HK, Jeong JY, Park MH, Kim W, Bang WC, Seong YK. A deep learning framework for supporting the classification of breast lesions in ultrasound images. Phys Med Biol 2017;62:7714-28. [Crossref] [PubMed]
  9. He F, Wang Y, Xiu Y, Zhang Y, Chen L. Artificial Intelligence in Prenatal Ultrasound Diagnosis. Front Med (Lausanne) 2021;8:729978. [Crossref] [PubMed]
  10. Ramirez Zegarra R, Ghi T. Use of artificial intelligence and deep learning in fetal ultrasound imaging. Ultrasound Obstet Gynecol 2023;62:185-94. [Crossref] [PubMed]
  11. Tonni G, Peixoto AB, Werner H, Grisolia G, Ruano R, Sepulveda F, Sepulveda W, Araujo Júnior E. Ultrasound and fetal magnetic resonance imaging: Clinical performance in the prenatal diagnosis of orofacial clefts and mandibular abnormalities. J Clin Ultrasound 2023;51:346-61. [Crossref] [PubMed]
  12. Salomon LJ, Alfirevic Z, Berghella V, Bilardo CM, Chalouhi GE, Da Silva Costa F, Hernandez-Andrade E, Malinger G, Munoz H, Paladini D, Prefumo F, Sotiriadis A, Toi A, Lee W. ISUOG Practice Guidelines (updated): performance of the routine mid-trimester fetal ultrasound scan. Ultrasound Obstet Gynecol 2022;59:840-56. [Crossref] [PubMed]
  13. Fuchs F, Grosjean F, Captier G, Faure JM. The 2D axial transverse views of the fetal face: A new technique to visualize the fetal hard palate; methodology description and feasibility. Prenat Diagn 2017;37:1353-9. [Crossref] [PubMed]
  14. Zhang L, Chen S, Chin CT, Wang T, Li S. Intelligent scanning: automated standard plane selection and biometric measurement of early gestational sac in routine ultrasound examination. Med Phys 2012;39:5015-27. [Crossref] [PubMed]
  15. Yang X, Yu L, Li S, Wen H, Luo D, Bian C, Qin J, Ni D, Heng PA. Towards Automated Semantic Segmentation in Prenatal Volumetric Ultrasound. IEEE Trans Med Imaging 2019;38:180-93. [Crossref] [PubMed]
  16. Ryou H, Yaqub M, Cavallaro A, Papageorghiou AT, Alison Noble J. Automated 3D ultrasound image analysis for first trimester assessment of fetal health. Phys Med Biol 2019;64:185010. [Crossref] [PubMed]
  17. Lee YB, Kim MJ, Kim MH. Robust border enhancement and detection for measurement of fetal nuchal translucency in ultrasound images. Med Biol Eng Comput 2007;45:1143-52. [Crossref] [PubMed]
  18. Deng Y, Wang Y, Chen P, Yu J. A hierarchical model for automatic nuchal translucency detection from ultrasound images. Comput Biol Med 2012;42:706-13. [Crossref] [PubMed]
  19. Nie S, Yu J, Chen P, Wang Y, Zhang JQ. Automatic Detection of Standard Sagittal Plane in the First Trimester of Pregnancy Using 3-D Ultrasound Data. Ultrasound Med Biol 2017;43:286-300. [Crossref] [PubMed]
  20. Li J, Wang Y, Lei B, Cheng JZ, Qin J, Wang T, Li S, Ni D. Automatic Fetal Head Circumference Measurement in Ultrasound Using Random Forest and Fast Ellipse Fitting. IEEE J Biomed Health Inform 2018;22:215-23. [Crossref] [PubMed]
  21. van den Heuvel TLA, Petros H, Santini S, de Korte CL, van Ginneken B. Automated Fetal Head Detection and Circumference Estimation from Free-Hand Ultrasound Sweeps Using Deep Learning in Resource-Limited Countries. Ultrasound Med Biol 2019;45:773-85. [Crossref] [PubMed]
  22. Frisova V, Cojocaru L, Turan S. A new two-dimensional sonographic approach to the assessment of the fetal hard and soft palates. J Clin Ultrasound 2021;49:8-11. [Crossref] [PubMed]
  23. Yu Zhen, Ni Dong, Chen Siping, Li Shengli, Wang Tianfu, Lei Baiying. Fetal facial standard plane recognition via very deep convolutional networks. Annu Int Conf IEEE Eng Med Biol Soc 2016;2016:627-30. [Crossref] [PubMed]
  24. Lei B, Tan EL, Chen S, Zhuo L, Li S, Ni D, Wang T. Automatic Recognition of Fetal Facial Standard Plane in Ultrasound Image via Fisher Vector. PLoS One 2015;10:e0121838. [Crossref] [PubMed]
  25. Lin M, He X, Guo H, He M, Zhang L, Xian J, Lei T, Xu Q, Zheng J, Feng J, Hao C, Yang Y, Wang N, Xie H. Use of real-time artificial intelligence in detection of abnormal image patterns in standard sonographic reference planes in screening for fetal intracranial malformations. Ultrasound Obstet Gynecol 2022;59:304-16. [Crossref] [PubMed]
  26. Gong Y, Zhang Y, Zhu H, Lv J, Cheng Q, Zhang H, He Y, Wang S. Fetal Congenital Heart Disease Echocardiogram Screening Based on DGACNN: Adversarial One-Class Classification Combined with Video Transfer Learning. IEEE Trans Med Imaging 2020;39:1206-22. [Crossref] [PubMed]
Cite this article as: Jiang W, Pang H, Deng S, Chen C. Artificial intelligence-based prenatal ultrasound for diagnosing fetal lip and palate during the second trimester. Quant Imaging Med Surg 2025;15(8):7497-7509. doi: 10.21037/qims-24-1454

Download Citation