3D reconstruction of the target tooth occlusal surface from a single-view image for root canal surgical navigation
Original Article

3D reconstruction of the target tooth occlusal surface from a single-view image for root canal surgical navigation

Guoliang Chen, Zhi Tang

School of Mechanical and Electronic Engineering, Wuhan University of Technology, Wuhan, China

Contributions: (I) Conception and design: Both authors; (II) Administrative support: G Chen; (III) Provision of study materials or patients: Both authors; (IV) Collection and assembly of data: Z Tang; (V) Data analysis and interpretation: Both authors; (VI) Manuscript writing: Both authors; (VII) Final approval of manuscript: Both authors.

Correspondence to: Guoliang Chen, PhD. School of Mechanical and Electronic Engineering, Wuhan University of Technology, No. 122 Luoshi Road, Hongshan District, Wuhan 430070, China. Email: glchen@whut.edu.cn.

Background: In robotic endodontic navigation, preoperative cone-beam computed tomography (CBCT) provides root canal information, whereas intraoperative vision usually provides only single-view tooth surface images. This study aimed to reconstruct the target tooth occlusal surface from a single visible-light image as a geometric interface for CBCT registration and root canal information mapping.

Methods: A statistical shape model (SSM) was constructed from 50 homologous tooth intraoral scan (IOS) models using curvature-adaptive non-rigid iterative closest point (NICP) registration. For each single-view image, the target tooth region was segmented, illumination interference was suppressed, and fissure features were enhanced. The extracted texture cues were converted into a bounded pseudo-height field and injected into the statistical template under top-surface gating and vertex stability weighting.

Results: On the occlusal region of interest (ROI), the proposed method achieved a root mean square distance (RMSD) of 0.432 mm, average symmetric surface distance (ASSD) of 0.308 mm, Hausdorff distance (HD) of 2.540 mm, Chamfer distance (CD) of 0.466 mm2, and Dice similarity coefficient (DSC) of 0.340. It outperformed the selected single-view comparison method in surface error metrics. Ablation experiments further showed that illumination normalization, stability weighting, and accurate segmentation contributed to improved reconstruction stability and detail recovery.

Conclusions: The proposed method reconstructs an anatomically constrained occlusal surface from a single-view image and provides a practical geometric interface for subsequent CBCT registration and root canal information mapping. It does not directly reconstruct the root canal anatomy or replace CBCT-based canal identification. Current validation remains limited to dental models, limited sample diversity, and intact or nearly intact occlusal morphology; altered clinical crowns require further evaluation.

Keywords: Dental contour extraction; 3D reconstruction; statistical shape model (SSM); cone-beam computed tomography (CBCT); surgical navigation


Submitted Mar 12, 2026. Accepted for publication Jul 16, 2026. Published online Aug 11, 2026.

doi: 10.21037/qims-2026-0602


Introduction

In robotic-assisted root canal surgery, the clinically relevant root canal trajectory is usually identified from preoperative cone-beam computed tomography (CBCT), whereas intraoperative visual sensing mainly provides two-dimensional visible-light images of the tooth surface. Therefore, the central problem is not to directly reconstruct the root canal structure from a photograph, but to construct a reliable three-dimensional (3D) tooth surface reference that can be registered with the tooth and root canal model derived from CBCT. After this registration, the planned access point and canal direction can be mapped into the robot coordinate system, as shown in Figure 1. Among the visible tooth regions, the occlusal surface is particularly suitable for reconstruction because cusps, grooves, and fissures provide rich geometric landmarks for image-to-CBCT alignment. However, 3D reconstruction in the intraoral environment remains challenging because of the confined workspace, restricted viewing direction, and optical phenomena such as specular reflection, diffuse reflection, and subsurface scattering on enamel and dentin surfaces (1). These factors weaken stable texture cues and may amplify local reconstruction errors.

Figure 1 Surgical procedure of the root canal treatment robot. 3D, three dimensional; CBCT, cone-beam computed tomography.

In recent years, advances in digital dentistry and dental artificial intelligence (AI) have further strengthened the roles of CBCT, intraoral scan (IOS), computer-aided design/computer-aided manufacturing (CAD/CAM), surgical navigation, and AI-based image interpretation in treatment planning and education (2-5). However, most existing dental AI studies have focused on disease diagnosis, image segmentation, restoration detection, educational assistance, or multi-view reconstruction, whereas relatively limited attention has been paid to low-cost single-view 3D reconstruction for intraoperative registration in endodontic navigation. This study focuses on a clinically motivated geometric interface: recovering the visible occlusal surface from a single-view image with sufficient geometric stability to support subsequent CBCT registration and root canal information mapping. However, altered occlusal surfaces caused by caries, restorations, fractures, or full-coverage crowns remain challenging for SSM-constrained reconstruction; therefore, the present study should be interpreted as a proof-of-concept investigation for cases with sufficiently visible occlusal landmarks.

In conventional vision-based approaches, Ali and colleagues (6) reconstructed sparse point clouds of the full dentition using photogrammetric principles, multi-view image matching, and the structure-from-motion (SfM) algorithm. However, such methods depend heavily on highly overlapping images and are susceptible to fragmented reconstruction due to saliva-induced specular highlights and weak surface texture. To reduce the number of required views, deep learning methods incorporating template priors have been proposed. Chen et al. (7) developed an automated framework that uses U-Net to extract contours from five standard orthodontic intraoral photographs and subsequently deforms a parametric statistical shape model (SSM) to fit the tooth boundaries. Similarly, Wirtz et al. (1) used five photographs and optimized a mean shape model to fit the segmented contours under the assumption of predefined fixed viewpoints. Although these methods reduce the number of input views, their contour-driven nature limits their ability to recover critical anatomical details, such as fissure depth. Mohamed et al. (8) further proposed an end-to-end deep learning framework in which EfficientNet was used to extract multi-view features, which were then fused by a long short-term memory (LSTM) network to regress a 3D voxel model. However, voxel-based representations generally involve a trade-off between resolution and computational cost. Moreover, the absence of explicit topological and local geometric constraints makes it difficult to represent high-frequency structures, such as sharp cusps and fine fissures, in a stable manner. Another line of research has focused on directly regressing 3D tooth morphology from a single dental image. Liang et al. (9) and Song et al. (10) employed graph convolutional and point-cloud generation networks, respectively, to directly regress 3D meshes or point clouds from a single panoramic radiograph. To further improve reconstruction accuracy, Li et al. (11) proposed the three-dimensional pathological eXpert (3DPX) framework, which adopts a hybrid multilayer perceptron (MLP)-convolutional neural network (CNN) architecture for progressive high-resolution reconstruction from two-dimensional X-ray images to 3D volumetric data. Nevertheless, such methods suffer from an inherent cross-modal information gap: X-ray images cannot capture the optical texture or color of the tooth surface, and the resulting models retain only statistical shape information, making them unsuitable as registration references for single-view visual navigation. Qian et al. (12) adopted a “physical stitching” strategy that rigidly registers a high-precision crown acquired by an IOS with the tooth root reconstructed from CBCT. Although this strategy achieves relatively high accuracy, it relies on additional expensive hardware, namely the IOS, thereby increasing both procedural time and equipment cost. Consequently, it does not satisfy the engineering requirements of root canal treatment for minimally invasive operation and simplified instrumentation. Dental-specific SSM studies further demonstrate that crown or surface information can be used to infer non-visible root-related structures for implant planning, but these studies also reveal that SSM inference introduces measurable uncertainty when unseen anatomy is estimated from surface scans (13,14). This observation is directly relevant to the present work: the proposed method uses the SSM as a constrained prior for the visible crown surface and does not claim patient-specific reconstruction of the root canal itself.

To address the common challenges faced by existing methods under single-view imaging, such as severe illumination interference and low-cost constraints, a dual-driven framework for occlusal surface reconstruction is proposed that integrates a single-view image with a tooth sample library. In the offline preparation stage, all homologous tooth samples in the library are first normalized in pose and scale to achieve canonical alignment. An improved non-rigid iterative closest point (NICP) algorithm is then employed to establish high-precision topological correspondences among the samples. Based on these correspondences, an SSM is constructed from the entire sample set as a prior model. In addition, allowable deformation weights and regions in the shape space are statistically estimated to define the global topological reference and the deformation range for reconstruction. In the real-time reconstruction stage, the segment anything model (SAM) is first used for interactive segmentation to extract the tooth mask and identify the region of interest (ROI) for reconstruction. Subsequently, photometric enhancement is performed within the ROI, and high-frequency texture features, such as cusps and fissures, are extracted and mapped to a geometric height field. Finally, under the constraints of top-surface gating and stability weights, controlled vertex deformation is achieved through energy minimization, thereby accurately injecting local details into the statistical template and generating a 3D occlusal surface model with realistic morphology and anatomically accurate features. The overall pipeline is illustrated in Figure 2, and the proposed framework provides an accurate geometric reference for subsequent CBCT registration and root canal mapping. The novelty of the proposed method lies in three aspects:

  • A curvature-adaptive NICP strategy is used to construct topology-consistent occlusal templates from homologous teeth;
  • Single-view fissure and cusp cues are converted into a bounded pseudo-height field rather than being treated only as 2D contours;
  • Anatomical gating and vertex stability weights constrain deformation to the occlusal region, thereby preserving sidewall stability while improving local surface detail.
Figure 2 Single-view image- and statistical template-based framework for 3D occlusal surface reconstruction. 3D, three dimensional; ROI, region of interest; SAM, segment anything model.

Compared with multi-view reconstruction, the method sacrifices completeness but reduces intraoperative imaging requirements; compared with generic image-to-3D models, it prioritizes metric stability and anatomical plausibility in a dental navigation scenario.


Methods

Template construction

To introduce stable geometric priors under single-view conditions with restricted viewpoints, a template was first established for each tooth type, and an IOS sample library was constructed accordingly for each tooth category. For simplicity, the two-digit tooth numbering system defined by the International Organization for Standardization (ISO) was adopted, as shown in Figure 3.

Figure 3 Tooth numbering and coordinate system definition.

To construct a statistical template from multiple homologous tooth samples, each sample in the library was first normalized within a standardized coordinate system. Following the method described in (7), an anatomically consistent right-handed coordinate system was established. Specifically, the centroid of the 14 teeth (excluding the wisdom teeth) was defined as the global origin, denoted by O. The vector from O to the midpoint A of the line connecting the centroids of the left and right central incisors was defined as the positive Z-axis. The vector CB, determined by the centroids B and C of the left and right second molars, was used to assist in defining the horizontal plane, and the positive directions of the X-axes and Y-axes were then determined according to the right-hand rule. Within this global reference frame, the origin of the local coordinate system for each tooth was set at its centroid, while the directions of the X-axes and Y-axes were aligned with those of the global coordinate system, as shown in Figure 3.

When constructing the homologous tooth sample library from segmented IOS data, all samples of the same tooth type were first aligned according to the coordinate system defined above to ensure a unified pose across samples. To eliminate size variations, the axis-aligned bounding box (AABB) of each sample was then computed in the unified coordinate system. Let the side lengths of the AABB along the x-axes, y-axes, and z-axes be denoted by Lx=xmaxxmin, Ly=ymaxymin, Lz=zmaxzmin. The corresponding diagonal length of the bounding box, denoted by L, is given by:

L=Lx2+Ly2+Lz2

Using the average diagonal length of the sample library, denoted by L0, as the reference scale, each original homologous tooth sample ti was isotropically scaled by a factor of s=L0/L, yielding the scale-normalized sample vi. Through this normalization procedure, which consists of pose alignment and scale normalization, homologous tooth samples from different individuals can be represented within a unified geometric reference frame and under a common size standard. This provides a consistent data foundation for subsequent topology-consistent registration and SSM.

To focus subsequent topology-consistent registration and detail reconstruction on the crown top-surface region most relevant to root canal navigation, the occlusal subset of sample ti was extracted along the Z-axis of the unified coordinate system. Specifically, the threshold Zth was defined as the Z-axis height at which the tooth cross-section formed a closed contour intersecting the Z-axis at only one point. Each tooth sample was then uniformly sampled into 20,000 points, and those with heights greater than or equal to Zth were retained as the candidate occlusal point-cloud region Pocc:

Pocc={pvipzzth}

Here, pz denotes the Z-coordinate of point p. This operation retains the upper portion of the point cloud with relatively high Z-values, as shown in Figure 4, thereby reducing interference from the sidewall and cervical regions in subsequent registration and SSM.

Figure 4 Sample normalization process. 3D, three dimensional.

(I) Original input scan with arbitrary orientation. (II) Pose and scale normalization based on the coordinate axes and bounding box. (III) Extraction of the ROI using a cropping threshold. (IV) Final normalized occlusal point set used for registration.

Due to issues such as uneven point density, holes, and topological inconsistencies after converting the IOS mesh into a point cloud, direct use for statistical modeling may result in missing point correspondences and distorted shape statistics. To address these problems, farthest point sampling (FPS) was employed to uniformly sample each Pocc into a point set P containing M=2000 points. Surface normals were then estimated for P, and a regularized triangular mesh was reconstructed using the Poisson reconstruction method (15).

To construct the topology-consistent training set with point-to-point correspondences required for SSM learning, an occlusal surface sample was first randomly selected from the sample library as the initial reference mesh Sref. An improved NICP algorithm was then used to register the reference mesh to the remaining samples. To mitigate the topological distortions caused by conventional NICP (16) in geometrically complex regions, an adaptive stiffness constraint based on surface curvature was introduced, and the resulting optimization objective E(X) was defined as follows:

E(X)=i=1Mdist2(Xivi,Pk)+γ{i,j}EwijXiXjF2

The energy function E(X) consists of a data term and an adaptive rigidity regularization term. Here, Xivi denotes the transformed position of vertex vi in Sref, where vi is obtained through resampling of the reference sample. Xi3×4 represents an independent affine transformation matrix associated with vertex vi, which is used to describe its local deformation. Pk denotes the target sample point cloud, and γ is a global stiffness coefficient that controls the smoothness of the deformation. ε denotes the set of mesh edges, representing all connected vertex pairs in the reference mesh, and F is the Frobenius norm used to quantify the difference between the transformation matrices of adjacent vertices. The improved weight wij is an adaptive factor based on surface curvature. By increasing the penalty in high-curvature regions, such as fissures and cusps, it enforces locally rigid deformation and thereby preserves the topology of anatomical tooth features. It is defined as follows:

wij=exp(βki+kj2)

Here, β>0 is a weighting coefficient, while ki and kj represent the surface curvatures at vertices vi and vj, respectively.

After the energy function converged, the reference mesh Sref was endowed with the geometric morphology of each sample while preserving the original vertex indices and connectivity. By extracting the vertex displacement vector of each sample relative to the reference template, δi(k)=Xivivi, a training set {Skreg}k=1K with strict point-to-point correspondence and consistent topology was obtained. On this basis, a SSM was constructed. Specifically, the vertex coordinates of each sample were concatenated into a high-dimensional vector xk, and principal component analysis (PCA) was applied to the sample set to compress the statistical variation in the high-dimensional space into an orthogonal basis matrix U composed of eigenvectors. Accordingly, an arbitrary tooth shape x(α) can be expressed as follows:

x(α)=μ+Uα

By adjusting the shape coefficients α, the SSM deforms the mean template μ along specific anatomical variation modes U, thereby representing individual morphological variability in a low-dimensional parameter space, as shown in Figure 5. To further quantify the stability of different anatomical regions, the standard deviation of vertex displacement across samples, denoted by σi, was computed as follows:

σi=1Kk=1Kvi(k)v¯i2

Figure 5 Construction of the statistical shape model and vertex displacement standard deviation. PCA, principal component analysis.

Here, K denotes the total number of samples in the training set, vi(k) represents the spatial coordinates of the iii-th vertex in the k-th topologically consistent sample, and v¯i corresponds to the vertex with the same index on the mean shape μ. The parameter σi characterizes the morphological variability of structures such as cusps, ridges, and fissures, thereby effectively compressing complex mesh deformations into a low-dimensional shape space. This representation not only achieves substantial dimensionality reduction but also incorporates a spatial stability prior, providing a well-constrained solution space for subsequent high-precision reconstruction under limited-view conditions.

Surgical tooth segmentation

The key to real-time reconstruction lies in accurately aligning the statistical shape prior provided by the SSM with the two-dimensional dental features observed from a single view. To this end, an interactive segmentation framework based on the SAM (17) was constructed to precisely extract the target tooth region from a single-view image. SAM is a large-scale vision foundation model built on the Transformer architecture, whose main advantages lie in its strong zero-shot generalization capability and flexible prompt engineering mechanism. The model consists of an image encoder, a prompt encoder, and a mask decoder, and it integrates image features with interactive prompts through a bidirectional cross-attention mechanism, as shown in Figure 6.

Figure 6 SAM model framework. SAM, segment anything model.

Based on this architecture, an interactive segmentation pipeline for single-view images was developed, comprising three stages: image feature encoding, human-computer prompt encoding, and rule-constrained mask generation, as illustrated in Figure 7. First, the image encoder Eϕ maps the input dental image IH×W×C to a low-resolution, high-dimensional feature embedding z, where H×W and C denote the image resolution and number of channels, respectively. The embedding z encodes the global topological structure of the dentition together with local texture details, and is computed as follows:

z=Eϕ(I)h×w×d

Figure 7 Tooth segmentation and surgical tooth selection.

Here, h×w denotes the spatial dimensions of the feature map, and d denotes the channel dimension. The operator then provides a single-point click prompt P within the target crown region to localize the target tooth. This prompt is encoded by Eϕ into a positional embedding vector ep, which is fused with the image feature z by the mask decoder Dθ through a bidirectional cross-attention mechanism. The decoder outputs a set of candidate masks, M={m1,m2,m3}, together with their predicted IoU-based confidence scores. To resolve ambiguity, the optimal target instance is automatically selected according to the IoU confidence and the prompt-point inclusion constraint (18), yielding the tooth mask SSS:

S^=argmax(Score(mk)I(pmk))

Here, Score(mk) denotes the model-predicted IoU score for the candidate mask mk, and I() is the indicator function. For mkM, I() takes the value 1 when the mask mkm_kmk spatially contains the prompt point p, and 0 otherwise. Finally, the selected mask S^ is thresholded to obtain a binary mask, whose pixel-wise mathematical representation is given as follows:

S^i,j={1,ifSigmoid(S^i,j)>τ0,otherwise

Here, (i, j) denotes the pixel coordinates, and τ is the binarization threshold. This mask precisely delineates the anatomical boundary of the target tooth, and the mask generation process is illustrated in Figure 7. This boundary is then used as a strong spatial prior to constrain the shape coefficients α and the rigid transformation parameters defined in section “Template construction”, thereby enabling optimal fitting between image features and the statistical prior during the 3D reconstruction process.

Image feature enhancement and height field

To stably extract local priors from a single-view image for geometric refinement, occlusal surface textures were enhanced under the constraint of the tooth mask M, and a fissure intensity map together with a pseudo-height field was constructed, as shown in Figure 8.

Figure 8 Comparison of the original image, highlight-removed image, and fissure intensity map.

To eliminate the interference caused by specular reflections from saliva and uneven illumination in the oral cavity, highlight removal and normalization techniques were first applied to suppress non-structural bright spots (19). Subsequently, contrast enhancement and multi-scale gradient operators were used to extract fissure and cusp features, thereby constructing the intensity map V.

To prevent differences in dynamic range and semantic inversion across images from adversely affecting subsequent reconstruction, V was first processed to enforce semantic consistency and then mapped to a stable range using percentile-based robust normalization. In conjunction with the binary mask S^i,j obtained in section “Surgical tooth segmentation”, the enhanced result was further constrained to the tooth region, thereby generating a height field H that can directly drive controlled mesh displacement. This height field serves as a local prior for fine-grained details, providing stable input for subsequent template alignment and occlusal surface refinement. To convert the two-dimensional features into 3D geometric constraints, the enhanced fissure intensity map Ivalley was transformed at pixel coordinate (u, v) into a pseudo-height field through a nonlinear mapping function f(), as follows:

Himg(u,v)=f(Ivalley(u,v))

Here, (u, v) denote the pixel coordinates, Ivalley represents the fissure intensity distribution after enhancement and mask-based constraint, and f() denotes a nonlinear mapping function used to convert pixel intensity into an equivalent geometric depth offset.

In the implementation, the nonlinear mapping was defined as:

f(I)=dmaxtanh(αInorm)

where Inorm denotes the robustly normalized fissure intensity in the range [0, 1], α was set to 1.5, and dmax was set to 0.60 mm according to the observed occlusal relief range of the IOS training samples. The mapping therefore does not infer absolute depth from image brightness alone; instead, it provides a bounded local displacement cue that is further constrained by the SSM, top-surface gating, and vertex stability weights.

Image-template alignment and occlusal surface reconstruction

Establishing a spatial mapping between image pixels and the 3D template is fundamental to detail injection. The reconstruction process is divided into three stages: spatial alignment, height-field-driven deformation, and global optimization.

First, the mean shape template μ constructed in section “Template construction” was initially aligned with the single-view image. The purpose of this alignment was to optimize the similarity transformation parameters T={R, t, s}, such that the projected boundary of the 3D template maximally overlapped with the boundary of the two-dimensional mask. A multi-scale grid search was then performed to discretely sample the rotation, translation, and scale parameter space. For each sampled pose, the intersection over union (IoU) between the projected template boundary and the two-dimensional mask S^ was computed, and the parameters yielding the maximum overlap were selected as the initial transformation Tinit (20). Subsequently, edge pixels were extracted from S^ to construct a Euclidean distance field D(u, v). Using the template boundary vertex set Vedge as a constraint, the transformation parameters were further optimized by minimizing the energy of the projected model boundary in the distance field, as follows:

Ealign(T)=iVedgeD(Project(vi;T,K))2

Here, K denotes the camera intrinsic matrix. The transformation parameters T were solved using the Gauss-Newton iterative method, while the scale factor s was simultaneously optimized to compensate for inter-individual scale variations. After obtaining the optimal transformation T*={R*, t*, s*}, each 3D vertex vi was projected onto a unique pixel coordinate, (ui,vi)=Project(vi;T*,K), thereby enabling the two-dimensional height field Himg(u,v) to be accurately associated with the vertices of the 3D template for subsequent normal-based detail injection. After image-template alignment, each vertex of the 3D template was projected onto the image plane, and the corresponding pseudo-height value was sampled from the two-dimensional height field to guide normal-direction displacement. The mapping relationship is illustrated in Figure 9.

Figure 9 Mapping relationship between the two-dimensional pseudo-height field and the three-dimensional template vertices.

After alignment, the two-dimensional height field Himg constructed in section “Image feature enhancement and height field” was sampled along the projection path onto the template mesh vertices to compute the local displacement magnitude ∆i. To ensure that deformation was applied only to the target anatomical region while preserving the stability of the lateral wall structure, a top-surface gating function giwas first introduced, allowing only vertices located in the occlusal region to participate in the deformation update.

gi={1,if(vi,z>zth)&(ninup>cosθ)0,otherwise

Here, vi,z denotes the z-coordinate of the i-th vertex, which represents the height in the unified coordinate system; zth is the height threshold for the top surface; ni is the unit normal vector of the i-th vertex on the template; nupis the unit vector along the positive Z-axis; and θ denotes the maximum allowable deviation angle. On this basis, a vertex weight wi was introduced and constructed from the cross-sample displacement standard deviation σi calculated by Eq. [6] in section “Template construction” to characterize the deformability of different regions. Regions with high anatomical consistency were assigned higher confidence, thereby enabling detailed fitting guided by anatomical priors. The final vertex update can be written as follows:

vi=vi+giwiΔini

To reduce the bias caused by scale differences among different tooth types, the stability weight was normalized separately for each tooth type. Specifically, the cross-sample vertex displacement standard deviation σiwas linearly mapped to the range of [0, 1] within the corresponding homologous tooth library. The vertex stability weight was then computed as:

wi=expσimedian(σ)

where median(σ) denotes the median value of the normalized cross-sample displacement standard deviations within the corresponding tooth library. In this way, anatomically stable regions with smaller inter-sample variations are assigned higher confidence during deformation. For teeth with atypical morphology, such as severe wear, restorations, caries, or fractures, this prior may suppress patient-specific deformation. Therefore, such samples were excluded from the current template-building dataset and are further discussed as a limitation.

This hierarchical constraint strategy achieves a balance between targeted refinement of the occlusal surface (21) and preservation of the lateral wall structure. The reconstruction problem is formulated as a coarse-to-fine energy minimization process, aiming to determine the optimal shape parameters α* and pose parameters T* such that the projected parametric model remains consistent with the image features extracted in section “Image feature enhancement and height field”. To constrain local deformation using the height field, a depth-consistency term is introduced:

Eheight=wizih(ui)2

Here, zi denotes the depth of the transformed vertex, and h(ui) represents the corresponding height-field value. The total energy function Etotalis defined as follows:

Etotal(α,T)=Eproj+λ1Eheight+λ2Ereg

Here, the projection-overlap term Eproj constrains the agreement between the projected contour of the model and the image mask, while Ereg denotes the shape regularization term. Through joint optimization under multiple constraints, this process enables the precise incorporation of local geometric details from the height field into the statistical template. Once the energy function converges, the system outputs a high-fidelity 3D occlusal surface mesh that integrates global topological stability with subject-specific local features, thereby providing an accurate geometric reference for subsequent root canal surgery navigation.

Ethical statement

The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. This study used anonymized dental models and IOS data for methodological validation. No identifiable patient information is presented in the manuscript. If additional clinical or in-mouth images are included in future work, ethical approval and informed consent will be obtained according to institutional requirements.


Results

The closed-loop evaluation framework for assessing accuracy, module contribution, and adaptability is illustrated in Figure 10. In the experiments, single-view images and high-precision IOS ground-truth models were acquired simultaneously, and adaptability was further validated through multi-view perturbations and cross-subject samples. In the core processing stage, ablation experiments were conducted to quantify the contributions of highlight removal, stability weighting, and the segmentation algorithm to suppressing geometric distortion. Subsequently, the reconstructed model was globally aligned with the ground-truth model using the scale factor s and the iterative closest point (ICP) algorithm. The occlusal surface ROI was then extracted based on the top-surface quantile threshold ztb, ensuring that the evaluation focused on the core anatomical functional region. Reconstruction quality was evaluated by computing the bidirectional surface distance distribution and using metrics such as the root mean square distance (RMSD), average symmetric surface distance (ASSD), Hausdorff distance (HD), Chamfer distance (CD), and Dice similarity coefficient (DSC), thereby establishing a geometric reference for root canal navigation.

Figure 10 Framework of the experimental workflow. 3D, three dimensional; ASSD, average symmetric surface distance; CD, Chamfer distance; DSC, Dice similarity coefficient; HD, Hausdorff distance; ICP, iterative closest point; IOS, intraoral scan; RMSD, root mean square distance; ROI, region of interest; SAM, segment anything model.

Experimental setup

The experimental platform consisted of a five-axis precision motion stage (HUAVEA31-60C and HUAVET23-60L), a PC-based image processing system, and a dental model, as shown in Figure 11. The single-view camera used in the experiments was an OV9734 endoscopic camera with a focal length of 22 mm and a maximum resolution of 1,280×720 pixels. The precision motion stage enabled angular adjustment with an accuracy of 0.02°. A sample library was constructed using 50 homologous tooth IOS models. The current dataset was mainly used for proof-of-concept validation under intact or nearly intact occlusal morphology, rather than for covering all clinical endodontic scenarios. To avoid introducing non-physiological deformation into the SSM, teeth with extensive restorations, severe fractures, large caries-related defects, full-coverage crowns, or indistinct occlusal morphology were excluded from template construction. All IOS meshes were anonymized, converted into point clouds, and normalized in pose and scale before correspondence construction. Subsequently, single-view images from different individuals were selected as reconstruction inputs for tooth segmentation, feature enhancement, and height-field generation.

Figure 11 Experimental platform.

Because the present method relies on an SSM prior and a visible occlusal texture-to-height cue, altered occlusal surfaces are expected to affect reconstruction in two ways: missing or filled fissures reduce the reliability of the pseudo-height field, while large defects or restorations may pull the image-template alignment toward an idealized mean anatomy. Accordingly, the current version should be interpreted as a navigation-interface method for cases with sufficiently visible occlusal landmarks, and not as a complete solution for severely destroyed crowns. Validation on clinically altered crowns is reserved for future work and is explicitly discussed as a limitation.

For segmentation, a single prompt point was placed inside the visible target crown region. The generated SAM masks were visually checked by two operators with experience in dental image processing. When the selected mask included adjacent teeth or gingival regions, the prompt was adjusted and the mask was regenerated. No manual boundary correction was used in the quantitative evaluation, so that the effect of segmentation quality on reconstruction could be assessed consistently. Images were resized to the camera resolution, undistorted using the calibrated intrinsic parameters, and processed using the same highlight removal, illumination normalization, gradient enhancement, and robust percentile normalization settings across all experiments.

The 50-sample library was not intended to represent the full morphological diversity of all ISO posterior tooth positions. Rather, it was used to construct proof-of-concept posterior occlusal templates under a unified preprocessing and correspondence workflow. Therefore, the generalization results should be interpreted as preliminary evidence under limited sample diversity. Larger tooth-position-specific libraries, including separate maxillary/mandibular premolar and molar templates, are required before broad population-level deployment.

The experiments strictly followed the reconstruction procedure described in section “Methods”, and all computations were performed within a unified coordinate system to reduce spatial deviations. Before data acquisition, the camera was calibrated using a planar calibration board, and fixed camera intrinsic parameters together with radial-tangential distortion coefficients were used throughout all experiments, as listed in Table 1. Image undistortion was performed prior to segmentation and template projection. During the experiments, the extrinsic transformation between the camera and the motion stage was kept fixed, while view perturbation experiments were generated through controlled rotation of the precision motion stage. Both template construction and reconstruction were conducted within the same coordinate frame, thereby preventing inter-sample differences in scale and orientation from affecting the final results.

Table 1

Local camera parameters

Parameter Transformation matrix
Local camera intrinsic matrix [719.350621.850718.53355.05001]
Distortion coefficients [−0.1965, −0.0352, 0.0023, −0.0039, 0.0324]

To achieve a balance between reconstruction detail and geometric stability, the key experimental parameters were uniformly specified as follows:

  • Top-surface percentile threshold zth: used to identify the occlusal surface region. It was defined as a fixed percentile threshold of template height along the superior direction in the coordinate system established in section “Template construction”, and only vertices above this threshold were included in the deformation.
  • Maximum displacement constraint dmax: used to bound the magnitude of height-field-driven vertex displacement, thereby preventing abnormal deformation caused by local noise or intense specular reflections.
  • Poisson reconstruction depth: used to standardize mesh topology and quality while ensuring recoverability of fissure structures and suppressing high-frequency noise.
  • Stability weight coefficient β: used to regulate reliance on the statistical prior and prevent non-physiological deformation induced by local noise.

These parameters were held constant throughout the main experiments to assess the stability and reproducibility of the proposed method across different sample conditions. The exact values used in the implementation are summarized in Table 2.

Table 2

Key hyperparameters and evaluation threshold used in the proposed reconstruction pipeline

Parameter Meaning Value range used
λ1 Projection-overlap term weight 1.00
λ2 Shape regularization term weight 0.08
β Curvature-adaptive stiffness coefficient 2.50
z_th Top-surface height percentile 70th percentile of template height
θ Maximum surface-normal deviation for top gating 45°
d_max Maximum height-field-driven displacement 0.60 mm
Poisson depth Poisson reconstruction depth 8
DSC threshold τ Distance tolerance for surface overlap 0.50 mm

DSC, Dice similarity coefficient.

Geometric error evaluation

The reconstructed model was spatially aligned with the IOS reference using automatic scaling, similarity transformation, and ICP. The occlusal surface point cloud was then extracted using the threshold zth for deviation analysis. Within the ROI, symmetric surface distances were computed between the two models, and reconstruction performance was evaluated using the following metrics derived from the bidirectional point-to-surface distances between the two surfaces:

  • Symmetric root mean square surface distance (RMSD): quantifies the overall deviation level and is sensitive to local extreme errors.

    RMSD(S,T)=12(1NSxSd(x,T)2+1NTyTd(y,S)2) Here, S and T denote the predicted and ground-truth sampled point sets within the ROI, respectively; xS and yT are sampled points; NS and NT are the numbers of points in S and T, respectively; and d(x, T) is the shortest Euclidean distance from point x to surface T.

  • ASSD: quantifies the average agreement between the two surfaces within the ROI.

    ASSD(S,T)=12(1NSxSd(x,T)+1NTyTd(y,S))

  • Symmetric HD: quantifies the maximum surface deviation under the worst-case condition.

    HD(S,T)=max(h(S,T),h(T,S)) Here, h(S, T) denotes the directed Hausdorff distance from S to T.

  • Symmetric CD: quantifies the overall bidirectional surface discrepancy.

    CD(S,T)=1NSxSd(x,T)2+1NTyTd(y,S)2

  • DSC: quantifies surface overlap consistency under a distance threshold.

    DSCτ=2PτRτPτ+Rτ

Here, τ denotes the distance threshold, while the two matching ratios represent the proportions of predicted and ground-truth surface points within this tolerance, respectively. In the main experiments, τ was set to 0.50 mm, which is close to the submillimeter tolerance expected for occlusal registration in a robotic navigation workflow. To evaluate the influence of this threshold, DSC was additionally recalculated at 0.30, 0.50, and 0.70 mm in the sensitivity analysis.

For fairness, all baseline methods were restricted to the same single-view input image. The method of (III) was implemented as a contour-driven parametric SSM fitting baseline without texture-to-height detail injection. The two large-scale image-to-3D baselines were instantiated using their publicly released inference pipelines or official demo settings, with the dental image used as the only visual input and no additional depth, multi-view, or IOS information provided. Their outputs were converted to meshes, aligned to the IOS reference using the same similarity transformation and ICP procedure, and evaluated within the same occlusal ROI. The results are shown in Table 3.

Table 3

Quantitative comparison of 3D reconstruction methods

Method RMSD (mm) ASSD (mm) HD (mm) CD (mm2) DSC
Statistical shape modeling
   Our 0.432117 0.307633 2.539716 0.465623 0.340257
   Method of (7) 0.526707 0.406348 3.318122 0.691364 0.252467
Large-scale model
   Method of (22) 0.736582 0.488608 3.785423 1.213353 0.270402
   Method of (23) 0.694209 0.456787 3.719028 1.073703 0.280457

The arrows indicate the change direction of the values of our method compared with the comparison methods: ↓, reduced; ↑, increased. 3D, three dimensional; ASSD, average symmetric surface distance; CD, Chamfer distance; DSC, Dice similarity coefficient; HD, Hausdorff distance; RMSD, root mean square distance.

As shown in Table 3, the proposed method outperformed existing methods across all evaluation metrics. It achieved the lowest RMSD and ASSD values, indicating superior reconstruction accuracy and geometric consistency. By contrast, general-purpose large models were more prone to producing meshes that appeared visually plausible yet suffered from geometric drift under intraoral conditions involving strong specular reflections, weak texture, and limited viewpoints, as shown in Figure 12.

Figure 12 Qualitative comparison of the reconstructed 3D tooth with other models in terms of shape completeness and geometric detail (3,15,16). 3D, three dimensional.

Ablation experiments

The ablation experiments showed that removing illumination normalization distorted the driving signal, causing pseudo-textures induced by saliva reflections to be incorrectly mapped into the geometric height field. This led to anatomically inaccurate deformation and significant increases in HD and CD, indicating that illumination preprocessing is essential for reliable detail recovery in single-view reconstruction. Removing the vertex-weight constraint caused deformation leakage, whereby sidewalls and boundaries absorbed more errors in the absence of stability constraints, resulting in greater extreme deviations and lower coverage. Comparative experiments using UNet-ASPP instead of SAM further demonstrated that ROI segmentation quality directly affects registration stability. Blurred occlusal boundaries introduced sidewall points into the evaluation region, broadening the error distribution and reducing the matching ratio. Overall, illumination normalization, stability weighting, and high-quality segmentation together ensured the geometric stability and anatomical fidelity of the reconstructed results.

Figure 13 presents the quantitative results of three ablation settings and the reference scheme on the occlusal ROI. Removing any key component resulted in a significant increase in error, and the degradation was not confined to a single average-error metric. Instead, RMSD, ASSD, HD, and CD all deteriorated simultaneously, indicating that these modules do not simply improve visual appearance, but are essential for maintaining the geometric usability of the reconstruction. Overall, the proposed method provides a more balanced performance in overall fitting accuracy, local extreme-error control, and within-tolerance coverage.

Figure 13 Quantitative ablation results for 3D tooth reconstruction. 1, removal of illumination normalization; 2, removal of weight constraints; 3, replacing SAM with Unet; 4, ours. 3D, three dimensional; ASSD, average symmetric surface distance; CD, Chamfer distance; DSC, Dice similarity coefficient; HD, Hausdorff distance; RMSD, root mean square distance; SAM, segment anything model.

Generalization validation

To verify adaptability under view perturbation, images of the target tooth from different viewpoints and homologous tooth images from different individuals were collected as independent single-view inputs. The current online reconstruction pipeline does not natively perform multi-view fusion or joint multi-view optimization. The initial position of the camera in the world coordinate system was [0, 20, 35]. View 2 was obtained by rotating the precision stage clockwise by 5 degrees about the X-, Z-, and Y-axes, respectively. The corresponding 4×4 homogeneous transformation matrix is given by:

T=[0.9924030.0940890.0792560.8921940.0871550.9924030.08682422.8869200.0868240.0792560.99306533.1721690001]

Using a unified parameter setting, each view was reconstructed independently to evaluate geometric consistency of the occlusal surface and the suppression of abnormal deformation, as shown in Figure 14.

Figure 14 Validation of model adaptability.

Model adaptability was evaluated through independent view-perturbation, cross-individual, and cross-tooth-position tests, with particular emphasis on the repeatability of occlusal surface details and the physiological plausibility of sidewall morphology. As shown in Table 4, relatively consistent error levels were maintained under cross-individual conditions, indicating that the statistical shape prior and explicit constraints can accommodate inter-individual variation to some extent. The corresponding visual results are presented in Figure 14. By constraining deformation within an allowable range, the proposed method maintains a relatively consistent geometric reference despite individual differences, thereby providing a more reliable basis for root canal structure mapping and subsequent navigation.

Table 4

Comparison of occlusal surface reconstruction accuracy across views, individuals, and tooth positions

Input image RMSD (mm) ASSD (mm) HD (mm) CD (mm2) DSC
Subject 1, tooth #26, view 1 0.432117 0.307633 2.539716 0.465623 0.340257
Subject 1, tooth #26, view 2 0.496425 0.297164 2.975975 0.605744 0.420969
Subject 2, tooth #26 0.466792 0.265467 2.890592 0.533729 0.498739
Subject 1, tooth #45 0.505066 0.286491 2.802227 0.635408 0.333443

ASSD, average symmetric surface distance; CD, Chamfer distance; DSC, Dice similarity coefficient; HD, Hausdorff distance; RMSD, root mean square distance.


Discussion

To address the lack of a reliable geometric reference for registration between visible-light images and CBCT in root canal treatment navigation, this study proposed a single-view occlusal surface reconstruction method. The reconstructed object should be interpreted as the visible crown/occlusal surface used for registration, not as the root canal anatomy itself. Compared with conventional multi-view geometric approaches, the proposed method reduces the dependence on complete intraoral image coverage and enhances fissure observability through illumination preprocessing and bounded height-field injection. Compared with purely generative image-to-3D approaches, the SSM-based constraint improves metric stability in a dental-specific setting. However, this benefit is obtained by prioritizing anatomical plausibility and deformation stability over unrestricted patient-specific completion.

The main limitation arises from the same SSM prior that makes the method stable. Regions that are not visible or are structurally altered in the single-view image are inferred mainly from the mean shape and deformation modes of the homologous tooth library. This can introduce systematic bias in cases with severe wear, restorations, caries, fractures, full-coverage crowns, developmental anomalies, or atypical cusp-fissure morphology. In such cases, the SSM may over-regularize the reconstructed surface toward an idealized crown, suppress patient-specific defects, or reduce the convergence stability of image-template alignment. Therefore, the current method is most appropriate when sufficient occlusal landmarks remain visible. It should not be used alone for severely destroyed crowns without additional imaging, scanner data, or defect-aware template adaptation.

The limited number of IOS samples also constrains the morphological coverage of the SSM. Although the current library was adequate for preliminary methodological validation, it is insufficient to represent the full population variability across all posterior tooth positions, age-related wear patterns, restorations, and pathological morphologies. Future work will expand the library in a tooth-position-specific manner and evaluate whether separate templates or mixture priors improve robustness for altered clinical crowns.

Future work can be advanced along four directions. First, multi-view fusion can be introduced to enforce cross-view consistency, compensate for invisible regions, and improve the reliability of the height field. Second, end-to-end learnable initialization and deformation modules can replace part of the current heuristic search while retaining anatomical constraints. Third, occlusal ROI detection, image-quality assessment, and uncertainty estimation should be automated to reduce operator interaction and improve reproducibility. Fourth, additional validation should be performed using realistic in-mouth images with saliva, blood, soft-tissue occlusion, restricted viewing angles, and patient motion, followed by CBCT registration and access-targeting experiments using clinically meaningful FRE, TRE, angular error, and safety-margin indicators.


Conclusions

This study proposed a single-view occlusal surface reconstruction method that integrates statistical shape priors with local texture-to-height detail injection. The method reconstructs an anatomically constrained 3D occlusal surface from one visible-light image and uses the reconstructed surface as a geometric interface for subsequent CBCT registration and root canal information mapping. Quantitative evaluation showed improved RMSD, ASSD, HD, and CD compared with selected single-view baselines, and ablation experiments confirmed the contributions of illumination normalization, segmentation quality, and stability weighting. The clinical value of the method lies in reducing the intraoperative sensing burden for robotic endodontic navigation while preserving a surface reference that can be aligned with preoperative CBCT. However, the approach does not replace CBCT-based canal identification, and its current evidence remains based on dental models, limited sample diversity, and crowns with intact or nearly intact occlusal morphology. Future studies should include larger tooth-position-specific datasets, altered clinical crowns, in-mouth imaging, downstream CBCT registration, and access-targeting accuracy tests before clinical deployment.


Acknowledgments

None.


Footnote

Data Sharing Statement: Available at https://qims.amegroups.com/article/view/10.21037/qims-2026-0602/dss

Funding: None.

Conflicts of Interest: Both authors have completed the ICMJE uniform disclosure form (available at https://qims.amegroups.com/article/view/10.21037/qims-2026-0602/coif). The authors have no conflicts of interest to declare.

Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments.

Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.


References

  1. Wirtz A, Jung F, Noll M, Wang A, Wesarg S. Automatic model-based 3-D reconstruction of the teeth from five photographs with predefined viewing directions. In: Landman BA, Išgum I, editors. Medical Imaging 2021: Image Processing. SPIE; 2021:11601-21.
  2. Sirirat W, Rewthamrongsris P, Bespinyowong K, Khurshid Z, Samaranayake L, Kaewkamnerdpong I, Osathanon T. A Comparative Study of Generative Artificial Intelligence Tools for Human Bone Learning. Clin Anat 2026;39:636-44. [Crossref] [PubMed]
  3. Khurshid Z, Faridoon F, Silkosessak OC, Trachoo V, Waqas M, Hasan S, Porntaveetus T. Multi-Regional deep learning models for identifying dental restorations and prosthesis in panoramic radiographs. BMC Oral Health 2025;25:1750. [Crossref] [PubMed]
  4. Khurshid Z. Digital Dentistry: Transformation of Oral Health and Dental Education with Technology. Eur J Dent 2023;17:943-4. [Crossref] [PubMed]
  5. Khurshid Z, Osathanon T, Shire MA, Schwendicke F, Samaranayake L. Artificial Intelligence in Dentistry: A Concise Review of Reporting Checklists and Guidelines. Int Dent J 2026;76:109322. [Crossref] [PubMed]
  6. Ali FI, Tarik Al-dahan Z. Teeth model reconstruction based on multiple view image capture. IOP Conf Ser: Mater Sci Eng 2020;978:12009.
  7. Chen Y, Gao S, Tu P, Chen X. Automatic 3D Teeth Reconstruction From Five Intra-Oral Photos Using Parametric Teeth Model. IEEE Trans Vis Comput Graph 2024;30:4780-91. [Crossref] [PubMed]
  8. Mohamed W, Nader N, Alsakar YM, Elazab N, Ezzat M, Elmogy M. 3D reconstruction from 2D multi-view dental 2D images based on EfficientNetB0 model. Sci Rep 2025;15:28775. [Crossref] [PubMed]
  9. Liang Y, Song W, Yang J, Qiu L, Wang K, He L. X2Teeth: 3D teeth reconstruction from a single panoramic radiograph. In: Martel AL, Abolmaesumi P, Stoyanov D, Mateus D, Zuluaga MA, Zhou SK, Racoceanu D, Jaskowicz L. editors. Medical Image Computing and Computer Assisted Intervention – MICCAI 2020. Cham: Springer International Publishing; 2020:400-9. doi: 10.1007/978-3-030-59713-9_39.
  10. Song W, Liang Y, Yang J, Wang K, He L. Oral-3D: reconstructing the 3D structure of oral cavity from panoramic X-ray. Proc AAAI Conf Artif Intell 2021;35:566-73.
  11. Li X, Meng M, Huang Z, et al. 3DPX: progressive 2D-to-3D oral image reconstruction with hybrid MLP-CNN networks. arXiv. Preprint posted online August 2, 2024: arXiv:2408.01292. doi: 10.48550/arXiv.2408.01292.
  12. Qian J, Lu S, Gao Y, Tao Y, Lin J, Lin H. An automatic tooth reconstruction method based on multimodal data. J Vis 2021;24:205-21.
  13. Brandenburg LS, Berger L, Schwarz SJ, Meine H, Weingart JV, Steybe D, Spies BC, Burkhardt F, Schlager S, Metzger MC. Reconstruction of dental roots for implant planning purposes: a feasibility study. Int J Comput Assist Radiol Surg 2022;17:1957-68. [Crossref] [PubMed]
  14. Brandenburg LS, Georgii J, Schmelzeisen R, Spies BC, Burkhardt F, Fuessinger MA, Rothweiler RM, Gross C, Schlager S, Metzger MC. Reconstruction of dental roots for implant planning purposes: a retrospective computational and radiographic assessment of single-implant cases. Int J Comput Assist Radiol Surg 2024;19:591-9. [Crossref] [PubMed]
  15. Kazhdan M, Hoppe H. Screened poisson surface reconstruction. ACM Trans Graph 2013;32:1-13.
  16. Zhang R, Sun Z, Hu J, Liu X. Curvature-adaptive deformation graph for 3D point cloud non-rigid registration under multi-geometric constraints. Journal of Computer-Aided Design & Computer Graphics 2023;35:991-9.
  17. Zhang Y, Shen Z, Jiao R. Segment anything model for medical image segmentation: Current applications and future directions. Comput Biol Med 2024;171:108238. [Crossref] [PubMed]
  18. Li Z, Tang W, Gao S, Wang Y, Wang S. Adapting SAM2 Model from Natural Images for Tooth Segmentation in Dental Panoramic X-Ray Images. Entropy (Basel) 2024;26:1059. [Crossref] [PubMed]
  19. Feng W, Sun J, Liu Q, Li X, Liu D, Zhai Z. Specular highlight removal of light field image combining dichromatic reflection with exemplar patch filling. Opt Lasers Eng 2024;178:108175.
  20. Mendoza CS, Safdar N, Okada K, Myers E, Rogers GF, Linguraru MG. Personalized assessment of craniosynostosis via statistical shape modeling. Med Image Anal 2014;18:635-46. [Crossref] [PubMed]
  21. Xiang J, Chen X, Xu S, et al. Native and compact structured latents for 3D generation. arXiv. Preprint posted online December 16, 2025: arXiv:2512.14692. doi: 10.48550/arXiv.2512.14692.
  22. Cai S, Ye Y, Cao J, Chen Z. FACE: feature-preserving CAD model surface reconstruction. Graphical Models 2024;136:101230.
  23. Hunyuan3D T, Yang S, Yang M, et al. Hunyuan3D 2.1: from images to high-fidelity 3D assets with production-ready PBR material. arXiv. Preprint posted online June 18, 2025:arXiv:2506.15442. doi: 10.48550/arXiv.2506.1544.
Cite this article as: Chen G, Tang Z. 3D reconstruction of the target tooth occlusal surface from a single-view image for root canal surgical navigation. Quant Imaging Med Surg 2026;16(9):709. doi: 10.21037/qims-2026-0602

Download Citation