Dry eye detection based on multidimensional temporal features of fluorescein tear film videos
Introduction
Dry eye disease (DED) is a common multifactorial chronic ocular surface disorder characterized by tear film instability or ocular surface microenvironment imbalance, resulting from abnormalities in tear quality and dynamics (1). With the widespread use of visual display terminals, its prevalence has exceeded 30%. Due to the complex etiology and pathogenesis of dry eye, the prevalence varies significantly across different regions (2-6). Tear film instability, elevated osmolarity accompanied by ocular surface inflammation and damage, and corneal nerve damage cause a wide range of ocular discomfort and visual dysfunction, which seriously affects patients’ quality of life. Recent advances have elucidated the core mechanisms and characteristics of DED, culminating in a formalized classification system that designates tear film stability as a pivotal criterion for diagnosis. The tear film structure prevents tear evaporation and provides a smooth optical surface for the cornea. So, evaluating tear film stability is crucial for detecting and categorizing dry eye. The diagnosis of dry eye is primarily based on a comprehensive analysis of subjective symptoms, risk factors, and ocular examinations such as tear film breakup time and tear secretion levels (7-10). However, current diagnostic methods have various issues, including examiner subjectivity, which leads to discrepancies in results, and the complexity of the tests.
The most commonly used clinical assessments of tear film stability include image analysis based on Placido ring projection (concentric circles) and non-contact fluorescein staining high-resolution video imaging, which detects dry eye by measuring tear film breakup time. The former method does not capture other important indicators of dry eye, such as the morphology of tear film breakup. It also requires the patient to keep their eyes open for extended periods, making the results susceptible to noise from eyelashes and other obstructions. If the patient blinks or makes eye movements, the recording must be restarted. Additionally, the high cost of the equipment limits its widespread use (11-16). Similarly, there are issues with analyzing tear film videos post-staining. For example, multiple blinks are recorded in the saved video, along with frames of the eye-opening process, which means a single test video contains multiple segments for analysis. Classification models are used to identify open and closed eye states, but frames with partially open eyes, brief eye movements, and noise from eyelashes affect the classification accuracy (ACC) and segmentation precision, ultimately impacting the final diagnostic ACC. Moreover, current video analysis methods require post-processing of the video and manual segmentation of test segments, which prevents real-time analysis and requires frame alignment, making it difficult to accommodate various types of test videos.
There is almost no research on automated dry eye detection using non-contact fluorescein video imaging (17-26). Remeseiro et al. utilized color and texture information to represent lipid layer images, selecting features to achieve dry eye classification (22). Su et al. proposed a technique requiring extensive preprocessing to segment frames before detecting tear film breakup (24). Yedidya et al. developed a method for identifying breakup areas, but they used manual feature extraction to detect the occurrence of tear breakup (25,26).
In recent years, the progress of deep learning in the medical field has led to success in medical image recognition and classification. Convolutional neural networks (CNNs) are the most versatile models, replacing manual feature extraction with hierarchical feature extraction and achieving good performance. Numerous studies have applied deep learning to develop computer-aided diagnostic systems for the automated diagnosis of chronic diseases (27-32). Several DED detection techniques have also been reported. These methods primarily focus on identifying areas indicative of DED, with a significant portion relying on fluorescein breakup time (FBUT) for detection (33-40). Vyas et al. and Shimizu et al. used deep learning to calculate tear film breakup time for simple dry eye screening. The use of unsegmented full-eye images for tear film breakup pattern recognition introduces substantial noise and is associated with compromised ACC (34). Similarly, and of notable interest, an approach employing full-sequence ocular video frames has been developed: a deep transfer learning model for the direct identification of DED from ocular surface videos. This approach, however, suffers from a key limitation validated in our experiments: the results are prone to interference from strong, extraneous features like eyelids and eyelashes, leading to suboptimal robustness (36). In addition, these approaches predominantly rely on unmodified CNN architectures. Furthermore, some of the extracted frames from the same video exhibited a high degree of similarity, which consequently constrained the feature selection space and led to suboptimal concept learning. The foundation of precision medicine within intelligent healthcare lies in the development of integrated therapeutic strategies tailored to individual or specific population profiles. However, prevailing clinical diagnostic and therapeutic approaches for DED remain largely subjective. They often depend on complex yet underutilized diagnostic data, leading to suboptimal efficacy. This limitation consequently hinders the precise stratification and management of DED based on its underlying mechanisms and subsequent severity levels (41).
Therefore, to solve these challenges, we apply deep learning to analyze tear film breakup dynamics from high-resolution fluorescein-stained video sequences, with a fully automatic framework for the precise characterization and analysis of DED. We proposed a system based on tear film region segmentation and spatio-temporal feature recognition of tear film breakup. Ophthalmologists need only operate a slit lamp to capture video and can use the system to assist in detecting dry eye. This system consists of three parts. The first part is the tear film region segmentation based on a CNN network. The second part is fully open-eye video segmentation and extracting tear film features. The third part is a transformer network, which compresses and recognizes the tear film texture features, visual-spatial features, and temporal features of each frame, thereby achieving efficient and accurate dry eye detection. Our main contributions are summarized as follows:
We constructed a dataset containing 248 fluorescein tear film video datasets based on slit lamp photography, with annotated DED and non-DED by clinical evaluations. It also contains 2,543 video frames with annotated tear film regions.
We proposed a novel MASN for tear film segmentation, which applies a channel and position attention module for deep feature fusion and utilizes a cross-attention module at skip connections to integrate shallow and deep features.
We developed a method to effectively identify the video segments with full eye opening, which calculates the percentage of tear film area in each frame and uses statistical methods to determine the upper and lower thresholds of the fully open eye.
We proposed an innovative cross-attention transformer classification network (CATCN). This network includes a dual-encoder module and a lightweight cross-modal attention interaction module (LCMAIM). The dual-encoder module supports both texture features and image data input. The LCMAIM was utilized to fuse the different types of representations, which used the texture as query (Q) to enhance the attention of tear film features. At last, the multilayer perceptron (MLP) was used to classify the fused features.
We validated our method on our annotated datasets, showing state-to-art tear film region segmentation and DED detection accuracies.
Methods
Data acquisition and proposed framework
Participants
This is a single-center, cross-sectional, and retrospective study; 148 DED patients [mean ± standard deviation (SD), age 45.6±8.7 years] and 100 non-DED adults (age 45.4±2.9 years) were recruited for this study. The inclusion criteria for DED subjects were consistent with the TFOS DEWS II (2) and Chinese Expert Consensus on Dry Eye (3). All patients and volunteers completed the Ocular Surface Disease Index (OSDI) questionnaire and dry eye examination, such as FBUT, Schirmer’s test I (ST-I), non-invasive breakup time (NIBUT), and corneal/conjunctival fluorescein staining (CFS). The inclusion and exclusion criteria are shown in Table 1.
Table 1
| Criteria | DED patients | Healthy volunteers |
|---|---|---|
| Inclusion | Meet one of the conditions: | (I) Voluntary participation and age >18 years |
| (I) OSDI ≥13 and (FBUT ≤5 s or ST-I ≤5 mm/5 min or NIBUT ≤10 s) | (II) No history or symptoms of dry eye, and FBUT >5 s, ST-I >5 mm/5 min, and negative CFS | |
| (II) OSDI ≥13 and (5 s < FBUT ≤10 s or 10 s < NIBUT ≤12 s or 5 mm/5 min < ST-I ≤10 mm/5 min) and positive CFS (≥5 dots) | ||
| Exclusion | (I) Active ocular disease, history of ocular surgery, or use of ocular anti-inflammatory medications (e.g., immunosuppressants, corticosteroids, NSAIDs) within the preceding 2 weeks | (I) Active ocular disease, history of ocular surgery, or use of anti-inflammatory medications within the past 2 weeks |
| (II) Conditions such as depression, anxiety, or other psychiatric disorders that could compromise cooperation with examinations | (II) Systemic connective tissue diseases or circulatory disorders | |
| (III) Systemic diseases associated with DED (e.g., connective tissue diseases, diabetes, autoimmune diseases) | (III) Pregnancy or poor compliance | |
| (IV) Severe systemic comorbidities including malignancy, active infection, circulatory disorders, chronic alcoholism, fluid/electrolyte imbalance, or cardiac/hepatic/renal/pulmonary insufficiency | ||
| (V) Any change in medication regimen within the past 3 months | ||
| (VI) Pregnancy or poor compliance |
CFS, corneal/conjunctival fluorescein staining; DED, dry eye disease; FBUT, fluorescein breakup time; NIBUT, noninvasive breakup time; NSAID, non-steroidal anti-inflammatory drug; OSDI, Ocular Surface Disease Index; ST-I, Schirmer’s test I.
The dataset is a consecutive series of patients at Zhongshan Ophthalmic Center, Sun Yat-sen University, collected between March 2023 and November 2024. This study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study was approved by the Ethics Committee of Zhongshan Ophthalmic Center, Sun Yat-sen University (approval No. 2023KYPJ251). Before the experiment, all participants gave informed consent and completed a questionnaire of subject information.
Data collection
The dataset was collected from patients attending the Ophthalmology Center. The videos were captured using a CANON EOS 60D camera connected to a slit lamp, with 16 frames per second, and stored in Audio Video Interleave (AVI) format. The videos were evaluated to confirm that the participants blinked at least three times to observe the stability of the tear film. The video with severe corneal defects or corneal epitheliitis, as well as images with poor quality, were excluded. The video acquisition was non-invasive, and no adverse events occurred. The dataset contains 248 videos from 248 eyes with a resolution of 1,056×704 pixels, including 148 DED and 100 non-DED videos. These videos were divided into 138,885 static images. A total of 200 videos (108,915 frames) were used for training the dry eye recognition model (including tear film segmentation and feature extraction), and 48 videos (29,970 frames) were reserved for testing. An additional 31 videos were included as an external validation set, which is used to assess the correlation between the system and actual clinical diagnoses. Each video was supplemented with basic clinical information, including sex, age, and slit-lamp microscopy findings. Two senior ophthalmologists independently reviewed and annotated the videos. When the cases had inconsistent diagnoses, the third senior ophthalmologist performed the final annotation. The third senior ophthalmologist’s annotation corresponds to the most chosen type. Each frame was annotated as normal, ruptured, or blink. The two senior ophthalmologists demonstrated substantial agreement for frame-level tear film rupture detection with Cohen’s k=0.85, and almost perfect agreement for video-level DED classification with Cohen’s k=0.92. The example frames of the dataset are shown in Figure 1.
For the segmentation of the tear film, we manually annotated the segmentation dataset of frames with different eye states, including fully open, non-fully open, and closed, using LabelMe software. The dataset includes 2,000 frames for training from the training videos dataset and 543 frames for testing from the test videos dataset. The dataset was enhanced by removing uneven lighting and horizontally flipping. In order to improve training efficiency, we resized the original image from 1,056×704 to 480×270 pixels.
The image processing workstation was equipped with an 18-core Intel Xeon Processor at 2.10 GHz, 64 GB RAM, a 512 GB NVMe SSD, and an RTX3090 with 24 GB RAM delivering 35.6 TFLOPS (FP32).
Proposed framework
The proposed method aligns with the clinical practice of diagnosing DED by visually observing dynamic changes in the tear film over time. We first employed the MASN network to pre-segment the tear film region, thereby restricting feature recognition to this area. Subsequently, the curve was plotted based on the proportion of the tear film region, and Gaussian distribution statistics were applied to determine the upper and lower thresholds for fully open frames, and then the continuous video frames with fully open frames. Within these segments, one frame was extracted every 0.1 seconds. In every frame, the tear film region and its texture features were extracted as input to the classification network. Finally, the CATCN network was utilized to detect DED. It included two vision transformer (ViT) encode modules, LCMAIM, a feed-forward network (FFN), and MLP. The two transformer (42) encode modules were used to encode the spatiotemporal and texture features along with the tear film region images. The LCMAIM was used to fuse the two output features of the transformer. The MLP was used to make the classification of the DED. The overall dry eye classification and recognition system is illustrated in Figure 2.
Tear film region segmentation
Network architecture
Fluorescein tear film video acquisition captures the complete ocular image and its characteristics, including both the tear film area and non-tear film areas. The strong features are predominantly concentrated in non-tear film regions, such as the upper and lower eyelids, eyelashes, and conjunctiva. In contrast, the patterns of tear film diffusion and breakup manifest as weak features characterized by subtle local structural changes, indistinct boundaries, and random diffusion. The strong features with their high frequency and saliency are easily captured by convolutional kernels and get high activation values in CNN. Conversely, weak features with low frequency and weak activation responses are prone to being overlooked by the neural network (43,44). Training CNN on entire video frames from a limited dataset may yield favorable classification and recognition results. However, the model can easily learn strong features unrelated to tear film breakup morphology while neglecting changes in the tear film region, thereby compromising the robustness of classification. Therefore, segmenting the tear film region for the classification of tear film breakup characteristics can effectively enhance the robustness of the classification results.
Based on the encoder-decoder U-shaped network (45) framework, we proposed a multi-attention segmentation network (MASN) to enhance the precise segmentation of the tear film area (Figure 3). First, we used the average pooling module at the downsampling part to improve the generalization capability for the tear film region features. Then, at the bottom of the encoder, we incorporated channel and positional attention modules to perform multi-dimensional attention fusion and enhancement on multi-level features, thereby boosting the network’s ability to recognize features across different dimensions. Finally, in the skip connections, we introduced a cross-attention module to strengthen the cross-attention between encoded and decoded features, consequently enhancing the network’s overall target recognition capability.
Cross-attention module
In U-shaped architecture networks, skip connections facilitate the fusion of deep and shallow features. They enhance target features by applying attention mechanisms to the compressed feature representations. Unlike other attention modules, this study proposes a cross-attention module that achieves bidirectional feature cross-attention between the encoder and decoder, followed by feature fusion. Through bilinear cross-attention (Figure 3), the contextual features from the encoder and decoder are cross-enhanced and deeply integrated. This approach accelerates the convergence speed of network training and improves the ACC of target segmentation.
Positional and channel attention module
The positional and channel attention modules can integrate features correlated across spatial positions and vertical channels, respectively. The positional module is designed to integrate cross-spatial information from different local feature maps. The input feature P is processed through two convolutional layers for horizontal and vertical filtering, thereby enhancing spatial position awareness. Subsequently, reshaping and transposition operations are applied to generate two new feature maps, G, and a feature matrix, H. The spatial weight feature matrix is derived through a matrix multiplication between features G and H, followed by a Softmax operation.
Similarly, the channel attention module extracts horizontal and vertical features via convolutional kernels. Finally, it fuses these horizontal and vertical features into the characteristics across high-dimensional channels through matrix multiplication.
Data preprocessing
To improve the segmentation performance of the tear film, we first performed consistent preprocessing on the input images. The image resolution was adjusted to 480×270, followed by the application of Contrast Limited Adaptive Histogram Equalization (CLAHE) to enhance the contrast between pixels. Furthermore, to resolve the issue of limited sample size, we employed data augmentation techniques, including horizontal flipping, to expand and diversify the training dataset.
Network training
The network was trained on our annotated dataset with the following hyperparameters: 100 epochs, a batch size of 10, a learning rate of 0.002, using the RMSprop gradient optimization algorithm and a binary cross-entropy loss function. Training was performed on an RTX 3090 GPU with 24 GB of memory. After training, the model achieved a mean pixel segmentation ACC of 0.963 on the validation set.
Tear film frame selection
FBUT is a standard diagnostic criterion for DED. Its measurement refers to the time interval from the last blink to the appearance of dark spots or dry spots on the cornea. Tear film rupture is correlated with changes in tear film thickness and stability. Therefore, identifying the period of fully opened eyes is a prerequisite for analyzing DED.
Fluorescein tear film video acquisition primarily focuses on the tear film, recording the morphological changes of the tear film over time during multiple eye-opening and closing cycles of the subject. Traditional methods for determining the fully-opened-eye video segments rely on a fixed threshold based on eye-opening parameters. However, due to individual variations in age, eye morphology, acquisition devices, and environmental factors, using a fixed threshold for identification has inherent limitations.
Based on this, we proposed a dynamic threshold video segmentation method. For each segmented tear film region, we calculated its area ratio relative to the entire image. Using statistical methods, we analyzed the distribution of this area ratio across all frames. The results show that the distribution essentially follows a Gaussian distribution characteristic (Figure 4). Based on the empirical rule of statistics, the upper and lower threshold ranges of eyes fully open can be determined by the mean and SD. The resulting histogram approximately follows the Gaussian distribution, as described by the following function:
represents the percentage of the tear film area in the input image, represents the mean, and σ represents the SD.
For each acquired video, we segmented the tear film percentage in all frames and calculated its mean and SD , thereby determining the upper and lower thresholds for the tear film percentage during complete eye opening for each sample video frame. The specific formula is as follows:
Where represents the upper limit of the tear film percentage, represents the lower limit of the tear film percentage.
To further validate the effectiveness of this method, we plotted a time-series curve of the tear film area ratio (Figure 5) and annotated frames corresponding to full eye opening and not fully opening near the determined thresholds. In the annotated figures, frames classified as not fully open show eyes in a semi-open state where the tear film area is severely occluded by eyelashes. In contrast, frames classified as fully open display the majority of the tear film area with no eyelash occlusion. The adaptive eyes-fully-open threshold overcomes the errors of fixed thresholds by providing individualized thresholds tailored to different subjects’ eye-opening states, leading to more robust recognition performance.
Video segments extraction
After identifying the fully open eye video segments, we selected one frame every 0.1 seconds starting from the fully open state for subsequent analysis. A minimum of 20 frames will be selected, and video segments not meeting the frame count will be filtered out. A maximum of 100 frames will be selected, corresponding to a maximum analysis time of 10 seconds along the temporal dimension. This protocol adheres to the clinical diagnostic standards for FBUT, ensuring that the network can learn both spatial tear film features and their temporal evolution, thereby enhancing the robustness of DED recognition.
Texture feature extraction
Texture Features play an exceptionally important role in medical image analysis, particularly in tasks such as pathological pattern recognition and classification analysis. They can more effectively capture microscopic changes in tissue structure, thereby enhancing the model’s sensitivity (SE) to early-stage lesions and the ACC of disease subtyping.
We integrated texture features into the network architecture. After embedding positional information, these texture features served as one of the input tokens to capture microscopic variations in tissue structure. This approach enhances the model’s ability to perceive local structural changes over time, thereby improving classification ACC for non-large-scale datasets as well as SE to early-stage lesions. We utilized the PyRadiomics library to extract texture features from the tear film region, thereby transforming the image data into high-dimensional data suitable for analysis. We extracted 94 feature parameters in total. Encompass the following categories: first-order statistics (19 features), gray level co-occurrence matrix (GLCM) (24 features), gray level run length matrix (GLRLM) (16 features), gray level size zone matrix (GLSZM) (16 features), neighboring gray tone difference matrix (NGTDM) (5 features), and gray level dependence matrix (GLDM) (14 features).
To further validate the utility of the extracted tear film region texture features, we visualized several GLCM-based texture features specific to the tear film (Figure 6). The visualization illustrates the variations in texture features corresponding to different tear film breakup morphologies within the region.
Dry eye classification
Network architecture
The design of this network is based on the transformer module. Since the video feature input consists of multiple image frames rather than a single image, the features of one frame and its position are compressed and served as input tokens for the transformer network. Based on this, we proposed a network with a hybrid decoding and multi-view attention fusion mechanism that enables it to accept input data from two distinct viewpoints S1 and S2 (Figure 7). Two independent transformer encoders extract feature representations H1 and H2 from different data sources, respectively. These are then processed through two standard multi-head decoding paths to preserve the unique features T1 and T2 of each viewpoint. Following this, LCMAIM is used to achieve multi-granular and deep information fusion . Finally, MLP was used for the classification of the fused features. The input to CATCN is the encoder part, and the computation formula is as follows:
Where and are the input features, and are the compressed features. After obtaining two different classification outputs, CATCN uses LCMAIM to achieve feature fusion and classification of different decoded features. Its attention fusion calculation is as follows.
Where is the fused token using hybrid attention mechanism, , and represent query, key and value.
Tear film area feature compression
The compression of consecutive tear film video frames requires a large amount of computation, especially the compression of real-time image features. To resolve this, we used EfficientNetV2-S (46) for compressing the tear film video frame features. The model combines Fused-MBConv and MBConv modules, achieving optimal training speed and parameter efficiency. To enhance the feature compression performance of EfficientNetV2-S, we perform transfer learning and fine-tune the classification output. First, the model is pretrained on the ImageNet-1K dataset. Second, a fully connected layer is added after the pooling layer with 128 output features, and the final classification layer is changed to output two classes. At last, transfer learning was performed on the model using the tear film breakup dataset. The detailed hyperparameters of training are in Table 2. This approach enables superior recognition of tear film breakup features.
Table 2
| Hyperparameter | Value |
|---|---|
| Training images | 980 |
| Validation images | 294 (30% of training images) |
| Test images | 120 |
| Optimizer | Stochastic Gradient Descent with Momentum |
| Epochs | 20 |
| Learning rate | 0.0002 |
| Batch size | 10 |
| Layer retrained | 42 |
| GPU | Nvidia GeForce RTX3090 |
| Classifier | Softmax |
Transformer classification
To effectively integrate the complementary information from the global image features and the specialized texture features under the constraint of limited data, we proposed an LCMAIM. The core design philosophy was used to treat the texture features as a query guide, enabling them to dynamically retrieve and aggregate relevant contextual information from the image feature space. This guiding and retrieving mechanism offers superior interpretability and fusion efficiency compared to simple concatenation or averaging operations.
Given the global image feature vector and the texture feature vector , the interaction process proceeds as follows.
Feature projection & alignment:
First, both features are projected into a shared latent space of dimension d using separate linear transformations to facilitate interaction:
Where , are learnable weight matrices, and , are biases. This yields aligned representations .
Linear cross-attention with texture as query
We adopted a computationally efficient linear attention mechanism where the texture feature serves as the Q to retrieve information from the image feature acting as both key (K) and value (V). This mimics a diagnostic process of locating relevant visual contexts based on specific textural cues.
The query, key, and value are derived as:
Here, and are projection weights. To further reduce parameters and stabilize training, we employed multi-head linear attention with heads. The head dimension is .
The linear attention for each head is computed without the quadratic-cost Softmax. We used a feature map applied to the Key to ensure non-negativity and numerical stability:
Where are features for head h. The outputs of all heads are concatenated and projected to obtain the attention-refined feature .
Residual fusion and nonlinear transformation
The refined feature A is first combined with the original texture feature via a residual connection, followed by layer normalization (LN). This ensures the preservation of critical textural information throughout the interaction:
Subsequently, passes through a lightweight, two-layer FFN with Gaussian Error Linear Unit (GELU) activation and dropout for nonlinear transformation:
A final residual connection and LN produce the fused output feature:
Classification head
The final fused representation which integrates guided visual context with specific textural patterns, is fed into a simple classifier consisting of a linear layer followed by a Softmax function for tear film breakup classification.
Results
Validation of the tear film segment
To enhance the segmentation performance of the tear film, we first performed consistent preprocessing on the input images. The image resolution was adjusted to 480×270, followed by the application of CLAHE to enhance the contrast between pixels. Furthermore, to address the issue of limited sample size, we employed data augmentation techniques, including vertical flipping, horizontal flipping, and a combination of both, to expand and diversify the training dataset.
For network training, we used the preprocessed RGB images as input (with 3 input channels) and generated an output mask (with 1 output channel). The training batch size was set to 10. We utilized the RMSprop algorithm with a weight decay of 0.0002 and a learning rate of 0.0002 as the optimizer. This configuration helps adapt the global learning rate and improves training speed. The BCEWithLogitsLoss function was employed as the loss function to train the network. We monitored the loss on the validation set and implemented early stopping if no improvement was observed for 10 consecutive epochs. The segmentation results of the MASN are presented in Figure 8.
To validate the segmentation performance of the MASN, we conducted a comparative performance analysis against other improved U-Net-based segmentation networks, including Unet (45), ResUnet (47), PSPnet (48), and ResAgUnet. ResUnet is based on U-Net and replaces the convolution blocks with residual blocks. Its area under the curve (AUC) gets an improvement of 10% against U-Net. ResAgUnet extends ResUnet by adding attention modules at the skip connections and gets 1% increase in both ACC and AUC. Based on ResAgUnet, MASN incorporates cross-attention modules at the skip connections and positional & channel attention modules at the bottom. It further enhances performance, getting an additional 1% improvement in both ACC and AUC. Tear film segmentation performance metrics of different networks are shown in Table 3.
Table 3
| Methods | Parameter (M) | FLOPs (G) | ACC | Sensitivity | Specificity | AUC |
|---|---|---|---|---|---|---|
| U-Net | 32.91 | 28.8 | 0.95 | 0.74 | 0.98 | 0.86 |
| ResUnet | 34.25 | 26.1 | 0.94 | 0.77 | 0.97 | 0.96 |
| PSPNet | 46.72 | 25.3 | 0.95 | 0.75 | 0.98 | 0.95 |
| ResAgUnet | 34.92 | 27.2 | 0.96 | 0.81 | 0.98 | 0.97 |
| MASN | 13.46 | 23.7 | 0.97 | 0.82 | 0.98 | 0.98 |
ACC, accuracy; AUC, area under the curve; FLOPs, floating point operations; MASN, multi-attention segmentation network.
At the same time, MASN contains fewer parameters and lower floating point operations (FLOPs) consumption. MASN reduces the network parameters and convolution operations by incorporating cross-attention, position attention, and channel attention modules. The cross-attention module contains no learnable parameters. It achieves cross-connection feature integration through pooling, activation, element-wise multiplication, and addition, requiring low FLOPs. Similarly, the position and channel attention module are parameter-free. It replaces the parameter-heavy module at the bottom of UNet, thereby effectively reducing the total number of network parameters. Furthermore, by using matrix element-wise multiplication and addition to replace extensive convolution operations, it significantly reduces FLOPs consumption.
Validation of Tear film feature encoding
The performance of the tear film feature compression directly impacts the classification ACC of the classification module. We recorded EfficientV2-S’s loss and ACC during transfer learning (Figure 9). It achieved the lowest loss and the highest ACC at the 7th epoch. To validate EfficientNetV2-S’s capability for compressing and identifying tear film features, we extracted the output features from the last convolutional module. These features were then rendered onto their corresponding tear film region images to construct feature distribution heatmaps (Figure 10). We conducted a comparative analysis using heatmaps generated from both the full images and the segmented tear film regions. Observations of the feature distribution heatmaps on the segmented tear film indicate that the features characterizing tear film breakup are predominantly concentrated within the inner regions of the tear film. This visualization confirms the feature encoder module’s effectiveness in identifying discriminative features related to tear film changes.
Validation of the dry eyes classification
To demonstrate the DED classification performance of the proposed methods. We evaluated them on our validation dataset. We selected two distinct network architectures, which are CNN and transformer, to compare their performance. For CNNs, the stacked frames were used as their input. While the Transformer networks used the compressed features from segmented tear films as the input. The learning rate of the networks was set to 0.002. The training epochs were 100. The batch size was set to 10 for all network training. The RMSprop algorithm was trained as an optimizer to adapt the global learning rate and improve the training speed. The classification performance of our network was evaluated against other methods on the same dataset.
In order to validate the loss convergence speed and ACC of different networks during training, we collected the loss values and ACC of the validation dataset during the training process and plotted their curves (Figure 11). In order to evaluate classification performance, we plotted receiver operating characteristic (ROC) curves for each network and calculated the AUC (Figure 11).
Our approach shows the best performance on tear film video databases through joint evaluation of precision, SE, specificity (SP), ACC, and AUC. The performance comparison of the networks is shown in Table 4. Although the CNNs achieved promising results in classification performance, with an ACC of 0.89 and an AUC of 0.92. But its primary discriminative features were from non-tear-film regions. We have validated this in Figure 9. For transformer networks, we used the features from the tear film region as input. The classification performance got an improvement over CNNs, with an increase in ACC >1%. Based on the transformer architecture, CATCN supports both compressed tear film features and textural features as inputs. It can fuse these multidimensional temporal features to achieve superior performance compared to other transformer networks. Its ACC is up to 0.92, and AUC is up to 0.96.
Table 4
| Model | Parameter (M) | FLOPs (G) | ACC | Precision | SE | SP | AUC |
|---|---|---|---|---|---|---|---|
| ResNet50V2 | 23.6 | ~15.8 | 0.85 | 0.83 | 0.73 | 0.95 | 0.86 |
| EfficientNetV2-S | 21.4 | ~6.5 | 0.89 | 0.90 | 0.84 | 0.95 | 0.92 |
| ViT | 151.1 | ~15.1 | 0.90 | 0.91 | 0.83 | 0.92 | 0.93 |
| ViT (EfficientNetV2-S) | 5.1 | ~5.1 | 0.91 | 0.90 | 0.85 | 0.91 | 0.94 |
| ViT (Texture) | 5.1 | ~5.1 | 0.91 | 0.91 | 0.84 | 0.93 | 0.95 |
| CATCN (Texture+EfficientNetV2-S) | 12.3 | ~10.4 | 0.92 | 0.92 | 0.86 | 0.96 | 0.96 |
ACC, accuracy; AUC, area under the curve; FLOPs, floating point operations; SE, sensitivity; SP, specificity.
ResNet50V2 (36), EfficientNetV2-S (46), and ViT used original frames as input, while ViT(EfficientNetV2-S), ViT(Texture), and CATCN used the tear film region as input, which was segmented by MASN. For the CNN networks and ViT, the input max size is 100×1,296. While 100 is the number of frames, which are extended from the video every 0.1 second, 1,296 is the scaled length of the frames (48*27). The ViT(EfficientNetV2-S) and ViT(Texture) input token dimensions are 94×2. They require only 5.1 GFLOPs with 100 tokens. But achieves better ACC than traditional CNNs and ViT. The precision and ACC of ViT(Texture) improved by 1% compared to ViT(EfficientNetV2-S). The CATCN (Text+EfficientNetV2-S) integrates ViT(EfficientNetV2-S) and ViT(Texture), employing LCMAIM for feature attention fusion and a residual connection with ViT(Texture). Its total parameter count and FLOPs consumption are more than double those of ViT(Texture). By incorporating two sets of input features, CATCN(Texture+EfficientNetV2-S) achieves an improvement of 1% in ACC, precision, and AUC.
To verify the actual clinical detection performance, we prepared 31 additional external fluorescein tear film videos for independent performance evaluation. This set of 31 videos included 14 normal images and 17 dry eye images. This effectively verifies whether the frames from training videos, used for feature extraction training, affect the system’s performance. The performance validation results are shown in Figure 12. The performance of the external dataset validation is basically consistent with the test dataset, with an ACC of 0.93, a precision of 0.94, and an SE of 0.94, showing slight improvements compared to the test set. SP is 0.92, and AUC is 0.9542, slightly lower than the test set. Overall, the validation performance is basically consistent with the test performance.
Error analysis and clinical interpretability
Although CATCN (Texture+EfficientNetV2-S) achieved an ACC of 0.92. But there is still an error rate of 8%. After reviewing and analysing the test data, several factors were found affecting the errors. The first one is uneven fluorescein staining of the tear film. The second one is significant obscuration by eyelashes. The last is the diffuse rupture of the tear film, which is a weak morphological feature. The typical cases of false positives (FPs) and false negatives (FNs) are shown in Figure 13.
Discussion
Dry eye detection using fluorescein-staining videos has become the expert consensus (2,3). Manual dry eye detection relies on the ophthalmologist observing screens, accurately classifying video frames, and timing the FBUT. It is a tedious and highly challenging task. Subsequently, Abdelmotaal et al. developed an AI-assisted detection method using ResNet50V2, DenseNet121, InceptionV3, and SVM to classify dry eye video frames, achieving an ACC of 0.9 (36). Based on frame-level recognition, Vyas and Shimizu et al. proposed GoogLeNet and Swin Transformer to identify tear film frames with breakup and calculate TBUP, offering a novel AI-based approach to dry eye detection (34,38), and achieving an ACC of 0.91 and 0.79. However, these methods perform tear film breakup detection using unsegmented, full-eye video frames. This introduces significant irrelevant features, such as eyelids, eyelashes, and conjunctiva, thereby reducing detection robustness. Most of these methods use unmodified traditional network architectures, although some methods utilize transfer learning to enhance hyperparameter learning. But further optimization is still required for the model. Finally, current dry eye detection methods rely on frame-level classification and TBUP calculation. There are no fully automated detection methods available at the video level.
In this study, we proposed a fully automated framework for DED detection based on the video-level. The system comprises three components. The first is the segmentation of the tear film region. The second is frame selection based on the dynamic threshold, followed by compressing the tear film features and extracting texture from the selected frames. The last is the integration of temporal, spatial, and textural features from selected tear film video frames to achieve precise DED classification. Unlike Vyas et al. and Shimizu et al. using full eye frames (34,38), our methods focus on characteristic changes within the tear film region. We developed the MASN for tear film segmentation to enhance the robustness of tear film feature recognition and ensure DED detection focuses on localized changes within the tear film, and obtained an ACC of 0.97 and an AUC of 0.98.
DED detection relies on identifying fully open-eye frames and then determining whether a rupture is present. In conventional approaches, identifying fully open eyes relies on fixed thresholds. However, due to physiological variations in eye morphology across individuals, fixed thresholds can lead to recognition errors. To resolve this, we developed a dynamic threshold. It enables more adaptive and accurate detection of fully open eyes across different subjects. And extracts video segments that meet the analysis requirements (Figure 5). Tear film breakup patterns and morphologies vary widely and are influenced by image acquisition equipment and environmental factors, often manifesting as blurred boundaries, image defocusing, and random diffusion. Traditional CNNs used for DED detection are difficult to distinguish these features from noise, thereby hindering effective learning of tear film breakup features (34). To overcome this problem, we proposed the CATCN network based on the ViT architecture for DED detection. This network uses extracted tear film features, textures, and spatio-temporal features as inputs to enhance the ability of capturing temporal changes within the tear film region. These will effectively improve the robustness of classification results. The CATCN supports the input of an arbitrary number of video frames as tokens. It effectively overcomes the limitation of CNNs, which require a fixed number of video frames as input channels. Within the CATCN network, we used a lightweight cross-attention interaction module. In this design, the encoded tear film features act as an information repository (K, V), while more pathologically specific texture features serve as Q. Through a linear attention mechanism, the model can dynamically retrieve and perform attention aggregation of relevant information from the global image context based on local abnormalities indicated by the texture features. This Q and K, V fusion is more effective in achieving information complementarity and suppressing interference from irrelevant backgrounds compared to simple feature concatenation or average pooling, and the process is more interpretable. Performance comparisons are shown in Table 4. The CATCN attained an ACC of 0.92 and an AUC of 0.96, representing more than a 1% improvement over competing networks. The CATCN achieved an ACC of 0.93, a precision of 0.94, and an SE of 0.94 in the external clinical validation set. The assessment has reached the level of professional ophthalmologists.
There are some potential limitations of the automated dry eye recognition system. First, the combined model architecture increases processing complexity and computational load compared to single-model methods. Although the combined module was designed to be lightweight, deploying it on resource-constrained edge devices, such as portable diagnostic instruments, will require further research into model compression, quantization, or distillation. Second, the datasets used for training and validation are relatively small and may not cover all possible scenarios, such as cases involving different acquisition devices, dry eye subtypes, or patient populations. Consequently, continuous validation and dataset optimization remain necessary to enhance the system’s coverage, reliability, and stability. Third, the dataset annotated by only three clinical ophthalmologists may contain subjective errors.
Conclusions
The fully automated dry eye detection system achieves excellent detection performance through precise tear film segmentation and multi-view tear film feature fusion recognition. It offers a convenient new method for large-scale screening of DED. Furthermore, CATCN supports multi-view and multi-modal data inputs. In the future, this architecture may be explored for its application in visual tasks requiring the fusion of heterogeneous information, such as the analysis of tissue slides, color fundus photographs, and OCT images.
Acknowledgments
None.
Footnote
Reporting Checklist: The authors have completed the STARD-AI reporting checklist. Available at https://qims.amegroups.com/article/view/10.21037/qims-2026-0950/rc
Data Sharing Statement: Available at https://qims.amegroups.com/article/view/10.21037/qims-2026-0950/dss
Funding: This study was supported by
Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://qims.amegroups.com/article/view/10.21037/qims-2026-0950/coif). All authors report that this study was supported by Guangzhou Science and Technology Program (Nos. 2025A04J4527 and 2025A04J4247), University Characteristic Innovation Project of Guangdong Province (No. 2025KTSCX083), and Key Program of the National Natural Science Foundation (No. 82230033). The authors have no other conflicts of interest to declare.
Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. This study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study was approved by the Ethics Committee of Zhongshan Ophthalmic Center, Sun Yat-sen University (approval No. 2023KYPJ251). Written informed consent was obtained from all subjects.
Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.
References
- Lemp MA, Foulks GN. The definition and classification of dry eye disease: Report of the Definition and Classification Subcommittee of the International Dry Eye Workshop. Ocul Surf 2007;5:75-92. [Crossref] [PubMed]
- Stapleton F, Alves M, Bunya VY, Jalbert I, Lekhanont K, Malet F, Na KS, Schaumberg D, Uchino M, Vehof J, Viso E, Vitale S, Jones L. TFOS DEWS II Epidemiology Report. Ocul Surf 2017;15:334-65. [Crossref] [PubMed]
- Asia Dry Eye Society China Branch, Ocular Surface and Tear Film Disease Group of the Ophthalmology Committee of the Cross-Strait Medical and Health Exchange Association, Ocular Surface and Dry Eye Group of the Chinese Ophthalmologist Association. Chinese expert consensus on dry eye: examination and diagnosis (2020). Chin J Ophthalmol 2020;56:741-7.
- Yu Z, Wu X, Zhu J, Jin J, Zhao Y, Yu L. Trends in Topical Prescriptional Therapy for Old Patients With Dry Eye Disease in Six Major Areas of China: 2013-2019. Front Pharmacol 2021;12:690640. [Crossref] [PubMed]
- Yang W, Luo Y, Wu S, Niu X, Yan Y, Qiao C, Ming W, Zhang Y, Wang H, Chen D, Qi M, Ke L, Wang Y, Li L, Li S, Zeng Q. Estimated Annual Economic Burden of Dry Eye Disease Based on a Multi-Center Analysis in China: A Retrospective Study. Front Med (Lausanne) 2021;8:771352. [Crossref] [PubMed]
- Tsubota K, Yokoi N, Shimazaki J, Watanabe H, Dogru M, Yamada M, Kinoshita S, Kim HM, Tchah HW, Hyon JY, Yoon KC, Seo KY, Sun X, Chen W, Liang L, Li M, Liu ZAsia Dry Eye Society. New Perspectives on Dry Eye Definition and Diagnosis: A Consensus Report by the Asia Dry Eye Society. Ocul Surf 2017;15:65-76. [Crossref] [PubMed]
- Weinreb RN, Moghimi S. Advances in Ocular Imaging. Asia Pac J Ophthalmol (Phila) 2019;8:97-8. [Crossref] [PubMed]
- Villani E, Nucci P, Benitez-Del-Castillo JM, Dahlmann-Noor A, Lagrèze WA, Bremond-Gignac D. PeDED Delphi Group. Expert consensus on pediatric dry eye: Insights from a European Delphi study. Ocul Surf 2025;37:189-97. [Crossref] [PubMed]
- Deng Y, Chen W, Xiao P, Jiang H, Wang J, Chen W, Li S, Zhong J, Peng L, Wang Q, Yuan J. Conjunctival microvascular responses to anti-inflammatory treatment in patients with dry eye. Microvasc Res 2020;131:104033. [Crossref] [PubMed]
- Deng Y, Wang Q, Luo Z, Li S, Wang B, Zhong J, Peng L, Xiao P, Yuan J. Quantitative analysis of morphological and functional features in Meibography for Meibomian Gland Dysfunction: Diagnosis and Grading. EClinicalMedicine 2021;40:101132. [Crossref] [PubMed]
- Bron AJ, Evans VE, Smith JA. Grading of corneal and conjunctival staining in the context of other dry eye tests. Cornea 2003;22:640-50. [Crossref] [PubMed]
- Titiyal JS, Falera RC, Kaur M, Sharma V, Sharma N. Prevalence and risk factors of dry eye disease in North India: Ocular surface disease index-based cross-sectional hospital study. Indian J Ophthalmol 2018;66:207-11. [Crossref] [PubMed]
- Galveia JN, Travassos A, Quadros FA, da Silva Cruz LA. Computer aided diagnosis in ophthalmology: deep learning applications. In: Classification in BioApps: Automation of Decision Making. Cham: Springer International Publishing; 2017:263-93.
- Wolffsohn JS, Arita R, Chalmers R, Djalilian A, Dogru M, Dumbleton K, Gupta PK, Karpecki P, Lazreg S, Pult H, Sullivan BD, Tomlinson A, Tong L, Villani E, Yoon KC, Jones L, Craig JP. TFOS DEWS II Diagnostic Methodology report. Ocul Surf 2017;15:539-74. [Crossref] [PubMed]
- Lin H, Yiu SC. Dry eye disease: A review of diagnostic approaches and treatments. Saudi J Ophthalmol 2014;28:173-81. [Crossref] [PubMed]
- Zhang X, Chen Q, Chen W, Cui L, Ma H, Lu F. Tear dynamics and corneal confocal microscopy of subjects with mild self-reported office dry eye. Ophthalmology 2011;118:902-7. [Crossref] [PubMed]
- Storås AM, Strümke I, Riegler MA, Grauslund J, Hammer HL, Yazidi A, Halvorsen P, Gundersen KG, Utheim TP, Jackson CJ. Artificial intelligence in dry eye disease. Ocul Surf 2022;23:74-86. [Crossref] [PubMed]
- Wang MH, Xing L, Pan Y, Gu F, Fang J, Yu X, Pang CP, Chong KKL, Cheung CYL, Liao X, Fang X. AI-based advanced approaches and dry eye disease detection based on multi-source evidence: cases, applications, issues, and future directions. Big Data Min Anal 2024;7:445-84.
- Di Cello L, Pellegrini M, Vagge A, Borselli M, Ferro Desideri L, Scorcia V, Traverso CE, Giannaccare G. Advances in the noninvasive diagnosis of dry eye disease. Appl Sci 2021;11:10384.
- Yedidya T, Hartley R, Guillon JP, Kanagasingam Y. Automatic dry eye detection. In: International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI). Berlin: Springer; 2007:792-9.
- Zeev MS, Miller DD, Latkany R. Diagnosis of dry eye disease and emerging technologies. Clin Ophthalmol 2014;8:581-90. [Crossref] [PubMed]
- Remeseiro B, Bolon-Canedo V, Peteiro-Barral D, Alonso-Betanzos A, Guijarro-Berdiñas B, Mosquera A, Penedo MG, Sánchez-Maroño N. A methodology for improving tear film lipid layer classification. IEEE J Biomed Health Inform 2014;18:1485-93. [Crossref] [PubMed]
- Vicnesh J, Oh SL, Wei JKE, Ciaccio EJ, Chua KC, Tong L, Acharya UR. Thoughts concerning the application of thermogram images for automated diagnosis of dry eye – a review. Infrared Phys Technol 2020;106:103271.
- Su TY, Liu ZY, Chen DY. Tear film break-up time measurement using deep convolutional neural networks for screening dry eye disease. IEEE Sensors Journal 2018;18:6857-62.
- Yedidya T, Carr P, Hartley R, Guillon JP. Enforcing monotonic temporal evolution in dry eye images. In: International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI). Berlin: Springer; 2009:976-84.
- Yedidya T, Hartley R, Guillon JP. Automatic detection of pre-ocular tear film break-up sequence in dry eyes. In: 2008 Digital Image Computing: Techniques and Applications (DICTA). IEEE; 2008:442-8.
- He K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 2016:770-8.
- Tajbakhsh N, Shin JY, Gurudu SR, Hurst RT, Kendall CB, Gotway MB, Liang Jianming. Convolutional Neural Networks for Medical Image Analysis: Full Training or Fine Tuning? IEEE Trans Med Imaging 2016;35:1299-312. [Crossref] [PubMed]
- Zheng Q, Wang L, Wen H, Ren Y, Huang S, Bai F, Li N, Craig JP, Tong L, Chen W. Impact of Incomplete Blinking Analyzed Using a Deep Learning Model With the Keratograph 5M in Dry Eye Disease. Transl Vis Sci Technol 2022;11:38. [Crossref] [PubMed]
- Xie X, Niu J, Liu X, Chen Z, Tang S, Yu S. A survey on incorporating domain knowledge into deep learning for medical image analysis. Med Image Anal 2021;69:101985. [Crossref] [PubMed]
- Bilkhu P, Sivardeen Z, Chen C, Craig JP, Mann K, Wang MTM, Jivraj S, Mohamed-Noriega K, Charles-Cantú DE, Wolffsohn JS. Patient-reported experience of dry eye management: An international multicentre survey. Cont Lens Anterior Eye 2022;45:101450. [Crossref] [PubMed]
- Ramos L, Barreira N, Pena-Verdeal H, Giráldez MJ, Yebra-Pimentel E. Computational approach for tear film assessment based on break-up dynamics. Biosyst Eng 2015;138:90-103.
- Ramos L, Barreira N, Mosquera A, Pena-Verdeal H, Yebra-Pimentel E. Break-up analysis of the tear film based on time, location, size and shape of the rupture area. In: International Conference on Image Analysis and Recognition (ICIAR). Berlin: Springer; 2013:695-702.
- Vyas AH, Mehta MA, Kotecha K, Pandya S, Alazab M, Gadekallu TR. Tear film breakup time-based dry eye disease detection using convolutional neural network. Neural Comput Appl 2024;36:143-61.
- Remeseiro B, Barreira N, Sánchez-Brea L, Ramos L, Mosquera A. Machine learning applied to optometry data. In: Advances in Biomedical Informatics. Cham: Springer; 2017:123-60.
- Abdelmotaal H, Hazarbasanov R, Taneri S, Al-Timemy A, Lavric A, Takahashi H, Yousefi S. Detecting dry eye from ocular surface videos based on deep learning. Ocul Surf 2023;28:90-8. [Crossref] [PubMed]
- Cebreiro E, Ramos L, Mosquera A, Barreira N, Penedo MF. Automation of the tear film break-up time test. In: Proceedings of the 4th International Symposium on Applied Sciences in Biomedical and Communication Technologies; 2011:1-5.
- Shimizu E, Ishikawa T, Tanji M, Agata N, Nakayama S, Nakahara Y, Yokoiwa R, Sato S, Hanyuda A, Ogawa Y, Hirayama M, Tsubota K, Sato Y, Shimazaki J, Negishi K. Artificial intelligence to estimate the tear film breakup time and diagnose dry eye disease. Sci Rep 2023;13:5822. [Crossref] [PubMed]
- Remeseiro B, Mosquera A, Penedo MG. CASDES: A Computer-Aided System to Support Dry Eye Diagnosis Based on Tear Film Maps. IEEE J Biomed Health Inform 2016;20:936-43. [Crossref] [PubMed]
- El Barche FZ, Benyoussef AA, El Habib Daho M, Lamard A, Quellec G, Cochener B, Lamard M. Automated tear film break-up time measurement for dry eye diagnosis using deep learning. Sci Rep 2024;14:11723. [Crossref] [PubMed]
- Yang HK, Che SA, Hyon JY, Han SB. Integration of Artificial Intelligence into the Approach for Diagnosis and Monitoring of Dry Eye Disease. Diagnostics (Basel) 2022;12:3167. [Crossref] [PubMed]
- Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ł, Polosukhin I. Attention is all you need. Adv Neural Inf Process Syst 2017;30:5998-6008.
- Verma R, Kumar N, Patil A, Kurian NC, Rane S, Graham S, et al. MoNuSAC2020: A Multi-Organ Nuclei Segmentation and Classification Challenge. IEEE Trans Med Imaging 2021;40:3413-23. [Crossref] [PubMed]
- Wu L, Huang Y, Lv T, Xiao C, Wang Y, Zhao S. Advances in AI-assisted quantification of dry eye indicators. Front Med. 2025;12:1628311.
- Ronneberger O, Fischer P, Brox T. U-Net: convolutional networks for biomedical image segmentation. In: Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015. Cham: Springer; 2015:234-41.
- Tan M, Le QV. EfficientNetV2: smaller models and faster training. Proceedings of the 38th International Conference on Machine Learning, PMLR; 2021;139:10096-106.
- Rahman H, Bukht TFN, Imran A, Tariq J, Tu S, Alzahrani A. A Deep Learning Approach for Liver and Tumor Segmentation in CT Images Using ResUNet. Bioengineering (Basel) 2022;9:368. [Crossref] [PubMed]
- Zhao H, Shi J, Qi X, Wang X, Jia J. Pyramid scene parsing network. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2017:2881-90.

