Data Set Paper
Data Set Paper
*Correspondence:
Shashi Kant Abstract
drshashikant38@[Link] Oral cancer poses a critical global health challenge, with early detection significantly
1
University of the Cumberlands,
Williamsburg, USA improving patient survival rates and treatment outcomes. This study proposes an
2
Trine University, Angola, USA advanced deep learning-based diagnostic model, LightSE-MobileViT, specifically
3
Pacific States University, Los designed to classify oral cancer using medical imaging. The Oral Cancer Classification
Angeles, USA
4
Bule Hora University, Hagere dataset used in this study comprises clinically validated lip and tongue images
Maryam, Ethiopia collected from various ENT hospitals in Ahmedabad. The original dataset consisted
of 131 images (87 cancerous and 44 non-cancerous). To address class imbalance
and enhance model generalizability, data augmentation techniques were employed,
expanding the dataset to 981 images with equal distribution across both classes. Our
proposed model, LightSE-MobileViT, integrates a lightweight convolutional neural
network (CNN) backbone consisting of sequential convolutional layers enhanced
with batch normalization and rectified linear unit activations. To further enrich feature
representation and spatial attention, a Squeeze-and-Excitation block is embedded
after the third convolutional layer. Subsequently, a MobileViT transformer encoder
is employed, effectively capturing global contextual information through efficient
multi-headed self-attention mechanisms. Experimental evaluations revealed that
LightSE-MobileViT achieved superior diagnostic performance, attaining an accuracy
of 98.39%, precision and recall values approaching 1.00 for both cancerous and
non-cancerous categories, a macro F1-score of 0.98, and an ROC-AUC of 1.00.
Comparative analysis demonstrated notable improvements over benchmark models,
including CST-CNN (98% accuracy), MobileNetV2 (97% accuracy), DenseNet121 (97%
accuracy), and InceptionV3 (90% accuracy). The exceptional performance of LightSE-
MobileViT underscores its robust capability and clinical applicability, suggesting
significant potential for deployment in automated oral cancer screening, thus
facilitating early detection and timely intervention.
Keywords Oral cancer detection, Lightweight deep learning, MobileViT transformer,
Squeeze-and-excitation (SE) module, Medical imaging
© The Author(s) 2025. Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use,
sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the
source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article
are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the
article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to
obtain permission directly from the copyright holder. To view a copy of this licence, visit [Link]
Kabir et al. Discover Artificial Intelligence (2025) 5:173 Page 2 of 21
1 Introduction
Oral cancer represents a major global health burden, ranking among the most preva-
lent and fatal forms of cancer worldwide. According to the World Health Organization
(WHO), approximately 377,713 new cases and 177,757 deaths are reported annually,
underscoring its substantial morbidity and mortality rates. Early and accurate detec-
tion remains critical to improving survival outcomes, as late-stage diagnosis significantly
diminishes treatment efficacy and prognosis.
Conventional diagnostic methods for oral cancer primarily include visual inspec-
tion, biopsy, histopathological analysis, and imaging modalities such as CT and MRI.
Although widely used, these approaches are often invasive, expensive, time-consuming,
and prone to observer variability, potentially delaying timely diagnosis and intervention.
This highlights the urgent need for reliable, non-invasive, and efficient diagnostic tools
to facilitate early detection and improve patient management. Advances in artificial
intelligence, particularly in DL, have revolutionized computer-aided diagnosis (CAD)
systems, offering powerful tools for medical imaging analysis. Among these, CNN-based
models have demonstrated outstanding performance in classification and segmentation
tasks by automatically learning hierarchical representations from medical images. How-
ever, standard CNN architectures often struggle to capture complex global relationships
within images, limiting their diagnostic [Link] overcome these challenges, hybrid
CNN-Transformer models have emerged, combining the strong local feature extraction
capabilities of CNNs with the global contextual understanding of Transformer networks.
Motivated by these advancements, this study proposes a novel hybrid model, LightSE-
MobileViT, specifically designed for efficient and precise classification of oral cancer
images. The architecture integrates lightweight CNN layers, SE blocks for channel-
wise feature recalibration, and MobileViT modules to capture long-range dependencies
alongside local textures, thereby enhancing feature [Link] primary objec-
tive of this research is to develop and evaluate the LightSE-MobileViT model for dis-
tinguishing between cancerous and non-cancerous images of the oral cavity. To ensure
robust and balanced training, we utilized an augmented version of the publicly available
Oral Cancer Classification (OCI) dataset, originally consisting of 131 clinically validated
images (87 cancerous and 44 non-cancerous), collected from ENT hospitals in Ahmed-
abad. Through data augmentation, the dataset was expanded to 981 images, with equal
representation of both categories. This enhanced dataset enabled a comprehensive eval-
uation of the proposed model’s effectiveness in addressing the diagnostic challenges of
oral cancer.
The contributions of this paper include:
This work underscores the transformative potential of advanced deep learning tech-
niques in medical diagnostics, paving the way for more accurate, accessible, and timely
cancer care.
2 Related work
Sampath et al. [1] proposed a deep learning fusion strategy integrating ResNet-50 fea-
tures with logistic regression to classify oral cancer from lips and tongue images, achiev-
ing a high accuracy of 97.8% through hierarchical feature fusion and SGD optimization.
Bansal et al. [2] performed a comparative analysis of machine learning classifiers includ-
ing Neural Networks, KNN Ensemble, and SVM on oral cancer datasets, highlighting
the effectiveness of segmentation and feature extraction in enhancing classification
accuracy. Mannelli et al. [3] introduced a defect-based classification system for tongue
cancer resection, offering a structured surgical approach to optimize functional and
oncologic outcomes through stratified defect types and tailored reconstructions.
Warnakulasuriya et al. [4] presented a WHO-endorsed consensus updating the clas-
sification of oral potentially malignant disorders (OPMDs), incorporating new entities
and emphasizing standardized nomenclature for improved diagnosis and risk assess-
ment. Alabi et al. [5] compared machine learning models to predict locoregional recur-
rence in early-stage oral tongue cancer, with the Boosted Decision Tree model achieving
81% accuracy, outperforming traditional DOI-based prognosis. Muller and Tilakaratne
[6] reviewed WHO’s 2022 updates on oral cavity and mobile tongue tumors, detailing
restructured sections and new inclusions like non-neoplastic lesions, enhancing the pre-
cision of clinical and pathological diagnoses. Shamim et al. [7] demonstrated that trans-
fer learning with deep convolutional neural networks can classify five types of tongue
lesions with near-human performance, supporting automated pre-screening of oral can-
cer in low-resource settings.
Ansarin et al. [8] proposed a five-tier classification system for glossectomy procedures
based on surgical anatomy and tumor spread, establishing a unified terminology to
improve surgical communication and research consistency. Bansal et al. [9] introduced
a CNN-based model named Oral_Cancer_Detection, trained on a small Kaggle dataset,
achieving 94% accuracy in classifying lips and tongue images as cancerous or non-can-
cerous with low complexity. Dwivedi et al. [10] developed a fusion-based deep learning
model using augmented lip and tongue images, achieving 94.62% accuracy and outper-
forming several state-of-the-art models, demonstrating its reliability for early-stage oral
cancer detection. Rajaguru and Prabhakar [11] compared GMM, MLP, and ELM clas-
sifiers for oral cancer classification using TNM staging data, with GMM showing the
highest average accuracy of 94.18%, emphasizing its clinical diagnostic potential.Müller
[12] reviewed the 2017 WHO classification highlighting the exclusion of oropharynx,
new entries like schwannoma, and the streamlined categorization of oral tumors, aid-
ing clarity in pathological diagnosis. Zini et al. [13] analyzed 36 years of national cancer
data, identifying site-specific incidence and survival trends; tongue and gums had the
lowest survival rates, reinforcing the need for targeted early screening and site-based
risk assessment. Bagan et al. [14] detailed the clinical presentation of OSCC, noting its
typical painless onset in high-risk sites like the tongue and floor of the mouth, and the
need for differential diagnosis from other aggressive oral malignancies. Speight and Far-
thing [15]discussed pathological distinctions between oral and oropharyngeal cancers,
Kabir et al. Discover Artificial Intelligence (2025) 5:173 Page 4 of 21
highlighting HPV’s role and prognostic factors, which are essential for developing tar-
geted therapeutic strategies. Sciubba [16] emphasised the epidemiological distinctions
between lip and intra-oral cancers, noticing a decline in the age-adjusted incidence of lip
cancer in contrast to a continuous increase in intra-oral cancer rates. The study empha-
sises the necessity of a distinct classification system to facilitate precise diagnosis and
targeted treatment strategies. Sciubba [16] underscored the importance of early diag-
nosis in enhancing the survival rates of oral cancer, focussing on key etiological factors
such as tobacco, alcohol, betel nut use, and HPV. He elucidated the progression from
dysplasia to invasive squamous cell carcinoma and emphasised that surgery is the pri-
mary treatment. Welikala et al. [17] utilised ResNet-101 and Faster R-CNN to create
deep learning models for the automated detection and classification of oral lesions as
part of the MeMoSA project. Their system demonstrated the feasibility of low-cost, AI-
assisted early oral cancer screening by achieving F1 scores of 87.07% for lesion identifi-
cation and 78.30% for referral classification.
Fig. 1 Architecture of the proposed LightSE-MobileViT model for oral cancer classification
1
H ∑
∑ W
Zc = SF 3 (c) (i, j)(4)
H ∗W
I=1 J=1
(c) (c)
S ∧ F3 = δc ∗ SF3 (6)
To address long-spatial dependencies, the feature map SF4 is passed through a Mobi-
leViT block,which replaces the high-cost AMSA block in the original LWENet. This
encoder is designed with:
Let F ∈ RH∗W ∗C be the output of Con4. Then the encoding process is:
1
H ∑
∑ W
Favg = Z2 (i, j)(10)
H ∗W
i=1 j=1
eFavg
Pout = ∑C (11)
j=1 eFavgj
Positional Encoding:
Learnable positional embeddings are added to the input tokens before the attention
mechanism to retain spatial information lost during flattening.
Normalization and Residuals:Both the multi-head attention and the FFN layers are
followed by residual connections and Layer Normalization, facilitating stable gradient
flow and improving convergence speed [36, 37].After transformer processing, the output
tokens are reshaped back to the spatial domain and concatenated with the original con-
volutional feature maps, thus preserving and integrating both local textures and global
contexts for robust [Link] final aggregated feature map is passed through a
Global Average Pooling (GAP) layer to reduce spatial dimensions, followed by two dense
layers with 128 and 64 neurons, each applying ReLU activation and dropout regulariza-
tion (dropout rate = 0.5) to prevent overfitting. Finally, a fully connected output layer
with softmax activation produces the probability scores for the two classes: cancerous
and [Link] hybrid architecture, combining lightweight convolutional pro-
cessing and transformer-based contextual learning, offers significant potential for clini-
cally deployable, interpretable, and efficient oral cancer detection (Fig. 2).
1 ∑
N
LBCE = − [xi log (yi ) + (1 − xi ) log (1 − yi )](12)
N
i=1
Contrastive Loss:
1 ∑ ∑[ ]
N N
2
LCL = xij d2ij + (1 − xij ) max (0, m − dij ) (13)
2N
i=1 j=1
where:xi is the true label, yi is the predicted probability,dij is the distance between sam-
ples i and j,m is the contrastive margin (e.g., 1).We empirically set α=β=0.5.
Final Loss:
3.6 CST-CNN
We introduce a novel convolutional neural network architecture named Multi-Scale
Residual Fusion CNN (MSRF-CNN) for accurate classification of oral cancer Fig. 3. The
model is tailored to extract discriminative features from tongue and lip images by lever-
aging multi-scale spatial representation, residual-inspired parallel paths, and deep fusion
mechanisms. This model is built upon principles similar to the CST-YOLO backbone
but customized for image classification.
Stage 1: Initial Convolutional Block,the first Convolutional block applies 32 filters of
size 3*3 followed by ReLU activation and max pooling, producing feature map SF1:
Fig. 3 Architecture of the proposed CST-CNN model for oral cancer classification
Kabir et al. Discover Artificial Intelligence (2025) 5:173 Page 9 of 21
These Outputs are concatenated along the channel axis to yield a fused multi-scale fea-
ture map SF2:
Stage 3: Deep Feature Extraction to learn more abstract representations, we apply two
additional convolutional layers:
Stage 4: Classification Head,the final feature map SF4 is flattened and passed through
fully connected layers with dropout regularization:Here, C is the number of classes, set
to 2 (cancer, non-cancer).
1 ∑
N
LSCCE = − log (pi,yi )(26)
N
i=1
were used for visualisation. Scikit-learn was used to calculate evaluation metrics. Table 2
below provides a summary of the parameters that were meticulously chosen for the
training process:
5 Comparative study
We conducted a thorough comparative analysis of our proposed LightSE-MobileViT
architecture against established baseline deep learning models, including MobileNetV2,
DenseNet121, and InceptionV3, as well as our custom-designed CST-CNN, to evaluate
diagnostic performance. To ensure fairness and consistency, all models were trained and
validated using the same preprocessed and augmented version of the Oral Cancer Clas-
sification dataset.
Table 3 summarizes the comparative results across key metrics such as accuracy, pre-
cision, recall, F1-score, and ROC-AUC. As shown in Fig. 4, LightSE-MobileViT achieved
the highest overall performance, with an accuracy of 98.39%, macro F1-score of 0.98,
and an ROC-AUC of 1.00. In the cancerous category, it obtained a precision of 1.00 and
recall of 0.97, while for the non-cancerous category, it achieved a precision of 0.96 and
recall of 1.00. These results highlight the model’s robust ability to distinguish between
classes with high [Link] superior performance is attributed to its hybrid archi-
tecture, which integrates lightweight convolutional layers, Squeeze-and-Excitation (SE)
blocks for channel-wise attention refinement, and a MobileViT transformer encoder
for modeling global contextual dependencies. The MobileViT component effectively
captures long-range spatial relationships that traditional CNNs may miss, while the SE
blocks enhance feature prioritization by emphasizing the most informative channels.
Our custom CST-CNN also demonstrated strong performance, achieving an accuracy
of 98% and a macro F1-score of 0.98. Its multi-scale design—with parallel convolutional
paths using 3 × 3, 5 × 5, and 7 × 7 kernels—allowed it to capture diverse receptive fields
and spatial textures. However, its lack of a transformer-based global attention mecha-
nism slightly limited its ability to model complex contextual interactions compared to
[Link] the baseline models, MobileNetV2 attained an accuracy of
97%, a macro F1-score of 0.96, and an ROC-AUC of 0.97, balancing efficiency with high
diagnostic performance. DenseNet121 performed similarly, with slightly higher recall
and ROC-AUC, reflecting its strength in gradient propagation via dense connectivity.
In contrast, InceptionV3 showed the lowest overall performance, with 90% accuracy, a
macro F1-score of 0.88, and a ROC-AUC of 0.95. Although it achieved a perfect recall of
1.00 in the cancerous class, it underperformed in the non-cancerous class (recall = 0.70),
indicating sensitivity to class [Link], these results validate the clinical appli-
cability and architectural advantages of the LightSE-MobileViT model, supporting its
potential use in real-world oral cancer screening and early detection (Figs. Figures 5, 6).
Fig. 4 Prediction results of the proposed LightSE-MobileViT model on oral cavity images
decisions and build trust in deployment. The proposed LightSE-MobileViT model dem-
onstrates strong potential for practical clinical use, offering a reliable, efficient, and accu-
rate solution for early oral cancer detection and decision support.
Fig. 5 ROC curve comparison of deep learning models for oral cancer classification receiver operating charac-
teristic (ROC) curves illustrating the classification performance of six models: a CST-CNN, b LightSE-MobileViT, c
InceptionV3, d MobileNetV2, and e DenseNet121
Fig. 6 Comparative Confusion Matrices of Deep Learning Models for Oral Cancer Classification Confusion matri-
ces for six models evaluated on the oral cancer classification task: a CST-CNN, b LightSE-MobileViT, c MobileNetV2,
d InceptionV3, e VGG16, and f DenseNet121
Kabir et al. Discover Artificial Intelligence (2025) 5:173 Page 17 of 21
Fig. 7 Training and Validation Accuracy and Loss Curves of Deep Learning Models for Oral Cancer Classification a
CST-CNN b VGG16 c InceptionV3 d MobileNetV2 e DenseNet121 f LightSE-MobileVi
Fig. 7 (continued)
research should focus on curating or accessing larger, multicenter datasets that encom-
pass a wider demographic spectrum to ensure broader applicability. The current model
lacks explainability features that are essential for clinician trust and regulatory approval
in medical AI systems. While the integration of SE blocks and the MobileViT encoder
enhances model interpretability at the feature level, the study does not currently incor-
porate explainable AI (XAI) techniques such as Grad-CAM, LIME, or SHAP to visu-
ally validate predictions. Incorporating such interpretability tools in future work would
allow for better transparency and human-AI collaboration in diagnostic settings.
Kabir et al. Discover Artificial Intelligence (2025) 5:173 Page 19 of 21
Author contributions
Md Firoz Kabir contributed to model optimization, baseline comparisons, and technical validation of results. Md Yousuf
Ahmad assisted with statistical analysis, performance evaluation, and drafting the results and discussion section.
Shashi Kant supported the design and implementation of key model components, contributed to literature review,
helped refine the methodology section, and handled manuscript correspondence and submission. Dr. Martin Cordero
supervised the overall research work, provided guidance on methodology, and critically reviewed the manuscript to
ensure academic rigor. Roise Uddin contributed to the review coordination and technical proofreading of the final draft.
Funding
This research received no external funding.
Data availability
The dataset used in this study is not publicly available due to privacy and data-sharing constraints but can be provided
by the corresponding author upon reasonable request. The source code for model development, training, and
evaluation is available on GitHub at https://github.com/Ir fanSadiqRahat/Oral-Cancer-Lips-and-tongue-classification.
Declarations
Ethics and consent to participate
Not applicable.
Consent for publication
Not applicable.
Competing Interests
The authors declare no competing interests.
References
1. Sampath P, Pradeepa J, Suganya R, Revathi R. Oralnet: Deep learning fusion for oral cancer identification from lips
and tongue images using stochastic gradient based logistic regression. Netw Model Anal Health Inform Bioinform.
2024;13(1):24.
Kabir et al. Discover Artificial Intelligence (2025) 5:173 Page 20 of 21
2. Bansal S, Jadon RS, & Gupta SK Features extraction and classification using machine learning classifiers for the recognition
of lips and tongue cancer. Available at SSRN 4719018; 2024.
3. Mannelli G, Arcuri F, Agostini T, Innocenti M, Raffaini M, Spinelli G. Classification of tongue cancer resection and treatment
algorithm. J Surg Oncol. 2018;117(5):1092–9.
4. Warnakulasuriya S, Kujan O, Aguirre-Urizar JM, Bagan JV, González-Moles MÁ, Kerr AR, Lodi G, et al. Oral potentially malig-
nant disorders: a consensus report from an international seminar on nomenclature and classification, convened by the
WHO collaborating centre for oral cancer. Oral Dis. 2021;27(8):1862–80.
5. Alabi RO, Elmusrati M, Sawazaki-Calone I, Kowalski LP, Haglund C, Coletta RD, Leivo I, et al. Comparison of supervised
machine learning classification techniques in prediction of locoregional recurrences in early oral tongue cancer. Int J Med
Inform. 2020;136: 104068.
6. Müller S, Tilakaratne WM. Update from the 5th edition of the World Health Organization classification of head and neck
tumors: tumours of the oral cavity and mobile tongue. Head Neck Pathol. 2022;16(1):54–62.
7. Shamim MZM, Syed S, Shiblee M, Usman M, Ali SJ, Hussein HS, Farrag M. Automated detection of oral pre-cancerous
tongue lesions using deep learning for early diagnosis of oral cavity cancer. Comput J. 2022;65(1):91–104.
8. Ansarin M, Bruschini R, Navach V, Giugliano G, Calabrese L, Chiesa F, Shah JP, et al. Classification of glossectomies: proposal
for tongue cancer resections. Head Neck. 2019;41(3):821–7.
9. Bansal S, Jadon RS, Gupta SK. Lips and tongue cancer classification using deep learning neural network. In: 2023 6th
international conference on information systems and computer networks (ISCON). IEEE; 2023, pp 1–3.
10. Dwivedi K, Patel K, Pandey JP, Garg P. An automatic robust deep learning and feature fusion-based classification method
for early diagnosis of oral cancer using lip and tongue images. In: 2024 2nd international conference on disruptive tech-
nologies (ICDT). IEEE; 2024, pp. 391–395.
11. Rajaguru H, Prabhakar SK. Performance comparison of oral cancer classification with Gaussian mixture measures and multi
layer perceptron. In: The 16th international conference on biomedical engineering: ICBME 2016. Singapore: Springer;
2017. p. 123–9.
12. Müller S. Update from the 4th edition of the World Health Organization of head and neck tumours: tumours of the oral
cavity and mobile tongue. Head Neck Pathol. 2017;11:33–40.
13. Zini A, Czerninski R, Sgan-Cohen HD. Oral cancer over four decades: epidemiology, trends, histology, and survival by
anatomical sites. J Oral Pathol Med. 2010;39(4):299–305.
14. Bagan J, Sarrion G, Jimenez Y. Oral cancer: clinical features. Oral Oncol. 2010;46(6):414–7.
15. Speight PM, Farthing PM. The pathology of oral cancer. Br Dent J. 2018;225(9):841–7.
16. Sciubba JJ. Oral cancer: the importance of early diagnosis and treatment. Am J Clin Dermatol. 2001;2:239–51.
17. Welikala RA, Remagnino P, Lim JH, Chan CS, Rajendran S, Kallarakkal TG, Yap MH, et al. Automated detection and classifica-
tion of oral lesions using deep learning for early detection of oral cancer. IEEE Access. 2020;8:132677–93.
18. Panigrahi S, Nanda BS, Bhuyan R, et al. Classifying histopathological images of oral squamous cell carcinoma using deep
transfer learning. Heliyon. 2023;9:e13444.
19. Smirnov EA, Timoshenko DM, Andrianov SN. Comparison of regularization methods for ImageNet classification with deep
convolutional neural networks. AASRI Procedia. 2014;6:89–94.
20. Krizhevsky A, Sutskever I, Hinton GE. Imagenet classification with deep convolutional neural networks. Commun ACM.
2017;60(6):84–90.
21. Szegedy C, Liu W, Jia Y, Sermanet P, Reed S, Anguelov D, Rabinovich A et al. Going deeper with convolutions. In: Proceed-
ings of the IEEE conference on computer vision and pattern recognition. 2015, pp. 1–9.
22. Chollet F. Xception: deep learning with depthwise separable convolutions. In: Proceedings of the IEEE conference on
computer vision and pattern recognition. 2017, pp. 1251–1258.
23. Simonyan K, Zisserman A. Very deep convolutional networks for large-scale image recognition. arXiv preprint
arXiv:1409.1556. 2014.
24. He K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. In: Proceedings of the IEEE conference on
computer vision and pattern recognition. 2016, pp. 770–778.
25. Howard AG, Zhu M, Chen B, Kalenichenko D, Wang W, Weyand T, Adam H et al. MobileNets: efficient convolutional neural
networks for mobile vision applications. arXiv preprint arXiv:1704.04861. 2017.
26. Litjens G, Kooi T, Bejnordi BE, Setio AAA, Ciompi F, Ghafoorian M, Sánchez CI, et al. A survey on deep learning in medical
image analysis. Med Image Anal. 2017;42:60–88.
27. Roth HR, Lu L, Liu J, Yao J, Seff A, Cherry K, Summers RM, et al. Improving computer-aided detection using convolutional
neural networks and random view aggregation. IEEE Trans Med Imaging. 2016;35(5):1170–81.
28. Shin HC, Roth HR, Gao M, Lu L, Xu Z, Nogues I, Summers RM, et al. Deep convolutional neural networks for com-
puter-aided detection: CNN architectures, dataset characteristics and transfer learning. IEEE Trans Med Imaging.
2016;35(5):1285–98.
29. Szegedy C, Vanhoucke V, Ioffe S, Shlens J, Wojna Z. Rethinking the Inception architecture for computer vision. In: Proceed-
ings of the IEEE conference on computer vision and pattern recognition. 2016, pp. 2818–2826.
30. Ker J, Wang L, Rao J, Lim T. Deep learning applications in medical image analysis. IEEE Access. 2017;6:9375–89.
31. Mehmood A, Iqbal M, Mehmood Z, Irtaza A, Nawaz M, Nazir T, Masood M. Prediction of heart disease using deep convolu-
tional neural networks. Arab J Sci Eng. 2020;46:1–14.
32. Shamim MZM, Syed S, Shiblee M, Usman M, Ali S. Automated detection of oral pre-cancerous tongue lesions using deep
learning for early diagnosis of oral cavity cancer. arXiv preprint arXiv:1909.08987. 2019.
33. Kim DW, Lee S, Kwon S, Nam W, Cha IH, Kim HJ. Deep learning-based survival prediction of oral cancer patients. Sci Rep.
2019;9(1):1–10.
34. Mohd F, Noor N, Bakar ZA, Rajion ZA. Analysis of oral cancer prediction using features selection with machine learning. In:
The 7th international conference on information technology (ICIT 2015). 2015, pp. 283–288.
35. Krishnan MMR, Chakraborty C, Ray AK. Wavelet based texture classification of oral histopathological sections. Int J Microsc
Sci Technol Appl Educ. 2010;2(4):897–906.
36. Krishnan MMR, Acharya U, Chakraborty C, Ray A. Automated diagnosis of oral cancer using higher order spectra features
and local binary pattern: a comparative study. Technol Cancer Res Treat. 2011;10(5):443–55.
Kabir et al. Discover Artificial Intelligence (2025) 5:173 Page 21 of 21
37. Patra R, Chakraborty C, Chatterjee J. Textural analysis of spinous layer for grading oral submucous fibrosis. Int J Comput
Appl. 2012;47:975–8887.
38. Krishnan MMR, Shah P, Choudhary A, Chakraborty C, Paul RR, Ray AK. Textural characterization of histopathological images
for oral sub-mucous fibrosis detection. Tissue Cell. 2011;43(5):318–30.
39. Rahman T, Mahanta L, Chakraborty C, Das A, Sarma J. Textural pattern classification for oral squamous cell carcinoma. J
Microsc. 2018;269(1):85–93.
40. Rahman TY, Mahanta LB, Das AK, Sarma JD. Automated oral squamous cell carcinoma identification using shape, texture
and color features of whole image strips. Tissue Cell. 2020;63: 101322.
41. Rahman TY. A histopathological image repository of normal epithelium of oral cavity and oral squamous cell carcinoma.
Mendeley Data. 2019. [Link]
42. Reinhard E, Adhikhmin M, Gooch B, Shirley P. Color transfer between images. IEEE Comput Graph Appl. 2001;21(5):34–41.
43. Shih FY. Image processing and pattern recognition: fundamentals and techniques. New Jersey: Wiley; 2010.
44. Krishnan MMR, Chakraborty C, Paul RR, Ray AK. Hybrid segmentation, characterization and classification of basal cell nuclei
from histopathological images of normal oral mucosa and oral submucous fibrosis. Expert Syst Appl. 2012;39(1):1062–77.
45. Mikołajczyk A, Grochowski M. Data augmentation for improving deep learning in image classification problem. In: 2018
international interdisciplinary PhD workshop (IIPhDW). IEEE; 2018, pp. 117–122.
46. Nair V, Hinton GE. Rectified linear units improve restricted Boltzmann machines. In: Proceedings of the 27th international
conference on machine learning (ICML). 2010.
47. Srivastava N, Hinton G, Krizhevsky A, Sutskever I, Salakhutdinov R. Dropout: a simple way to prevent neural networks from
overfitting. J Mach Learn Res. 2014;15(1):1929–58.
48. Arslan S, Kaya MK, Tasci B, Kaya S, Tasci G, Ozsoy F, Dogan S, Tuncer T. Attention turkernext: investigations into bipolar
disorder detection using OCT images. Diagnostics. 2023;13:3422. [Link]
49. Smith MQP, Ruxton GD. Effective use of the McNemar test. Behav Ecol Sociobiol. 2020;74(11):1–9.
50. Thomas B, Kumar V, Saini S. Texture analysis based segmentation and classification of oral cancer lesions in color images
using ANN. In: 2013 IEEE international conference on signal processing, computing and control (ISPCC). IEEE; 2013, pp.
1–5.
51. Coşgun Baybars S, Talu MH, Danacı Ç, Tuncer SA. Artificial intelligence in oral diagnosis: detecting coated tongue with
convolutional neural networks. Diagnostics. 2025;15(8):1024.
52. Dwivedi K, Chugh B, Srivastava A, Pandey JP. Transens-network: an optimized light-weight transformer and feature fusion
based approach of deep learning models for the classification of oral cancer. Int J Comput Model Appl. 2024;1(1):32–44.
53. Sampath P, Sasikaladevi N, Vimal S, Kaliappan M. Oralnet: deep learning fusion for oral cancer identification from lips and
tongue images using stochastic gradient based logistic regression. Netw Model Anal Health Inform Bioinform. 2024.
[Link]
54. Bansal S, Jadon RS, Gupta SK. Lips and tongue cancer classification using deep learning neural networks. Gwalior: Depart-
ment of Computer Science & Applications Jiwaji University; 2023.
Publisher's Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.