Deep Learning for Cervical Cancer Detection
Deep Learning for Cervical Cancer Detection
[Link]
Abstract
The conventional approach of cervical cancer classification mostly depends on pathologists’
expertise, which is typically less precise. Over the last 75 years, the incidence and death
rates of cervical cance6r have decreased dramatically thanks in large part to the widespread
use of colposcopy, an essential component of cervical cancer prevention. On the other hand,
misdiagnoses and a decline in diagnostic efficacy have resulted from the increasing workload
in visual screening. In deep learning, cervical cancer type classification has shown improved
performance using medical image processing and Convolutional Neural Network (ConvNet)
models, namely the ResNet50 model and Cervix Ensemble Network (CervixNET). In order
to automatically identify cervical cancer from colposcopy images, this research presents
two ConvNet architectures. In one architecture, ResNet50 functions as a transfer learning
(TL) model, and for classification, a new model called CervixNET is created. Both models’
sensitivity, specificity, and accuracy are evaluated. With an accuracy of 82.67%, ResNet50
produces findings that are quite excellent. For ResNet50, the kappa value suggests a moderate
categorization. CervixNET, however, performs very well; its sensitivity, specificity, and kappa
score are 99.58%, 99.63%, and 99.12%, respectively, in the experimental data. Notably, the
CervixNET model’s classification accuracy of 99.23% marks a notable increase over the
ResNet50 (TL) model by 16.56%.
1 Introduction
Women in the medical field consider cervical cancer as the second most life-threatening
ailment after breast cancer. Previously, a diagnosis of late-stage cervical cancer has always
instilled a loss of hope. New innovations have however massively improved the rates of
success in detection of medical diseases aided by machines imaging [1]. According to WHO,
Cervical cancer is one of the top four most prevalent diseases in the world with nearly 570,000
new cases recorded in 2018 and representing 7.5% of the total cancer deaths among women
worldwide. Approximately 3,11,000 deaths caused by cervical cancer are reported every
year, and around 85% of such occurrences are in low- and middle-income countries [2].
In 2022, cervical cancer accounted for an estimated 6,62,000 new cases and about 3,50,000
123
Content courtesy of Springer Nature, terms of use apply. Rights reserved.
65 Page 2 of 27 S. Dayalane et al.
deaths worldwide. Despite preventive measures, such as HPV vaccines and regular screening,
low- and middle-income countries still bear a significant portion of the burden, with about
90% of cases and deaths occurring in these regions. India and China were the countries
with the highest incidence rates, with India also leading in cervical cancer-related deaths
[3]. In 2023, cervical cancer is still one of the major global health issues, though there are
variations related to income and region. According to the recent WHO statistics, up to 90%
of cervical cancer cases and deaths now occur in developing countries which have limited
access to HPV vaccination and screening. The HPV vaccine is a well-known intervention that
has been effective in decreasing incidences of cervical cancer, provided it is given at an age
before sexual intercourse begins. The data from vaccinated females in the US are reported
to show a 65% reduction in the incidence of cervical cancer among women who are aged
twenty to twenty-four years [4].
Early screening may have the potential to help save lives. Cervical cancer is said to
have occurred because of HIV in only 5% of the cases. Some factors such as availability
of equipment, standardization of screening protocols, adequate follow up, lesion detection
and treatment have an impact on scanning technology’s capability. Despite significant devel-
opments in the field of science and medicine, it is still not possible to cure some diseases
when they have become chronic especially those chronic diseases that were diagnosed many
years back [5]. So, to win the battle against the cervical cancer, firstly screening programs
and secondly role of prevention has to be underlined. The following techniques Pap smear,
Colposcopy, and HPV testing often take place during the screening procedure. Several addi-
tional technologies have been developed and patented to improve this procedure and make
it cheaper, easier and more effective. However, there is a problem with PAP smear imaging
in cervical cancer treatment because it is dependent on several histology examinations, it is
time-consuming and requires skilled personnel [6].
From a potential perspective, there are also possibilities that positive instances might be
missed while screening. HPV tests and PAP smears are typically not very sensitive, and
protrude a high cost. In less developed countries however, colposcopy therapy is widely
practiced. Colposcopy screening meets the disadvantages that PAP smear and the imaging
as well as the detection of HPV has [7]. Cancers of the cervix and others that are picked
up at an early stage are associated with a better prognosis, but the problem is that currently,
there are no symptoms to aid early diagnosis. Effective screening programs may help reduce
morbidity and mortality from mortality associated with cervical cancer. Unfortunately, the
shortage of trained medical personnel and limited funding for screening programs lead to a
low coverage of cervical cancer screening in developing and middle-income countries [8].
Colposcopy is a simple surgical procedure that is readily available and is done in an effort
to avert cervical cancer. Since timely detection and diagnosis of this type of cancer, the next
level of clinical management of the patient can be greatly enhanced. Details concerning the
images have been collected via digital colposcopy through various efforts that have been
undertaken [9].
The basic idea is to provide instruments that can assist medical professionals, regardless of
how skilled they are, to conduct colposcopy examinations. Previous studies have focused on
the conception of a computer aid for numerous actions, including but not limited to, enhance-
ment and assessment of image quality, segmentation of regions of interest, detection of the
patterns and of the unstable areas, classification of transition zone (TZ) and the probability
of carcinogenesis [10]. Computer-aided design (CAD) systems which see broad application
in the field of colposcopy automation aid in analysing images acquired through cervical
colposcopy and locating areas of interest or specific anomaly. Even though these methods
relieve the physicians in making decisions concerning the diagnosis it is however important
123
Content courtesy of Springer Nature, terms of use apply. Rights reserved.
Cervical Cancer Classification Using Deep Learning Approach Using … Page 3 of 27 65
that these doctors possess the necessary knowledge and skill as far as the diagnosis is con-
cerned. The existence of pathological regions may suggest the presence of neoplasms, and in
this regard, it is quite important to identify such lesions during the colposcopy examination
[11].
In these cases, there is also the establishment of atypical regions such as acetowhite
lesions, regard of abnormally vascular structure, abnormal mosaic patterning, and punctation
that require targeted focus. Most literature evaluations seem to agree that a method should
be designed to locate the abnormal areas in the colposcopy pictures [12]. Despite the drastic
reduction in cervical cancer deaths that have been reported in western nations, almost 90%
of the said deaths occur in less and middle-income countries. Western countries are currently
busy enjoying the benefits of AI but most countries in sub-Saharan Africa are light years
behind. In Uganda, cervical cancer detection by cytologists using Pap smear images is a
manual process that is time consuming, labour intensive and prone to errors as it is highly
reliant on human judgment [13].
Elevated diseases are leading the advance for better health, for which classification of
digital images has helped the health industry. According to various studies, the pap smear
exams that classify healthy cervical cells from the abnormal ones are quite tedious, expensive
and error-prone. The aim is to employ a ConvNet-based algorithm to classify images of
cervical cells featured in the SIPaKMeD dataset, which includes five types of cells such as
superficial-intermediate, parabasal, koilocytotic, metaplastic, and dyskeratotic [14].
In Pap smear screenings, it takes days for a pathologist to examine the millions of cells
in detail as accurate cell identification and classification are difficult seeing the overlap
of the cells in the images. DL models have been applied in the recognition of cells and
materials in pap smear images, but these aspects render classification complicated. Another
issue adding to the complexity is the lack of sufficient annotated data in this domain. In an
attempt to overcome these hurdles, a new approach based on DL for screening of cervical
cancer based on the colposcopy images has been proposed. Colposcopy is characterized by
non-invasiveness and relative ease of use which can improve the acceptance of this method
by patients. However, the amount of colposcopy exam datasets is small when compared to
other screening tests. It has the potential of reducing the number of colposcopy examinations
through the quick determination of whether further diagnostic examinations are required.
This paper presents a model for predicting cervical cancer based on colposcopy images and
intends to help speed up the cervical cancer screening process. The main inputs of this study
consist of:
123
Content courtesy of Springer Nature, terms of use apply. Rights reserved.
65 Page 4 of 27 S. Dayalane et al.
screening and has potential to improve diagnosis accuracy and facilitate efficient mass
screening.
The structure of this study is organized as follows: In Sect. 2, previous studies in relation
to cervical screening are reviewed in order to provide an understanding and rationale for
the study. Section 3 gives the concepts and development of the model CervixNET that is
intended to enhance cervical cancer screening processes. Section 4 discusses the findings from
the implementation of CervixNET in terms of its effectiveness as well as its shortcomings.
Lastly, Sect. 5 draws the study to a close by summarizing the most important outputs of the
study as well as suggesting some areas for follow up work and research in this area.
2 Related Work
Ishak et al. [15] proposed a diagnostic system utilizing ConvNets and vision transform-
ers (ViTs). The study incorporated data augmentation and ensemble learning techniques to
improve the diversity of the dataset and enhance model accuracy. A thorough comparison
was conducted on the SIPaKMeD pap-smear dataset, evaluating 40 ConvNet-based models
and over 20 ViT-based models. The findings highlighted that data augmentation significantly
improved the performance of ViT models, which outperformed their ConvNet counterparts.
Sher et al. [16] aimed to design DL models for the automated segmentation of cervi-
cal cancer, thereby bypassing the need for large datasets and traditional classifiers. The
research utilized pre-trained deep neural networks applied to pap smear images from the
Herlev database. Among the 13 models tested, DenseNet-201 emerged as the most accurate,
showcasing its potential in cervical cancer segmentation tasks.
Xia et al. [17] introduced a global contextual aware module implemented via the Region
Proposal Network (RPN) to enhance the spatial relationship between background and fore-
ground elements in cervical cancer images. Their deformable and global context aware
(DGCA) RCNN model was evaluated using the ‘Digital Human Body’ Vision Challenge
dataset. The experiments demonstrated a notable improvement in mean average precision
(mAP) by 6–9% compared to previous methods.
Mamunur et al. [18] developed DeepCervix, a hybrid deep feature fusion system, to
enhance manual screening processes. Using the SIPaKMeD dataset for training and testing,
DeepCervix achieved impressive accuracy rates for binary, three-class, and five-class classifi-
cations at 99.85%, 99.38%, and 99.14%, respectively. This work underscored the robustness
and reliability of hybrid approaches in cervical cancer diagnosis.
Mohammed et al. [19] presented a ConvNet-based framework for classifying cervix cells
into categories such as Normal, Abnormal, and Normal with cancerous changes. The study
emphasized early detection of cervical cancer, achieving high accuracy and reduced testing
time. The research highlighted the efficiency and reliability of ConvNet models in medical
diagnostics.
Wei et al. [20] introduced the 3cDe-Net, a cervical cancer cell detection network that com-
bined a backbone network and a detection head to enhance resolution and feature extraction.
Utilizing group convolution and dilated convolution for multiscale feature extraction, the
network was able to accurately detect small cells. The detection head also generated adaptive
cancer cell anchors using unsupervised clustering, achieving a mean average accuracy of
50.4%.
Yuta et al. [21] employed the YOLOv4 algorithm for abnormal cell detection and ResNeSt
for classification, using 919 cell images from liquid-based cervical cytological samples. The
123
Content courtesy of Springer Nature, terms of use apply. Rights reserved.
Cervical Cancer Classification Using Deep Learning Approach Using … Page 5 of 27 65
Bethesda system was used for annotations, and label smoothing was applied to prepare the
data. The model achieved a mean accuracy of 90.5% and an F-measure of 70.5%, demon-
strating its ability to accurately classify atypical cells.
Anant R. Bhatt et al. [22] explored the use of TL and ConvNets for multiclass classification
of cervical cells captured from whole-slide images (WSI). The application of progressive
resizing in ConvNet model training led to significant improvements. The approach achieved
outstanding accuracy of 99.70% and precision of 99.76% on the SIPaKMeD dataset, as well
as strong performance on the Herlev dataset.
Elakkiya et al. [23] proposed the FSOD-GAN framework for automating cervical spot
identification and lesion classification. The FR-ConvNet model was employed to hierarchi-
cally classify three types of cervical cancer lesions. Among the 1,993 participants involved,
the system achieved a prediction accuracy of 99%, underscoring its effectiveness as a cost-
efficient screening method.
Yoon et al. [24] assessed the performance of DL architectures in cervical cancer classi-
fication using original images, acetowhite mask images, and RGB channel superposition.
The model combining original and acetowhite mask images recorded the highest accuracy of
81.31% and an AUC of 0.817, demonstrating the benefits of image processing in improving
classification outcomes.
Shahin et al. [25] implemented an ensemble classifier that integrated salient features
extracted from multiple classifiers, including Decision Trees, Random Forest, Gaussian Naïve
Bayes, and Support Vector Machines. The classifier achieved accuracy rates of 98.06% and
95.45% on two datasets, highlighting its robustness in cross-validation experiments. The
study also reported high AUC scores, demonstrating the ensemble approach’s effectiveness.
Nur et al. [26] leveraged segmentation algorithms to separate the cytoplasm and nucleus
in histological cell images. Using a graph convolutional network (GCN) built from strongly
linked characteristics, the study achieved accuracies of 99.11% and 98.18% on the SIPaKMeD
and Herlev datasets, respectively. This work demonstrated the efficacy of ranking diagnostic
features to enhance classification accuracy.
Meenu et al. [27] proposed a four-step cancer neural network system comprising prepro-
cessing, outlier removal, dimensionality reduction via PCA, and classification. The approach
was tested on high-dimensional datasets, demonstrating strong performance in classifying
normal and abnormal data. Metrics such as accuracy and precision validated the robustness
of the method.
Table 1 demonstrated that the proposed model addresses several limitations of previous
studies by combining ResNet50 for TL and a novel architecture called CervixNET. ResNet50
reduces dependency on large, dataset-specific training data by leveraging pre-trained features,
improving generalization across diverse datasets. CervixNET, tailored specifically for cervi-
cal cancer classification.
An image from a colposcopy is essential for early cancer diagnosis. The TZ must be examined
under a microscope in order to determine whether patients with abnormal cytology should
be sent for further testing. Characterising the TZ is crucial to our investigation. Research on
observer heterogeneity in TZ form appraisal and squamocolumnar junction (SCJ) visibility,
as well as quantitative measurement of intra- and interobserver agreement in TZ contour
tracing [31], are still lacking despite the recognition of intra-and interobserver variability
123
Content courtesy of Springer Nature, terms of use apply. Rights reserved.
65 Page 6 of 27 S. Dayalane et al.
[15] ViT, ConvNets, Ensemble learning Heavy reliance on physician’s skill for Pap
smear tests, data quality heterogeneity
[16] DenseNet-201 Limited to a specific dataset, requires
transfer learning
[17] DGCA-RConvNet Requires complex convolution techniques,
dataset specificity
[18] DeepCervix Prone to false positives in manual testing,
requires large amounts of data for training
[19] Ensemble DL Requires heavy computational resources
for WSI analysis
[20] 3cDe-Net Moderate accuracy (50.4%), complex setup
with unsupervised clustering
[21] YOLOv4, ResNeSt Moderate accuracy and F-measure (90.5%,
70.5%), somewhat limited scope
[22] ConvNet Limited to Herlev and SIPaKMeD datasets,
may not generalize across all datasets
[23] FSOD-GAN, FR-ConvNet Dataset specificity, may not generalize well
across other populations
[24] DL Relatively low accuracy compared to other
models (81.31%)
[25] Ensemble Classifier Requires extensive pre-processing, model
complexity for handling multiple
classifiers
[26] GCN Requires manual segmentation, feature
engineering complexity
[27] Cancer Neural Network, PCA May not be as effective for smaller
datasets, limited to certain types of data
123
Content courtesy of Springer Nature, terms of use apply. Rights reserved.
Cervical Cancer Classification Using Deep Learning Approach Using … Page 7 of 27 65
to determine the new border that separates the squamous epithelium from the columnar
epithelium. Figure 2 illustrates several examples of images taken from the dataset that have
been classified as type 1, type 2, and type 3. Green light is used to capture the core image in
green, which helps to improve visibility of the cervical region.
3.1 Materials
The dataset includes a total of 5,679 colposcopy images obtained from Smartphone ODT and
Intel’s cervical screening data collection initiatives. The diagnostic images are categorised
according to how visible the transition zone is in each image. All sensitive data pertaining to
the cases has been eliminated before analysis. Using diagnostic records, the data is first split
into three categories: type 1, type 2, and type 3. MATLAB image labeller apps help identify
the area of interest (ROI) within the cervical images by referencing a pretrained dataset [32].
This ROI, also known as the TZ of the clinic, includes the centre region where lesions are
usually seen. The ROI is used to annotate, mark, and outline the initial images in preparation
for further analysis. The collection includes 691 instances of type 1, and 3126 cases of type
2, and 1862 cases of type 3. The dataset is unbalanced, as shown by an investigation, mainly
because of the uneven distribution of images across the different kinds. In model training,
this mismatch increases the danger of overfitting. In order to solve this, an oversampling
approach is used, which includes copying type 1 and type 3 images in order to equal the
quantity of type 2 images. This yields a total of 9378 images over time.
Furthermore, techniques for data augmentation are used in order to maximise the amount
of training data. By using methods like data augmentation—rotation, brightness modifica-
tion, cropping, and randomization, for example—the model’s resilience is increased and the
danger of overfitting is reduced. The total number of images has increased to 11,266 after the
procedure of image enhancement has been completed. The subsequent step involves resizing
all of the processed image data to a size of 227 × 227 in order to conform to the input require-
ments of the ConvNet model. It is possible to divide the dataset into three distinct categories:
training data, which includes 7,498 images; validation data, which includes 1,884 images;
123
Content courtesy of Springer Nature, terms of use apply. Rights reserved.
65 Page 8 of 27 S. Dayalane et al.
Fig. 2 Dataset sample images Intel & MobileODT Cervical Cancer Screening
and testing data, which includes 1,884 images. Visual representations of the implementation
of the data augmentation approach are shown in Fig. 3, respectively. This technique involves
enhancing the dataset before it is used for the training of neural networks.
ConvNet models have achieved considerable acceptance across a variety of varied disci-
plines, with their effectiveness being proved most prominently in applications related to
medical imaging. In the field of computer vision, one of these crucial tasks is the analysis of
colposcopy images for the purpose of detecting cervical cancer. This task is quite challenging.
As far as distinguishing various types of cervical carcinoma, for instance type 1, type 2 and
type 3, ConvNet approaches are better than the traditional methods of extracting features.
In evaluating the capabilities of ConvNet in conducting diagnosis of cervical lesions, we
calibrate the proposed TL ResNet50 model by selectively freezing its upper layers and then
subjecting it to an extensive cervical image dataset. With the introduction of the CervixNET
architecture, we are able to take use of the inherent benefits of depth and parallel convolu-
tional filtering in order to enhance the extraction of discrete cervical cancer characteristics
from colposcopy images.
123
Content courtesy of Springer Nature, terms of use apply. Rights reserved.
Cervical Cancer Classification Using Deep Learning Approach Using … Page 9 of 27 65
Fig. 3 Input images from dataset and performed image augmentation technique
This allows us to advance the state of the art in this very important medical sector. In the
proposed model, separate convolutional layers are used, such as the standard convolutional
layer that sits at the very beginning of the network after a single convolution filter and
multiple convolution layers which are aimed at obtaining different features from same input.
The elimination of biassed components is accomplished by the use of numerous convolutional
filters, which in turn helps to mitigate the phenomena of overfitting occurs. The construction
of this model is based on three primary phases: (1) the preprocessing of data, (2) the training
of the ConvNet model, and (3) the categorization of the findings. The CervixNET model
comprises a total of 15 convolutional layers alongside 14 activation layers, 5 max pooling
layers and 4 cross channel normalization layers. To analyze the parameters generated from
the trained model, the test data are first computed within the model. With the design of
GoogleNet serving as a source of inspiration, the first strata are comprised of numerous levels
for functional manipulation, which are then followed by two completely linked layers that
include Softmax classification, as shown in Fig. 4. The differences in filter size that are present
inside the parallel convolutional block are shown in Table 2, which has a comprehensive
network description of the CervixNET model. This description includes the convolution and
max-pooling layers.
123
Content courtesy of Springer Nature, terms of use apply. Rights reserved.
65 Page 10 of 27 S. Dayalane et al.
This study focuses on the use of a dual DL algorithm in the evaluation of colposcopy images
as well as highlighting its importance when screening for cervical anomalies. Due to TL, the
performance of the previously trained ResNet50 model gets better in the process of detection.
Furthermore, the original CervixNET architecture is painstakingly designed from the bottom
up, built exclusively to solve the complexities of cervical anomaly identification, therefore
offering fresh insights to the field.
3.2.2 CervixNET
The conventional neural network architecture employs a solitary variant of the ConvNet filter,
and input data dimensions range from 1 × 1 to 5 × 5. Convolution with the input data generates
a thorough input data map that encompasses features that are relevant to the problem at hand.
The motivation for implementing these multilayer convolutional filters is the fundamental
idea of combining many convolutional filters. This technique aims to selectively focus on the
progressively increasing levels of the architecture designed for the extraction of features that
depend on classification. Through the integration of these filters, the network is capable of
efficiently capturing and representing complex patterns and distinctions present in the input
data. As a result, its capacity to learn and generalise across various tasks and domains is
significantly improved. The technique includes expanding the clusters that are embedded in
the data in order to enhance the ability of the model to recognize subtle differences in the
data.
During training, it is achieved by simultaneously using three kernel sizes: 1 × 1, 3 × 3 and
5 × 5 which are designed to obtain such features that are useful for accurate classification.
Model parameters remain constant throughout the proposed CervixNET architecture: 50
123
Content courtesy of Springer Nature, terms of use apply. Rights reserved.
Cervical Cancer Classification Using Deep Learning Approach Using … Page 11 of 27 65
epochs, 64 batches, and the Adam optimisation algorithm with a preset learning rate of
0.0001 illustrated in Table 3. Furthermore, in order to fine-tune the performance of the model,
a decaying learning rate strategy is implemented, whereby the learning rate is decreased to
0.01 every ten epochs utilising the piecewise technique. Before the training steps start, at every
iteration, the dataset is always shuffled to ensure an equal representation and distribution. Such
an arrangement eases the task of proper normalizing throughout the training process. The
structure includes a number of convolutional layers that slowly build up active characteristics
for the model increasing the ability of the model to predict and improving competitiveness
in the generation of accurate predictions. The activation map, which is depicted in Fig. 5, has
been carefully constructed to illustrate the complex details that are unique to type 1 cases.
Epochs 50
Batch size 64
Optimization algorithm Adam
Learning rate 0.0001
123
Content courtesy of Springer Nature, terms of use apply. Rights reserved.
65 Page 12 of 27 S. Dayalane et al.
The generated map represents the result of employing a solitary filter nested within the
initial convolutional layer to process the input data. On the other hand, a more comprehensive
viewpoint is presented in Fig. 5, which displays the activation map produced by combining
insights from all 64 filters functioning in the identical convolutional layer, with a particular
emphasis on instances of type 1. The main goal of this research work is to investigate deeper
any underlying facets within our ConvNet model. We wish to analyse these activation maps in
order to understand what features are important for making a correct classification of specific
classes in the dataset.
Activation functions are essential for determining the behaviour and performance of neural
networks. These functions, which are simply mathematical expressions, serve as gatekeep-
ers inside each neuron. They determine whether the neuron should activate depending on
the input’s relevance in relation to the model’s predictions. An extensively used activation
function is ReLU, which is an abbreviation for ReLU. The ReLU function functions as a
piecewise linear function, where it directly outputs the input if it is positive, and outputs zero
otherwise. The main reason is that ReLU is easy to use and nonlinearity as such there are
several benefits of ReLU usage, for example it assists in speeding up the convergence process
in the course of the training stage. The fast convergence can be understood by looking at the
one-sided linear argument of ReLU for positive arguments which aids during the gradient
calculations. In addition, ReLU reduces the saturation issues observed with other activation
units such as logistic regression and hyperbolic tangent functions. Unlike logistic regression
and hyperbolic tangent, ReLU doesn’t have the problem of a ‘vanishing gradient’ which is
when the gradient gets so small that learning hardly takes place. In addition, the ReLU may
enable the outputs of the neurons to exceed the value of 1 which would be quite interesting
123
Content courtesy of Springer Nature, terms of use apply. Rights reserved.
Cervical Cancer Classification Using Deep Learning Approach Using … Page 13 of 27 65
due to enabling more activation potential. It is for a good reason that ReLU is used in nearly
all hidden layers of the neural networks because of the numerous benefits it provides. The
fact that it is now so widely used means that it works to enhance the performance of the net-
work and to simplify the training processes. The equation for the ReLU activation function
commonly used in all hidden layers of neural networks can be expressed as:
where, x represents the input to the neuron, and f (x) represents the output. This equation
essentially outputs the input directly if it’s positive, and zero otherwise, making it a simple
yet effective activation function. A concatenation layer performs the important function of
fusing features obtained from different kernels in a neural network architecture. The layer
acts as an element-wise vector by invoking concatenation across many features which are
indicative of the same underlying concept. Thereupon the becoming of the understanding of
the network is richer and also more complex patterns and more complex connections can be
learnt. So, after concatenation normalization is introduced as an important procedure. This
type of normalization is applied on a per channel basis and aims at scaling the activations
functions of several channels correctly. Since the risk of extremely large activations that
could lead to chaotic training dynamics is reduced, the application of activation normalization
in each channel enhances the resilience of the network against overfitting. Local response
normalisation may be accomplished in two main modes: intra-channel and inter-channel.
The main task of internal normalization in a channel is to keep the mean or average
of each feature map’s activations to a given value. In other words, that internal mean is
fixed simply by doing the computation within the feature map. However, if normalization
needs to be done for several channels using activations obtained from the same channels,
the range within which the normalization is done is broader. To be precise, a cross-channel
normalisation technique is adopted in this case for local response normalisation. With this
technique, multiple channels within the same layer can be normalised, therefore normalising
it down to the level of pixels becomes a possibility. This makes it easier to understand the
input data and thus improves the performance of the network in seeing other data and learning
more generalized representations of the given input data.
yi
yi (2)
(l + (β j y 2j ))β
The Eq. (2) encapsulates the process of local response normalization, a technique fre-
quently employed in neural network architectures. Here, yi represents the ith element of the
output vector y, while l serves as a small positive constant preventing division by zero and
ensuring stability in the normalization process. The parameter ββ regulates the intensity of
the normalization, determining how strongly the magnitudes of activations are adjusted. The
β
denominator term (l + (β j y 2j )) sums the squares of all elements in y, effectively capturing
the overall activation level. By raising this sum to the power of β and adding l, the equation
dynamically scales each activation based on its relative magnitude within the context of the
entire vector y. This pixel-wise normalization aids in controlling the overall dynamic range
of activations, fostering more stable training dynamics and reducing the risk of overfitting in
neural networks. The purpose of the max-pooling layer is to reduce the dimensionality of the
features obtained from the convolutional layer, which in turn decreases the computational
complexity of the model. It does this by selecting and keeping only the highest pixel values
within predetermined kernel sizes, usually 2 × 2. After the fifth max-pooling layer, a fully
connected layer called fully connected layer 1 is added. This layer has 128 output nodes
123
Content courtesy of Springer Nature, terms of use apply. Rights reserved.
65 Page 14 of 27 S. Dayalane et al.
and has a dropout ratio of 0.5 to reduce overfitting. After FC1, a fully connected layer 2 is
created with three output nodes and a dropout ratio of 0.3% is used to mitigate overfitting.
The softmax layer produces probability distributions for each class by using the training and
validation data, which are in agreement with the ground truth labels. The output classes in
the colposcopy images, namely type 1, type 2, and type 3, are limited to three in order to
simplify the model complexity. This approach avoids the need of a larger number of nodes,
which may range from 100 to 1000. The softmax activation function is defined as an integral
component of this procedure.
g zi
f i (y) z (3)
hg
h
The expanded Eq. (3) represents the computation of the ith component f (y) within a
softmax function, where g is the base of the exponential function and zi is the input to the ith
component. This equation reflects the process of transforming raw scores or logits zizi into
probabilities. Firstly, each raw score zi is exponentiated using the base g, producing a value
that represents the likelihood of the ith class. Then, these exponentiated scores are normalized
by dividing each by the sum of all exponentiated scores across all classes. This normalization
ensures that the resulting probabilities sum up to one, allowing them to be interpreted as
probabilities of belonging to each class. Thus, the softmax function effectively maps raw
scores to a probability distribution over multiple classes, facilitating decision-making in
classification tasks. The cost function, which is the categorical crossentropy function, is
employed to evaluate the discrepancy between predicted and actual classes. The function
represented by Eq. (4) measures the extent to which class predictions are erroneous.
M
Jq ( p) − xi )
xi log( (4)
i1
Our research investigated two models, specifically CervixNET and ResNet50, in accor-
dance with the methodology we proposed. CervixNET was developed incrementally from
its inception, whereas ResNet50 was investigated through the implementation of TL. The
training objective for both models was to accurately classify the various types of cervical
cancer that are illustrated in colposcopy images.
The study employs MATLAB r2021b running on a workstation having Intel i9 CPU along
with AMD Radeon RX 7000 GPU equipped with 24 VRAM. This research adopts colposcopy
cervical cancer baseline dataset acquired from Kaggle. The dataset is split as follows: 80% for
training purposes, 10% for validation and the remaining 10% set aside for testing. The training
set contains approximately 7498 images whereas the validation set has 1884 images. Bayesian
optimization is used in the selection of parameters that include L2 regularisation, layer depth,
initial learning rate, optimizer, and momentum value. The training process includes fifty
epochs and uses a multi-GPU setup, 64 batches, and a base learning rate of 0.0001. CervixNET
and fine-tuned ResNet50 were trained with the same set of datasets and constant parameters.
Performance analysis is very broad, for example, it includes Cohen’s kappa score, precision,
sensitivity and specificity among others.
To measure multiclass classification’s performance, a confusion matrix is defined in the
exact subSect. 3. Utilization of the confusion matrix is the most important device displayed
123
Content courtesy of Springer Nature, terms of use apply. Rights reserved.
Cervical Cancer Classification Using Deep Learning Approach Using … Page 15 of 27 65
in Fig. 6 which helps in the assessment of the performance of the model of classification.
Examining the performance index in-depth also helps to assess the accuracy of the model
in terms of classifying the data instances and this is the purposed method of the model.
The aim of this task is to evaluate with some level of concern the training accuracy of two
models, that is the ResNet50 method that has been proposed and the CervixNET model.
Both models are put through rigorous training procedures that extend over 50 epochs during
which, knowledge contained in the training dataset is greatly exploited. Such a comprehensive
analysis seeks to understand the specific learning mechanisms of the models, their ability
to learn complex structures and their susceptibility or tolerance to different scenarios. This
helps in appreciating the models in terms of their relative advantages and disadvantages.
The performance of the models scaled new heights as training progressed through the
epochs until the best performance of 85% with the ResNet50 (TL) model and 98.7% with the
CervixNET model was achieved. The Fig. 7 depicts the training and validation performance
of both the proposed CervixNET and the carefully designed ResNet50 through all the epochs
respectively. It was very apparent that as the number of epochs increased, the model accuracy
improved. The CervixNET model had a noticeable validation accuracy of 94.1% as early as
the 30th epoch. Its counterpart, the ResNet50 model, which was more advanced than the other,
had its accuracy decline at first due to a very low learning rate of 0.0001 but it managed to
recover and had a final validation score of 70.2%. From the aforementioned data, it is clear
that the CervixNET model is superior to the ResNet50 model in cervical screening using
colposcopy images. This is mainly because of the robust and simple architecture of the
CervixNET model.
Figure 8 presents the visual depictions of the training and validation loss curves for the
ResNet50 and CervixNET models, respectively. Through a meticulous analysis of the val-
idation loss curve’s movement, the convergence pattern of the CervixNET model under
consideration can be discerned. CervixNET exhibits a remarkable rate of convergence, cul-
minating in a remarkably minimal loss value of 0.2982. ResNet50, on the contrary, converges
123
Content courtesy of Springer Nature, terms of use apply. Rights reserved.
65 Page 16 of 27 S. Dayalane et al.
with a significantly greater loss value of 0.9885. A noteworthy observation pertains to the
instability that is evident in the ResNet50 validation model, as opposed to the consistent
stability that is observed in the CervixNET model. The achieved stability is visually repre-
sented by a more refined loss curve, which provides additional evidence of the effectiveness
of CervixNET’s training procedure. Figure 9 depicts the training and validation loss of pro-
posed model. Figures 8 and 9 illustrate the training and validation loss curves, respectively,
123
Content courtesy of Springer Nature, terms of use apply. Rights reserved.
Cervical Cancer Classification Using Deep Learning Approach Using … Page 17 of 27 65
for the CervixNET and ResNet50 models that have been proposed. For the verification of the
convergence behavior of the proposed network the changes in the validation loss curve are
being examined. CervixNET achieves more rapid convergence than ResNet50, obtaining the
value of 0.2982 in terms of loss index and whereas ResNet50 only achieves the loss value
of 0.9885. In addition, it can be noted that the validation model of ResNet50 fluctuates but
CervixNET is able to hold its level of stability, leading in turn to a less erratic loss curve
(Fig. 10).
The provided confusion matrix summarizes the classification results for Type-1, Type-2,
and Type-3, with each row representing the true class labels and each column representing the
predicted class labels. For Type-1, the model correctly classified 645 instances, misclassified
5 instances as Type-2, and 2 instances as Type-3. For Type-2, 640 instances were correctly
classified, while 3 were misclassified as Type-1 and 5 as Type-3. For Type-3, the model
correctly classified 577 instances, misclassified 6 instances as Type-2, and 1 as Type-1. This
breakdown provides a comprehensive overview of the model’s performance, highlighting
areas of accurate classification and opportunities for improvement where errors occurred.
Figure 11 depicts the proposed model ROC curve and each curve represents a different class
(Type-1, Type-2, Type-3), with corresponding AUC values.
The confusion matrix of the proposed model to test data is depicted in Fig. 10. The
confusion matrix data illustrates the total number of accurately classified images whose labels
correspond to their predicted counterparts. Metrics including true positive ([Link]), false
positive ([Link]), true negative ([Link]), and false negative ([Link]) are delineated
within the CervixNET confusion matrix. An exposition of evaluation metrics is presented in
Table 3, encompassing positive predicted value (NPV), accuracy, sensitivity, and specificity.
It is worth mentioning that in the context of medical imaging, sensitivity and specificity,
which are derived from the confusion matrix, are the most dependable metrics for evaluating
the performance of classifiers. In model evaluation, Eq. (5), which is derived from the sum
of true positive ([Link]) and true negative ([Link]) values divided by the number of
123
Content courtesy of Springer Nature, terms of use apply. Rights reserved.
65 Page 18 of 27 S. Dayalane et al.
123
Content courtesy of Springer Nature, terms of use apply. Rights reserved.
Cervical Cancer Classification Using Deep Learning Approach Using … Page 19 of 27 65
samples, calculates the accuracy. Equation (6) evaluates the specificity of the model in terms
of the correctly identified negative instances (T. negative) in relation all the actual negative
instances (T. negative + F. positive) Typically false. Specificity is finding those negatives
which the model was supposed to find. Equation (7) deals with sensitivity assessing it through
calculating the proportion of true positives (T. positive) to the sum of true positives and false
negatives (F. negative). Sensitivity is the measure of the model to classify the true positives
correctly. Equation (8) defines the positive predicted value (PPV) which is the fraction of true
positives to the sum of true positives and false positives (T. negative + F. positive). PPV is
the measure of the number of correct positive predictions over the positive predictions made
by the model. Finally, Eq. (9) finds the NPV as the fraction of true negatives (T. negative) to
the sum of true negatives and false negatives (F. negative). This helps in knowing how good
the model is in predicting negative cases.
[Link] + [Link]
Accuracy (5)
[Link] + [Link] + [Link] + [Link]
[Link]
Specificity (6)
[Link] + [Link]
[Link]
Sensitivity (7)
[Link] + [Link]
[Link]
Positive predicted value (8)
[Link] + [Link]
[Link]
Negative predicted value (9)
[Link] + [Link]
The performance evaluations for two models, ResNet50 with TL and the proposed model,
are provided in Table 4. These evaluation criteria are sensitivity (Se), specificity (Sp), accuracy
(Acc), positive predictive value (PP value), negative predictive value (NP value), and Cohen’s
kappa coefficient (Kappa). As far as ResNet50 with TL is concerned, the empirical figures
for sensitivity stands at 68.21%, specificity at 81.13%, accuracy at 82.67%, PP value at
71.82%, NP value at 89.17%, and Kappa at 72.08%. On the other end, the suggested model
registers much higher performance on the corresponding metrics with sensitivity of 99.58%,
specificity of 99.63%, accuracy of 99.23%, PP value of 99.25%, NP value of 99.78%, and
Kappa of 99.12%. The efficacy of this model over ResNet50 with TL in respect of correctly
classifying the instances and reducing the prediction error.
The Table 5 provides a thorough summary of the classification accuracy of cervical cancer
attained by several methods. While Sami Azam [16] used a Random Forest (RF) model to get
a better accuracy of 99.05%, Ishak Pacal’s study [15] used a ConvNet to achieve an accuracy
of 97.8%. Sher Lyn Tan [17] also used a ConvNet, and she got a 96.57% accuracy rate. With an
123
Content courtesy of Springer Nature, terms of use apply. Rights reserved.
65 Page 20 of 27 S. Dayalane et al.
accuracy of 98.77%, Xia Li [18] introduced DGCA-RCNN (Deep Global and Context-Aware
Recurrent ConvNet), and Mamunur Rahaman [19] produced DeepCervix, which achieved
98.32%. Wei Wang [21] introduced 3cDe-Net (3-channel DenseNet), which achieved an
accuracy of 97.13%, whereas Mohammed Alsalatie [20] used a ConvNet and achieved an
accuracy of 99.02%. ResNet50 and CervixNET were used in the suggested investigation,
and the results showed a noteworthy high accuracy of 99.23%. A number of factors, such
as the efficient use of DL architectures like ResNet50 and CervixNET, strong preprocessing
methods, a sizable and varied dataset, and possibly novel features or methodologies, could be
responsible for the proposed model’s superior performance. Moreover, the model’s capacity
for generalisation may have been strengthened by the addition of advanced regularisation
techniques or ensemble learning strategies, which would have led to the increased accuracy
noted. Figure 12 depicts the classification accuracy comparison of proposed and other state-
of-the-art methods, and the comparison outcome shows that our proposed study outperforms
well than the other existing classification models.
The effectiveness of ResNet50 and CervixNET models can be understood from a number
of aspects, which are not the same as the studies conducted in those areas. First, because
the ResNet50 base network was pretrained on large datasets, it was easier to perform feature
extraction even with a relatively small number of cervical cancer images due to the use of TL.
Consequently, that pre-training assisted the model in acquiring rich and more generalizable
features which improved its performance. However, the most remarkable achievement was
the development of the CervixNET that was proposed for cervical image analysis. To handle
compound and diverse datasets of cervical cancer images, CervixNET employs multi Con-
vNet architectures within an ensemble method to strengthen the classification tool. The model
was able to achieve sensitivity of 99.58%, specificity of 99.63% and kappa score of 99.12%.
123
Content courtesy of Springer Nature, terms of use apply. Rights reserved.
Cervical Cancer Classification Using Deep Learning Approach Using … Page 21 of 27 65
This shows its ability to detect and classify most types of cervical cancer lesions. Addition-
ally, ResNet50 has lower accuracy compared to CervixNET’s 99.23% accuracy which is a
considerable 16.56% rise, marking it as a better model. When put side by side with simi-
lar previous works, CervixNET models outperform the previously available methods using
better accuracy and reliability in classification of cervical cancer.
The expected positive and negative values (PPV) produced by the CervixNET and
ResNet50 models are depicted in Fig. 13. As illustrated in Fig. 13b, the positive and nega-
tive predicted values of the CervixNET model are calculated with sensitivity and specificity
rates of 99.58% and 99.63%, respectively, while the infection probability remains constant
at 0.05. In contrast, as illustrated in Fig. 13a, the ResNet50 (TL) model attains sensitivity and
specificity rates of 81.13% and 68.21%, respectively, when applied to infections of varying
probabilities. The information gleaned from the incidence graph assists physicians in clas-
sifying patients into distinct categories according to their previous risk of being diagnosed
with cervical cancer.
The execution time for the proposed CervixNET model is 3 min and 8 s, whereas ResNet50
requires 4 min and 35 s, with each model utilising a batch size of 64. The CervixNET model is
composed of a grand total of 8,465,376 parameters, while ResNet50 comprises 44,549,160
parameters. Figure 13a provides the effectiveness metrics of the ResNet50 model, which
includes the Positive Predictive Value (PPV) as well as the Negative Predictive Value (NPV),
across different cut-off points. The lowest allowed threshold of 0 implies that both PPV
and NPV begin at 0 and 1 respectively; this represents total failure to positively forecast
events as well as perfect negativity predictions. It can be noted that both PPV and NPV
values improve progressively as the probability threshold is increased, which indicates that
the model can be better able to discriminate between positive and negative cases when more
source information is provided. The progressive improvement in performance emphasises the
predictive power of the models with regard to distinguishing between positive and negative
cases over a range of probability threshold values, which makes them quite useful in terms of
their overall reliability and predictive accuracy. Figure 13b shows the metric scores obtained
123
Content courtesy of Springer Nature, terms of use apply. Rights reserved.
65 Page 22 of 27 S. Dayalane et al.
123
Content courtesy of Springer Nature, terms of use apply. Rights reserved.
Cervical Cancer Classification Using Deep Learning Approach Using … Page 23 of 27 65
by the CervixNET model at various thresholds, however, focus is on the positive and negative
predictive values (use or replace PPV and NPV). With the increase of probability threshold
towards one, PPV moves from one at probability threshold five to zero probability. The
situation which has been explained seems to indicate that the more the probability threshold
is increased the more the model is able to correctly predict cases that are positive. Inversely,
NPV falls at 1 and slightly drops as probability metrics goes high which somehow implies
the model is less efficient in identifying negative cases as the thresholds go high. The metrics
in question also help to understand how efficiently the model is able to separate positive and
negative cases at different probability thresholds.
While the proposed CervixNET model demonstrates high accuracy in cervical cancer clas-
sification, certain limitations must be acknowledged. First, despite applying oversampling
techniques, class imbalance may still introduce bias, potentially affecting predictions for
underrepresented categories. Additionally, the model was trained on a specific dataset
(Intel ODT dataset), and its generalizability to other datasets with different imaging con-
ditions remains uncertain. Another limitation is the computational complexity; deploying
CervixNET in real-time clinical settings may require optimization to ensure efficiency on
resource-limited hardware. Furthermore, although activation maps provide some level of
interpretability, deep learning models inherently function as black boxes, which can impact
clinical trust and adoption. Finally, the high accuracy raises the possibility of overfitting, and
while data augmentation and cross-validation were used to mitigate this, further validation
on external datasets is necessary to confirm robustness. Future work will focus on addressing
these challenges by incorporating additional datasets and optimizing the model for practical
deployment.
5 Conclusion
The novel DL framework known as CervixNET was developed with the specific purpose of
classifying diverse forms of cervical cancer by analysing images acquired during colposcopy.
The implementation of oversampling techniques serves the purpose of preserving balance
within the image collection, thereby culminating in enhanced classification outcomes. This
research study introduces two unique models: one incorporates TL with the ResNet50 archi-
tecture, and the other introduces a completely new model called CervixNET. CervixNET
has been purposefully developed to distinguish various types of cervical cancer through
the utilisation of the ODT colposcopy image collection. The comparison study between
the suggested model and ResNet50. ResNet50 produced results with an overall accuracy
(Acc) of 82.67%, specificity (Sp) of 81.13%, and sensitivity (Se) of 68.21%. Although these
findings are noteworthy, the suggested model performs much better. The suggested model
performs remarkably well in properly recognising positive and negative situations, resulting
in a greater accuracy rate, with a Se of 99.58%, Sp of 99.63%, and Acc of 99.23%. Further-
more, the suggested model’s positive predictive value (PP value) is 99.25%, and its negative
predictive value (NP value) is 99.78%, demonstrating its effectiveness in correctly projecting
both positive and negative events. Moreover, the model proposed has a Kappa of 99.12%
which means there is a good level of demonstration that is greater than simply chance selec-
tion thus confirming the reliability and accuracy of the model in the classification of the
123
Content courtesy of Springer Nature, terms of use apply. Rights reserved.
65 Page 24 of 27 S. Dayalane et al.
various types of cervical cancer. These results underscore the significance of the proposed
model and point out the possibility of using it as a supplement in the detection of various
forms of cervical cancers in a clinical practice.
The theoretical DL model will be tested in future endeavours using a variety of datasets to
determine its resilience and generalizability in a range of settings. Additionally, the strat-
egy may be improved by combining ConvNet models with advanced image processing
approaches. With this combination it is possible to create a sophisticated diagnostic sys-
tem for the detection of cervical pre-cancerous states. It would also enable a better analysis
of new data in the future. The diagnostic system might improve the capabilities and the accu-
racy of the diagnosis by integrating ConvNets and state of art image processing techniques.
This might ensure more efficient management of cervical precancerous lesion screening and
diagnosis.
Author Contributions Author Contributions: All authors contributed equally in this work. All authors have
read this work.
Data Availability No datasets were generated or analysed during the current study.
Declarations
References
1. Mudawi NA, Alazeb A (2022) A model for predicting cervical cancer using ML algorithms. Sensors
22(11):1–21
2. Singh SK, Goyal A (2020) Performance analysis of ML algorithms for cervical cancer detection. Int J
Healthc Inf Syst Inf 15(2):1–21
3. International Agency for Research on Cancer (2022) Cervical Cancer, [Link]
type/cervical-cancer/
4. McDowell S (2023) Incidence drops for cervical cancer but rises for prostate cancer. American Cancer
Society, [Link]
5. Chauhan R, Goel A, Alankar B, Kaur H (2024) Predictive modelling and web-based tool for cervical
cancer risk assessment: a comparative study of ML models. MethodsX 12(1):102653
6. Chandran V, Sumithra MG, Karthick A, George T, Deivakani M, Elakkiya B, Subramaniam U, Manoharan
S (2021) Diagnosis of cervical cancer based on ensemble DL network using colposcopy images. Biomed
Res Int, PMCID: PMC8112909
7. Ali S, Miah S, Haque J, Rahman M, Islam K (2021) An enhanced technique of skin cancer classification
using deep convolutional neural network with transfer learning models. ML with Appl 5(1):100036
123
Content courtesy of Springer Nature, terms of use apply. Rights reserved.
Cervical Cancer Classification Using Deep Learning Approach Using … Page 25 of 27 65
8. Ghoneim A, Ghulam Muhammad M, Hossain S (2020) Cervical cancer classification using convolutional
neural networks and extreme learning machines. Futur Gener Comput Syst 102(1):643–649
9. Kalbhor M, Shinde S, Popescu DE, Jude Hemanth D (2023) Hybridization of DL pre-trained models with
ML classifiers and fuzzy min-max neural network for cervical cancer diagnosis. Diagnostics 13(7):1363
10. Liming H, Bell D, Antani S, Xue Z, Kai Y, Horning MP, Gachuhi N, Wilson B, Jaiswal MS, Befano B
(2019) An observational study of DL and automated evaluation of cervical images for cancer screening.
JNCI J Natl Cancer Inst 111(9):923–932
11. Chitra B, Kumar SS, Subbulekshmi D (2024) Prediction models applying convolutional neural network
based DL to cervical cancer outcomes. IETE J Res. [Link]
12. Sahay A, Gopakumar G, Gokulan S, Subham D, Thakur A (2024) Applying ML algorithms to investigate
cervical cancer, In: International conference on intelligent and innovative technologies in computing,
electrical and electronics (IITCEE). [Link]
13. Ahishakiye E, Kanobe F (2024) Optimizing cervical cancer classification using transfer learning with
deep gaussian processes and support vector machines. Discov Artif Intell 73(4):1–16
14. Alsubai S, Alqahtani A, Sha M, Almadhor A, Abbas S, Mughal H, Gregus M (2023) Privacy preserved
cervical cancer detection using convolutional neural networks applied to pap smear images. Comput Math
Methods Med 9676206:1–8
15. Pacal I, Kılıcarslan S (2023) DL-based approaches for robust classification of cervical cancer. Neural
Comput Appl 35(1):18813–18828
16. Tan SL, Selvachandran G, Ding W, Paramesran R, Kotecha K (2023) Cervical cancer classification from
pap smear images using deep convolutional neural network models. Interdiscip Sci Comput Life Sci
16:16–38
17. Li X, Zhenhao Xu, Shen Xi, Zhou Y, Xiao B, Li T-Q (2021) Detection of cervical cancer cells in whole
slide images using deformable and Global context aware faster RCNN-FPN. Curr Oncol 28(1):3585–3601
18. Rahaman M, Li C, Yao Y, Kulwa F, Xiangchen Wu, Li X, Wan Q (2021) DeepCervix: a DL-based
framework for the classification of cervical cells using hybrid deep feature fusion techniques. Comput
Biol Med 136:104649
19. Alsalatie M, Alquran H, Mustafa WA, Yacob YM, Alayed AA (2022) Analysis of cytology pap smear
images based on ensemble DL approach. Diagnostics 12(1):2756
20. Wang W, Tian Y, Yang Xu, Zhang X-X, Li Y-S, Zhao S-F, Bai Y-H (2022) 3cDe-Net: a cervical cancer
cell detection network based on an improved backbone network and multiscale feature fusion. BMC Med
Imag 22:130
21. Nambu Y, Mariya T, Shinkai S, Umemoto M, Asanuma H, Sato I, Hirohashi Y (2022) A screening
assistance system for cervical cytology of squamous cell atypia based on a two-step combined CNN
algorithm with label smoothing. Cancer Med 11:520–529
22. Bhatt AR, Ganatra A, Kotecha K (2021) Cervical cancer detection in pap smearwhole slide images using
convNet with transfer learning and progressive resizing. PeerJ Comput Sci 7:e348. [Link]
7717/peerj-cs.348
23. Elakkiya R, Subramaniyaswamy V, Vijayakumar V, Mahanti A (2022) Cervical cancer diagnostics
healthcare system using hybrid object detection adversarial networks. IEEE J Biomed Health Inform
26(4):1464–1471
24. Kim YJ, Woong Ju, Nam KH, Kim SN, Kim YJ (2022) RGB channel superposition algorithm with
acetowhite mask images in a cervical cancer classification DL model. Sensors 22(9):1–10
25. Ali S, Hossain M, Kona MA, Nowrin KR, Islam K (2024) An ensemble classification approach for cervical
cancer prediction using behavioral risk factors. Healthc Anal 5(1):100324
26. Fahad NM, Azam S, Sidratul Montaha M, Mukta SH (2024) Enhancing cervical cancer diagnosis with
graph convolution network: AI-powered segmentation, feature analysis, and classification for early detec-
tion. Multimed Tools Appl 83(30):75343–75367. [Link]
27. Meenu Kumari C, Bhavani R, Padmashree S, Priya R (2024) Automated cervical cancer classification
using deep neural network classifier. Int J Model Simul Sci Comput 15(01):1–14
28. Xu T, Zhang H, Huang X, Zhang S, Metaxas DN (2016) Multimodal DL for cervical dysplasia diagnosis
29. Medical Image Computing and Computer-Assisted Intervention—MICCAI, lecture notes in computer
science, pp 115–123, Springer, Cham
30. Plissiti ME, Tripoliti EE, Charchanti A, Krikoni O, Fotiadis DI (2009) Automated detection of cell nuclei
in pap stained cervical smear images using fuzzy clustering. IFMBE Proc 22(1):637–641
31. Shi J, Wang R, Zheng Y, Jiang Z, Zhang H, Yu L (2021) Cervical cell classification with graph convolu-
tional network. Comput Methods Programs Biomed 198(1):105807
32. Ksi˛ażek W, Hammad M, Pławiak P, Acharya UR, Tadeusiewicz R (2020) Development of novel ensemble
model using stacking learning and evolutionary computation techniques for automated hepatocellular
carcinoma detection. Biocybern Biomed Eng 40(4):1512–1524
123
Content courtesy of Springer Nature, terms of use apply. Rights reserved.
65 Page 26 of 27 S. Dayalane et al.
33. Su H, Yu Y, Du Q, Du P (2020) Ensemble learning for hyperspectral image classification using tangent
collaborative representation. IEEE Trans Geosci Remote Sens 58(6):3778–3790
34. Fahad NM, Sakib S, Raiaan MAK, Mukta SH (2023) SkinNet-8: an efficient cnn architecture for clas-
sifying skin cancer on an imbalanced dataset. In: 2023 International conference on electrical, computer
and communication engineering (ECCE), [Link]
35. Schwaiger C, Aruda M, Lacoursiere S, Rubin R (2012) Current guidelines for cervical cancer screening.
J Am Acad Nurse Pract 24(7):417–424
36. Shakil R, Islam S, Akter B (2024) A precise ML model: detecting cervical cancer using feature selection
and explainable AI. J Pathol Inform 15(1):100398
37. Khanarsa P, Kitsiranuwat S (2024) DL-based ensemble approach for conventional pap smear image
classification. ECTI Trans CIT 18(1):101–111
38. Pacal I (2024) MaxCerVixT: a novel lightweight vision transformer-based approach for precise cervical
cancer detection. Knowl Based Syst 289:1–15
39. Hong Z, Xiong J, Yang H, Mo YK (2024) Lightweight low-rank adaptation vision transformer framework
for cervical cancer detection and cervix type classification. Bioengineering 11(5):468
Publisher’s Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and
institutional affiliations.
123
Content courtesy of Springer Nature, terms of use apply. Rights reserved.
Cervical Cancer Classification Using Deep Learning Approach Using … Page 27 of 27 65
8 University Centre for Research & Development, Chandigarh University, Gharuan, Mohali,
Punjab 140413, India
123
Content courtesy of Springer Nature, terms of use apply. Rights reserved.
Terms and Conditions
Springer Nature journal content, brought to you courtesy of Springer Nature Customer Service Center
GmbH (“Springer Nature”).
Springer Nature supports a reasonable amount of sharing of research papers by authors, subscribers
and authorised users (“Users”), for small-scale personal, non-commercial use provided that all
copyright, trade and service marks and other proprietary notices are maintained. By accessing,
sharing, receiving or otherwise using the Springer Nature journal content you agree to these terms of
use (“Terms”). For these purposes, Springer Nature considers academic use (by researchers and
students) to be non-commercial.
These Terms are supplementary and will apply in addition to any applicable website terms and
conditions, a relevant site licence or a personal subscription. These Terms will prevail over any
conflict or ambiguity with regards to the relevant terms, a site licence or a personal subscription (to
the extent of the conflict or ambiguity only). For Creative Commons-licensed articles, the terms of
the Creative Commons license used will apply.
We collect and use personal data to provide access to the Springer Nature journal content. We may
also use these personal data internally within ResearchGate and Springer Nature and as agreed share
it, in an anonymised way, for purposes of tracking, analysis and reporting. We will not otherwise
disclose your personal data outside the ResearchGate or the Springer Nature group of companies
unless we have your permission as detailed in the Privacy Policy.
While Users may use the Springer Nature journal content for small scale, personal non-commercial
use, it is important to note that Users may not:
1. use such content for the purpose of providing other users with access on a regular or large scale
basis or as a means to circumvent access control;
2. use such content where to do so would be considered a criminal or statutory offence in any
jurisdiction, or gives rise to civil liability, or is otherwise unlawful;
3. falsely or misleadingly imply or suggest endorsement, approval , sponsorship, or association
unless explicitly agreed to by Springer Nature in writing;
4. use bots or other automated methods to access the content or redirect messages
5. override any security feature or exclusionary protocol; or
6. share the content in order to create substitute for Springer Nature products or services or a
systematic database of Springer Nature journal content.
In line with the restriction against commercial use, Springer Nature does not permit the creation of a
product or service that creates revenue, royalties, rent or income from our content or its inclusion as
part of a paid for service or for other commercial gain. Springer Nature journal content cannot be
used for inter-library loans and librarians may not upload Springer Nature journal content on a large
scale into their, or any other, institutional repository.
These terms of use are reviewed regularly and may be amended at any time. Springer Nature is not
obligated to publish any information or content on this website and may remove it or features or
functionality at our sole discretion, at any time with or without notice. Springer Nature may revoke
this licence to you at any time and remove access to any copies of the Springer Nature journal content
which have been saved.
To the fullest extent permitted by law, Springer Nature makes no warranties, representations or
guarantees to Users, either express or implied with respect to the Springer nature journal content and
all parties disclaim and waive any implied warranties or warranties imposed by law, including
merchantability or fitness for any particular purpose.
Please note that these rights do not automatically extend to content, data or other material published
by Springer Nature that may be licensed from third parties.
If you would like to use or distribute our Springer Nature journal content to a wider audience or on a
regular basis or in any other manner not expressly permitted by these Terms, please contact Springer
Nature at
onlineservice@[Link]
Different studies employed various strategies to improve the accuracy of cervical cancer diagnostic models. Anant R. Bhatt et al. utilized transfer learning and ConvNets, which led to a precision of 99.76% on the SIPaKMeD dataset . Elakkiya et al. proposed the FSOD-GAN framework, achieving prediction accuracy of 99% through hierarchical classification . Yoon et al. demonstrated that combining original images with image processing like acetowhite mask images improved accuracy to 81.31% . Another study leveraged segmentation algorithms alongside a GCN to achieve high accuracy on datasets . Notably, the development of CervixNET with multi ConvNet architectures achieved sensitivity of 99.58% and specificity of 99.63% . Additionally, Bayesian optimization in parameter selection and intensive training across epochs were used to enhance model performance .
The CervixNET model outperforms the ResNet50 model in both training and validation performance for cervical cancer classification. CervixNET achieved a validation accuracy of 94.1% by the 30th epoch, while ResNet50 had a final validation score of 70.2% . CervixNET's robustness was further demonstrated through its minimal loss value of 0.2982, compared to ResNet50's suboptimal performance at various probabilities . CervixNET's architecture, simpler yet robust, allows it to manage complex datasets more effectively than the heavier ResNet50 with its greater parameter count .
CervixNET incorporates innovative methodological approaches that enable it to outperform predecessor models in cervical cancer classification. It uses a multi ConvNet architecture within an ensemble framework, effectively strengthening its classification capabilities by combining diverse model strengths . Bayesian optimization in parameter selection enhances its learning efficiency, optimizing key factors like L2 regularization and learning rate . Moreover, its architecture is specifically tailored for cervical cancer datasets, unlike generic models, allowing it to achieve a sensitivity of 99.58% and specificity of 99.63% . These innovations collectively contribute to its superior performance. .
Transfer learning (TL) plays a crucial role in enhancing model performance for cervical cancer classification. This technique enables models like ResNet50 to perform effective feature extraction even with limited data by leveraging pre-trained features from larger datasets . By doing so, TL helps in acquiring rich, generalizable features, improving the overall accuracy and reliability of the models, although in this context, CervixNET surpassed ResNet50 despite its reliance on TL .
Image processing significantly contributes to enhancing cervical cancer classification models. Yoon et al. highlighted that the use of acetowhite mask images alongside original images improved classification accuracy to 81.31% and an AUC of 0.817 . This approach demonstrates how processed images can highlight features essential for accurate classification, augmenting raw data to provide clearer insights for deep learning architectures .
The ensemble classifier approach enhances cervical cancer classification by integrating multiple classifiers, thus leveraging diverse predictive strengths. Shahin et al. demonstrated that combining algorithms like Decision Trees, Random Forest, Gaussian Naïve Bayes, and Support Vector Machines into an ensemble classifier yielded high accuracy rates of 98.06% and 95.45% on different datasets . This multi-classifier method improves robustness by capturing varied aspects of the data, leading to superior cross-validation results and higher AUC scores compared to individual classifiers .
The CervixNET model has several advantages over other models when handling diverse cervical cancer datasets. It employs multi ConvNet architectures within an ensemble method, strengthening its classification capabilities . This allows CervixNET to achieve high sensitivity (99.58%) and specificity (99.63%) even in compound datasets . Additionally, its simpler architecture compared to ResNet50 ensures faster execution times (3 min 8 s for CervixNET vs. 4 min 35 s for ResNet50). Its ability to outperform previous methods in precision and reliability further demonstrates its superior handling of complex datasets .
The application of transfer learning and ConvNets significantly enhances performance on the SIPaKMeD dataset by allowing the model to leverage pre-trained knowledge from larger datasets to handle small or specific datasets more effectively . This approach enables the extraction of rich, generalizable features which improves classification precision to 99.76% on the SIPaKMeD dataset . Progressive resizing during ConvNet training further augments model adaptability and accuracy, ensuring comprehensive learning of detailed cellular structures .
The comparison of validation loss curves between CervixNET and ResNet50 reveals significant insights into their learning efficiency. CervixNET showcases a rapid rate of convergence, ultimately achieving a minimal loss value of 0.2982 . In contrast, ResNet50's convergence is less efficient, with fluctuations in its learning trajectory, partly due to its larger parameter size and initial lower learning rate . The stability and lower loss of CervixNET indicate its effective learning capability and adaptability to data patterns, highlighting its superiority in achieving efficient and robust convergence .
Sensitivity and specificity are critical metrics in evaluating cervical cancer classification models as they reflect the model's ability to accurately detect positive cases and correctly reject negative cases, respectively. CervixNET achieved high sensitivity (99.58%) and specificity (99.63%), indicating its reliability in both identifying cervical cancer lesions and minimizing false positives . These metrics underscore its effectiveness in correctly classifying diverse and complex datasets, making it a robust tool for clinical applications .