0% found this document useful (0 votes)
4 views32 pages

Chapter 1 Introduction

This document provides an in-depth overview of RNA, its structure, types, and functions, highlighting the significance of RNA modifications like methylation in gene regulation and disease. It discusses the roles of different RNA types, including mRNA, tRNA, and rRNA, and their involvement in protein synthesis and cellular processes. Additionally, it emphasizes the need for advanced computational methods to identify 5-methylcytosine sites in RNA, addressing the limitations of traditional experimental approaches.

Uploaded by

onemardan
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views32 pages

Chapter 1 Introduction

This document provides an in-depth overview of RNA, its structure, types, and functions, highlighting the significance of RNA modifications like methylation in gene regulation and disease. It discusses the roles of different RNA types, including mRNA, tRNA, and rRNA, and their involvement in protein synthesis and cellular processes. Additionally, it emphasizes the need for advanced computational methods to identify 5-methylcytosine sites in RNA, addressing the limitations of traditional experimental approaches.

Uploaded by

onemardan
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Chapter 1: INTRODUCTION

Introduction:
Despite its many biological roles, RNA is a complicated molecule that controls gene
transcription and regulation [1]. It can be personalized with several kinds of chemical changes.
RNA modification is the procedure of changing an RNA molecule's chemical makeup [2].
The main reason DNA became the favored genetic material in most species is thought to be the
chemical upgrading process of RNA compared to DNA, which has a stable (OH) group in a
comparable spot on the methyl group (deoxyribose).[3], [4]. The structure of the RNA molecule
was discovered in 1965 by R.W. Holley [5]. In reality, the sugar known as ribose in RNA is five-
carbon cyclical structure made up of hydrogen and oxygen. [6]. Although the final carbon group
of the ribose sugar molecule connects to a strongly reactive hydroxyl (OH) group, RNA is easily
hydrolyzed. [7]. In recent years, modifications to RNA have also been found in extra RNA
classifications, such mRNA, tRNA, rRNA, and snRNA. N1-methyladenine exists in tRNA and
rRNA, 5-methylcytosine is present in the majority of RNA molecules, and N6-methylladenosine
is present in mRNA [8]. The nitrogenous bases of RNA are uracil, guanine, cytosine, and adenine.
RNA is made up of chains of ribose nucleotide comprised of nitrogenous bases united to ribose
sugars and joined by phosphate group bonds. The highly-compound RNA (ribonucleic acid)
helps produce protein substrates and performs as as a carrier of the genetic code
(deoxyribonucleic acid) in specific diseases [9], [10]. The nitrogenous bases that can be used to
create RNA (U) are adenine (A), guanine (G), cytosine (C), and uracil (U). Nucleotide polymers
called RNA molecules are connected by covalent bonds comprised of one nucleotide's phosphate
and another nucleotide's sugar.[11]. Adenine (A), guanine (G), cytosine (C), and uracil (U) are
nitrogenous bases that can be used in RNA (U). RNA molecules are nucleotide polymers
connected by covalent bonds formed by the phosphate of one nucleotide and the sugar of another
nucleotide. [1]
Table 1 RNA Bases

A Adenine
G Guanine
C Cytosine
U Uracil

It can fold into complicated three-dimensional shapes and create hairpin loops. Hairpin loops are
typical in Messenger RNA (mRNA) and transition RNA are two types of RNA molecules
(tRNA). Adenine and uracil (A-U) are paired together, and guanine and cytosine are paired
together (G-C). The nitrogenous bases connect to one another when this happens. RNA is not
always linear, despite being single-stranded [2]

1.1 Structure of RNA


It has been showed that the RNA fraction of a cellular RNP can, at most, once, carry out a
biochemical reaction—a function that is typically restricted to proteins. The RNA strand's self-
complementary sequences regulate intrachain base pairing and ribonucleotide chain bending into
a variety of cellular forms, include bulges and hydroxyl groups. [13, 14].
After a mutation at position 58 of the tRNA chain, a methyl-free transfer RNA (tRNA) molecule
that operates as an initiator (tRNAiMet) becomes dysfunctional and subsequently non-functional;
the nonfunctional chain has been disrupted by cellular tRNA supply chain activities.[15]
Molecules which have the ability to produce clusters with RNAs are identified as ribonucleic
proteins (RNPs). Imperfect stabilization and modifications to structure cause molecules to break
easily. The stability and function of RNA are contingent upon its three-dimensional structure,
and this helps biological enzymes in binding organic molecules (such methyl groups) to the
chain and changing the ribose sugar and nitrogenous bas es in multiple ways.[16], [17].
Figure 1. 1 RNA Structure1

Although purines comprise a bicyclic structure while pyrimidines have a single ring, the size of
the pyramidal is smaller and the space between molecule can be preserved at 2 nm, producing
purines larger than pyrimidine compounds Purines and pyrimidines are naturally occurring
substances that participate in RNA binding, earning them the title "components of genetic
structure RNA." In RNA, these nitrogen bases make up two separate nucleic acids. Pyrimidines
cytosine and thymine contain one carbon azacyclic base, but purines both of which have two.
[18], [19].

An imidazole molecules ring and the ring of pyrimidine interact to produce purine, a naturally
generated carboxylic acid molecule. Most usual spawn are guanine and adenine. The molecular
structure is composed of two hydrogen carbon rings and nitrogenous base protons. The point at
which it softens of purine is 214°C. Oxidation results in the production of uric acid. Pyrimidine
is a polycyclic aromatic chemical family based on carbon and hydrogen. There are actually two
nucleobases: uracil and cytosine. It has a hydrogen carbon ring and two nitrogen molecules. The
melting point of pyridinium is 20–22°C. During biosynthesis, carbon dioxide, amino acids, and
undesirable salts are produced [20], [21]

1
[Link]
Figure 1. 2 Pyrimidines and Purines2

1.2 Recombination of RNA

The exchange of genetic data between two non-divided RNA genomes is often referred to as
recombination in RNA infections, and it is particularly evident in infections with shattered
genomes. Since all aspects of RNA recombination incorporate polymerase bouncing in RNA
binding—which may differ from the synthesis of DI RNA—they are likened to the imperfect
integrator DI RNA era. It might be a reasonably common miracle related to RNA infection. The
properties of homologous RNA recombine are as follows: in the early 1960s, Hirst 1962 and
Delink 1963 discovered two exact distinctions regarding the poliovirus. RNA recombination is
so distant, as various RNA is a infections have shown. transmission of the same RNA segments
among regions 22 and 24.

2
Figure 1. 3RNA recombination process3

1.3 RNA Types and Functions

Some, on the flip side, play a variety of control roles in cells. This type of RNA primarily
participates in biochemical operations, along with another like enzymes. Because they are
plentiful, have a variety of functions, and are involved with numerous regulatory processes,
RNAs have significance for both transcriptional regulation and cancer. All mammals have
messenger RNA (mRNA), transfer RNA (tRNA), and ribosomal RNA (rRNA), which are the

3
[Link]
three forms of RNA that have been thoroughly examined. The cellular protein subunits are
encoded by rRNA and formed by the nucleolus.[25], [26].

rRNA and protein are the two distinct elements which make up a ribosome. mRNA carries
genetic information from the nucleus to ribosomes, that are cytoplasmic locations where protein
translation occurs, during the procedure of protein synthesis. A group of three nitrogenous bases
in mRNA establish which amino acid is present in the protein's chain. Upon reaching the
cytoplasm, they proceed to function as translation regulators by a process called "reading" the
mRNA-carried code. Several kinds of amino acids are carried to ribosomes by the reporter gene
messenger tRNA, also referred to as receptive or activator RNA, where they come together to
produce proteins. [27, 28].

Maintenance and regulatory ncRNAs (tRNA and rRNA) are the two different kinds of non-
coding RNAs (ncRNAs), which are further separated by number. Small ncRNAs (sncRNAs)
have even fewer nucleotides than long ncRNAs (lncRNAs), which have more than 200
nucleotides. MicroRNA (miRNA), small nucleotide RNA (snRNA), small-interfering RNA
(siRNA), small nucleolar RNA (snoRNA), and PIWI-interacting RNA (piRNA) are examples of
these tiny non-coding RNAs (ncRNAs). RNAs can be classified as either coding (cRNA) or
noncoding (nRNA), in contrast to mRNA, tRNA, and rRNA (ncRNA). Through their binding to
target mRNA, they restrict communication, so silenced (inhibiting) the expression of genes and
inhibiting the formation of protein aggregates. [29, 30].

miRNAs usually have an important impact on how disease like cancer grow. One cannot stress
the importance of microRNAs. We now have an estimated length of roughly 22 nucleotides
common to all eukaryotes, and we are important for gene expression. In influencing particular
specific genes, cancer cell flash hider and tumor suppression (also known as cancer-initiating)
miRNAs, for example, may govern the growth and genesis of tumors.[31]

1.4 RNA in illness


Like other RNA types, tRNAs may connect to caspases, which contain specialized proteins
involved in apoptosis, also known as programmed cell death).
As mentioned before, certain miRNAs can regulate leukemia genes in ways that promote the
development of cancer. Non-coding RNAs called tRNA-derived fragments (tRFs) are
additionally believed to play a role in leukemia. Thanks to advancements in techniques such as
sequence analysis, increased levels of MALAT1 (transcript related to lung cancer metastases)
have been found in many malignant tissues and have been tied to tumor cell proliferation and
dissemination (spread)[2], [3], [27]. Remarkable discoveries have also been revealed about the
relationship between RNA and human illness. Cancer can be defined by neurons' propensity to
evade procedures and operations, and tRNAs inhibit mortality by attaching to phosphorylation
proteins. Several neurological diseases, most notably Alzheimer's, have been connected to
downregulated miRNA oxidation.[32]. These findings are predicted to be boosted by a greater
understanding of RNA and its functions, ongoing developments in methodologies for
sequencing, and initiatives to identify RNA and RBPs as potential targets for targeted
therapeutics. Multiple human illnesses have been connected to RBP issue, instability, and
evolution. Wondered to be instrumental in developing of foci or collates in brain tissues, a family
of RNAs that contained tandem repeats that sequester RNA-binding protein (RBPs). Those
particles possess an impact on neurological disorders such as myotonic dystrophy (MD) and
motor neuron disease (ALS). It is expected that additional links between RNA and illness will be
identified in the future. [33]–[35].

1.5 5mRNA Methylcytosine

Although messenger RNA (mRNA) performs as an interface among DNA and proteins, it is
necessary for the production of instructions. Transforming genetic information into proteins—
which are required for biological processes—is this vital function. Apart from its essential
function, mRNA experiences other chemical changes that impact its stability, translation
effectiveness, and localization. [36] The methylation of cytosine residues, leading to in the
creation of 5-methylcytosine (m5C), is one of these modifications that is most significant. The
significant influence this particular methylation process has on mRNA metabolism and its more
general consequences for cellular function and illnesses like cancer have made it a focus of
research. Particular enzymes known as RNA methyltransferases play a role for the m5C
modification of mRNA, and NSUN2 is an essential aspect of this process. It has been shown that
m5C in mRNA enhances mRNA stability, prolonging its lifespan and promoting continuous
protein synthesis, particularly during times of stress. Additionally, the movement of mRNA from
the nucleus to the cytoplasm—a process essential to successful protein synthesis—is increased
by m5C.

Furthermore associated with the control of transformation m5C enhances mRNA binding and
handling by ribosomes. This is particularly crucial for quickly proliferating cells since these cells
require a lot of protein synthesis. A variety of tumors have been related to [37] m5C dysregulation
in mRNA, which makes it a potential target for treatment. The applications of mRNA
methylcytosine in identifying and treating diseases has become more and more obvious as
research into this molecule proceeds.

1.6 Structure of 5mRNA Methylcytosine

Messenger RNA, or mRNA, acts as an example for the manufacture of proteins and is a key
component of the process of gene expression. mRNA is a single-stranded molecule made up of
ribonucleotides, that are made up of one of four nitrogenous bases—adenine (A), guanine (G),
cytosine (C), and uracil (U)—as well as a phosphate group and a ribose sugar. The DNA coding
sequence, the 5' cap, the 5' untranslated region (5' UTR), the 3' UTR, and the poly(A) tail are the
several parts that make up the structure of mRNA.

Before the transcription process, a changed a substance called nucleotide known as the 5' cap is
appended to the 5' end of the mRNA molecule. This cap structure shields the mRNA from
exonucleolytic devastation, which is essential for the continued viability of the molecule.
Figure 1. 4 – Methylcitosine 5mC Structure4

[Link]

The 5' UTR modulates protein binding and scanning, which can affect translation effectiveness.
By affecting protein access to the primary coding order, upstream open reading frames (uORFs)
or other structures within the 5' UTR can further regulate translation [39].
Overall 3' untranslated region (3' UTR), additional non-coding area which is crucial for post-
transcriptional legislation, follows the coding sequences. A number of regulation components,
which include binding regions for RNA-binding proteins (RBPs) and microRNAs (miRNAs), are
present in the 3' UTR and can affect the ability to translate efficiency, equilibrium, and
placement of mRNA [40]. Some sequences inside the 3' UTR determine the half-life of the
mRNA molecule, and this region additionally plays a role in controlling the process of mRNA
breakdown.

Adenine nucleotides are attached to the 3' end of the mRNA molecule through transcription to
form the poly(A) tail. This tail aids in the control of translation and increases the stability of the
4
[Link]
presence-in-A_fig5_332385818
mRNA by shielding it from premature devastation. This tail aids in the management of synthesis
and enhances the integrity of the mRNA by protecting it from quick destruction. [41].

In summary, the accurate and successful interpretation of genetic data into proteins is guarantees
by the complicated structure of mRNA. The 5' cap, 5' UTR, coding sequence, 3' UTR, and

poly(A) tail of the mRNA molecule all have a role in the stability, the translation, and
distribution of the mRNA, which in turn controls the quantity and duration of protein production
in the cell.

1.7 Problem statement

Recognizing an assortment of biological functions and conditions depends upon the proper
identification of these locations within RNA sequences, given the critical function that 5-
methylcytosine (5mC) plays in molecular biology. Conventional experimental methods are
useful, but they are problematic for large-scale research since they are frequently costly and
time-consuming. As a result, the creation of mathematical methods for identifying 5mC sites has
grown in relevance. While previous models have used one-hot encoding and convolutional
neural networks (CNN) for representation of features to identify 5mC sites, such as the
iRhm5CNN put out by Ali et al., these methods are limited by the intrinsic complexity of RNA
molecules. It may be difficult for traditional machine learning algorithms to understand the
intricate structure and patterns required for precise 5mC site prediction. It could be challenging
for traditional machine learning algorithms to understand the complicated structures and patterns
required for exact 5mC site prediction. Researching strong learning algorithms, which includes
deep neural networks (DNN), in conjunction with dependable representation methods is crucial
in order to meet these problems. Such methods have the possibility to significantly enhance
prediction accuracy, making researchers much better and more reliable tool for identifying 5mC
sites in RNA segments.

1.8 Research Objective:

The purpose of this work is to develop an advanced computational method for accurately finding
5-methylcytidine (m5C) sites in mRNA sequences. m5C plays a role for both gene expression
and RNA stability, but since RNA is so complex, it can be hard to discover. The desire for
successful computational techniques comes from the time and cost contribution of traditional
testing methods.

Deep learning methods, especially deep neural networks (DNNs), will be utilized during the goal
to improve the predictions accuracy of m5C sites. We will investigate a range of machine
learning models, such as Recurrent neural networks (RN (RNNs), specifically Long Short-Term
Memory (LSTM) networks, which are ideal for analyzing the sequential nature of RNA, and
Convolutional Neural Networks (CNNs), which are able to identify sequence motifs and
structures suggestive of m5C modifications. Furthermore, long-range dependencies in RNA
sequences will be represented by transformer models, like BERT, offering an improved
understanding of the sequence the environment around m5C sites.

Additionally, the research will focus on powerful feature representation methods that quickly
transform unprocessed RNA sequence data into a machine learning-ready format, including
embedding methods, k-mer representation, and one-hot encoding. These methods seek to extract
the small details required for effective m5C site prediction.

The current study aims to create a computational tool that significantly improves the
identification of m5C sites by integrating cutting-edge machine learning methods with advanced
representation of characteristics techniques. This computational tool will provide a scalable and
effective solution for large-scale research in epitranscriptomics and molecular biology.

1.9 Research Methodology

Recombination of gene expression is an important process that promotes evolution and


contributes to variation in genes. Identifying and predicting recombination hotspots can provide
important information about genetic variation and pattern of inheritance. The systematic
procedure for developing and evaluating a computer model for predicting gene expression
recombination sites is described in this work. The approach comprises multiple key stages:
1 Theoretical analysis:

Theoretical examination of the current mathematical frameworks for forecasting gene expression
combination hotspots is conducted as part of the research. This entails evaluating the benefits
and drawbacks of the existing strategies and locating areas in need of change. The goal is to
understand these models' workings and lay the groundwork for creating a predictor that is more
accurate and efficient.

2: Dataset Selection:

Selecting a reliable benchmark dataset with biological sequences pertinent to gene expression
recombination is an essential first step. This dataset ought to be thorough, well annotated, and
indicative of the biological processes under research. The dataset's quality and applicability are
essential for both developing and evaluating the model's predictive capabilities.

3: Feature Extraction:

Selecting and utilizing a suitable feature extraction method is the subsequent stage. Using this
procedure, unstructured biological sequence data can be converted through meaningful features
that machine learning algorithms may use. Feature extraction techniques can involve counting k-
mers, analyzing the makeup of nucleotides, or identifying secondary structures, based on the
nature of the data and the research objectives.

4: Feature Selection:

A strict feature selection procedure is used to make sure that the model contains the most
pertinent and instructive features. This involves identifying which features are most important
and choosing a subset that improves the model's functionality. To find and save the most
important features, strategies including decreasing dimensionality, statistical testing, and
wrapping approaches can be used.

5: Model Training and Classification:


To construct the model for prediction, an efficient machine learning algorithm is used as a
classification engine. The selected algorithm will be developed utilizing the features that were
obtained and chosen to identify patterns linked to recombination sites. According to how well
suited they are for the job, algorithms including ensemble techniques, support vector machines,
and deep neural networks may be used.

6: Evaluation of Performance:

Using the benchmark dataset, an extensive evaluation of the suggested predictor's performance is
conducted in the last phase. Evaluation metrics are used to measure how well the model predicts
recombination hotspots, including accuracy, precision, recall, and F1-score. This stage
guarantees the reliability and precision of the model and offers useful details about its
effectiveness.

The research project aims to improve the comprehension of genetic variation and evolutionary
processes by developing a computational model that enhances the prediction of gene expression
recombination sites through the implementation of this structured methodology.

1.1 0 Structure of Thesis

This thesis' remaining sections are arranged as follows:

Chapter 2: Literature Review:

This chapter presents a thorough analysis of related literature on predictors for recombine point
detection. It looks at many designs and methods that have been put out and talks about their
benefits and drawbacks.

Chapter 3: Methods and Materials:

This chapter describes the materials and procedures that went into creating the prediction
models. The study involves multiple components such as feature selection strategies, hybrid
operations, sequence representation techniques, benchmark dataset selection, and the algorithms
used for classification used.
Chapter 4: Results and Discussion

Numerous performance metrics will be addressed in this chapter, such as Matthew's correlation
coefficient, sensitivity, specificity, and accuracy. The chapter includes grid search methods for
hyper-parameter optimization and model validation procedures. Using the benchmark dataset,
the efficacy of various learning algorithms is tested, and the results of the suggested classifier is
assessed using a variety of feature extraction methods. The chapter concludes with an assessment
of the suggested model's outcome with current predictors.

Chapter 5: Conclusion and Future Work:

This last chapter provides a summary of the main conclusions derived from the study as well as
possible paths for further investigation.

1.11 Summery of Chapiter 1

In the Chapter 1, we highlighted its arrangement and essential functions in biological functions
like protein synthesis and gene control. We talked about RNA-related problems that result from
errors in the above procedures and RNA the process of re a crucial mechanism that fosters
variation in genes. The chapter also discussed various types of RNA, such as mRNA, tRNA,
rRNA, and non-coding RNAs, to highlight the range of uses for this flexible molecule.

The idea of 5-methylcytosine (5mC), a key environmental change found in RNA, occupied an
important part of the chapter. We studied the impact of 5mC methylcytosine on cell function and
its role in regulating gene expression. The chapter provided a quick overview of the many
cytosine alterations, emphasizing the significance of 5mC in biological processes such as disease
progression and development.

Furthermore, we contrasted deep learning techniques with conventional machine learning


algorithms and showed the latter's proficiency in analyzing intricate datasets. knowledge the
purpose of the study we are doing, which is to anticipate 5mC sites in RNA sequences using
sophisticated computational tools, requires a knowledge of this similarity. A summary of the
research goals and methods provided in the chapter's conclusion provided the foundation for the
in-depth analysis that would come ahead.

Chapter 3: METHODS AND MATERIALS


3.1 Benchmark dataset

A dependable and accurate benchmark dataset, as recently utilized by Jia et al. [1], is crucial
when developing a robust and efficient prediction model. The benchmark dataset typically
consists of two key components: a testing dataset and a reference dataset. The testing dataset is
used to validate the output of the proposed model, while the learning model is trained using the
reference dataset. In this study, the primary benchmark 5mC dataset consists of both positive and
negative examples. Mathematically, the benchmark dataset is represented by the following
equation (1):

−¿¿
+¿∪ R 1 ¿
R1=R 1 (1)

+¿ ¿ −¿¿
where R1 represents the total RNA sequences, R1 denotes the positive 5mC sequences, and R1
, indicates the non-5mC sequences. The symbol ∪ represents the union of these two subsets. To
ensure the individuality of sequences, a CD-HIT tool was employed to remove sequences with
more than 20% similarity. As a result, a benchmark dataset comprising 58,160 total sequences
was generated, including 29,080 5mC samples and 29,080 non-5mC samples. The independent
benchmark dataset is mathematically formulated using the following equation (2):

−¿¿
+¿∪R 2 ¿
R2=R 2 (2)
+¿ ¿
Where R1 refers to the independent set of RNA sequences, R2 represents the positive RNA
−¿¿
5mC sequences, and R2 represents the non-5mC sequences. The training and validation set
includes 46,528 samples, while the independent set includes 11,632 samples. It is important to
note that the results from the independent tests were stored distinctly and were not used in the
learning or parameter optimization developments.

3.2 Support Vector Machine, or SVM

Support Vector Machine, or SVM, is a robust and flexible supervised machine learning
algorithm primarily used for classification, though it can also handle regression tasks. The main
goal of SVM is to find the optimal boundaries, known as a hyperplane, that separate different
data classes. In an n-dimensional space, where n is the number of features, each data point is
represented as a point, and the algorithm seeks to determine the hyperplane that best separates
these points.[41]

Figure 3. 1 svm
[Link] -

The hyperplane is a decision-making boundary that maximizes the margin between itself and the
closest data points from each class, known as support vectors. These support vectors are crucial
as they influence the orientation and position of the hyperplane, ensuring the model's robustness.
Originally designed for binary classification, SVM has been extended to multiclass problems by
breaking them down into binary classifications.

One of SVM's key features is the "kernel trick," which allows it to perform non-linear
classification. Kernel functions map non-linear data into a higher-dimensional space where it
becomes linearly separable. [42] Two key parameters in SVM are C, which controls the trade-off
between classification error and margin size, and gamma, which determines the influence of
individual training examples. SVM is particularly effective in high-dimensional spaces and is
generally robust against overfitting.

3.3 Randum Forest (RF)

Another common kind of collaborative learning method for handling both regression and
classification problems are the Random Forest (RF) algorithm. It operates by creating several
autonomous decision trees, each of which is trained using a different subset of the training set
produced by bootstrap sampling. [43] By ensuring that every decision tree is trained on a little
different data, this method—known as bagging—promotes variation among the trees and lowers
the possibility of excess fitting. Particularly appreciated for its simplicity and efficacy, Random
Forest frequently yields good results without requiring a great deal of hyper-parameter
modification.

The Random Forest algorithm is essentially a collective process. As the name indicates, it
consists of numerous independent decision trees that collaborate to generate projections. In task
classification, each tree in the forest makes a separate forecasting about the class; in regression
tasks, it makes an independent prediction about the outcome. The Random Forest model's final
forecast is the result of combining these individual forecasts The result of the model is
categorized based to which class the trees vote for the most. Regression yields a final result by
averaging the predictions made by each tree. By combining the data, variation reduces and the
capacity of the model to generalize to new data is enhanced.

Random Forest's ability to manage huge data sets with multiple characteristics, which makes it
appropriate for a variety of applications, is one of its key benefits. Furthermore, even with a large
number of trees in the forest, the resulting model is normally robust against the overfitting
because the trees in a Random Forest are constructed using random subsets of features and data.
The algorithm's utility is further improved by its ability to provide feature significance scores,
which provide insights into the variables that have the greatest influence on projections. When
all factors considered, Random Forest is a strong and adaptable algorithm that finds broad
application in many different domains because of reliability, efficiency, and simplicity of
operation.

3.4 Decision Tree (DT):

A popular and flexible supervised learning method for both regression and classification
applications is decision trees (DT). The algorithm divisions a dataset according to various
characteristics and then develops a tree-like model of decisions. Because it doesn't assume a
specific format for the data, this non-parametric approach is flexible and helpful for a wide range
of issues.[44]

By applying multiple characteristics to separate the data into subsets, a decision tree constructs a
structure that resembles a tree. Every node in the tree represents a choice made in response to a
feature, and the branches indicate the possible results of that choice. Until the data are divided
into pure subsets or satisfy a stopping criterion, the process is repeated.

Decision trees are intuitive and easy to understand due to their visual representation, which
includes expanding paths and terminal nodes. Their clear and simple decision-making process
makes them popular in real-world applications like risk assessment, medical diagnosis, and
segmentation of consumers. Decision Trees are straightforward tools that can handle complex
information sets and offer lucid insights into the process of making decisions.[45]
3.5 KNN K-Nearest Neighbor:

Although it is mainly employed for issues related to classification, the K-Nearest Neighbors
(KNN) algorithm is a supervised machine learning technique that is also used for regression
tasks. It operates on the assumption that equivalent objects are found in proximity to one another
in feature space. KNN essentially assesses newly discovered, unlabeled data points according to
how close they are to already-existing, labeled data points. Essentially, similar objects tend to
group together.[46]

Given that KNN uses the entire dataset in the process of prediction phase rather than a separate
training phase, it is referred to as a "lazy" learning algorithm. This means that KNN uses distance
metrics for contrasting new data points with existing ones in order to make choices based on the
whole data set. It is frequently used to measure the nearness of data points using the distance
calculated by Euclidean geometry.

The most important variable in the KNN algorithm's performance is the parameter K. While a
larger K value can smooth out projections and increase stability, it may also reduce the model's
ability to capture intricate patterns. A smaller K value may result in a less stable model that is
highly sensitive to noise. KNN can become longer as the dataset grows, which reduces its
efficiency when dealing with very large datasets. Situations where data is regularly updated or
changes as time passes are best for this algorithm.[47].

3.6 ADA BOOST

Adaptive Boosting, or AdaBoost, is a popular ensemble learning method that improves machine
learning model accomplishment by fusing together several weak learners to create a strong
learner. AdaBoost, [48] first presented by Freund and Schapire in 1997, is a technique which
trains an order of weak classifiers one after the other by emphasizing on the errors created by the
preceding classifiers. Iteratively assigning a higher weight to misclassified data points forces
subsequent classifiers to focus more on these difficult cases.

AdaBoost modifies the weights of sample training data in each iteration in response to the errors
of the before classifier. [49] The new classifiers are now able to rectify the errors made by earlier
models through to this modification. AdaBoost collects the classifiers' estimations after they
have been thoroughly trained, weighing the contributions of each classifier according to its
accuracy. By lowering bias and variance, this combination enhances the effectiveness of the
model as a whole.[50]

Thanks to its ease of use and efficiency, AdaBoost is highly appreciated in a wide range of
applications, such as bioinformatics, text classification, and recognition of images. It is a flexible
machine learning tool because it is capable of turning poor performers into a robust model.

3.7 Gradient Boosting

Gradient Boosting is an ensemble learning method that uses iterative improvement to create
predictive models. To build a strong model, it combines the predictions from multiple weak
learners, usually decision trees.[51] The main goal of gradient boosting is to reduce the model's
loss function by gradually fitting new models to the residuals of the previous ones. This strategy
increases prediction accuracy by concentrating on the mistakes made by earlier models.

Every fresh algorithm in gradient boosting gets instruction to correct any errors made by
previous models. The algorithm begins with a basic model and then builds upon it by adding new
models to fix the errors made by the original ensemble. Gradient descent is used in the learning
process to reduce a loss function, such as the mean squared error for regression or the log loss for
classification. All of the separate models' predictions have been combined and weighted based on
accuracy to create a final model.[52]
Applications such as machine learning matches, healthcare, and finance all benefit greatly from
gradient boosting. Its versatility for handling diverse data kinds and the ability for improving the
accuracy of models through iterative improvement make it an appealing choice for intricate
prediction assignments.[53]

3.8 Logistic Regression

Another statistical technique for issues with binary classification is logistic regression. By fitting
data to a logistic function, which is translated predictions to probabilities between 0 and 1, it
calculates the likelihood that a given input belongs to a specific class. When the outcome
variable is categorical, specifically binary, a logistic regression model is utilized as opposed to
linear regression, which forecasts a continuous output.[54]

The basic idea of logistic regression is to use the logistic function, sometimes referred to as a
sigmoid function, to model the association between the input features and the binary outcome. A
probability value, or the chance that the input will belong to the positive class, is generated by
this function from the linear combination of the input features. Using the estimation of maximum
likelihood, the model is trained by discovering the perfect coefficients that reduce the disparity
between the actual results and the estimated probabilities.

Because of its efficiency, interpretability, [55]and ease of use, logistic regression is used
extensively. It provides a simple probabilistic framework for classification, which makes it
appropriate for a number of uses, such as risk prevention, marketing, and medical diagnosis.

3.9 Naive Bayes

The Naive Bayes algorithm is a statistical classification technique that relies on assuming that
features are independent of one another based on the class label. It is derived from the Bayes
Theorem. Naive Bayes is surprisingly effective at many tasks, especially text classification and
spam filtering, considering how simple it is.[56]
Based on the features of the input data, the algorithm establishes the probability of each class and
chooses the class with the highest probability. It makes use of the Bayes Theorem, which asserts
that the likelihood of the features given the class and the prior probability of the class increase to
determine the chance of a class given the features. The "naive" an element stems from the belief
that the features are unrelated to one another, which makes calculation easier but may not always
indicate dependence in the real world.[57][58]

Since Naive Bayes scales well with the total amount of features and instances, it is especially
useful when working with massive databases and high-dimensional data. It is a well-liked
solution for issues like document classification and sentiment analysis given its performance,
which continues to be solid even when the assumption of independence is not fully satisfied.

Here's a concise mathematical proof of the Naive Bayes classifier:

1. **Bayes' Theorem: **

\[

P(C|X) = \frac{P(X|C) \cdot P(C)} {P(X)}

\]

2. **Naive Bayes Assumption: ** Assuming feature independence, the likelihood \(P(X|C) \)


becomes:

\[

P(X|C) = \prod_{i=1} ^{n} P(X_i|C)

\]

3. **Substitute into Bayes' Theorem: **

\[
P(C|X) = \frac{P(C) \cdot \prod_{i=1} ^{n} P(X_i|C)} {P(X)}

\]

4. **Classification Decision: ** Since \(P(X) \) is constant, maximize:

\[

\hat{C} = \arg\max_C \left(P(C) \cdot \prod_{i=1} ^{n} P(X_i|C) \right)

\].

---------------------------------------------------------------------------------------------------------------------

3.10 XGBOOST:

XGBoost, which stands for extreme Gradient Boosting, is a complex gradient boosting machine
application that prioritizes both speed and effectiveness. Its scalability and flexibility have made
it highly utilized in the machine learning belonging, especially for tasks involving structured
data. For example, for XGBoost to function, weak learners—usually decision trees—are added
one after the other while their mistakes from previous versions are fixed. The model works
especially well with complicated data sets because it reduces a loss function and includes a term
of regularization to avoid overfitting.[59][60]

XGBoost's ability to successfully deal with sparse data and values that are missing is one of its
primary benefits. It makes use of a cutting-edge sparsity-aware algorithm that can process and
maximize such data with a negligible achievement loss. Furthermore, XGBoost is capable of
performing large-scale datasets due to its supports distributed and parallel processing. Due to its
adaptability, it performs better than other models based on machine learning in a range of use
cases, such as regression, classification, and ranking tasks.[61][62]

A major factor in XGBoost's success is its extensive parameter tuning options, which provide a
fine-grained control over the amount of detail and performance of the model. The model's status
as a preferred tool for data researchers and scientists across different disciplines has been further
strengthened by its capacity to manage non-linear relationships and interactions between features
without the need for important preprocessing.
References

[1] J. Brosius and C. A. Raabe, “What is an RNA? A top layer for RNA classification,” RNA
Biol., vol. 13, no. 2, pp. 140–144, Feb. 2016, doi: 10.1080/15476286.2015.1128064.

[2] P. S. Thomas, “Hybridization of denatured RNA and small DNA fragments transferred to
nitrocellulose.,” Proc. Natl. Acad. Sci. U. S. A., vol. 77, no. 9, pp. 5201–5205, 1980, doi:
10.1073/pnas.77.9.5201.

[3] O. Preisig, N. Moleleki, W. A. Smit, B. D. Wingfield, and M. J. Wingfield, “A novel


RNA mycovirus in a hypovirulent isolate of the plant pathogen Diaporthe ambigua,” J.
Gen. Virol., vol. 81, no. 12, pp. 3107–3114, Dec. 2000, doi: 10.1099/0022-1317-81-12-
3107.

[4] B. Mateescu et al., “Obstacles and opportunities in the functional analysis of extracellular
vesicle RNA – an ISEV position paper,” J. Extracell. Vesicles, vol. 6, no. 1, p. 1286095,
Dec. 2017, doi: 10.1080/20013078.2017.1286095.

[5] L. Liu et al., “A method for extracting high-quality total RNA from plant rich in
polysaccharides and polyphenols using Dendrobium huoshanense,” PLoS One, vol. 13,
no. 5, p. e0196592, May 2018, doi: 10.1371/[Link].0196592.

[6] L. Deng, W. Yang, and H. Liu, “PredPRBA: Prediction of Protein-RNA Binding Affinity
Using Gradient Boosted Regression Trees,” Front. Genet., vol. 10, no. August, pp. 1–11,
Aug. 2019, doi: 10.3389/fgene.2019.00637.

[7] J. Lan et al., “Functional role of Tet-mediated RNA hydroxymethylcytosine in mouse ES


cells and during differentiation,” Nat. Commun., vol. 11, no. 1, pp. 1–15, 2020, doi:
10.1038/s41467-020-18729-6.

[8] D. Haussecker, Y. Huang, A. Lau, P. Parameswaran, A. Z. Fire, and M. A. Kay, “Human


tRNA-derived small RNAs in the global regulation of RNA silencing,” Rna, vol. 16, no. 4,
pp. 673–695, 2010, doi: 10.1261/rna.2000810.

[9] N. G. Nguyen et al., “DNA Sequence Classification by Convolutional Neural Network,”


J. Biomed. Sci. Eng., vol. 09, no. 05, pp. 280–286, 2016, doi: 10.4236/jbise.2016.95021.

[10] S. Linjawi, “What is nucleic acid ?,” pp. 1–15.

[11] G. A. Soukup and R. R. Breaker, “Relationship between internucleotide linkage geometry


and the stability of RNA,” Rna, vol. 5, no. 10, pp. 1308–1325, 1999, doi:
10.1017/S1355838299990891.

[12] N. B. Leontis and E. Westhof, “Geometric nomenclature and classification of RNA base
pairs,” Rna, vol. 7, no. 4, pp. 499–512, 2001, doi: 10.1017/S1355838201002515.

[13] Y. Zeng and B. R. Cullen, “Sequence requirements for micro RNA processing and
function in human cells,” Rna, vol. 9, no. 1, pp. 112–123, 2003, doi:
10.1261/rna.2780503.

[14] D. E. Draper, “A guide to ions and RNA structure,” Rna, vol. 10, no. 3, pp. 335–343,
2004, doi: 10.1261/rna.5205404.

[15] G. Meister, M. Landthaler, Y. Dorsett, and T. Tuschl, “Sequence-specific inhibition of


microRNA-and siRNA-induced RNA silencing,” Rna, vol. 10, no. 3, pp. 544–550, 2004,
doi: 10.1261/rna.5235104.

[16] L. Dölken et al., “High-resolution gene expression profiling for simultaneous kinetic
parameter analysis of RNA synthesis and decay,” Rna, vol. 14, no. 9, pp. 1959–1972,
2008, doi: 10.1261/rna.1136108.

[17] H. J. Peltier and G. J. Latham, “Nc,” Rna, vol. 14, no. 5, pp. 844–852, 2008, doi:
10.1261/rna.939908.

[18] P. T. Lang et al., “DOCK 6: Combining techniques to model RNA-small molecule


complexes,” Rna, vol. 15, no. 6, pp. 1219–1230, 2009, doi: 10.1261/rna.1563609.

[19] P. J. Shepard, E. A. Choi, J. Lu, L. A. Flanagan, K. J. Hertel, and Y. Shi, “Complex and
dynamic landscape of RNA polyadenylation revealed by PAS-Seq,” Rna, vol. 17, no. 4,
pp. 761–772, 2011, doi: 10.1261/rna.2581711.

[20] B. F. Medics, “RNA ( Ribonucleic acid ).”

[21] “ Nucleic acids are biopolymers , or small biomolecules , essential to all known forms of
life . They are composed of nucleotides , which are monomers made of three components :
a 5-carbon sugar , a phosphate group and a nitrogenous base . If the sugar is.”

[22] N. J. Schurch et al., “How many biological replicates are needed in an RNA-seq
experiment and which differential expression tool should you use?,” Rna, vol. 22, no. 6,
pp. 839–851, 2016, doi: 10.1261/rna.053959.115.

[23] D.-Q. Shi, I. Ali, J. Tang, and W.-C. Yang, “New Insights into 5hmC DNA Modification:
Generation, Distribution and Function,” Front. Genet., vol. 8, no. JUL, pp. 1–11, Jul.
2017, doi: 10.3389/fgene.2017.00100.

[24] W. Chen, P. Feng, X. Song, H. Lv, and H. Lin, “iRNA-m7G: Identifying N7-
methylguanosine Sites by Fusing Multiple Features,” Mol. Ther. - Nucleic Acids, vol. 18,
pp. 269–274, Dec. 2019, doi: 10.1016/[Link].2019.08.022.

[25] X. Liu, P. He, W. Chen, and J. Gao, “Multi-Task Deep Neural Networks for Natural
Language Understanding,” 2019, doi: 10.18653/v1/p19-1441.

[26] P. J. Batista, “The RNA Modification N 6 -methyladenosine and Its Implications in


Human Disease,” Genomics. Proteomics Bioinformatics, vol. 15, no. 3, pp. 154–163, Jun.
2017, doi: 10.1016/[Link].2017.03.002.

[27] Y. Uemura, A. Hasegawa, S. Kobayashi, and T. Yokomori, “Tree adjoining grammars for
RNA structure prediction,” Theor. Comput. Sci., vol. 210, no. 2, pp. 277–303, Jan. 1999,
doi: 10.1016/S0304-3975(98)00090-5.

[28] B. S. Cheriyedath and M. Sc, “Types of RNA : mRNA , rRNA and tRNA,” pp. 1–5.

[29] X.-Q. Gao et al., “The piRNA CHAPIR regulates cardiac hypertrophy by controlling
METTL3-dependent N6-methyladenosine methylation of Parp10 mRNA,” Nat. Cell Biol.,
vol. 22, no. 11, pp. 1319–1331, Nov. 2020, doi: 10.1038/s41556-020-0576-y.

[30] Y. P. Shi, S. Thouta, Y. M. Cheng, and T. W. Claydon, “Extracellular protons accelerate


hERG channel deactivation by destabilizing voltage sensor relaxation,” J. Gen. Physiol.,
vol. 151, no. 2, pp. 231–246, Feb. 2019, doi: 10.1085/jgp.201812137.

[31] W. Chen, H. Tang, J. Ye, H. Lin, and K. C. Chou, “iRNA-PseU: Identifying RNA
pseudouridine sites,” Mol. Ther. - Nucleic Acids, vol. 5, no. March, p. e332, 2016, doi:
10.1038/mtna.2016.37.

[32] V. Thiel, J. Herold, B. Schelle, and S. G. Siddell, “Infectious RNA transcribed in vitro
from a cDNA copy of the human coronavirus genome cloned in vaccinia virus,” J. Gen.
Virol., vol. 82, no. 6, pp. 1273–1281, Jun. 2001, doi: 10.1099/0022-1317-82-6-1273.

[33] T. Iwamoto, K. Mise, K. Mori, M. Arimoto, T. Nakai, and T. Okuno, “Establishment of an


infectious RNA transcription system for Striped jack nervous necrosis virus, the type
species of the betanodaviruses,” J. Gen. Virol., vol. 82, no. 11, pp. 2653–2662, Nov. 2001,
doi: 10.1099/0022-1317-82-11-2653.

[34] K. Abe and N. Konomi, “Hepatitis C Virus RNA in Dried Serum Spotted onto Filter Paper
Is Stable at Room TemperatKatz 2002ure,” J. Clin. Microbiol., vol. 36, no. 10, pp. 3070–
3072, 1998, doi: 10.1128/JCM.36.10.3070-3072.1998.

[35] R. S. Katz, M. Premenko-Lanier, M. B. McChesney, P. A. Rota, and W. J. Bellini,


“Detection of measles virus RNA in whole blood stored on filter paper,” J. Med. Virol.,
vol. 67, no. 4, pp. 596–602, Aug. 2002, doi: 10.1002/jmv.101

[36] G. A. Soukup and R. R. Breaker, “Relationship between internucleotide linkage geometry and the
stability of RNA,” Rna, vol. 5, no. 10, pp. 1308–1325, 1999, doi: 10.1017/S1355838299990891.

[37] N. B. Leontis and E. Westhof, “Geometric nomenclature and classification of RNA base pairs,” Rna, vol.
7, no. 4, pp. 499–512, 2001, doi: 10.1017/S1355838201002515.

[38] M. Li et al., “5-methylcytosine RNA methyltransferases and their potential roles in cancer,” Journal of
Translational Medicine 2022 20:1, vol. 20, no. 1, pp. 1–16, May 2022, doi: 10.1186/S12967-022-03427-
2.
[39] L. Shen, C. X. Song, C. He, and Y. Zhang, “Mechanism and function of oxidative reversal of DNA and
RNA methylation,” Annu Rev Biochem, vol. 83, no. Volume 83, 2014, pp. 585–614, Jun. 2014, doi:
10.1146/ANNUREV-BIOCHEM-060713-035513/CITE/REFWORKS.

[40] Muthukrishnan, S., Both, G. W., Furuichi, Y., & Shatkin, A. J. (1975). 5'-Terminal 7-
methylguanosine in eukaryotic mRNA is required for translation. Nature, 255(5503), 33-37.
doi:10.1038/255033a0

[41] B. Manavalan, T. H. Shin, and G. Lee, “PVP-SVM: Sequence-based prediction of phage


virion proteins using a support vector machine,” Front. Microbiol., vol. 9, no. MAR, pp.
1–10, 2018, doi: 10.3389/fmicb.2018.00476.

[42] T. Joachims, “1 Making Large-Scale SVM Learning Practical,” 1998, [Online]. Available:
http:::[Link].

[43] “Mendeley Reference Manager.” Accessed: Aug. 18, 2024. [Online]. Available:
[Link]

[44] Q. Li, Z. Wen, and B. He, “Practical federated gradient boosting decision trees,” arXiv,
2019, doi: 10.1609/aaai.v34i04.5895.

[45] A. Blanco-Justicia, J. Domingo-Ferrer, S. Martínez, and D. Sánchez, “Machine learning


explainability via microaggregation and shallow decision trees,” Knowledge-Based Syst.,
vol. 194, no. xxxx, p. 105532, 2020, doi: 10.1016/[Link].2020.105532.

[46] R. T. Sataloff, M. M. Johns, and K. M. Kost, “No 主観的健康感を中心とした在宅高齢者における 健康関連指


標に関する共分散構造分析 Title.”

[47] M. L. Zhang and Z. H. Zhou, “ML-KNN: A lazy learning approach to multi-label


learning,” Pattern Recognit., 2007, doi: 10.1016/[Link].2006.12.019.

[48] Freund, Y., & Schapire, R. E. (1997). A Decision-Theoretic Generalization of On-Line


Learning and an Application to Boosting. Journal of Computer and System Sciences,
55(1), 119-139.

[49] Schapire, R. E. (2003). The Boosting Approach to Machine Learning: An Overview.


MSRI Publications, 71, 149-171.

[50] Zhang, T., & Yu, B. (2005). Boosting with Early Stopping: Convergence and
Consistency. The Annals of Statistics, 33(4), 1538-1557

[51] riedman, J. H. (2001). Greedy Function Approximation: A Gradient Boosting


Machine. The Annals of Statistics, 29(5), 1189-1232.

[52] Chen, T., & Guestrin, C. (2016). XGBoost: A Scalable Tree Boosting System.
Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery
and Data Mining, 785-794.

[53] Schapire, R. E., & Freund, Y. (2012). Boosting: Foundations and Algorithms. MIT
Press

[54] Cox, D. R. (1958). The Regression Analysis of Binary Sequences. Journal of the
Royal Statistical Society: Series B (Methodological), 20(2), 215-242.

[55] Hosmer, D. W., & Lemeshow, S. (2000). Applied Logistic Regression. Wiley.

[56] Bayes, T. (1763). An Essay towards Solving a Problem in the Doctrine of Chances.
Philosophical Transactions of the Royal Society of London, 53, 370-418.

[57] Johns, A., & Langley, P. (1994). Estimating Continuous Distributions in Bayesian
Classifiers. Proceedings of the 11th International Conference on Machine Learning, 338-345.

[58] Rennie, J. D. M., Shih, L., Teevan, J., & Edwards, S. (2003). Tackling the Poor
Assumptions of Naive Bayes Text Classifiers. Proceedings of the 20th International Conference
on Machine Learning, 616-623.

[59] Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system.
Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and
Data Mining, 785-794.
[60] Tianqi Chen, & Tong He. (2015). Higgs Boson Discovery with Boosted Trees.
International Conference on Machine Learning.

[61] Zhang, Y., & Jiang, W. (2020). A review of XGBoost as a new machine learning
approach in financial risk management. Journal of Risk and Financial Management, 13(7), 166.

[62] Brownlee, J. (2016). XGBoost with Python. Machine Learning Mastery.

You might also like