0% found this document useful (0 votes)
11 views6 pages

Sentiment Analysis of COVID-19 Booster Vaccine

This document summarizes a research paper that analyzed sentiment in Indonesia about the COVID-19 booster vaccine. The researchers collected 681 tweets with the hashtag "vaksinbooster" and labeled them as either positive or negative sentiment. Most tweets (554) expressed positive sentiment about the booster vaccine, while 127 expressed negative sentiment. The researchers tested various machine learning algorithms to classify sentiment and found that a support vector machine (SVM) algorithm using a polynomial kernel achieved the highest accuracy of 79.22% according to a 10-fold cross validation test. The sentiment analysis provides useful information for the government on public views of COVID-19 vaccination efforts.

Uploaded by

Kasiyem
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
11 views6 pages

Sentiment Analysis of COVID-19 Booster Vaccine

This document summarizes a research paper that analyzed sentiment in Indonesia about the COVID-19 booster vaccine. The researchers collected 681 tweets with the hashtag "vaksinbooster" and labeled them as either positive or negative sentiment. Most tweets (554) expressed positive sentiment about the booster vaccine, while 127 expressed negative sentiment. The researchers tested various machine learning algorithms to classify sentiment and found that a support vector machine (SVM) algorithm using a polynomial kernel achieved the highest accuracy of 79.22% according to a 10-fold cross validation test. The sentiment analysis provides useful information for the government on public views of COVID-19 vaccination efforts.

Uploaded by

Kasiyem
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

EN-274 JURNAL NASIONAL TEKNIK ELEKTRO DAN TEKNOLOGI INFORMASI

p-ISSN 2301–4156 | e-ISSN 2460–5719

Indonesian Society’s Sentiment Analysis Against


the COVID-19 Booster Vaccine
Dionisia Bhisetya Rarasati1, Angelina Pramana Thenata2, Afiyah Salsabila Arief3
1,2,3
Informatics Study Program, Faculty of Technology and Design, Universitas Bunda Mulia, Tangerang 15143 INDONESIA (tel. : 021 80821428; email: [Link]@[Link],
2
[Link]@[Link], 3afiiyahsarief@[Link])

[Received: 10 April 2023, Revised: 4 August 2023]


Corresponding Author: Dionisia Bhisetya Rarasati

ABSTRACT — The COVID-19 pandemic is still occurring in various countries, including Indonesia. This pandemic is
caused by the coronavirus, which has mutated into multiple virus variants, such as Delta and Omicron. As of 9 February
2022, 4,626,936 people were confirmed positive for COVID-19 in Indonesia. This number continues to rise. The Indonesian
government has prevented the spread of these virus variants by introducing booster vaccines to the public. However, this
vaccination program has caused various sentiments among Indonesians. To optimize efforts to combat COVID-19, the
government needs to know these sentiments immediately. Based on these problems, the researcher proposes the application
of machine learning technology to develop a system that can analyze the sentiments of the Indonesians toward the booster
vaccine. This research has several stages: data collection, data labeling, text preprocessing, feature extraction, and application
of the support vector machine (SVM) algorithm using various kernels, namely the linear kernel, Gaussian radial basis
function (RBF) kernel, and polynomial kernel. Furthermore, the results of the system were tested for accuracy using a 10-
fold cross validation and confusion matrix. The dataset used was 681 tweets with the hashtag “vaksinbooster.” The dataset
consists of two classes: negative (0) and positive (1). The results showed that the data were positive for the booster vaccine,
as evidenced by the higher number of positive tweets, with 554 data, compared to 127 negative tweets. In addition, the
dataset was divided into training data of 545 and testing data of 136. In addition, the test results of this study revealed that
the SVM algorithm with the polynomial kernel, which was evaluated with 10-fold cross validation, yielded the highest level
of accuracy, namely 79.22%.

KEYWORDS — COVID-19, Booster Vaccine, Sentiment Analysis, Support Vector Machine.

I. INTRODUCTION expected to be input for the government or interested parties to


The coronavirus, commonly known as COVID-19, is still decide on the next steps to suppress the spread of the COVID-
spreading almost worldwide, including in Indonesia. Various 19.
variants of the COVID-19 spreading in Indonesia include This paper is organized as such. Section II discusses
Alpha, Delta, Beta, Kappa, and Omicron. As of 9 February research related to this study and introduce the methods used.
2022, a total of 4,626,936 Indonesians were tested positive for Section III introduce the methodology that is formulated to
COVID-19 and this number is increasing [1]. It has then achieve the research goal. Examples of some methodology
prompted the government to put a greater effort to suppress the processes are also presented in this section. Then, Section IV
spread of the COVID-19’s various variants by conducting presents the result of the research, along with the
massive vaccination for Indonesians. implementation of the formulated methodology. Finally,
In Indonesia, the COVID-19 vaccination itself consists of Section V is the conclusion of the research conducted.
two mandatory doses, and the most recent of which is the
booster vaccination. The Indonesian government is striving to II. SENTIMENT ANALYSIS
socialize and invite the Indonesians to have booster Sentiment analysis is conducted to know public/user
vaccinations in order to suppress the surge in COVID-19 cases. opinions on a topic or platform. The sentiment analysis process
However, these government’s efforts have induced various generally uses data from the internet and social
sentiments among the Indonesians. These sentiments must be media/platforms. A study whose data was sourced from
immediately known by the government to optimize efforts to Facebook examines public sentiment analysis towards
combat the numerous variants of the COVID-19. These various Indonesia’s presidential and vice-presidential candidates in
sentiments are usually poured into social media such as Twitter, 2019. The research resulted in the presidential and vice-
where users can exchange information through text, images, presidential candidates, Jokowi-Ma’ruf received 56.76%
and video [2]. positive sentiment and 43.24% negative sentiment, while
Because of these issues, this research aims to develop a Prabowo-Sandi received 24.21% positive sentiment and 75.79%
system that can analyze the sentiments of Indonesians on negative sentiment. The said results were obtained using the
Twitter toward the booster vaccine. This analysis will classify naïve Bayes algorithm [6].
positive and negative groups using the support vector machine A sentiment analysis study on COVID-19 was done in 2019.
(SVM) algorithm. The SVM algorithm will characterize it by With data sourced from Twitter, the researcher used the k-
developing an N-dimensional hyperplane that isolates nearest neighbors (KNN) and naïve Bayes algorithm to attain
information into two types of classification (positive and results. The results of the sentiment test obtained an accuracy
negative) [3]. Furthermore, the results of the system will be rate of 63.21% for the naïve Bayes and 58.10% for the KNN.
tested for accuracy using a 10-fold cross validation [4] and Then, the precision obtained was 59.11% for the naïve Bayes
confusion matrix [5]. The sentiment analysis results are and 53.10% for the KNN. In addition, it was found that the

Volume 12 Number 4 November 2023 Dionisia Bhisetya Rarasati: Indonesian Society's Sentiment Analysis …
JURNAL NASIONAL TEKNIK ELEKTRO DAN TEKNOLOGI INFORMASI EN-275
p-ISSN 2301–4156 | e-ISSN 2460–5719

tendency of public sentiment on Twitter was positive. It is


evidenced by the fact that there were 610 positive sentiments
and 488 negative sentiments [7].
Research done in 2020 compared the SVM, random forest
(RF), and stochastic gradient descent (SGD) algorithms to
classify the performance (good or bad) of programmers during
social media activities. It obtained accurate results with cross-
validation of SVM (81.3%), RF (74.4%), and SGD (80.1%).
These results indicate that the SVM algorithm performs better
than the other two algorithms in classifying programmer
performance (good or bad) during social media activities [8].
Furthermore, most related research descriptions only examine
sentiment analysis of public opinion regarding the presidential
election and the COVID-19 outbreak using the naïve Bayes
algorithm.
However, public opinion on the COVID-19 booster
vaccination, which is an effort from the government to suppress
the surge in COVID-19 cases, has yet to be investigated. In
addition, the SVM algorithm has good performance in
Figure 1. SVM hyperplane.
grouping datasets. Thus, using the SVM algorithm, this study
applies machine learning technology to analyze public 𝑇
𝜅(𝑥𝑖 , 𝑥𝑗 ) = 𝑥 𝑥𝑗 . (3)
sentiment on the COVID-19 booster vaccination from Twitter. 𝑖
An appropriate kernel function, such as the nonlinear Gaussian
III. MACHINE LEARNING
Machine learning is a part of artificial intelligence that is RBF, can enable the SVM if there is an occurrence of problems
widely used to solve various problems through learning which cannot be separated linearly. The kernel equation is
through data in various forms, such as texts, numbers, images, shown in (4), where 𝜎 signifies the kernels width [13].
and videos. The learning process from these data are obtained 2
‖𝑥𝑖 𝑥𝑗 ‖
through two stages: training and testing [9]. The discovery of 𝜅(𝑥𝑖 , 𝑥𝑗 ) = exp (− ). (4)
2𝜎 2
interesting knowledge is obtained from data in the form of texts.
However, in the Gaussian RBF kernel with the σ parameter
Machine learning works with various algorithms such as
as the kernel width, the SVM will be overfitting in all training
decision trees, naïve Bayes, KNN, RF, and SVM [10].
instances when the parameter is close to zero. Setting a more
IV. SUPPORT VECTOR MACHINE significant value to σ may result in underfitting, in which all
The SVM algorithm aims to find the best hyperplane that instances are classified into one class. Because of that, the
can perfectly separate two classes with the widest margins. correct value must be selected for the kernel. The polynomial
Margins are described as the distance between the said kernel degrees control higher degrees, allowing more flexible
hyperplane with the nearest support vectors of each class, while decision limits than linear limits and the flexibility of the
support vectors can be described as the furthest data point of classifier width. The kernel equation is shown in (5),
each class in the hyperplane [5]. The SVM theory has been where p signifies the degree of the polynomial kernel [13].
evolving since the 1960s, but it was not introduced until 1992 𝑇 𝑝
𝜅(𝑥𝑖 , 𝑥𝑗 ) = (1 + 𝑥 𝑥𝑗 ) . (5)
by Vapnik, Boser, and Guyon. The SVM functions as a method 𝑖
for linearly generating a hyperplane from a dataset into two Meanwhile, when training data using the SVM and
classes. In Figure 1, the hyperplane is a general term for all selecting kernel functions, several decisions must be made
dimensions [11]. For example, for a one-dimensional data set, during the data preparation, among others, by labeling and
the hyperplane can manifest as a point; if the set is in the form setting SVM parameters to provide optimal results.
of two dimensions, the hyperplane is a straight line [12].
Figure 1 shows that a pair of parallel hyperplanes can V. ACCURACY
separate two classes. The first boundary plane becomes the Below is the implementation of the testing stages.
boundary of the first class. In contrast, the second boundary A. CROSS VALIDATION
plane is the boundary of the second class, so (1) and (2) are Cross validation is one of the methods that could be used to
obtained where w is the normal plane and b is the position of search for the validity of machine learning models. This
the plane relative to the coordinate center [11]. method works by partitioning learning set into k-subsets and
𝑥𝑖 . 𝑤 + 𝑏 ≥ +1, if 𝑦𝑖 = +1 (1) thus, have knns. Where, in each fold, as much as (k-1)-subsets
are used as training set and the rest is used as the validation set.
𝑥𝑖 . 𝑤 + 𝑏 ≤ −1, if 𝑦𝑖 = −1. (2) This procedure is repeated until the entire subset becomes a
validation set. It is recommended to use 10 as the best value for
As for this algorithm, the most widely used kernel learning k when trying to validate a model, or it is widely known as 10-
includes linear kernels, Gaussian RBF, and polynomials. The fold cross validation [4].
SVM algorithm with this kernel finds hyperplanes by data
mapping from feature space to higher dimensional kernel space. B. CONFUSION MATRIX
This way leads to achieving nonlinear separation in kernel The confusion matrix test is a size N  N matrix, where N is
space. In addition, the linear kernel can be expressed as (3) [13]. the number of classes used in the classification. This matrix

Dionisia Bhisetya Rarasati: Indonesian Society's Sentiment Analysis … Volume 12 Number 4 November 2023
EN-276 JURNAL NASIONAL TEKNIK ELEKTRO DAN TEKNOLOGI INFORMASI
p-ISSN 2301–4156 | e-ISSN 2460–5719

TABLE I TABLE II
CONFUSION MATRIX 2  2 TOKENIZING STAGE PROCESS

Stages Results
Predicted Actual Values DPP Projo mengadakan vaksin booster
Values
Positive Negative Initial Tweet Covid- gratis untuk rakyat selama lima hari
Positive True positive (TP) False positive (FP) untuk masyarakat DKI Jakarta
Negative False negative (FN) True Negative (TN) dpp projo mengadakan vaksin booster
Lowercasing covid- gratis untuk rakyat selama lima hari
untuk masyarakat dki jakarta
dpp projo mengadakan vaksin booster
Punctuation
covid gratis untuk rakyat selama lima hari
removal
untuk masyarakat dki jakarta
“dpp” “projo” “mengadakan” “vaksin”
Words separation
“booster” “covid” “gratis” “untuk”
as individual
“rakyat” “selama” “lima” “hari” “untuk”
token
“masyarakat” “dki” “jakarta”
TABLE III
STOPWORD STAGE PROCESS

Stages Results
“dpp” “projo” “mengadakan” “vaksin”
Before
“booster” “covid” “gratis” “untuk” “rakyat”
stopword
“selama” “lima” “hari” “untuk”
removal
“masyarakat” “dki” “jakarta”
“dpp” “projo” “mengadakan” “vaksin”
After stopword
“booster” “covid” “gratis” “rakyat” “lima”
removal
“masyarakat” “dki” “jakarta”
Figure 2. Research flow. TABLE IV
STEMMING STAGE PROCESS
compares the value predicted by the model with the actual
value. The confusion matrix used in this study is a 2  2 matrix, Stages Examples
as shown in Table I [5]. Token mengadakan
Affixes meng–ada-kan
Table I shows that the true positive (TP) value is obtained
Stemming Results ada
when the predicted and actual values are positive. On the other
hand, the true negative (TN) value is obtained when the noise in tweets, such as URLs and usernames. In addition, the
predicted value and actual value are negative. Furthermore, the system will also convert nonformal Indonesian words or
false positive (FP) value is obtained if the predicted value is abbreviation into words that adhere to the Great Dictionary of
positive, but the actual value is negative. The false negative Indonesian Language (Kamus Besar Bahasa Indonesia, KBBI)
(FN) value is obtained if the predicted value is negative, but the and extract words that start with hashtags.
actual value is positive. Meanwhile, the accuracy formula in
this test can be seen in (6). Based on the equation, the greater 1) TOKENIZING
the value of TN and TP, the higher the level of accuracy [5]. The first stage in preprocessing is tokenizing [12]. Several
𝑇𝑃 + 𝑇𝑁
processes are carried out at this stage: lowercasing, punctuation
𝐴𝑐𝑐𝑢𝑟𝑎𝑐𝑦 = × 100%. (6) removal, and breaking a sentence into separate words [14]. The
𝑇𝑃 + 𝑇𝑁 + 𝐹𝑃 + 𝐹𝑁
process can be seen in Table II.
VI. METHODOLOGY
The research is conducted with the research flow depicted 2) STOPWORD
in Figure 2. The second stage is the stopword. This stage is carried by
removing high-frequency words in the texts that have no
A. DATA COLLECTION STAGE AND LABELING STAGE special or insignificant meaning. Examples of stopword in
The data used in this research are Indonesian tweets sourced Indonesian are “dan,” “atau,” “sebuah,” and “adalah” [15]. This
from Twitter. The use of these data aligns with the research’s process can be seen in Table III. As seen in Table III, stopword
goal of developing a system that can analyze the sentiments of will be deleted once it is identified.
Indonesian Twitter users regarding the booster vaccine. The
aforementioned data were collected from the data source using 3) STEMMING
Python script with “vaksinbooster” as the keyword. The data The third stage is stemming, a process that eliminates
comprised 681 tweet data with the hashtag “vaccination affixes in order to find the root words. This stage aims to reduce
booster,” gathered from 12 January 2022 until 12 April 2022. the number of words processed in text mining with the purpose
Then, the step proceeded to the labeling stage, during which the of reducing the processing time and minimizing the memory
data were assigned a class label based on two classes (positive used [16]. The results of the stemming process can be seen in
and negative). Table IV.
B. TEXT PREPROCESSING STAGE 4) SYNONYM MERGING
This stage prepares data so the machine can analyze them The final stage in preprocessing is combining synonyms or
easily [10]. There are several stages: tokenizing, removing language forms with similar meanings. At this stage, different
stopword, and stemming. These steps are carried out to remove words with similar meanings will be merged. Examples of

Volume 12 Number 4 November 2023 Dionisia Bhisetya Rarasati: Indonesian Society's Sentiment Analysis …
JURNAL NASIONAL TEKNIK ELEKTRO DAN TEKNOLOGI INFORMASI EN-277
p-ISSN 2301–4156 | e-ISSN 2460–5719

Figure 3. confusion matrix accuracy with linear kernel. Figure 5. Confusion matrix accuracy with the Gaussian RBF kernel.

Figure 6. 10-fold cross validation accuracy with the Gaussian RBF kernel.

Figure 4. Ten-fold cross validation accuracy with linear kernel.


On (8), Z represents the z-score normalization, x represents the
value from data, 𝑥̅ is the mean of the data, and s is the standard
words with the same meaning are “saya” and “aku.” This stage deviation.
aims to minimize the number of words in the system while
maintaining the number of frequencies [3]. D. SUPPORT VECTOR MACHINE (SVM) STAGE
The system groups tweets into two clusters, namely positive
C. FEATURE EXTRACTION STAGE and negative. A hyperplane groups each tweet. When using
Feature extraction is the subsequent stage after the SVM algorithm, there are three kernels that could be used to
preprocessing stage. Since the textual data (in this case is tweets) find the best hyperplane, namely the Gaussian RBF kernel,
are public opinion, the writing structure is often not structured. linear kernel, and polynomial kernel [13]. Thus, during the
Therefore, it is necessary to transform textual data into research the three kernels’ accuracies were compared with each
structured data so that the machine learning algorithm can work other, and the best kernel found was used as a result of
immediately [17]. There are two processes in this stage, namely sentiment analysis based on their accuracy results.
weighting and z-score normalization. E. ACCURACY STAGE
1) WEIGHTING The system results were tested for accuracy using the 10-
fold cross validation and confusion matrix. The test results of
Weighting is a stage that reflects how important a word is
the two methods were then compared, and the best results was
in the document. The method used to do weighting in text
utilized to test the accuracy of sentiment analysis.
mining is the term frequency-inverse document frequency (TF-
IDF), where the formula can be seen in (7). The weight of the VII. RESULTS AND DISCUSSION
terms in the text will increase as their occurrences rise. The dataset collected on Twitter with the hashtag
However, this increase is also in line with the frequency of “vaksinbooster” is 681 tweets. These data consisted of two
terms’ occurrences that are included within the research classes: negative (0) and positive (1). In addition, the data
domain of interest [18]. showed that public sentiment towards the booster vaccine
tended to be positive, as indicated by the 554 positive and 127
𝑊𝑡,𝑑 = 𝑡𝑓𝑡,𝑑 𝑥 𝑖𝑑𝑓𝑡 . (7) negative tweets. The dataset was divided into 545 training data
and 136 testing data. Furthermore, text preprocessing was
where Wt,d represent the weight, tft,d represent the term
applied to the dataset, resulting in an array of significant words
frequency (TF) of the word, and idft represent the inverse
from a tweet. Meanwhile, the weight values of each word from
document frequency (IDF).
the text preprocessing stage were obtained from conducting the
2) Z-SCORE NORMALIZATION feature extraction stage. The weight value from extraction
The z-score normalization process is a step that is carried results were afterwards processed as inputs using the SVM
out after obtaining the weight value. Due to the significant algorithm. The SVM algorithm was then applied using three
difference in the range value that will impact the classification kernels: the linear kernel (see Figure 3 and Figure 4), Gaussian
problems later, normalization must be carried out [19]. In RBF kernel (see Figure 4), and polynomial kernel (see Table V
addition, (8) can be used to perform z-score normalization [20]. and Table VI).
𝑥−𝑥̅ Figure 3 describes the results of testing the SVM algorithm
𝑍= . (8) with a linear kernel on the sentiment analysis of booster vaccine.
𝑠

Dionisia Bhisetya Rarasati: Indonesian Society's Sentiment Analysis … Volume 12 Number 4 November 2023
EN-278 JURNAL NASIONAL TEKNIK ELEKTRO DAN TEKNOLOGI INFORMASI
p-ISSN 2301–4156 | e-ISSN 2460–5719

TABLE V In 2020, research was conducted to investigate the


CONFUSION MATRIX RESULTS sentiment analysis of OVO services on Twitter using the SVM
n = 136 Actual: No Actual: Yes Total algorithm. The research found that using a polynomial kernel
Predicted: No 0 0 0
in the SVM process had a high accuracy rate of 89.09% [22].
In addition, research on sentiment analysis of the large-scale
Predicted: Yes 37 99 136
social restrictions (pembatasan sosial berskala besar,
37 99
PSBB) policies on Twitter found that the SVM algorithm with
TABLE VI a polynomial kernel had an accuracy rate of 94.55% [23].
10-FOLD CROSS VALIDATION RESULTS Based on the results of this study, the SVM algorithm with a
Fold Accuracy Training Testing polynomial kernel is the best method for analyzing booster
k-1 88.41% 612 69 vaccine sentiment on Twitter social media.
k -2 82.61% 612 69
VIII. CONCLUSION
k -3 81.16% 612 69 In Indonesia, the COVID-19 vaccination entails the
k -4 81.16% 612 69 administration of two compulsory doses, with the most recent
k -5 75.36% 612 69 dosage being the booster vaccination. However, the booster
k -6 78.26% 612 69 vaccines have rose various sentiments among the Indonesians.
k -7 86.96% 612 69 Therefore, the researcher proposes the application of machine
k -8 75.36% 612 69 learning technology to develop a system that can analyze the
k -9 69.57% 612 69 sentiments of the Indonesian towards the booster vaccine using
k -10 73.33% 621 60 the SVM algorithm. This study gathered 681 Indonesian
Twitter data, specifically focusing on the hashtag
Testing with the confusion matrix method yielded positive “vaksinbooster.” The data were subsequently categorized into
predictive data following positive facts in as many as 92 and negative (0) and positive (1). In addition, the dataset was
negative predictive data following negative reality in as many divided into training data consisting of 545 data and a testing
as eight positive predictive data. Still, negative reality found as data consisting of 136 data. The data shows that public
many as 29 data and negative predictive data, but positive sentiment toward the booster vaccine tends to be positive, as
reality found as many as 7 data. Therefore, the results of the evidenced by the positive tweets of 554 data and the negative
accuracy level with the confusion matrix test method were 74%. tweets of 127.
On the other hand, as seen in Figure 4, testing with the 10-fold The system testing using the 10-fold cross validation found
cross validation method yielded a higher level of accuracy than that the SVM algorithm using the linear kernel achieved an
the confusion matrix, which was 76.07%. accuracy rate of 76.07%. At the same time, the Gaussian RBF
Figure 5 shows the results of testing the SVM algorithm kernel obtained an accuracy rate of 78.2%. Then, the
with the Gaussian RBF kernel for sentiment analysis of booster polynomial kernel yielded an accuracy rate of 79.22%. The
vaccines. Testing with the confusion matrix method yielded system testing using the confusion matrix method found that
positive predictive data corresponding to positive reality as the SVM algorithm using the linear kernel resulted in an
much as 99 and negative predictive data corresponding to accuracy rate of 74%. Meanwhile, the Gaussian RBF kernel
negative reality as much as 0. Meanwhile, positive predictive
obtained an accuracy rate of 73%. The polynomial kernel
data with negative reality were found in 37 data, and negative
achieved an accuracy rate of 73%. Hence, the best kernel for
predictive data with positive reality were found in 0 data.
Therefore, the results of the confusion matrix testing yielded an the application of the SVM algorithm is the polynomial kernel,
accuracy level of 73%. On the other hand, the utilization of the which achieves the highest level of accuracy when tested with
10-fold cross validation method yielded a higher level of the 10-fold cross validation.
accuracy than the confusion matrix, which was 78.2% (see CONFLICT OF INTEREST
Figure 6). All authors state that there is no conflict of interest in this
In 2019, research on the human development index (HDI) study.
classification using the SVM with the Gaussian RBF kernel
achieved an accuracy of 98.1%, which was higher than the AUTHOR CONTRIBUTION
application of the linear kernel, which yielded an accuracy of Conceptualization, Dionisia Bhisetya Rarasati;
95.1% [21]. methodology, Dionisia Bhisetya Rarasati; writing—original
Table V and Table VI describes the results of testing the draft preparation, Angelina Pramana Thenata and Afiyah
SVM algorithm with polynomial kernel on the sentiment of Salsabila Arief; writing—review and editing, Dionisia
booster vaccine analysis. As shown on Table V, testing with Bhisetya Rarasati, Angelina Pramana Thenata, and Afiyah
the confusion matrix method obtained the same results as those Salsabila Arief.
achieved by the SVM algorithm with Gaussian RBF kernel,
which had an accuracy rate of 73%. However, as depicted on ACKNOWLEDGMENT
Table VI, testing using the 10-fold cross validation method Thanks to the Ministry of Research, Technology, and
yielded a higher level of accuracy than the confusion matrix, Higher Education who has supported this research through the
Beginner Lecturer Research Grant for the 2022 fiscal year with
which was averaged at 79.22%.
the contract number of 155/E5/[Link]/2022.
Thus, the application of the SVM algorithm with the best
kernel was found in the polynomial kernel, exhibiting the REFERENCES
highest level of accuracy through the implementation of using [1] (2022) “Peta Sebaran COVID-19,” [Online], [Link]
the 10-fold cross validation test method, which is 79.22%. sebaran#, access date: 14-Mar-2022.

Volume 12 Number 4 November 2023 Dionisia Bhisetya Rarasati: Indonesian Society's Sentiment Analysis …
JURNAL NASIONAL TEKNIK ELEKTRO DAN TEKNOLOGI INFORMASI EN-279
p-ISSN 2301–4156 | e-ISSN 2460–5719

[2] D.A. Agustina, S. Subanti, and E. Zukhronah, “Implementasi Text Machines,” Sens., Vol. 19, No. 23, pp. 1–16, Nov. 2019, doi:
Mining pada Analisis Sentimen Pengguna Twitter Terhadap Marketplace 10.3390/s19235219.
di Indonesia Menggunakan Algoritma Support Vector Machine,” Indones. [14] V.Y. Radygin et al., “Application of Text Mining Technologies in
J. Appl. Stat., Vol. 3, No. 2, pp. 109–122, Nov. 2020, doi: Russian language for Solving the Problems of Primary Financial
10.13057/ijas.v3i2.44337. Monitoring for Solving the Problems of Primary Financial Monitoring,”
[3] D.B. Rarasati, “A Grouping of Song-Lyric Themes Using K-Means Procedia Comput. Sci., Vol. 190, pp. 678–683, 2021, doi:
Clustering,” JISA (J. Inform., Sains), Vol. 3, No. 2, pp. 38–41, Dec. 2020, 10.1016/[Link].2021.06.078.
doi: 10.31326/jisa.v3i2.658. [15] H. Najjichah, A. Syukur, and H. Subagyo, “Pengaruh Text Preprocessing
[4] D. Berrar, “Cross-Validation,” in Encyclopedia of Bioinformatics and dan Kombinasinya pada Peringkas Dokumen Otomatis Teks Berbahasa
Computational Biology: ABC of Bioinformatics, vol. 1, S. Ranganathan, Indonesia,” J. Teknol. Inf. Cyberku, Vol. 15, No. 1, pp. 1–11, Jan. 2019.
K. Nakai, C. Schönbach, Eds., Amsterdam, Netherlands: Elsevier, 2019, [16] D. Sebastian, “Implementasi Algoritma K-Nearest Neighbor untuk
pp. 542–545, doi: 10.1016/B978-0-12-809633-8.20349-X. Melakukan Klasifikasi Produk dari Beberapa E-marketplace,” J. Tek.
[5] Y. Sari, P.B. Prakoso, and A.R. Baskara, “Road Crack Detection using Inform., Sist. Inf., Vol. 5, No. 1, pp. 51–61, Apr. 2019, doi:
Support Vector Machine (SVM) and OTSU Algorithm,” 2019 6th Int. 10.28932/jutisi.v5i1.1581.
Conf. Electr. Veh. Technol. (ICEVT), 2019, pp. 349–354, doi: [17] H. Liu, P. Burnap, W. Alorainy, and M. L. Williams, “A Fuzzy Approach
10.1109/ICEVT48285.2019.8993969. to Text Classification with Two-Stage Training for Ambiguous Instances,”
[6] B. Haryanto et al., “Facebook Analysis of Community Sentiment on 2019 IEEE Trans. Comput. Soc. Syst., Vol. 6, No. 2, pp. 227–240, Apr. 2019,
Indonesian Presidential Candidates From Facebook Opinion Data,” doi: 10.1109/TCSS.2019.2892037.
Procedia Comput. Sci., Vol. 161, pp. 715–722, 2019, doi: [18] I. Yahav, O. Shehory, and D. Schwartz, “Comments Mining With TF-
10.1016/[Link].2019.11.175. IDF: The Inherent Bias and Its Removal,” IEEE Trans. Knowl., Data Eng.,
[7] M. Syarifuddin, “Analisis Sentimen Opini Publik Mengenai COVID-19 Vol. 31, No. 3, pp. 437–450, Mar. 2019, doi:
pada Twitter Menggunakan Metode Naïve Bayes dan K-NN,” Inti Nusa 10.1109/TKDE.2018.2840127.
Mandiri, Vol. 15, No. 1, pp. 23–28, Aug. 2020, doi: [19] Henderi, W. Tri, and R. Efana, “Comparison of Min-Max Normalization
10.33480/inti.v15i1.1347. and Z-Score Normalization in the K-Nearest Neighbor (KNN) Algorithm
[8] R. Umar, I. Riadi, and Purwono, “Perbandingan Metode SVM, RF dan to Test the Accuracy of Types of Breast Cancer,” IJIIS Int. J. Inform., Inf.
SGD untuk Penentuan Model Klasifikasi Kinerja Programmer pada Syst., Vol. 4, No. 1, pp. 13–20, Mar. 2021, doi: 10.47738/ijiis.v4i1.73.
Aktivitas Media Sosial,” J. RESTI (Rekayasa Sist., Teknol. Inf.), Vol. 4, [20] M.A. Imron and B. Prasetiyo, “Improving Algorithm Accuracy K-
No. 2, pp. 329–335, Apr. 2020, doi: 10.29207/resti.v4i2.1770. Nearest Neighbor Using Z-Score Normalization and Particle Swarm
[9] A. Roihan, P.A. Sunarya, and A.S. Rafika, “Pemanfaatan Machine Optimization to Predict Customer Churn,” J. Soft Comput. Explor., Vol.
Learning dalam Berbagai Bidang: Review Paper,” IJCIT (Indones. J. 1, No. 1, pp. 56–62, Sep. 2020, doi: 10.52465/joscex.v1i1.7.
Comput. Inf. Technol., Vol. 5, No. 1, pp. 75–82, May 2020, [21] H. Al Azies, D. Trishnanti, and E. Mustikawati P.H, “Comparison of
doi:10.31294/ijcit.v5i1.7951. Kernel Support Vector Machine (SVM) in Classification of Human
[10] A.P. Thenata, “Text Mining Literature Review on Indonesian Social Development Index (HDI),” IPTEK J. Proc. Ser., No. 6, pp. 53–57, Nov.
Media,” J. Eduk., Penelit. Inform., Vol. 7, No. 2, pp. 226–232, Aug. 2021, 2019, doi: 10.12962/j23546026.y2019i6.6394.
doi: 10.26418/jp.v7i2.47975. [22] F. Romadoni, Y. Umaidah, and B.N. Sari, “Text Mining untuk Analisis
[11] F.S. Jumeilah, “Penerapan Support Vector Machine (SVM) untuk Sentimen Pelanggan Terhadap Layanan Uang Elektronik Menggunakan
Pengkategorian Penelitian,” J. RESTI (Rekayasa Sist., Teknol. Inf.), Vol. Algoritma Support Vector Machine,” J. Sisfokom (Sist. Inf., Komput.),
1, No. 1, pp. 19–25, Apr. 2017, doi: 10.29207/resti.v1i1.11. Vol. 9, No. 2, pp. 247–253, Jul. 2020, doi: 10.32736/sisfokom.v9i2.903.
[12] S. Symeonidis, D. Effrosynidis, and A. Arampatzis, “A Comparative [23] H.P.P. Zuriel and A. Fahrurozi, “Implementasi Algoritma Klasifikasi
Evaluation of Pre-Processing Techniques and Their Interactions for Support Vector Machine untuk Analisa Sentimen Pengguna Twitter
Twitter Sentiment Analysis,” Expert Syst. Appl., Vol. 110, pp. 298–310, Terhadap Kebijakan PABB,” J. Ilm. Inform. Komput., Vol. 26, No. 2, pp.
Nov. 2018, doi: 10.1016/[Link].2018.06.022. 149–162, Aug. 2021, doi: 10.35760/ik.2021.v26i2.4289.
[13] C. Savas and F. Dovis, “The Impact of Different Kernel Functions on the
Performance of Scintillation Detection Based on Support Vector

Dionisia Bhisetya Rarasati: Indonesian Society's Sentiment Analysis … Volume 12 Number 4 November 2023

Common questions

Powered by AI

The researchers analyzed sentiments by collecting 681 tweets with the hashtag 'vaksinbooster' and categorizing them as negative (0) or positive (1). They applied machine learning techniques, specifically the Support Vector Machine (SVM) algorithm, with various kernels to classify these sentiments. The effectiveness of different kernels—linear, Gaussian RBF, and polynomial—was assessed using a 10-fold cross-validation and confusion matrix, with the polynomial kernel achieving the highest accuracy of 79.22% .

The research revealed that the SVM algorithm with different kernels produced varying accuracy levels in sentiment classification. The linear kernel achieved an accuracy rate of 76.07%, the Gaussian RBF kernel had an accuracy rate of 78.2%, and the polynomial kernel resulted in the highest accuracy rate of 79.22% using 10-fold cross-validation . These results indicate that the polynomial kernel was the most effective for this dataset .

The accuracy of the SVM algorithm has significant implications for deploying sentiment analysis in real-world applications. High accuracy suggests the method's reliability and effectiveness in correctly interpreting public sentiment, which is critical for guiding decisions, especially in public health and policy-making contexts such as vaccine campaigns. The study's 79.22% accuracy with the polynomial kernel indicates its potential utility in informing government strategies and public outreach efforts .

Data preprocessing and feature extraction were vital in transforming raw tweet data into a structured format suitable for machine learning models. Preprocessing involved cleaning and normalizing text, while feature extraction quantified the textual information using term frequency-inverse document frequency (TF-IDF). This resulting structured data was then used as input for the SVM algorithm, enabling effective sentiment classification .

Sentiment analysis outcomes using SVM help inform strategic decision-making by providing insights into public sentiment regarding health interventions such as vaccines. Understanding public opinion allows policymakers to identify misperceptions, design effective communication strategies, and manage public concerns, thereby enhancing the effectiveness of health campaigns and ensuring higher compliance and support among the populace .

Social media data provides a rich, real-time repository of public opinion that can be analyzed to understand prevalent sentiments. In this study, Twitter served as a platform enabling the collection of sentiments related to the COVID-19 booster vaccine, reflecting diverse public emotions and thoughts on this issue. This data offers actionable insights into public opinion, which can guide policy and communication strategies .

Challenges in using Twitter data for sentiment analysis include handling unstructured text, interpreting context from limited characters, and managing biases inherent in social media platforms. Additionally, data preprocessing is complex, as it requires filtering out noise, emojis, and irrelevant information. The study's reliance on a specific hashtag may also limit the generalizability of findings, as it might not capture all relevant sentiments .

Understanding sentiments towards the COVID-19 booster vaccine is crucial for the Indonesian government to optimize their efforts in combating the spread of the virus. By knowing public perceptions, the government can better address concerns, improve vaccination campaigns, and tailor communication strategies that promote vaccine acceptance .

The SVM algorithm classifies sentiments by creating a hyperplane that separates data into two categories: positive and negative sentiments. Kernels are critical in determining the shape and dimensions of the hyperplane. Different kernels—linear, Gaussian RBF, and polynomial—offer distinct methods for mapping input data into a higher-dimensional space, which affects classification performance. In this research, the choice of kernel significantly impacted accuracy, with the polynomial kernel providing the best results .

The polynomial kernel achieved superior accuracy in this study because it effectively mapped the input data into a higher-dimensional space where a clear separation between sentiment classes was attainable. This suggests that polynomial kernels may offer advantages in certain datasets where non-linear relationships are present, encouraging future studies to consider kernel selection carefully based on data characteristics to maximize classification performance .

You might also like