Sentiment Analysis of Maxim App Reviews
Sentiment Analysis of Maxim App Reviews
1. INTRODUCTION
The continuous advancement of technology has had a significant impact, leading to more and more young generations
continually innovating to simplify daily activities. One such innovation is the emergence of online transportation
service providers. Online transportation has become an integral part of daily life in Indonesia. According to a survey
conducted by the Association of Indonesian Internet Service Providers (Asosiasi Penyelenggara Jasa Internet
Indonesia or APJII), online transportation ranks 16th among the 22 reasons for using the internet in Indonesia. This
indicates that using online transportation services has become a habit for Indonesian society [1]. Online transportation
service providers in Indonesia are currently in fierce competition. The Chairman of the Online Driver Association
(Asosiasi Driver Online or ADO) revealed that in recent years, more than 10 online transportation companies have
been forced to shut down because they couldn't capture and dominate the ride-hailing market in Indonesia. Currently,
Gojek, Grab, and Maxim are the top online transportation services based on the number of users [2][3].
Maxim is one of the online transportation service providers operating in Indonesia since 2018. In addition to
providing pick-up and delivery services, Maxim also offers food delivery and freight transport services. Maxim has
consistently improved its technology and business processes with the main mission of enhancing user interaction and
assisting many people in reaching their destinations [4].
According to a survey conducted by the Ministry of Transportation's Research and Development Agency
(Balitbang Kementerian Perhubungan), the most widely used online transportation service in the Jabodetabek area is
Gojek, with 59.13% of the total respondents using it. In second place is Grab, with 32.24% of users, while Maxim
ranks third with 6.93% of users [5]. Therefore, it can be concluded that there may still be areas for improvement. One
way to assess Maxim's performance and quality is by understanding user experiences, which can be obtained from
reviews and comments on the Google Play Store. User reviews can reflect their views on the application experience,
including service quality, user-friendliness, and more.
As Maxim's popularity grows, the number of users and feedback on Google Play Store has also increased.
Based on data recorded on Google Play Store, the Maxim application has an overall rating of 4.7 out of 5 and has
received 2.85 million product reviews to date. Achieving high ratings and reviews is crucial for assessing an
application. New users often consider ratings and reviews as references and considerations before using an app.
However, rating acquisition is not always accurate and may not provide in-depth information about user experiences.
Rating analysis is easy to understand and aids in decision-making by looking at the average ratings provided by users.
However, rating analysis has some drawbacks, such as the inability to capture complex emotions that may be present
in textual reviews and susceptibility to bias and manipulation, such as ratings without explanatory reviews and
mismatched reviews and ratings.
Another way to assess Maxim's performance and quality is by understanding user experiences through reviews
and comments on the Google Play Store. User reviews can reflect their experiences with the app, including service
quality, ease of use, and more [6]. Therefore, sentiment analysis of user reviews is necessary. Sentiment analysis
provides deeper insights compared to rating analysis alone. The advantage of sentiment analysis over rating analysis
is its ability to identify the emotions and context conveyed in user-written text [7]. Sentiment analysis categorizes data
into two classes: positive sentiment and negative sentiment, providing valuable insights from the text. With numerous
user reviews, sentiment analysis offers a faster and more efficient way to process review data.
In this study, the Support Vector Machine (SVM) algorithm is used as the research method. Previous research
has shown that the SVM algorithm outperforms other algorithms in review data processing. In previous studies, SVM
achieved better accuracy than Naïve Bayes, with SVM accuracy at 77.00% and Naïve Bayes at 70.40%. The study
also suggested that SVM could be used effectively for further research with the same data characteristics [8].
Furthermore, a study comparing sentiment analysis of review data between SVM and Decision Tree found that SVM
outperformed with an accuracy of 90.20% compared to Decision Tree's 89.80% [9]. Another study comparing Naïve
Bayes and SVM for sentiment analysis found that SVM achieved the best accuracy at 82.48%, while Naïve Bayes
achieved 76.56% [10]. Additionally, a study conducting sentiment analysis with Naïve Bayes reported an accuracy of
74.37% and an AUC of 0.659, while SVM performed better with an accuracy of 81.22% and an AUC of 0.886 [11].
This research uses the Support Vector Machine (SVM) algorithm for text classification into positive and
negative sentiments. Proper preprocessing, processing, and evaluation steps are followed to obtain more accurate
results. The goal of this research is to implement the Support Vector Machine (SVM) algorithm for sentiment analysis
and to compare sentiment results for the Maxim application. The results of this research, which include positive and
negative sentiment results, can serve as an evaluation and reference for improving Maxim's performance and services.
2. RESEARCH METHODOLOGY
2.1 Research Stages
The research conducted focuses on the user environment of the Maxim online transportation service application, where
users provide opinions in the form of comments or reviews about the Maxim application on Google Play Store. The
method used in this research is text data classification using Support Vector Machine (SVM). The problem-solving
process in this research consists of four stages: Problem and Solution Identification Stage, Preprocessing Stage,
Modeling Stage, and Conclusion Stage.
𝑻𝑷
𝑹𝒆𝒄𝒂𝒍𝒍 = 𝑻𝑷+𝑭𝑵 (5)
𝑻𝑷
𝑷𝒓𝒆𝒄𝒊𝒔𝒊𝒐𝒏 = 𝑻𝑷+𝑭𝑷
(6)
𝟐(𝑹𝒆𝒄𝒂𝒍𝒍 𝒙 𝑷𝒓𝒆𝒄𝒊𝒔𝒊𝒐𝒏)
𝑭𝟏 𝑺𝒄𝒐𝒓𝒆 = (7)
𝑹𝒆𝒄𝒂𝒍𝒍+𝑷𝒓𝒆𝒄𝒊𝒔𝒊𝒐𝒏
The preprocessing stages carried out include Cleaning Text to remove meaningless characters such as
punctuation marks, numbers, specific symbols, and convert the text to lowercase, Tokenization to separate each word
from a review sentence, Stop Word Removal to eliminate words that have no meaning, and Stemming to transform
words with affixes into their base form. The following is an example of the preprocessing results as shown in Table
4.
Table 4. Preprocessing Stages Results
Preprocessing Stages Results
Data Layanan mantap terima kasih sudah mengantar saya sampai tujuan.
(Great service, thank you for taking me to my destination.)
Cleaning Text layanan mantap terima kasih sudah mengantar saya sampai tujuan
(Great service, thank you for taking me to my destination)
Tokenization ['layanan', 'mantap', 'terima', 'kasih', 'sudah', 'mengantar', 'saya', 'sampai', 'tujuan']
(['service', 'steady', 'accept', 'love', 'already', 'delivered', 'me', 'arrived', 'destination'])
Stop Word Removal ['layanan', 'mantap', 'terima', 'kasih', 'mengantar', 'tujuan']
(['service', 'steady', 'accept', 'love', 'deliver', 'destination'])
Stemming ['layan', 'mantap', 'terima', 'kasih', 'antar', 'tuju']
(['serve', 'steady', 'receive', 'love', 'deliver', 'go'])
Table 4 shows that the initial data consists of reviews with punctuation marks and initial capitalization in
sentences. During the Text Cleaning, punctuation marks are removed, and sentences are converted to lowercase. Then,
in the Tokenization, words within sentences are separated for ease of processing. In the Stop Word Removal,
meaningless words are eliminated as they do not play a role in the processing. Finally, in the Stemming, words
containing affixes are converted to their base forms.
3.3 Data Labeling
In this stage, labels are assigned to the data based on the content and meaning of the review words. There are two
label groups: positive and negative labels. In this context, a positive label represents a review or comment that contains
indicators of satisfaction from Maxim's online transportation users expressed in various words or sentences. Examples
of reviews falling into the positive label category include words of praise and expressions of gratitude. On the other
hand, negative labels represent reviews that contain complaints or user dissatisfaction, and there are many reviews
mentioning shortcomings in the Maxim application or services. An example of data labeling results, as found in Table
5, is as follows.
Table 5. Data Labelling Results
Review Dataset Label
['maxim', 'jalan', 'mudah', 'terima', 'kasih', 'maxim', 'sukses'] POSITIF
(['maxim', 'way', 'easy', 'accept', 'love', 'maxim', 'success'])
['driver', 'nya', 'nakal', 'apk', 'ribu', 'suruh', 'bayar', 'ribu', 'alas', 'sesuai', 'aplikasi', 'pake', 'maxim', NEGATIF
'drivernya', 'nakal', 'curang']
(['driver', 'its', 'naughty', 'apk', 'thousand', 'tell', 'pay', 'thousand', 'base', 'appropriate', 'app', 'use',
'maxim', 'driver', 'naughty', 'cheat'])
['layan', 'ramah', 'nyaman', 'kasih', 'bintang'] POSITIF
(['serve', 'friendly', 'cozy', 'love', 'star'])
['tolong', 'lokasi', 'kurang', 'antar', 'susah', 'akurat', 'tuju'] NEGATIF
(['please', 'location', 'less', 'deliver', 'difficult', 'accurate', 'go'])
In Table 5 above, there are examples of review data that have undergone preprocessing and manual labeling
stages to prevent potential sentiment misinterpretation. For instance, the first data entry is labeled as positive because
it contains positive sentiments, such as appreciation and expressions of gratitude from a user for the user-friendly
Maxim service. The second data indicates that the user expressed disappointment in their experience with the Maxim
application, as evidenced by words like 'nakal' (naughty) and 'curang' (cheat).
3.3 Modeling
The modeling stage, also known as the modeling phase, is the stage where a model is created to represent a system or
problem that needs to be solved. The steps in this stage consist of data preparation, including dividing the dataset into
two parts: testing data and training data, assigning weights to all data using TF-IDF, forming a model by selecting a
classification algorithm and the parameters to be used, and finally evaluating the trained model.
3.3.1 Splitting the Data into Training Data and Testing Data.
The next step is data splitting, where the data will be divided into two categories: training data and testing data.
Training data is the portion of data needed by the machine to learn the characteristics of the data in order to generate
a model for making predictions on new data. Meanwhile, testing data is the new data that will be predicted using the
trained model [20]. There are three data splitting ratios used in this research: a 60:40 ratio, which means dividing the
data into 60% training data and 40% testing data; a 70:30 ratio, which means dividing the data into 70% training data
and 30% testing data; and an 80:20 ratio, which means dividing the data into 80% training data and 20% testing data.
Each of these ratios will undergo the classification process to test accuracy and determine the best accuracy value.
The details of the data distribution or split, as shown in Table 6 below, are as follows.
Table 6. Split Data Training and Testing
Data Split
Data Split Ratio Total
Training Testing
60:40 10.945 7.297
70:30 12.769 5.473 18.242
80:20 14.593 3.649
As seen in the Table 6 above, the total overall data before the split amounted to 18,242 review data. In the
60:40 data split ratio, there were 10,945 data for training and 7,297 data for testing. In the 70:30 ratio, there were
12,769 data for training and 5,473 data for testing. Finally, in the 80:20 ratio, there were 14,593 data for training and
3,649 data for testing.
3.3.2 TF-IDF Weighting
Before being used in classification, text data needs to be transformed into a numeric form that represents the weight
of each word. In Table 7 below, examples of word weighting using TF-IDF are presented for three sample documents.
Table 7. TF-IDF Weighting Examples
TFIDF
Term
D1 D2 D3
bagus 0 0 0,46735
cepat 0,40204 0,51785 0
jalan 0 0 0,46735
jemput 0 0,68091 0
motor 0 0 0,46735
mudah 0,52863 0 0
murah 0,52863 0 0
nyaman 0 0 0,46735
ramah 0 0,51785 0,35543
terimakasih 0,52863 0 0
In the classification process, during model training, SVM uses training model vectors (training data) in the
form of TF-IDF weight values as input in the equation to calculate the best hyperplane that can effectively separate
the two classes. After the model is trained, the TF-IDF weight value vectors from the testing data are used as input to
predict the correct class for each document. The testing data is projected into the same feature space used in the
training data, and then the predicted class is determined based on the position or coordinates of the vector relative to
the hyperplane that was defined during the training.
3.3.3 Model Formation
In determining the parameters of the classification model, this research uses the GridSearch method to find the best
combination of kernel and parameter C that can optimize the performance of the classification model. The kernel used
in this research is the Linear kernel. The parameter C was tested with values of [0.01, 0.1, 1, 100]. After parameter
optimization, the best combination was found to be the Linear kernel with a parameter C value of 1. These parameters
can be used to create a classification model with SVM and perform training on the training data before testing the
model on the testing data. The training and testing of the best model are conducted using the syntax in the following
Table 8.
Table 8. Classification Model
svm_model = [Link](kernel='linear', C=1)
svm_model.fit(tf_train, y_train)
predicted = svm_model.predict(tf_test)
The syntax above indicates that a linear kernel SVM is used with a parameter C = 1. Next, the model is 'fit' or
applied to the pre-weighted training data. After the training process with the training data, prediction is carried out
using new data, which is the testing data, to predict its labels or sentiments. Then, to measure how accurate or good
the classification model's performance is, evaluation is performed in the next stage.
3.3.4 Model Formation
In the SVM classification process, accuracy and evaluation results were obtained for three data splitting scenarios:
60:40, 70:30, and 80:20, as shown in Table 9 below.
Table 9. Model Performance Evaluation Results
Rasio Akurasi Precision Recall F1 Score AUC
60:40 89.02% 92.17% 93.48% 92.82% 0.8419
70:30 89.82% 92.69% 94.09% 93.38% 0.8505
80:20 88.98% 92.26% 93.39% 92.82% 0.8408
Based on the model evaluation conducted, the model with a data ratio of 60:40 achieved an accuracy of 92.17%,
the 70:30 ratio had an accuracy of 89.82%, and the 80:20 ratio had an accuracy of 88.98%. The model with the best
accuracy was the 70:30 ratio, which resulted in a precision of 92.69%, recall of 94.09%, and an F1 score of 93.38%
with a confusion matrix as shown in Table 10 below.
Table 10. Confusion Matrix Result
Prediction
Confusion Matrix
Negative Positive
Negative 982 310
Actual
Positive 247 3934
The precision value of 92.69% indicates that the model can correctly identify 92.69% of all data declared as
positive by the model. The recall value of 94.09% shows that the model can correctly identify 94.09% of all positive
data. The F1 score value of 93.38% indicates that the model has a good balance between recall and precision values.
In addition to evaluation with the confusion matrix, another technique that can be used to evaluate model performance
is by considering the area under the ROC curve or AUC (Area Under Curve). In Table 9 above, the best AUC value
was obtained by the model with a 70:30 ratio, with an AUC value of 0.8505. This value indicates that the model has
a fairly good ability to distinguish between positive and negative label classes and falls into the criteria of good
classification based on the criteria in Table 2 and the curve visualization as shown in the Figure 4 below.
The expression of positive sentiment can be depicted using various words. In the image above, there is a
visualization of frequently occurring words in the reviews. To get a clearer view of the sequence of words that appear
most dominantly in positive reviews, this research also includes a visualization with a bar chart to display the words
and their frequency of occurrence.
4. CONCLUSION
The conclusions drawn from the conducted research include the following, The implementation of the Support Vector
Machine algorithm in sentiment analysis of online transportation application reviews, specifically Maxim, resulted in
a reasonably good model performance. Performance evaluation using the confusion matrix revealed the highest
accuracy achieved by the classification model with a 70:30 data splitting ratio, which reached 89.82%. It also achieved
precision of 92.17%, recall of 94.09%, and an F1 score of 93.38%. Additionally, this model obtained the best AUC
(Area Under Curve) value of 0.8505, indicating good classification performance. Based on sentiment analysis results,
the Maxim application tends to receive more positive comments from users, accounting for 76% of the feedback.
Users expressed satisfaction with Maxim's services, such as the friendly and prompt behavior of drivers. This is one
of Maxim's strengths that should be maintained. However, there were negative comments accounting for 24% of the
feedback, which also require attention from developers to improve the application's performance and functionality.
Some users reported complaints while using the application, such as malfunctioning features. This feedback can serve
as an evaluation for developers to enhance performance, including improvements to crucial features like location
accuracy. This research can also serve as a basis for future studies in sentiment analysis of application reviews,
especially for online transportation applications. Overall, the research findings provide valuable insights into user
sentiments and feedback regarding the Maxim application, which can guide developers in making improvements and
maintaining user satisfaction.
REFERENCES
[1] M. Arif and K. U. Apjii, “Profil Internet Indonesia 2022,” no. June, 2022.
[2] R. K. Hastuti, “Duh, 2 Tahun Terakhir Ada 10 Ojol yang Tergilas Grab & Gojek,” CNBC Indonesia, 2019.
[Link]
gojek
[3] Shopback, “Sering Membandingkan Harga Transportasi Online? Aplikasi Ini Akan Memudahkan Penggunanya,”
[Link], 2018. [Link]
[4] [Link], “Tentang Perusahaan.” [Link] (tanggal akses 10 Juli 2023)
[5] A. Mutia, “Survei: Publik Jabodetabek Paling Sering Pakai Gojek, Bagaimana Grab, Maxim, dan InDriver?,”
[Link], 2022. [Link]
sering-pakai-gojek-bagaimana-grab-maxim-dan-indriver
[6] P. A. Permatasari, L. Linawati, and L. Jasa, “Survei Tentang Analisis Sentimen Pada Media Sosial,” Maj. Ilm. Teknol.
Elektro, vol. 20, no. 2, p. 177, 2021, doi: 10.24843/mite.2021.v20i02.p01.
[7] S. W. Iriananda et al., “ANALISIS SENTIMEN DAN ANALISIS DATA EKSPLORATIF ULASAN,” no. Ciastech, pp.
473–482, 2021.
[8] A. Saepulrohman, S. Saepudin, and D. Gustian, “Analisis Sentimen Kepuasan Pengguna Aplikasi Whatsapp Menggunakan
Algoritma Naïve Bayes Dan Support Vector Machine,” is Best Account. Inf. Syst. Inf. Technol. Bus. Enterp. this is link
OJS usf@, vol. 6, no. 2, pp. 91–105, 2021, doi: 10.34010/aisthebest.v6i2.4919.
[9] K. A. Rokhman, B. Berlilana, and P. Arsi, “Perbandingan Metode Support Vector Machine Dan Decision Tree Untuk Analisis
Sentimen Review Komentar Pada Aplikasi Transportasi Online,” J. Inf. Syst. Manag., vol. 3, no. 1, pp. 1–7, 2021, doi:
10.24076/joism.2021v3i1.341.
[10] A. M. Rahat, A. Kahir, and A. K. M. Masum, “Comparison of Naive Bayes and SVM Algorithm based on Sentiment Analysis
Using Review Dataset,” Proc. 2019 8th Int. Conf. Syst. Model. Adv. Res. Trends, SMART 2019, pp. 266–270, 2020, doi:
10.1109/SMART46866.2019.9117512.
[11] N. Herlinawati, Y. Yuliani, S. Faizah, W. Gata, and S. Samudi, “Analisis Sentimen Zoom Cloud Meetings di Play Store
Menggunakan Naïve Bayes dan Support Vector Machine,” CESS (Journal Comput. Eng. Syst. Sci., vol. 5, no. 2, p. 293,
2020, doi: 10.24114/cess.v5i2.18186.
[12] S. A. Salloum, M. Al-emran, and A. A. Monem, “Using Text Mining Techniques for Extracting Information from Research
Using Text Mining Techniques for Extracting Information from Research Articles,” no. January, 2018, doi: 10.1007/978-3-
319-67056-0.
[13] S. Bhatia, M. Sharma, and K. K. Bhatia, “Sentiment Analysis and Mining of Opinions,” Stud. Big Data, vol. 30, no. May,
pp. 503–523, 2018, doi: 10.1007/978-3-319-60435-0_20.
[14] J. Han, J. Pei, and H. Tong, Data Mining : Concepts and Techniques. Morgan Kaufmann, 2022.
[15] K. X. Han, W. Chien, C. C. Chiu, and Y. T. Cheng, “Application of support vector machine (SVM) in the sentiment analysis
of twitter dataset,” Appl. Sci., vol. 10, no. 3, 2020, doi: 10.3390/app10031125.
[16] P. S. Saragih, D. Witarsyah, F. Hamami, and J. M. MacHado, “Sentiment Analysis of Social Media Twitter with Case of
Large Scale Social Restriction in Jakarta using Support Vector Machine Algorithm,” 2021 Int. Conf. Adv. Data Sci. E-
Learning Inf. Syst. ICADEIS 2021, vol. 19, no. January 2020, pp. 1–6, 2021, doi: 10.1109/ICADEIS52521.2021.9701961.
[17] S. S. Chaeikar, A. A. Manaf, A. A. Alarood, and M. Zamani, “PFW: Polygonal fuzzy weighted—an SVM kernel for the
classification of overlapping data groups,” Electron., vol. 9, no. 4, 2020, doi: 10.3390/electronics9040615.
[18] A. Zaiem and N. Charibaldi, “Komparasi Fungsi Kernel Metode Support Vector Machine untuk Analisis Sentimen Instagram
dan Twitter ( Studi Kasus : Komisi Pemberantasan Korupsi ),” vol. 9, no. 2, pp. 33–42, 2021.
[19] A. Pranata, E. Budianita, Yusra, and E. P. Cynthia, “Klasifikasi Sentimen Terhadap Maxim Menggunakan Algoritma SVM
Pada Media Sosial TwittAnggi Pranata, N. (2022). Klasifikasi Sentimen Terhadap Maxim Menggunakan Algoritma SVM
Pada Media Sosial Twitter. Klasifikasi Sentimen Terhadap Maxim Menggunakan Algorit,” Klasifikasi Sentimen Terhadap
Maxim Menggunakan Algoritm. SVM Pada Media Sos. Twitter, vol. 5, no. 3, pp. 332–341, 2022.
[20] W. Musu, A. Ibrahim, and Heriadi, “Pengaruh Komposisi Data Training dan Testing terhadap Akurasi Algoritma C4 . 5,”
Pros. Semin. Ilm. Sist. Inf. Dan Teknol. Inf., vol. X, no. 1, pp. 186–195, 2021.