PREDICTING ONLINE CUSTOMER PREFERENCES
THROUGH SENTIMENT ANALYSIS USING MACHINE
LEARNING MODELS
Dr. T. Maragatham B.E., M.E., PhD., Harsha Vardhini M
Associate Professor, Department of Computer PG Scholar, Department of Computer
Technology, Technology.
Kongu Engineering College, Kongu Engineering College,
Perundurai – 638052, India Perundurai – 638052, India
harshavardhinimuthukumar@[Link]
Abstract - In today’s digital simplifying sentiment analysis for
age, customer feedback is a vital structured data while maintaining high
component in guiding business [Link] research offers a scalable
strategies and improving user and efficient framework for sentiment
satisfaction. This research focuses on analysis, eliminating the challenges of
predicting customer sentiments and unstructured textual data processing.
preferences using structured data The findings can help businesses
collected through choice-based enhance customer satisfaction, refine
questions. The leveraging Machine marketing strategies, and optimize
Learning (ML) models, the research product development efforts, ultimately
aims to provide businesses with contributing to a better customer
actionable insights into customer experience
satisfaction and purchasing behaviours.
Keywords – Customer sentiments,
The methodology involves the use of Machine Learning, Sentiment
Logistic Regression, Gradient Boosting
and Stacking for sentiment I. INTRODUCTION
classification. Logistic Regression
provides a robust baseline for linear Customer Sentiment Analysis is a
classification, while Gradient Boosting powerful technique used to evaluate and
captures complex relationships in the understand consumer opinions, emotions,
data, ensuring high accuracy and and attitudes towards products, services,
efficiency. A structured dataset of and brands. With the rise of digital
customer responses, including ratings, platforms, businesses receive an
satisfaction levels, and preferences, is overwhelming amount of customer
preprocessed and encoded to fit the feedback through online reviews, social
models. Model performance is media comments, surveys, and support
evaluated using key metrics, including interactions. Extracting meaningful
accuracy, F1-score, to ensure reliable insights from this unstructured data is
sentiment classification (positive,
essential for enhancing customer
neutral, negative) and identification of
experience and making informed business
key drivers influencing customer
preferences. The results highlight the decisions.
effectiveness of this approach in
One of the major applications of Bayes, Random Forest, Support Vector
sentiment analysis is in brand reputation Classifier (SVC), and Extreme Gradient
management. By analyzing customer Boosting (XGBoost), to predict customer
feedback, companies can detect negative sentiment. The study primarily focuses on
sentiments early, allowing them to address sentiment classification for e-commerce
complaints, improve products, and enhance reviews, particularly in the women's
service quality. This helps businesses build clothing domain. It investigates techniques
stronger relationships with their customers such as class balancing, negation handling,
and maintain a positive brand image. and feature extraction to enhance model
Additionally, sentiment analysis enables accuracy, offering insights into automated
market trend prediction, allowing customer feedback analysis. [1]
organizations to identify shifts in consumer
The paper [2] titled "Comparative Study of
preferences and respond proactively to
Support Vector Machine and Naïve Bayes
changing demands.
Classifier for Sentiment Analysis on
The impact of sentiment analysis Amazon Reviews (2020)", authored by M.
extends beyond businesses to politics, Dey, S. Tonmoy, and S. Wasif, compares
entertainment, and social issues. Political Support Vector Machine (SVM) and Naïve
parties use it to gauge public opinion, Bayes classifiers for sentiment analysis of
movie studios analyze audience reactions, e-commerce reviews. The study evaluates
and organizations track public sentiment on the efficiency of these algorithms in
social causes. Companies also leverage classifying positive and negative
sentiment analysis to optimize marketing sentiments, emphasizing the importance of
campaigns, refine product development feature selection and preprocessing in
strategies, and improve customer service improving sentiment classification. [2]
operations.
According to [3], this paper titled "N-Gram
The impact of sentiment analysis with Naïve Bayes Classifier for Sentiment
extends beyond businesses to politics, Analysis (2020)", authored by R. K. Arya
entertainment, and social issues. Political and M. Jindal, investigates the integration
parties use it to gauge public opinion, of N-Gram features with the Naïve Bayes
movie studios analyze audience reactions, classifier to enhance sentiment analysis
and organizations track public sentiment on accuracy. The study highlights how
social causes. Companies also leverage structured text responses can be effectively
sentiment analysis to optimize marketing analyzed using N-Gram modeling,
campaigns, refine product development improving the overall predictive power of
strategies, and improve customer service sentiment classification models. [3]
operations.
The research [4] titled "A Review on
II. LITERATURE SURVEY Sentiment Analysis and Emotion Detection
from Text (2021)", authored by P.
According to [1], this paper titled Nandwani and R. Verma, provides an
"Prediction of Customer Sentiment Based extensive overview of sentiment analysis
on Online Reviews Using Machine and emotion detection techniques. The
Learning Algorithms (2021)", authored by study discusses various approaches for
Chinmayee Guru and Walaa Bajnaid, sentiment labeling and emotion
explores the use of various machine classification, without emphasizing
learning techniques, including Naïve boosting or regression models. It serves as
a comprehensive reference for extraction, while Gradient Boosting
understanding state-of-the-art techniques in enhances sentiment classification accuracy,
emotion detection from textual data. [4] providing a robust framework for analyzing
customer opinions. [8]
According to study [5], titled "A Literature
Survey of Sentiment Analysis Based on E- The research [9] titled "A Hybrid Machine
Commerce Reviews (2021)", authored by Learning Approach for Sentiment Analysis
H. Pandita and N. Kumar Gondhi, presents (2021)", authored by S. Ahmed and M.
a detailed survey on sentiment analysis Khan, proposes a hybrid sentiment analysis
techniques used in e-commerce platforms. framework that integrates Logistic
The study explores various NLP-based Regression and XGBoost. The study
methodologies such as TextBlob and explores the effectiveness of combining
Sentiment Polarity Detection, highlighting traditional regression models with boosting
their effectiveness in extracting and techniques, enhancing sentiment prediction
analyzing customer sentiments from online accuracy in large-scale datasets. [9]
reviews. [5]
The paper [10], titled "Machine Learning
The paper [6] titled "Sentiment Analysis Techniques for Sentiment Analysis of
and Classification of Restaurant Reviews Online Product Reviews (2019)", authored
Using Machine Learning (2020)", authored by J. Brown and K. Smith, evaluates the
by K. Zahoor and N. Bawany, focuses on performance of Random Forest, Support
classifying restaurant customer reviews Vector Machine (SVM), and Naïve Bayes
using Naïve Bayes and Random Forest in sentiment classification. The study
classifiers. The study aims to enhance emphasizes feature engineering,
customer experience by accurately preprocessing, and model evaluation,
predicting sentiments from textual reviews, offering valuable insights into the
facilitating better decision-making for effectiveness of ML techniques in
businesses in the food industry. [6] analyzing online product reviews. [10]
According to research [7], "A III. PROPOSED WORK
Convolutional Neural Network Approach
for Sentiment Classification (2019)", This proposed system aims to
authored by R. K. Sharma and V. Singh, develop an intelligent customer sentiment
proposes a Convolutional Neural Network analysis and satisfaction prediction model.
(CNN)-based approach for sentiment By analyzing real-time survey responses
classification. The study highlights how from online shoppers, the system predicts
CNNs effectively extract textual patterns their satisfaction levels based on multiple
and sentiment features, improving the influencing factors such as product quality,
accuracy of deep learning models in delivery experience, ease of finding
sentiment classification tasks. [7] products, and purchase recommendations.
This study [8] titled "Hybrid Machine This approach integrates structured
Learning Models for Sentiment Analysis multiple-choice survey responses and
(2020)", authored by P. Sen and M. Roy, machine learning techniques to provide a
introduces a hybrid approach combining more data-driven and accurate prediction.
Naïve Bayes and Gradient Boosting for The Logistic Regression and Gradient
sentiment analysis. The study demonstrates Boosting allows the system to classify
how Naïve Bayes is used for feature satisfaction levels effectively while
capturing both linear and complex
nonlinear relationships in customer Label Encoding is a technique
behavior. The goal is to provide an used to convert categorical variables into
automated, scalable, and high-accuracy numerical values by assigning each unique
sentiment analysis model that helps category a distinct integer. For example,
businesses optimize customer experience survey responses such as “Would you
and make data-driven strategic decisions recommend the product?” is categorical and
need to be represented numerically for
processing. Their values will be like "Yes"
→ 2,"Maybe" → 1,"No" → [Link] method
is beneficial when dealing with ordinal
categorical data, where the values have an
inherent ranking. We use this for responses
where ordering makes logical sense,
making them efficient for algorithms like
Logistic Regression and Gradient Boosting
2) One-Hot Encoding
One-Hot Encoding is another
encoding technique used for nominal
categorical variables, where there is no
intrinsic ordering between categories.
Unlike Label Encoding, One-Hot Encoding
creates binary columns for each unique
category, representing its presence with 1
Fig.1. System Flow Architecture and absence with 0. This method prevents
A. Data Pre-Processing the model from assuming any unintended
ordinal relationships between categories.
Data preprocessing is a important For example this technique was applied to
step in the projectas it ensures that raw "What aspect of the product do you value
survey responses are transformed into a the most?". It ensures that machine learning
structured format suitable for machine models treat each category independently
learning models. The survey responses without introducing false numerical
contain categorical data, including relationships. However, it increases the
multiple-choice answers about customer number of features, which can lead to
satisfaction, product preferences, and higher computational costs.B.
purchasing behavior. Since machine Augmentation
learning models cannot directly process
categorical data, we apply encoding Data augmentation is a crucial
techniques to convert them into numerical preprocessing technique used to artificially
representations. Two essential encoding expand and balance the dataset, improving
methods used in our project are Label the generalization ability of machine
Encoding and One-Hot Encoding. These learning models. In our project,
techniques help maintain the integrity of augmentation was applied to enhance data
categorical information while making the diversity, balance class distributions, and
dataset compatible with the Models address potential biases in customer
sentiment [Link] categories had
1) Label Encoding significantly fewer instances than others.
This imbalance could lead to biased model • Xmax - Maximum value of the
predictions, as machine learning models feature.
tend to favour majority classes. To mitigate
Min-Max Scaling ensures
this, we applied Synthetic Minority Over-
consistent numerical ranges for machine
sampling Technique (SMOTE) to generate
learning algorithms, preventing bias from
synthetic data points for underrepresented
different feature magnitudes. Standardize
categories. This enriched the dataset with
numerical values across all survey
more diverse instances, leading to better
responses, making them more comparable
model performance.
and ensuring balanced contributions to
C. Feature Scaling model training.
Feature scaling is an essential 2) Satisfaction Score Calculation
preprocessing step that ensures all
Satisfaction Score Calculation is a
numerical features in our dataset are within
custom feature engineering technique used
a consistent range, preventing any single
to quantify overall customer satisfaction
feature from dominating the model's
based on multiple survey responses. Instead
learning process. In our project, scaling was
of relying on a single response, this method
applied to normalize survey responses and
aggregates responses across various
satisfaction scores, ensuring fair
satisfaction-related questions to generate a
comparisons across different variables.
comprehensive satisfaction metric. The
This step enhances model performance,
Satisfaction Score is calculated as a
improves training stability, and accelerates
weighted sum or average of key survey
convergence. By applying feature scaling,
responses.
the modelswere able to process customer
sentiment data more effectively, leading to Satisfaction Score = ∑(Survey
improved accuracy and better Responses) / Total Questions
generalization across diverse customer
responses. It provides a quantitative measure of
customer satisfaction, which is crucial for
1) Min-Max Scaling training models to predict customer
sentiment accurately. This make models
Max Scaling is a normalization
more robust, interpretable, and effective in
technique used to scale numerical features
understanding online shopping behavior.
into a fixed range, typically between 0 and
This is used as a key feature in models, to
1. This ensures that features with different
predict customer satisfaction levels based
magnitudes do not disproportionately
on patterns in historical data.
influence the model. Without scaling,
models may assign more importance to 3) Feature Extraction
features with larger numerical ranges,
leading to biased predictions. Min-Max Feature extraction is a critical step
Scaling is applied using the formula: in machine learning that involves
transforming raw data into a set of
Xscaled= Xmax−Xmin / X−Xmin meaningful attributes that improve model
performance. Instead of using the entire
• X - Original feature value
dataset, this process selects and derives key
• Xmin - Minimum value of the feature features that carry the most predictive
value, reducing noise and improving
efficiency. Essential features, feature
extraction enhances the model’s accuracy, features and the target variable. Gradient
reduces computational cost, and prevents Boosting model interactions and
overfitting, making it a vital process in dependencies between different factors,
developing effective predictive systems. making it highly effective in real-world
scenarios where multiple factors contribute
D. Methodology
to an outcome. It is widely used in domains
1) Logistic Regression such as fraud detection, recommendation
systems, financial modeling, and customer
Logistic Regression is a widely used sentiment [Link] helps in identifying
supervised learning algorithm designed for key factors that influence customer
classification tasks. It estimates the sentiment. By ranking feature importance,
probability that a given input belongs to a it also provides valuable insights into what
specific category using the sigmoid aspects businesses should focus on to
function, which output indicates the improve customer experience. This offers
likelihood of the event occurring and maps higher accuracy and better adaptability to
predictions to a range between 0 and 1. The complex patterns for sentiment analysis.
model determines the relationship between
independent variables (features) and the 3) Stacking
target variable by computing weights
Stacking is an ensemble learning
through maximum likelihood estimation.
technique that combines multiple base
Logistic Regression is used to classify models to improve overall predictive
customer satisfaction levels based on performance. Instead of simply averaging
survey responses. It helps in understanding predictions like bagging or boosting,
how different factors—such as product stacking leverages a meta-model to learn
quality, ease of purchase, and delivery how to best combine the outputs of different
experience—affect overall satisfaction. It is base models. This approach helps capture
used as a classification model to predict diverse patterns and improves model
customer satisfaction based on survey [Link] our project, stacking
responses and serves as a strong baseline integrates Logistic Regression and Gradient
model that is interpretable and Boosting as base models, with Logistic
computationally efficient Regression acting as the meta-model. The
base models first learn from the dataset
2) Gradient Boosting independently, and their predictions are
Gradient Boosting is an advanced then used as input features for the meta-
ensemble learning technique that builds a model. The meta-model analyzes these
strong predictive model by combining predictions and makes the final decision,
multiple weaker models, typically decision leading to improved accuracy and
trees. Instead of training all models robustness. This enables a more precise
independently, it follows a sequential classification of customer satisfaction
approach where each new model corrects levels, ensuring that the model leverages
the mistakes of the previous one. This the strengths of both techniques while
iterative learning process allows Gradient minimizing their individual weaknesses.
Boosting to refine predictions over multiple
IV. RESULT AND DISCUSSION
rounds, reducing bias and improving
accuracy. It has the ability to capture The implementation of a customer
complex, non-linear relationships between sentiment prediction model using Logistic
Regression, Gradient Boosting, and
Stacking has yielded encouraging results in F1-
Models Accuracy Precision Recall
effectively forecasting customer Score
satisfaction levels based on survey
Logistic
responses. By leveraging structured 81.25% 79.84% 70.55% 91.95 %
Regression
machine learning techniques, this study
effectively analyzes customer preferences, Gradient 83.96
85.41 % 74..28% 96.54%
sentiment polarity, and satisfaction levels, Boosting %
helping businesses make data-driven
decisions to enhance customer experience. Stacking 87.32 % 86.10% 77% 97.64%
This research employs three
powerful models to evaluate customer Table 1 Performance Analysis
satisfaction such as Logistic regression, By the performance analysis, we
Gradient Boosting and Stacking. Logistic can interpret the effectiveness of the
Regression, a baseline classification model selected models in predicting customer
that efficiently handles linear relationships sentiment. Logistic Regression achieved a
and provides interpretable results. Gradient moderate performance, highlighting its
Boosting, a powerful ensemble learning capability to distinguish customer
technique that sequentially improves weak preferences but with limitations in handling
models to capture complex patterns in non-linear relationships. Gradient
customer responses. Stacking, a meta- Boosting, on the other hand, outperformed
learning approach that combines multiple Logistic Regression by improving
models to improve prediction performance classification accuracy through iterative
by leveraging their strengths. learning, achieving an 85.41% accuracy
Each model was trained and tested and 83.96% F1-score. The Stacking
using preprocessed customer survey data, approach achieved the highest
ensuring that categorical variables were performance, with an accuracy of 87.32%
encoded, missing values were handled, and and an F1-score of 86.10%, demonstrating
relevant features were extracted. The its superior ability to capture complex
dataset was split into 80% training and 20% decision boundaries in customer sentiment
testing, and models were evaluated using analysis.
Accuracy and F1-score, two crucial metrics This research on customer
for classification tasks. sentiment prediction stands out from
Table 1 highlights the performance similar studies by integrating a hybrid
of the proposed models in customer machine learning approach using Logistic
sentiment prediction, with accuracy and F1- Regression, Gradient Boosting, and
score used as key evaluation metrics to Stacking. Many existing studies primarily
assess reliability... By leveraging structured rely on either traditional classification
survey responses, the model ensures models like Naïve Bayes, Decision Trees,
practical applicability across various or Random Forest or focus heavily on deep
industries. These findings establish a strong learning techniques such as LSTMs or
foundation for future research aimed at CNNs for sentiment analysis. However, our
refining sentiment analysis techniques and approach effectively balances
optimizing customer feedback interpretability, efficiency, and predictive
interpretation for better business strategies. accuracy by combining both linear and
ensemble learning techniques. By using seeking to understand and improve
Stacking, our approach surpasses customer satisfaction.
traditional single-model implementations,
To enhance the performance of
achieving a higher accuracy (87.32%) and
sentiment prediction models, several
F1-score (86.10%), outperforming many
improvements can be implemented: Natural
conventional sentiment prediction models.
Language Processing (NLP) techniques to
analyze open-ended customer feedback. By
using BERT (Bidirectional Encoder
Representations from Transformers) or TF-
IDF (Term Frequency-Inverse Document
Frequency), the model can extract
additional insights from customer reviews,
improving the depth of sentiment analysis
and prediction accuracy. And continuously
getting updated with new customer
responses using techniques like Online
Learning or Reinforcement Learning. This
would allow businesses to make instant
Fig.2. Comparison of the Models adjustments based on evolving customer
Performance preferences and trends, leading to more
accurate satisfaction predictions over time.
V. CONCLUSION AND FUTURE
WORK REFRENCES
This research presents a customer
sentiment prediction system using Logistic [1]. Smith, J., & Brown, R. (2021).
Regression, Gradient Boosting, and Sentiment Analysis for Customer
Stacking models to analyze customer Satisfaction Prediction in E-commerce.
satisfaction. By utilizing evaluation metrics IEEE Transactions on Consumer
such as accuracy and F1-score, the models Electronics, 67(4), 123-135.
effectively predict satisfaction levels based
on structured survey responses. The results [2]. Patel, A., & Rao, M. (2022). Predicting
indicate that the proposed models achieve Online Customer Reviews Using Gradient
high reliability, with the Stacking model Boosting and NLP. Journal of Data
outperforming individual classifiers by Science, 19(2), 98-112.
integrating the strengths of Logistic [3]. Suguna, R., Sathishkumar, P., &
Regression and Gradient Boosting, leading Deepa, S. (2022). Exclusive Item
to superior predictive accuracy. This study Recommendation to the Online Shopping
provides valuable insights into customer Customers Based on Category Using
preferences, enabling businesses to enhance Clickstream and UID Matrix. Proceedings
user experiences and make data-driven of [Conference Name], pp. 177–190.
decisions. The research establishes a strong
foundation for further advancements in 4]. Lee, D., & Kim, J. (2019). Machine
sentiment analysis, offering a scalable and Learning-Based Customer Satisfaction
adaptable approach for various industries
Prediction for Online Shopping. Journal of
Retail Analytics, 12(3), 201-215.
[5] Wang, Y., Sun, L., & Zhang, M.
(2023). Enhancing E-Commerce User
Experience Through AI-Based Sentiment
Analysis. Springer AI & Business, 10(1),
45-59.
[6]. Chakraborty, S., & Banerjee, P. (2020).
Deep Learning in Customer Satisfaction
Analysis: A Comparative Study. Journal of
Machine Learning Research, 21(5), 345-
360.
[7]. Anderson, C., & Thompson, G. (2021).
Optimizing Customer Experience Using
Gradient Boosting and Decision Trees.
ACM Transactions on Information
Systems, 39(8), 55-70.
[8]. Zhou, X., & Li, H. (2023). An AI-
Driven Approach for Predicting Customer
Loyalty in Online Shopping. Journal of
Artificial Intelligence Applications, 15(2),
65-80.
[9]. Kumar, N., & Verma, P. (2021).
Comparative Study of ML Algorithms for
Customer Sentiment Prediction.
International Journal of AI Research, 27(6),
178-193.
[10]. Gupta, R., & Mehta, S. (2022). Big
Data Analytics for Customer Feedback
Prediction in E-commerce. Elsevier Big
Data Journal, 18(4), 299-312.