AI-Driven User Churn Prediction System
AI-Driven User Churn Prediction System
INTRODUCTION
1.1 Overview
learning (ML)–based methods capable of learning behavioural patterns from large user
datasets to predict churn risk. Parallel advances in Natural Language Processing (NLP)
have enabled sentiment extraction from user feedback, reviews, and social-media
discussions, allowing organizations to capture the emotional drivers behind customer
decisions. However, most ML systems remain black-box models - accurate but opaque -
leaving decision-makers uncertain about why particular users are predicted to churn.
To address this limitation, Explainable Artificial Intelligence (XAI) frameworks
such as SHAP and LIME have emerged. These techniques apply principles from
cooperative game theory to highlight the relative contribution of each feature to a
prediction, thereby restoring transparency and trust in AI systems. Recent academic
work, including hybrid churn-prediction architectures combining behavioural analytics
with sentiment analysis, has shown improved accuracy and interpretability.
Building on these developments, the Retention-AI project integrates ML, NLP,
and XAI into a unified platform that predicts churn, interprets its causes, and
recommends personalized retention strategies. By leveraging both structured behavioural
data and unstructured textual feedback, the system not only forecasts potential user drop-
offs but also provides clear, actionable insights for businesses to enhance engagement
and loyalty.
1.2 Objectives
1.3.1 Purpose
The primary purpose of this project is to help businesses understand and reduce
user churn by leveraging the capabilities of Artificial Intelligence. Many organizations
collect large volumes of user data but struggle to interpret it effectively to identify users
at risk of leaving. Retention-AI aims to bridge this gap by providing an intelligent,
interpretable, and proactive churn prediction system.
1.3.2 Scope
The scope of Retention-AI extends across the entire churn prediction and
prevention pipeline. It involves collecting and preprocessing both structured data (user
behavior logs, demographics, usage metrics) and unstructured data (user feedback,
reviews).
Two models are used - Model A (XGBoost) for structured data and Model B (Random
Forest with VADER) for unstructured data. After preprocessing and balancing with
SMOTE, churn probabilities from both models are combined to produce a final churn
score.
While the project effectively predicts and explains churn, its scope is limited by
data availability, the quality of unstructured feedback, and computational resources.
1.3.3 Applicability
The report is structured into nine chapters. Chapter 1 introduces the project along
with its background, objectives, and scope. Chapter 2 reviews related research work and
existing churn prediction systems. Chapter 3 outlines the functional and non-functional
requirements, while Chapter 4 discusses project planning and system architecture.
Chapter 5 presents the module decomposition, interface design, data structures, and
algorithm design. Chapter 6 focuses on the implementation process, and Chapter 7
explains the testing approaches used for system validation. Chapter 8 presents the results,
performance analysis, and user documentation, and Chapter 9 concludes the report with
applications, limitations, and future scope of the project.
2.1 Introduction
In any research-oriented project, reviewing existing literature is crucial for
understanding the current state of the field and identifying areas that require
improvement. For this project, the focus lies on user churn prediction and customer
retention, where businesses strive to anticipate and prevent user drop-offs using Artificial
Intelligence (AI) techniques. This chapter presents a detailed review of previous studies,
methodologies, and technologies applied in churn prediction, sentiment analysis, and
explainable AI (XAI). The goal of this literature survey is to examine the strengths and
limitations of existing systems, understand how recent advancements have improved
prediction accuracy and interpretability, and identify research gaps that led to the
development of the Retention-AI framework. By integrating findings from past studies,
this chapter provides the theoretical foundation for designing a system that combines
ML, NLP, and XAI to deliver transparent, data-driven churn insights and personalized
retention strategies.
Key Findings: The model effectively identifies at-risk users and highlights
behavioural factors contributing to churn, allowing targeted retention strategies. The
use of explainable AI improves transparency and trust in model outputs.
[2] Upma Singh, Sandeep Singh, Tripti Rathee, and Manav Vaish (2024), “Customer
Retention Modelling over the OTT Platform using Machine Learning,” Indian
Journal of Science and Technology, 2024.
duration, and content preferences. These models are trained to identify factors
influencing user retention and predict potential churners.
Key Findings: The models effectively predict customer churn and retention rates,
enabling OTT platforms to design targeted promotional strategies and improve user
satisfaction.
[3] Shuzlina Abdul-Rahman, Muhamad Faidi Akif Md Ali, Azuraliza Abu Bakar,
and Sofianita Mutalib (2024), “Enhancing Churn Forecasting with Sentiment
Analysis of Steam Reviews,” Social Network Analysis and Mining, Springer,
2024.
Key Findings: The integration of sentiment data improves the precision of churn
forecasts and provides deeper behavioural insights into player engagement and
dissatisfaction.
Limitations: The model focuses only on sentiment polarity, lacking emotion, intent,
and contextual understanding. Additionally, Steam reviews may not accurately
represent the overall player experience, limiting the generalizability of results.
Methodology: The study employs a fuzzy logic approach integrated with explainable
AI principles to analyze broadband usage patterns, network performance indicators,
and customer service interactions. The model generates churn predictions while
maintaining transparency through rule-based explanations.
[5] Kapil Arora, M. Lalitha, Poonam, Hemalatha Yadav J., and Biswo Ranjan
Mishra (2024), “Machine Learning for Customer Retention in E-Commerce
Healthcare Startups,” South Eastern European Journal of Public Health
(SEEJPH), TensorGate, 2024.
Objective: To apply machine learning techniques for predicting customer churn and
enhancing user retention in healthcare-focused e-commerce startups.
Key Findings: The model successfully improves retention rates by identifying high-
risk customers and providing data-driven recommendations for personalized
engagement strategies.
[6] Tamer Ahmed Ibrahim Abou El-Fotouh and Musliudeen Toheeb Akanbi (2024),
“Impact of Predictive Analytics and Machine Learning on Customer Retention and
Loyalty in Service-Oriented Businesses,” International Journal of Business
Intelligence and Big Data Analytics, 2024.
Methodology: The study applies decision tree classifiers and ensemble learning
methods to analyze customer transaction data and behavioural patterns. These
techniques are used to identify churn indicators and predict customer loyalty levels
across different service sectors.
[7] Ziru Liu, Shuchang Liu, Bin Yang, Zhenghai Xue, Qingpeng Cai, Xiangyu Zhao,
Zijian Zhang, Lantao Hu, Han Li, and Peng Jiang (2024), “Modelling User
Retention through Generative Flow Networks,” Proceedings of the ACM
SIGKDD Conference on Knowledge Discovery and Data Mining (KDD), 2024.
Limitations: Despite strong predictive performance, the model suffers from limited
interpretability, making it difficult for recommender systems to understand why users
churn. This lack of transparency restricts its applicability for actionable business
decision-making.
[8] Kapil Arora, M. Lalitha, Poonam, Hemalatha Yadav J., and Dr. Biswo Ranjan
Mishra (2024), “Machine Learning for Customer Retention in E-Commerce
Healthcare Startups,” South Eastern European Journal of Public Health
(SEEJPH), 2024.
[9] Suchita Sharma and Nishith Desai (2023), “Identifying Customer Churn Patterns
Using Machine Learning Predictive Analysis,” IEEE SMART GENCON, 2023.
Key Findings: The research demonstrates that machine learning can accurately
predict customer churn, providing businesses with early insights to design data-
driven retention strategies.
[10] Sabreen Abulhaija, Shyma Hattab, Ahmad Abdeen, and Wael Etaiwi (2022),
“Predicting Mobile Apps Performance using Machine Learning,” Journal of
System and Management Sciences, 2022.
Methodology: The study analyzes data from the Google Play Store, applying
algorithms such as Random Forest to predict app performance based on factors
including user ratings, downloads, and update frequency. The research focuses on
improving prediction generalizability through extensive feature extraction and dataset
expansion.
Key Findings: The model achieves strong performance in predicting app success,
demonstrating that integrating multiple app-related features enhances prediction
accuracy and robustness.
Limitations: Although the study uses a significantly large dataset of 2.24 million
records, it remains limited by feature diversity and a lack of qualitative insights such
as user sentiment or contextual engagement data.
[11] Manzura Jorayeva, Akhan Akbulut, Cagatay Catal, and Alok Mishra (2022),
“Machine Learning-Based Software Defect Prediction for Mobile Applications,”
Sensors, 2022.
Limitations: The study highlights a lack of unified frameworks for evaluating fault
prediction models across mobile platforms. It also reports issues such as data
imbalance, high computational cost, and the absence of real-world validation,
limiting the practical applicability of these models in large-scale mobile app
development.
[12] Carson K. Leung, Adam G. M. Pazdor, and Joglas Souza (2021), “Explainable
Artificial Intelligence for Data Science on Customer Churn,” IEEE International
Conference on Data Science and Advanced Analytics (DSAA), 2021.
[13] Zeynep Hilal Kilimci, Hasan Yörük, and Selim Akyokus (2020), “Sentiment
Analysis Based Churn Prediction in Mobile Games using Word Embedding Models
and Deep Learning Algorithms,” IEEE International Conference on Machine
Learning,2022.
Objective: To predict mobile game churn by applying deep learning models using
sentiment extracted from user reviews represented through advanced word
embedding models.
Methodology: The study employs CNN, LSTM, and RNN architectures to process
user reviews, utilizing Word2Vec, GloVe, and FastText embeddings to represent
semantic relationships between words. Sentiment scores are derived from these
Key Findings: The deep learning models outperform traditional churn prediction
approaches based solely on behavioral or demographic data. Integrating sentiment
representations significantly improves accuracy, emphasizing the emotional tone.
Limitations: The model focuses only on textual data, excluding behavioral and
demographic factors that could improve prediction reliability. Moreover, its
application is restricted to the gaming domain, limiting scalability to other industries.
Key Findings: Random Forest and Decision Tree models demonstrate high
predictive accuracy, revealing that usage frequency, contract length, and service
satisfaction are key churn determinants. The study highlights the effectiveness of
traditional ML models for churn prediction in SaaS contexts.
[15] Carl Yang, Xiaolin Shi, Jie Luo, and Jiawei Han (2019), “I Know You’ll Be
Back: Interpretable New User Clustering and Churn Prediction on a Mobile Social
Application,”2019.
Objective: To develop an interpretable churn prediction framework that identifies
churn-prone users in mobile social applications by analyzing behavioral patterns of
new users.
Methodology: The study proposes ClusChurn, a two-step framework combining
LSTM networks and attention mechanisms with k-means clustering to group new
users based on engagement behavior. It is applied on large-scale Snapchat user data
to predict churn and interpret user retention patterns.
Key Findings: The model effectively captures sequential user behavior and improves
churn prediction accuracy for new users. The clustering component enhances
interpretability by revealing behavioral segments associated with different churn
risks.
Limitations: Despite offering improved interpretability compared to traditional
models, ClusChurn still lacks comprehensive explainability and actionable insights
for targeted retention strategies. Its reliance on a single platform dataset also limits
generalizability to other applications.
Despite significant advancements in the fields of customer churn prediction and retention
analytics, existing systems still face several limitations that restrict their effectiveness,
scalability, and practical adoption.
4. Data Imbalance and Limited Real-Time Learning: Many systems suffer from
class imbalance, where churn instances are significantly fewer than non-churn
instances. This leads to biased predictions and lower sensitivity toward churn-
prone users. Additionally, most models are static and do not retrain dynamically
with new incoming data, making them unsuitable for real-time applications.
The proposed system, Retention-AI, aims to bridge the gaps observed in existing
churn prediction frameworks by integrating Machine Learning (ML), Natural Language
Processing (NLP), and Explainable AI (XAI) into a unified and interpretable customer
retention platform. The system focuses on providing both predictive accuracy and
actionable insights, enabling organizations to not only identify potential churners but also
understand the reasons behind churn and take proactive measures. FIN-NET employs
Artificial Neural Network (Multi layered Perceptron) for capturing complex patterns and
interdependencies.
2. Both models generate independent churn predictions, which are later combined to
produce a final aggregated churn score through averaging.
3. Balanced and Preprocessed Data Pipeline: The system uses data preprocessing
techniques to clean and normalize inputs. To address the class imbalance problem,
it applies SMOTE (Synthetic Minority Oversampling Technique) to balance churn
and non-churn samples, ensuring fair model training and better prediction
reliability.
6. Interactive Dashboard and Joblib Integration: The use of the Joblib library
ensures that while new computations are being processed, the dashboard remains
responsive and displays cached results to avoid user downtime.
7. Personalized Retention Strategies: Based on the churn score and SHAP
explanations, the system recommends personalized retention actions such as
targeted discounts, promotional offers, or re-engagement messages. These insights
are sent to the business owners, allowing them to make data-driven decisions to
retain high-risk customers.
REQUIREMENT ENGINEERING
Requirement Engineering is a crucial phase in the software development life cycle
that defines what the system should accomplish and how it should perform under
specified conditions. This chapter focuses on identifying and analyzing user needs,
system expectations, and transforming them into clearly defined functional and non-
functional requirements for the proposed system, Retention-AI. The objective of this
phase is to ensure that the churn prediction and retention system is technically feasible,
user-centric, and aligned with real-world business needs related to customer engagement
and retention. It also details the software and hardware tools used in the development
process, the conceptual and analytical modeling, and various UML diagrams such as use
case, sequence, activity, and state chart diagrams. This chapter forms the foundation for
the subsequent system design and implementation phases, ensuring that every functional
requirement is systematically aligned.
1. Processor
A mid-range processor is required to ensure smooth execution of machine
learning and natural language processing tasks. It supports efficient data
preprocessing, model training, and SHAP-based interpretability computations.
The processor also maintains responsiveness while handling simultaneous
operations such as structured and unstructured data analysis.
2. RAM
Adequate memory is essential for training ML models, processing large datasets,
Python libraries and enable seamless data preprocessing and model input
handling.
5. Version Control System: is used for local version tracking, while GitHub serves
as the remote repository for collaborative development, version management, and
code maintenance.
Fig
This diagram illustrates the working of the Retention-AI System, which is designed to
predict and prevent customer churn using artificial intelligence and data analytics. The
use case represents how different actors interact with the system to achieve the goal of
improving user retention.
Key Actors
User: Regular app users whose engagement, behavior, and feedback are
continuously monitored to assess churn risk.
Admin : Business stakeholders who access the system’s dashboard, interpret
churn insights, and implement retention strategies based on AI recommendations.
AI System:The intelligent computational engine that processes user data, predicts
churn risk, and provides explainable insights to the admin.
System Workflow
1. User Interaction Layer: Users interact with the app by performing regular
actions (such as browsing, purchasing, or engaging with content) and submitting
feedback or reviews. This generates both structured (user activity logs) and
unstructured (text feedback) data that feed into the AI system.
2. AI Analysis Engine: The AI system performs three major analyses:
3. Strategic Output: Based on the aggregated churn score, the system recommends
personalized retention strategies such as targeted offers, in-app notifications, or
content suggestions for at-risk users.
4. Business Intelligence: Admins can:
System Components
1. User - The end user of the application who interacts with the platform and
provides behavioral and feedback data.
2. Retention-AI System - The central orchestration unit that manages
communication among AI modules.
3. Behavior Analyzer (ML) - The machine learning component responsible for
analyzing user engagement patterns and activity data.
4. Sentiment Analyzer (NLP) -The natural language processing module that
extracts sentiment from user feedback using models such as VADER.
5. Churn Predictor - The machine learning model that aggregates structured and
unstructured data to predict churn probability.
6. XAI Engine - The Explainable AI layer utilizing SHAP values to interpret model
outputs and highlight key churn factors.
7. Admin / Data Analyst - The business user who views predictions, explanations,
and personalized retention recommendations on the dashboard.
Interaction Flow
In the initial phase, the User interacts with the application by performing regular
activities such as browsing, purchasing, or engaging with app features. These interactions
generate structured behavioral data in the form of activity logs, frequency patterns, and
engagement metrics. Simultaneously, the user may provide textual feedback or reviews.
The Retention-AI System captures this data in real time and forwards it to the Behavior
Analyzer (ML).
In this phase, the Retention-AI System sends the collected textual feedback to the
Sentiment Analyzer (NLP). The NLP module, implemented using the VADER sentiment
analysis model, evaluates the polarity of the feedback-categorizing it as positive,
negative, or neutral. The system interprets this sentiment data as a measure of user
satisfaction, allowing it to complement the behavioral data collected earlier.
Once both behavioral and sentiment data have been analyzed, the Retention-AI
System consolidates the results and forwards them to the Churn Predictor. This model
combines insights from both data types to calculate a churn probability score for each
user. The predictor determines the likelihood that a user will discontinue app usage based
on past engagement behavior, satisfaction levels, and historical patterns.
Phase 4: Explainability
To ensure transparency, the Churn Predictor interacts with the XAI Engine, which
provides an interpretable explanation of the model’s decision-making process. Using
SHAP (SHapley Additive exPlanations), the XAI Engine identifies key features such as
session frequency, review sentiment, or inactivity duration that contribute most to each
user’s churn risk. These explainable outputs enable business users to understand the
reasoning behind churn predictions, converting complex model outputs into actionable
business intelligence.
After obtaining churn scores and their corresponding explanations, the Retention-
AI System compiles the results and displays them through an interactive dashboard. The
Admin or Data Analyst accesses this dashboard to view churn rates, customer segments,
In the final phase, the Admin can request personalized retention strategies for
users identified as high-risk. The Retention-AI System then generates recommendations
such as promotional offers, loyalty rewards, or personalized engagement messages. These
strategies are designed to re-engage users and enhance their experience, ultimately
reducing churn. The system can also deliver these recommendations directly to users
through notifications or app-based messages, closing the loop between prediction and
proactive retention action.
Description of Activities
1. User Interaction: The workflow begins when a user engages with the application
by performing actions such as browsing, ordering, or using specific app features.
These interactions generate structured behavioral data like session counts, feature
usage, and time spent in the app. Additionally, users may provide written
feedback or reviews, contributing unstructured textual data that reflects their
experience and satisfaction.
2. Data Collection : Once user interaction occurs, the Retention-AI System initiates
its automated data collection pipeline. It gathers structured behavioral metrics and
unstructured sentiment data simultaneously. These datasets form the foundation
for the system’s analytical process, capturing both quantitative and qualitative
aspects of user engagement.
3. Behavioral Analysis (ML Model): The collected behavioral data is sent to the
Machine Learning module for processing. This model analyzes engagement
frequency, activity duration, and interaction diversity to identify key patterns and
anomalies that may signal a decline in user activity. The analysis output includes
summarized engagement trends and behavioral risk indicators.
4. Sentiment Analysis (NLP Model): Parallelly, the unstructured textual feedback
is processed by the Natural Language Processing module using the VADER
sentiment analysis model. The system interprets the tone and polarity of the
feedback—categorizing it as positive, negative, or neutral—and extracts
emotional cues and dissatisfaction points that influence user behavior. This
sentiment information complements behavioral insights to provide a holistic user
understanding.
5. Churn Prediction: Once both behavioral and sentiment insights are derived, the
system aggregates them and forwards the combined dataset to the Churn
Predictor. This model, based on algorithms like XGBoost and Random Forest,
computes a churn probability score for each user. The churn score quantifies the
likelihood that a user may stop using the app, helping businesses identify at-risk
users proactively.
6. Explainability (XAI Engine): Following churn prediction, the system engages
the Explainable AI (XAI) Engine, which uses SHAP (SHapley Additive
exPlanations) to interpret the model’s output. The engine identifies key
contributing factors—such as low session frequency, negative sentiment, or
feature inactivity—and visualizes their impact on the churn probability. This
ensures that the system remains transparent and interpretable for business
stakeholders.
7. Results Processing and Visualization: After obtaining churn scores and SHAP
explanations, the Retention-AI System consolidates the results. These outputs are
displayed through an interactive dashboard where the Admin or Data Analyst
can monitor churn rates, visualize user risk segments, and review feature-level
insights. This stage enables data-driven decision-making by providing actionable
intelligence in a user-friendly format.
8. Generate Personalized Retention Strategy: When retention strategies are
requested, the AI system generates tailored responses based on each user’s churn
drivers. These strategies may include personalized discounts, reward offers, re-
engagement messages, or feature-based recommendations. The system aligns the
recommendations with the factors identified by SHAP to ensure that interventions
are targeted and effective.
9. Deliver Retention Action to User: The final step involves delivering these
personalized strategies to the respective users through notifications or in-app
messages. This closes the feedback loop, where AI-driven insights translate into
proactive business actions aimed at improving user satisfaction and retention.
Summary:
The Retention-AI activity diagram represents a closed-loop system that continuously
learns from user behavior. It integrates automated data analysis, explainable predictions,
and human decision-making to enhance customer retention. By combining Machine
Learning, Natural Language Processing, and Explainable AI, the system ensures a
transparent, adaptive, and personalized approach to mitigating user churn.
Fig
The Data Flow Diagram (DFD) illustrates the complete flow of data within the Retention-
AI System, showing how user interactions and feedback are transformed into intelligent
The process begins when users interact with the application, generating behavioral and
feedback data. The system captures app usage patterns such as session time, frequency,
and feature interactions, along with textual feedback like reviews and complaints. This
data is categorized and stored in two distinct repositories: the User Behavior Database for
structured logs and the Feedback Database for unstructured text data. This bifurcation
ensures efficient processing and easy retrieval during later stages of analysis.
In this stage, behavioral data from the User Behavior Database is analyzed using Machine
Learning algorithms. Metrics such as user activity frequency, feature adoption, and
engagement duration are studied to identify declining trends or inactivity patterns. The
output of this stage is a detailed behavior analysis report that helps the system understand
user engagement levels and detect early signs of churn risk.
The Churn Prediction stage serves as the integration point where results from the ML and
NLP modules converge. The combined behavioral insights and sentiment scores are input
into predictive models like XGBoost and Random Forest to calculate a churn probability
for each user. The output is a churn risk score, indicating how likely a user is to
discontinue app usage. This predictive intelligence allows proactive identification of
high-risk customers.
To ensure transparency, the Explainable AI (XAI) Engine interprets the results of the
churn prediction model. Using SHAP (SHapley Additive exPlanations), it highlights key
contributing features such as reduced activity, low engagement, or negative sentiment
toward specific features. These interpretable results are then displayed to the admin,
providing clear reasoning behind each prediction. This interpretability builds trust and
supports data-driven decision-making.
This structured data flow allows the Retention-AI System to operate as a transparent,
adaptive, and intelligent customer retention solution that blends machine learning
precision with human insight to reduce churn and enhance user loyalty.
This section outlines the functional and non-functional requirements of the system
developed for the project titled Retention-AI. The primary objective of this system is to
help businesses identify potential customer churn and implement proactive strategies to
retain users. It processes both structured data (user behavior metrics) and unstructured
data (customer feedback) using advanced Artificial Intelligence techniques such as
Machine Learning (ML), Natural Language Processing (NLP), and Explainable AI
(XAI).The system predicts the churn risk of users based on their engagement behavior
i. User Requirements
The user shall be able to upload new user data (CSV format) or connect to a
live database for churn prediction.
The system shall provide a churn prediction output for each user, including
the churn probability score and status (e.g., At Risk or Retained).
The user shall be able to view SHAP-based explainable insights that highlight
the top contributing factors influencing churn predictions.
The user shall have the option to generate personalized retention strategies
for high-risk users directly from the dashboard.
predictions but also maintains transparency, scalability, and usability across diverse
business environments.
1. Functional Requirements
The system shall enable uploading, storage, and preprocessing of customer data,
including both structured (user activity logs, engagement metrics) and
unstructured (user feedback, reviews) datasets. Data preprocessing shall include
handling missing values, normalization, and encoding to ensure model readiness.
The system shall support training and evaluation of multiple machine learning
models, such as Random Forest, XGBoost, and Ensemble Classifiers, to identify
the model that provides the highest churn prediction accuracy.
The system shall be capable of accepting new user data inputs in CSV format or
through a live database connection for churn prediction in real time.
The system shall generate and return churn prediction results for each user,
categorizing them as ―At Risk‖ or ―Retained,‖ along with the churn probability
score.
The system shall incorporate sentiment analysis using NLP techniques like
VADER to extract and quantify emotions from user feedback, contributing to the
overall churn prediction process.
The system shall integrate an Explainable AI (XAI) module utilizing SHAP to
interpret model outputs, highlighting the top features that influenced each churn
prediction.
The system shall provide an interactive dashboard for the Admin or Data Analyst
to visualize churn metrics, review insights, and generate personalized retention
strategies.
The system shall support real-time retraining of models based on defined
thresholds (e.g., retraining after a set number of new users), ensuring adaptive and
updated churn predictions.
2. Non-Functional Requirements
Usability: The system shall provide an intuitive, visually appealing, and easy-
to-navigate user interface through the dashboard. Users should be able to
access model results, churn insights, and personalized strategy
recommendations without technical complexity.
Performance: The system shall process uploaded datasets, perform churn
predictions, and return interpretable results within 3 - 5 seconds under normal
operating conditions, ensuring a smooth real-time experience for end-users.
Scalability: The system shall be capable of handling increasing data volumes
and concurrent user requests as more customers and feedback data are
integrated into the platform. It should support future expansion to enterprise-
level datasets without performance degradation.
Reliability: The system shall maintain operational accuracy and stability at
least 99% of the time, ensuring that model predictions and dashboards remain
consistently available and accurate.
Security: The system shall validate all user inputs and ensure that sensitive
customer data is securely stored and protected from unauthorized access
through appropriate encryption and authentication mechanisms.
Maintainability: The system’s codebase shall be modular, well-documented,
and structured to allow easy debugging, model updates, and integration of new
AI components or additional features in the future.
PROJECT PLANNING
This chapter presents the structured plan and timeline followed for developing the
Retention-AI System. Proper project planning ensures smooth execution, efficient
resource utilization, and timely completion of all deliverables. The development process
was divided into systematic phases — beginning with problem identification and
literature survey, followed by data collection, preprocessing, model development,
integration of Explainable AI (XAI), dashboard implementation, and testing. Each phase
was strategically scheduled to ensure steady progress toward the project’s objectives and
timely milestone achievement.
minimizing delays and maintaining consistent workflow throughout the project lifecycle.
During this stage, the project scope, objectives, KPIs, and success metrics were defined.
The development team was assembled, roles were assigned, and resource allocation was
finalized. Data governance, ethical AI use, and privacy policies were also established to
ensure compliance and transparency.
The system architecture was designed based on modular integration between structured
and unstructured data pipelines. Data schemas for the User Behavior Database and
Feedback Database were defined. Appropriate ML and NLP algorithms were selected,
and API integration plans were outlined for dashboard connectivity and model
deployment.
Data pipelines were developed to capture both behavioral and feedback data. The system
tracked user engagement metrics and feature usage patterns while also recording textual
reviews through feedback forms. ETL (Extract, Transform, Load) mechanisms were built
to ensure seamless and structured data ingestion.
The collected data underwent rigorous preprocessing, including handling missing values,
detecting outliers, and normalizing user activity logs. Textual data from feedback was
cleaned, tokenized, and transformed into numerical embeddings for NLP analysis.
Training datasets were prepared for ML and sentiment models.
Machine learning models were trained for behavioral analysis, while NLP models
(VADER) were used for sentiment scoring. Both outputs were fused in a churn prediction
ensemble model using algorithms such as Random Forest and XGBoost. Performance
validation ensured consistency and minimized overfitting.
Models were optimized using hyperparameter tuning and cross-validation techniques. The
Explainable AI (XAI) layer, powered by SHAP, was integrated to provide interpretable
insights. The models were evaluated across precision, recall, and F1-score metrics to
benchmark performance.
Comprehensive testing was performed, including User Acceptance Testing (UAT) and
A/B testing of retention strategies. Scalability and stress tests were conducted to evaluate
system performance under large data loads. Results validated that the Retention-AI
system operated efficiently with reliable churn prediction accuracy.
In the final phase, system documentation, algorithm explanations, and dashboard user
manuals were compiled. The system was deployed for production, followed by post-
launch monitoring to ensure operational stability and continuous performance tracking.
The architecture of Retention-AI is organized into multiple layers that handle data
ingestion, preprocessing, model execution, explainability, and user interaction. It is
designed to support structured behavioural logs and unstructured textual feedback,
enabling comprehensive churn analysis.
System Layers
1. Data Input & Preprocessing Layer: Handles ingestion of user logs and feedback,
followed by cleaning, tokenization, feature extraction, and model-specific dataset
splitting.
2. ML & NLP Processing Layer: Uses XGBoost for structured data prediction and
VADER-based sentiment scoring for textual inputs. Outputs are combined for ensemble
prediction.
4. Insights & Recommendation Layer: Produces aggregated churn reports, user risk
segmentation, and personalized retention strategy recommendations.
1. Data Input Model : The Data Input Module serves as the entry point for all raw data.
It enables admins to upload structured logs such as session data, activity frequency, and
retention metrics, along with unstructured textual feedback from user reviews. The
module validates file formats, performs schema checks, and ensures the uploaded datasets
4. Churn Prediction Module: The Churn Prediction Module is the core intelligence
engine of Retention-AI. It deploys XGBoost for behavioural churn detection and Random
Forest for sentiment-augmented churn prediction. Each submodel generates a churn
probability based on its respective input data. The module then combines the two scores
using a weighted ensemble approach to produce a final, more reliable churn probability. It
also categorizes users into different churn risk segments such as Low Risk, Medium Risk,
and High Risk, enabling businesses to prioritize intervention.
5. XAI Utility Layer (Explainability Module): Once the churn probability is computed,
the XAI module provides full interpretability through SHAP values. It explains how each
feature contributed to the prediction outcome, ensuring transparency and trust in the
model. The module generates local explanations for individual users and global
explanations to highlight dominant churn factors across the user base. These explanations
6. Dashboard Output Layer: This module handles the graphical output and visualization
layer through Streamlit. It provides interactive charts, churn risk summaries, SHAP
visualizations, sentiment distributions, and data upload interfaces. The dashboard ensures
that all model results, explanations, and retention recommendations are displayed in a
clear and business-friendly manner. Admins can explore individual user profiles, filter
risk segments, and download analytical reports.
The component design defines the internal structure, class relationships, attributes,
and operational methods that form the core logic of Retention-AI. Each major module is
implemented as a dedicated class, encapsulating its functionality and ensuring modular,
object-oriented development. The following components define the backbone of the
system:
Class: DataProcessor
The DataProcessor class forms the foundation of the system’s data pipeline. It handles
raw data ingestion, cleaning, transformation, and preparation for machine learning and
NLP tasks. The class ensures dataset consistency across all operational branches,
preserving data quality and improving downstream model performance.
Attributes
Methods
Class: ModelTrainer
This class orchestrates the training and evaluation of ML models used in the prediction
pipeline. It supports multiple model architectures and handles balancing, splitting, and
validation.
Attributes
Methods
Class: PredictionEngine
Responsible for generating churn predictions and explainable insights. It bridges ML
outputs with XAI reasoning.
Attributes
Methods
4. Frontend Module
Attributes
Methods
Class: AutomationOrchestrator
Coordinates scheduling, automated triggers, and continuous model updates.
Attributes
Methods
ML churn scorer, XAI engine, and strategy engine. This modular panel-based structure
ensures that each operation is clearly segregated and easily accessible.
When an administrator wants to upload behavioural log data, they navigate to the
―Upload Behavior Logs‖ section from the dashboard. The UI Manager displays a file-
upload widget allowing the admin to choose a CSV file containing user activity metrics
such as screen time, usage frequency, last seen date, and session duration. Once the file is
selected, the admin clicks the ―Upload‖ button, which triggers the frontend’s
handleLogUpload() function. This function packages the selected file into a multipart form-
data request and sends it securely to the backend via the API Client.
Upon receiving the upload, the backend validates the file structure, saves the
dataset to the designated storage path, and registers the upload event. It then returns a
response summarizing the number of records successfully received and any detected
issues. Once the frontend receives the backend response, the UI Manager displays a
success message confirming that the behavioural logs were uploaded and queued for
preprocessing. This feedback loop ensures that admins have real-time awareness of
upload outcomes and can immediately proceed to other analytical operations.
For uploading user reviews or free-text feedback, the admin navigates to the
―Upload Feedback Text‖ section. The UI presents a file-upload component similar to
behavioural logs but expects text-based entries. When a text file or CSV containing user
reviews is selected and uploaded, the UI’s handleFeedbackUpload() function is triggered.
This function sends the file to the backend pipeline, where the NLP Preprocessing
Module processes it through tokenization, text cleaning, and sentiment extraction. After
processing, the backend responds with a summary including the count of reviews
processed and the associated sentiment scores. The frontend displays a confirmation
message and updates the dashboard to reflect that sentiment data is now available for
downstream churn prediction.
2. Component-to-Component Interaction
Automation & Pipeline Module to Data Processing Module: After registering the new
data, the Automation & Pipeline Module triggers the Data Processing Module. The
preprocessing component performs cleaning, missing-value imputation, outlier capping,
text preprocessing, and dataset segmentation. This ensures that the data is in the correct
format for both the ML churn model and the NLP sentiment engine.
Data Processing Module to Prediction & Analysis Module: Once data preparation is
complete, the cleaned dataset is transferred to the Prediction & Analysis Module. This
module loads the currently active churn models and computes churn probabilities, risk
labels, and combined ensemble scores for users. It generates user-level prediction objects
and compiles the results into structures that are ready for backend delivery.
Prediction & Analysis Module to Frontend Module: The Prediction & Analysis
Module sends the churn predictions, risk categories, and explainability outputs back to the
frontend via API responses. The UI immediately updates the corresponding dashboard
panels, showing key metrics such as number of high-risk users, average sentiment trends,
and the top SHAP features contributing to churn predictions. This enables admins to
interpret user behaviour trends quickly and take informed action.
Automation & Pipeline Module to Model Training Module: If the Automation &
Pipeline Module identifies that the number of newly uploaded records exceeds predefined
retraining thresholds (e.g., Model A retrains after 20 new logs, Model B after 10 new
feedback entries), it initiates a retraining cycle. The Model Training Module re-trains the
ML and NLP models on the updated dataset and stores the new model versions. These
updated models are then used for all subsequent predictions and explanations, ensuring
that the system continuously adapts to new user behaviour.
This structure represents each individual feedback entry. The review_id uniquely
identifies each review. The user_id links the feedback to the corresponding user profile.
The text field contains the full review content. The sentiment_score stores the VADER-
derived sentiment value, while timestamp records when the feedback was created.
spent_on_app: float
left_review: bool
rating: int
sentiment_score: float
churn_risk: enum("High", "Medium", "Low")
explainability_factors: dict<string, float>
The UserProfile structure is used to store combined behavioural and sentiment attributes
of each user. The profile includes both structured metrics such as screen_time and
spent_on_app and unstructured metrics such as sentiment_score. The churn_risk attribute
holds the model-generated risk level, while explainability_factors stores SHAP-derived
influence scores for each relevant feature.
class User:
def __init__(self, user_id, screen_time, spent_on_app, feedback_text):
self.user_id = user_id
self.screen_time = screen_time
self.spent_on_app = spent_on_app
self.feedback_text = feedback_text
self.sentiment_score = None
self.churn_probability = None
self.risk_level = None
When new data is uploaded and processed, the system can create a list of User objects.
The Sentiment Analysis Module updates the sentiment_score attribute, and the Churn
Prediction Module fills the churn_probability and risk_level attributes. This sequential
update mechanism makes it easy to track how each user progresses through the pipeline
and how their churn risk is computed.
The algorithm design of Retention-AI outlines the logical flow and computational
processes that enable the system to perform churn prediction, sentiment analysis,
explainable reasoning, and automated retraining. Each algorithm is structured to ensure
consistent behavior, reproducibility, and modular execution across all system
components. The system integrates multiple algorithmic pipelines, including behavioural
preprocessing, text preprocessing, churn prediction through ensemble learning, SHAP-
based explainability, and automated retraining. These algorithms collectively ensure that
Retention-AI can efficiently handle large datasets, extract relevant patterns, and produce
actionable insights for the admin. The following subsections describe the key algorithms
that drive the core functionalities of the platform.
This algorithm processes raw behavioural logs uploaded by the admin and prepares them
for input into the churn prediction model. It handles missing values, caps extreme
numeric values, normalizes key features, and splits the dataset into training and testing
subsets. This ensures that the data is clean, consistent, and suitable for ML operations.
Through this algorithm, Retention-AI ensures that the behavioural dataset remains free
from noise and anomalies, improving the stability of downstream model inference.
This algorithm transforms unstructured textual feedback into usable sentiment features. It
applies cleaning, tokenization, lexicon-based scoring, and sentiment aggregation.
This algorithm ensures that qualitative feedback is quantified into meaningful numerical
signals that contribute to churn prediction.
The churn prediction algorithm combines the outputs of two predictive models—Model A
(behaviour-based XGBoost) and Model B (sentiment-augmented Random Forest). The
ensemble mechanism produces a final churn probability that is more stable and accurate
than individual model predictions.
This algorithm uses churn risk levels and sentiment signals to assign personalized
retention strategies to users. It maps high-impact factors to specific business actions.
Notify the Prediction Engine to use updated models for future predictions.
IMPLEMENTATION
Implementation Plan
2. Data Integration and Preprocessing: Behavioural logs and text feedback were
uploaded through the frontend dashboard. These datasets were cleaned, normalized,
tokenized (for text), scored using VADER sentiment analysis, and stored in their
respective processed folders. Outlier capping, missing-value imputation, type
standardization, and text preprocessing were implemented to ensure high-quality model
inputs.
Both models were trained, grid-searched, evaluated, and saved into the models directory.
Prediction files were generated for each user and stored for ensemble processing.
6. Testing & Debugging: Each module was tested independently to ensure correctness—
upload operations, preprocessing steps, training accuracy, SHAP value correctness, API
responses, and dashboard functionality. Integration testing was performed to validate the
full pipeline.
FUNCTION upload(file):
IF file missing OR not CSV -> ERROR "Invalid Input"
SAVE file -> datasets/[Link]
df <- READ CSV
df <- CLEAN (missing/outliers) + NORMALIZE + ENCODE
SAVE -> datasets/[Link]
INCREMENT retrain_counters
RETURN message, records_processed
[Link] Model A
FUNCTION train_model_a():
df <- READ datasets/model_a_train.csv
DROP leakage cols [userid, churn_risk, ...]
(X_tr, X_te, y_tr, y_te) <- SPLIT 80/20 (stratify)
(X_trb, y_trb) <- SMOTE(X_tr, y_tr)
model_a <- XGBClassifier()
FIT model_a ON (X_trb, y_trb)
SAVE model_a -> models/model_a_latest.pkl
SAVE feature_names -> models/model_a_features.json
PREDICT probs (all users) ->datasets/model_a_predictions.csv
[Link] Model B
FUNCTION train_model_b():
df <- READ datasets/model_b_train.csv
text <- PREPROCESS_TEXT(df.review_text)
feats <- SENTIMENT_SCORES(text)
model_b <- XGBClassifier()
FIT model_b ON (feats, y)
SAVE -> models/model_b_latest.pkl
PREDICT probs -> datasets/model_b_predictions.csv
FUNCTION build_churn_predictions():
A <- READ model_a_predictions.csv (userid, a_prob)
B <- READ model_b_predictions.csv (userid, b_prob)
FOR each userid:
final_prob <- (a_prob + b_prob)/2 IF b_prob exists ELSE a_prob
risk <- High if final_prob>=80
Medium if final_prob>=50
Low otherwise
5. Explainable AI
FUNCTION explain_user(user_id):
IF cache outputs/explanations/user_{id}.json exists -> RETURN cache
model_a <- LOAD models/model_a_latest.pklfeats <- LOAD user features
shap_vals <- SHAP(model_a, feats)
top_features <- TOP_K(shap_vals, 10)
chart_path <- outputs/charts/user_{id}.png (Agg)
SAVE JSON {user_id, prob, risk, top_features, recommendations, chart_path}
RETURN JSON
Code Efficiency
TESTING
The testing of the Retention-AI system was carried out using both unit testing and
integration testing to ensure that every component, from churn prediction to explainable
AI and retention strategy generation, performs correctly and reliably. Unit testing focused
on validating individual modules such as data upload, preprocessing, prediction logic, and
explanation generation. Integration testing evaluated how these modules work together
across the full flow — from user actions in the UI to backend API responses and database
updates.
This combined approach ensured that both the independent modules and the entire end-to-
end pipeline behaved as expected under different scenarios, user inputs, and data
conditions. Testing confirmed that Retention-AI consistently produces accurate churn
predictions, handles invalid inputs, and displays results through a smooth and intuitive
user interface.
Unit testing was carried out to verify that each functional component of Retention-
AI behaves correctly in isolation. The main focus was on validating:
Each test case assessed how the system responded to different inputs such as valid files,
incorrect formats, missing values, invalid symbols, and corrupted explanation caches. The
tests confirmed that Retention-AI successfully handles proper inputs by generating
accurate churn probabilities, user-level insights, and correct SHAP-based explanations.
Some limitations were also identified. For instance, the system initially did not validate
unknown columns in uploaded CSV files and crashed when the explanation API was
called without a valid user ID. These issues highlight areas for stronger input validation
and error-handling mechanisms. Nevertheless, the overall unit testing phase demonstrated
that the core functionalities - including prediction accuracy, explanation generation, and
data processing - are functioning correctly while revealing important opportunities for
refinement.
The File Upload Module correctly handles valid files, invalid file formats, and empty
datasets. However, it fails when unknown columns are present, indicating the need for
stricter schema validation before processing.
The Explanation Module performs well in generating valid results, handling missing
users, and recovering from corrupted cache files. The failure occurs when the user_id
parameter is missing, showing that API input validation needs improvement.
System-level tests reveal strong handling of malformed requests, prediction accuracy, and
NLP processing. However, performance limitations exist for large file uploads and high-
load API usage, highlighting areas requiring optimization to ensure scalability.
Integration testing was conducted after successful unit testing to verify the
seamless interaction between all system components. The key modules tested together
include the React frontend, Flask API layer, XGBoost-based churn model, sentiment
analysis module, explainable AI engine, and the backend dataset storage.
The goal of integration testing was to ensure that data flows correctly across the
system -from file uploads and prediction requests to dashboard refreshes, user explanation
retrieval, and retention strategy generation. During testing, the system successfully
processed uploaded CSV files, refreshed churn predictions, generated updated
dashboards, and visualized user-specific insights through SHAP values. The
communication between the frontend and backend remained stable, confirming that the
APIs, data processing logic, and visualization components are well integrated.
2 Upload CSV missing Show ―Invalid Input‖; Error shown; no file Pass
required columns no DB update saved; predictions
unchanged
3 Upload valid CSV Trigger retrain; new Retrain logs present; Pass
exceeding retrain model saved; model updated
threshold predictions updated
4 Upload CSV during Upload request queued Upload blocked but UI Fail
active retraining or blocked gracefully freezes
This table validates the integration between the upload module, retraining logic, and
prediction pipeline. The system successfully handles valid/invalid uploads and triggers
retraining correctly when thresholds are met. However, when a retraining process is
active, the UI freezes instead of gracefully queuing or blocking the request, indicating a
concurrency-handling issue.
This table verifies the integration between the explanation engine, cache layer, and
dashboard visualization. The failure occurs when the system attempts to load a chart that
does not exist, resulting in an API exception instead of a controlled error message.
3 Server under high load (50 System should Significant delays; Fail
concurrent uploads + degrade partial hangs
explanations) gracefully
This table evaluates system-level interactions under stress and error conditions. The
system performs well under moderate load but struggles under heavy load, indicating
performance bottlenecks. The missing user_id case also fails, showing insufficient API-
level validation. These tests highlight areas requiring optimization for robustness and
scalability.
When valid CSV files containing complete user behavior logs and feedback were
uploaded, the system successfully processed and normalized the data.
Churn predictions were generated accurately, reflecting the correct probabilities
and risk levels.
The explainability module displayed correct SHAP insights, validating the
correctness of model inference.
On uploading unsupported file formats (TXT, JSON) or empty CSV files, the
system produced appropriate error messages such as ―Invalid Input‖ or ―Missing
required columns‖.
This confirms strong validation checks in the upload and ingestion pipeline.
The system provided clear visualizations for risk distribution, churn probability
distribution, and explanation charts.
Searching users, filtering by risk levels, and generating insights worked smoothly
across all test cycles
For every user ID requested, the system returned impact scores and top risk
factors.
Cached explanations were retrieved correctly, reducing computation time for
repeated users.
These findings collectively prove that RETENTION-AI is robust, accurate, and capable
of handling diverse churn-prediction scenarios with stability and clarity.
1. Dashboard Overview
The Dashboard provides a complete snapshot of user churn and retention metrics. Users
can view the total number of customers, churn rate, retention rate, and visual breakdowns
such as risk distribution and churn probability charts.
(Ref: Screenshot – Dashboard Overview)
2. Data Upload
In the Data Upload interface, admins can upload user behavior logs or feedback data via
CSV. The platform validates the file, preprocesses features, and updates the dataset. The
page also provides formatting instructions to ensure correct upload structure.
(Ref: Data Upload screenshot)
This page displays churn predictions for all users after processing the dataset. Each user is
assigned a churn probability and a corresponding risk category (Low, Medium, High).
The table supports searching, filtering, and exporting results.
(Ref: Churn Prediction screenshot)
RETENTION-AI analyzes customer feedback using NLP models and assigns sentiment
labels along with numerical sentiment scores. This helps businesses understand customer
emotions contributing to churn.
(Ref: Sentiment Analysis screenshot)
5. Explainable AI (XAI)
6. Retention Strategies
This module provides actionable recommendations tailored to each user’s risk level.
Admins can select a user, choose from generated suggestions, customize the message, and
send retention emails directly through the interface.
(Ref: Retention Strategies screenshot)
SNAPSHOTS
This screenshot highlights the data upload portal, where administrators can drag and drop
CSV files containing customer logs. It also provides required fields and format
guidelines.
This page lists users alongside their churn risk and predicted probability. Users can filter,
sort, and search through 1029 predictions efficiently.
The sentiment analysis module displays customer reviews and their sentiment labels. It
helps identify key emotional signals contributing to churn.
This interface presents recommended user-specific retention actions and allows sending
personalized emails to targeted users.
This figure shows SHAP explanations including key risk factors, bar plots, and a user
summary with churn risk justification.
9.2 Applications
RETENTION-AI can be applied across various industries where customer churn directly
impacts revenue and growth:
5. Digital Marketing & Customer Support: Marketing teams can leverage Retention-AI
for targeted outreach campaigns by identifying customers most at risk of leaving.
Customer support teams can use churn insights and sentiment data to address issues
before they escalate. Together, these insights help create personalized, data-driven
communication strategies that improve retention.
1. Static Dataset Usage: The current system relies on uploaded CSV files rather than
real-time data streams, which limits its ability to reflect sudden changes in user behavior.
Because predictions are based on periodic updates, continuous monitoring is not yet
possible. This reduces responsiveness in rapidly changing environments.
2. Limited Domain Adaptability: Although Retention-AI works well for general user
churn prediction, different industries may require customized features or behavior
metrics. Highly specialized domains like healthcare or finance might not fit the default
feature set. This necessitates domain-specific tuning and retraining of models.
3. Sentiment Model Constraints: The sentiment analysis component may struggle with
sarcasm, slang, or complex multilingual expressions. This can lead to misclassification of
user reviews or feedback. As a result, certain emotional cues may not be captured
accurately, affecting overall prediction quality.
1. Real-Time Data Streaming: Future versions can integrate real-time event tracking
pipelines using tools like Kafka or AWS Kinesis. This will allow instant churn updates
whenever user behaviour changes. Such real-time predictions significantly improve
responsiveness and retention accuracy.
4. Cloud Deployment for Scalability: Deploying the system on AWS, Azure, or GCP
will support large-scale churn monitoring across thousands of users in parallel. Cloud
services also offer automated scaling, faster processing, and high reliability.
[6] Suchita Sharma, Nishith Desai. ―Identifying Customer Churn Patterns Using
Machine Learning Predictive Analysis‖, Proceedings of the 3rd International
Conference on Smart Generation Computing, Communication and Networking
(SMART GENCON), IEEE, pp. 1–6, 2023.
[7] Kapil Arora, M. Lalitha, Poonam, Hemalatha Yadav J, Biswa Ranjan Mishra.
―Machine Learning for Customer Retention in E-Commerce Healthcare Startups‖,
South Eastern European Journal of Public Health (SEEJPH), Vol. XXV, Suppl.
2, pp. 4303–4308, 2024.
[8] Sabreen Abulhaija, Shyma Hattab, Ahmad Abdeen, Wael Etaiwi. ―Predicting
Mobile Apps Performance using Machine Learning‖, Journal of System and
Management Sciences, Vol. 12, No. 6, pp. 300–314, 2022.
[9] Manzura Jorayeva, Akhan Akbulut, Cagatay Catal, Alok Mishra. ―Machine
Learning-Based Software Defect Prediction for Mobile Applications: A
Systematic Literature Review‖, Sensors, Vol. 22, No. 7, pp. 1–17, 2022
[10] Zeynep Hilal Kilimci, Hasan Yörük, Selim Akyokus. ―Sentiment Analysis Based
Churn Prediction in Mobile Games using Word Embedding Models and Deep
Learning Algorithms‖, IEEE International Conference on Machine Learning and
Applications (ICMLA), pp. 1–8, 2020.
[11] Tamer Ahmed Ibrahim Abou El-Fotouh, Musliuende Toheeb Akanbi. ―Impact of
Predictive Analytics and Machine Learning on Customer Retention and Loyalty in
Service-Oriented Businesses‖, International Journal of Business Intelligence and
Big Data Analytics, Vol. 07, No. 03, pp. 1–11, 2024.
[12] Naragain Phumchusri, Phongsatorn Amornvetchayakul. ―Machine Learning
Models for Predicting Customer Churn: A Case Study in a Software-as-a-Service
Inventory Management Company‖, International Journal of Business Intelligence
and Data Mining, Vol. 24, No. 1, pp. 1–18, 2024.
[13] Carl Yang, Xiaolin Shi, Jie Luo, Jiawei Han. ―I Know You’ll Be Back:
Interpretable New User Clustering and Churn Prediction on a Mobile Social
Application‖, KDD, pp. 1–9, 2019.
[14] Ziru Liu, Shuchang Liu, Bin Yang, Zhenghai Xue, Qingpeng Cai, Xiangyu Zhao,
Zijian Zhang, Lantao Hu, Han Li, Peng Jiang. ―Modeling User Retention through
Generative Flow Networks‖, KDD, pp. 1–12, 2024.
[15] Mudassir Rafi, Md. Faiz Ahmad, Varshitha K, Siri Varsha T, Lahari K, Md.
Asadul Haque, Pavan Kumar Pagadala, Sushama Rani Dutta. ―Customer Churn
Prediction employing Ensemble Learning‖, IEEE 6th International Conference on
Cybernetics, Cognition and Machine Learning Applications (ICCMLA), pp. 207–
211, 2024.