Predictive Modeling for Loan Default Analysis
Predictive Modeling for Loan Default Analysis
As 2025 unfolds, financial institutions contend with persistent high interest rates (4.25%-4.50%),
stagflation Ary pressures (3% inflation, 1.4% GDP growth), tighter credit conditions, rising
delinquencies and charge-offs, declining small business lending, compressed net interest margins,
and increased competition from fintechs—all while navigating regulatory changes and economic
uncertainty that demand enhanced analytics, risk management, and modernization of lending
strategies.
Strategic Imperative for Customer Value Segmentation
Lenders must move beyond traditional risk-based models and adopt advanced analytics for
customer value segmentation—especially targeting high-value, long-term loan customers—to
optimize portfolios and drive sustainable growth, as research and industry practice show that
focusing on Customer Lifetime Value (CLV) enables banks to improve profitability, enhance
cross-sell opportunities, strengthen customer relationships, and achieve significant gains in
marketing effectiveness and retention, particularly in a challenging, high-rate, and competitive
lending environment
Environment: financial institutions must balance risk management with maximizing customer
lifetime value by deploying sophisticated segmentation strategies—such as segmenting borrowers
by loan term—to identify and target valuable customer segments, enabling tailored products,
optimized risk-adjusted returns, and stronger long-term relationships
Analytical: By leveraging advanced machine learning models—including logistic regression,
decision trees, random forests, and boosted forests—lenders can robustly identify the factors
driving long-term customer value and behavior, achieving both high predictive accuracy and
interpretability essential for regulatory compliance and effective stakeholder communication
Research Gap and Intended Outcomes
While current literature and industry practice primarily focus on default prediction and risk
management, they often neglect customer value segmentation based on loan term preferences,
creating a research gap that this project addresses by shifting from risk-only models to
comprehensive customer value identification through loan term prediction; the intended outcomes
are actionable insights for product design, risk management, and customer retention, as well as
empirical evidence from comparative machine learning analysis to support optimal, interpretable
customer segmentation strategies that enhance institutional capabilities and meet regulatory
requirements in the competitive digital lending landscape.
1
2. Problem Statement and Objectives
The problem is the inability of financial institutions to efficiently identify and segment valuable
customers based on their loan term preferences, resulting in suboptimal customer targeting,
resource allocation, and revenue optimization.
Challenges: Distinguishing between customers who prefer long-term versus short-term loan
products, which directly impacts their ability to develop targeted marketing strategies, optimize
product offerings, and maximize customer lifetime value. Current manual assessment processes
are time-consuming, inconsistent, and fail to leverage the full potential of available customer data
for predictive insights
2.2. Objective
To develop and compare multiple machine learning algorithms (Logistic Regression, Decision
Tree, Random Forest, and Boosted Forest) for predicting loan term preferences using customer
financial and demographic characteristics, thereby enabling financial institutions to identify
valuable customers and optimize loan product targeting strategies.
Model Development and Performance Optimization
Develop four distinct machine learning models to classify customers based on loan term
preferences with over 70% accuracy by implementing comprehensive data preprocessing—
including handling missing values, outlier detection, and feature scaling—and conducting rigorous
train-validation-test splits to ensure model generalizability and prevent overfitting.
Comparative Algorithm Analysis
Evaluate and compare the predictive performance of Logistic Regression, Decision Tree,
Random Forest, and Boosted Forest algorithms using standardized metrics (accuracy, precision,
recall, F1-score, and AUC-ROC), analyze their computational efficiency and training time to
assess practical feasibility, and determine the optimal algorithm by balancing predictive
accuracy, interpretability, and operational requirements.
Feature Importance and Variable Analysis
Identify and rank the relative importance of predictor variables—including financial capacity
indicators (annual income, debt-to-income ratio, installment amount, revolving balance), credit
profile metrics (grade, sub-grade, interest rate, revolving utilization, verification status), credit
history factors (public records, bankruptcy records, mortgage accounts, open accounts, total
accounts), and employment/housing stability (employment length, home ownership status, loan
purpose)—across all four modeling approaches; analyze interaction effects between predictors to
uncover synergistic relationships that enhance prediction accuracy; and validate feature
importance findings through cross-algorithm comparison to ensure robust and reliable variable
selection.
2
3. Literature Review
The quantification and management of customer value form the foundation of modern
marketing and financial strategy. Related theories have evolved from early models focused solely
on financial valuation to more sophisticated, multidimensional frameworks for managing
customer relationships.
3
behavior. Its methodological approach goes beyond traditional segmentation and valuation,
incorporating complex Expected Credit Loss (ECL) models and climate risk scenario analysis.
This reflects an integrated perspective in real-world business environments—one that
simultaneously considers customer value, credit risk, and macroeconomic conditions.
The existing literature adopts a wide range of research paradigms to investigate customer
value, spanning from quantitative modeling to conceptual integration. Each approach brings
distinct analytical priorities, contributing to a more nuanced and multidimensional understanding
of the construction. the following table is provided:
Table 3-1 Research Paradigms
Primary
Study Research Paradigm Description
Paradigm
Developed a mathematical model based on customer acquisition
Financial
rate, retention rate, profit margin, and discount rate to quantify
Valuation
Gupta et total customer equity. Empirically tested the model using public
Model &
al. (2004) financial data from five listed companies, demonstrating a strong
Empirical
correlation between model-based valuations and actual market
Analysis
capitalization.
Proposed a three-dimensional LTV model incorporating current
Theoretical
value, potential value, and customer loyalty. Applied the model to
Hwang et Model
real customer data from a major South Korean wireless telecom
al. (2004) Building &
company, validating its effectiveness in customer segmentation and
Case Study
high-value customer identification.
Conducted a systematic review and synthesis of literature in
Kumar & customer value, customer relationship management, and customer
Conceptual
Reinartz engagement. Constructed an integrative theoretical framework
Review
(2016) through critical analysis and conceptual innovation rather than
empirical testing.
Analyzed corporate reports and strategic documents to examine
how the institution applies data science to customer management
NatWest Corporate
and risk control. The report outlines the internal use of customer
Group Case
segmentation models, Expected Credit Loss (ECL) models, and
(2024) Analysis
climate risk assessments, showcasing the in-depth application of
theory in large financial institutions.
Accordingly, it is evident that the definition, measurement, and application of customer value
remain subjects of ongoing theoretical development and scholarly debate.
Against the backdrop of digital transformation in the global financial services industry, the
4
sector is undergoing a fundamental strategic shift—from a product-centric to a customer-centric
orientation. Financial institutions are no longer satisfied with offering standardized credit
products; instead, they increasingly seek to deliver highly personalized services through data-
driven insights. Leading institutions have adopted predictive analytics to enable smarter lending
decisions (Verbraken, Verbeke, & Baesens, 2013). These applications extend beyond conventional
credit risk assessments to encompass refined customer management approaches informed by
behavioral and psychographic profiling (Kumar & Reinartz, 2018).
Accurately predicting customer preferences—such as optimal loan duration—requires a
nuanced understanding of customer heterogeneity. Modern financial customer segmentation has
moved beyond traditional demographic-based models and evolved into multidimensional
frameworks that integrate behavioral patterns, psychological traits, and lifecycle stages. Such
approaches offer a more comprehensive perspective on customer preferences and enable
institutions to design more targeted financial strategies.
Academic research in this domain generally approaches the topic from three major angles:
(1) Behavioral Segmentation
This approach classifies customers based on their actual interactions with financial
institutions. For example, credit card users can be segmented into transactors (who pay their
balances in full each month), revolvers (who pay only the minimum and carry forward balances),
and inactive users (Verbraken, Verbeke, & Baesens, 2013). Such segmentation reflects customers’
financial habits and risk preferences in a direct and observable manner.
(2) Psychographic Segmentation
This paradigm incorporates insights from behavioral economics into customer profiling.
Traits such as loss aversion or present bias significantly influence decision-making processes—for
instance, how individuals trade off short-term high interest against long-term low monthly
payments, thereby shaping their loan term preferences (Kahneman & Tversky, 1979).
(3) Lifecycle Segmentation
This method dynamically segments customers according to their life stages and associated
financial needs. Empirical cases show that age- and segment-based classifications—such as youth,
retail, and affluent clients—can effectively predict product preferences across life events such as
education, marriage, and retirement. These patterns are particularly informative in forecasting
demand for loan products of varying maturities (Verbraken, Verbeke, & Baesens, 2013).
The ultimate goal of customer segmentation is to enable optimal allocation of resources, and
the Customer Lifetime Value (CLV) model provides a scientific and quantitative tool for this
purpose. CLV estimates the total profit a customer is expected to generate over the entire duration
of the business relationship, allowing financial institutions to allocate marketing and service
5
resources more efficiently toward high-value clients. Empirical evidence shows significant
disparities in CLV across different customer segments: while the average CLV of a private
banking client may amount to several tens of thousands of dollars, that of a basic retail account
holder might be only a few hundred (Kumar & Reinartz, 2018).
Traditional CLV models often rely on historical transactional data, which can be limiting in
low-frequency, sparse-data contexts such as banking. To address this, researchers have developed
dynamic models based on stochastic processes, such as Markov chains, which simulate the
probabilities of customers transitioning across segments over time. These models offer more
robust forecasts of future value, especially for products with infrequent purchase cycles
(Verbraken, Verbeke, & Baesens, 2013).
From the perspective of personalization strategy, CLV analysis also serves as a direct
foundation for product customization. For customer groups with high predicted CLV, financial
institutions have stronger incentives to develop and recommend tailored loan offerings that match
specific preferences—such as loan maturity or interest rate structure—thereby enhancing customer
satisfaction and long-term loyalty (Kumar & Reinartz, 2018).
However, current research on customer value has largely overlooked a critical and observable
behavioral dimension: customer preferences for loan maturity. Existing studies tend to adopt a
supply-side perspective, focusing on how financial institutions shorten loan durations to mitigate
risk under asymmetric information. Yet few have investigated, from the demand side, the micro-
level factors driving customers’ deliberate choices of specific repayment horizons. This oversight
has led to an understanding of customers that remains largely superficial—limited to financial and
demographic attributes—without fully capturing the deeper behavioral signals embedded in
customer decisions.
Loan maturity choice is not an isolated financial decision. It can serve as a comprehensive
proxy variable, reflecting an individual’s risk tolerance, time discounting behavior, financial
planning ability, and even broader indicators of financial health. Therefore, incorporating maturity
preference as a distinct analytical dimension not only addresses a meaningful gap in the existing
literature but also deepens and strengthens the conceptual and practical foundations of customer
value theory.
The theory of Enduring Customer Value proposed by Kumar and Reinartz (2016) emphasizes
a strategic shift from focusing solely on transactional contributions—such as Customer Lifetime
Value (CLV)—to building long-term relationships based on comprehensive customer engagement.
6
In this context, a customer’s loan maturity preference serves as a valuable leading indicator for
anticipating relationship longevity and estimating the customer’s potential enduring value.
When analyzed from a preference-based perspective, the choice of loan maturity may signal
divergent customer trajectories and risk profiles. Short-term preferences often correspond to
transaction-oriented, interest-sensitive customer segments that exhibit lower loyalty and higher
attrition risk. In contrast, long-term preferences may reflect a greater willingness to establish
durable relationships and a higher degree of trust in the financial institution. Such customers are
more likely to generate stable future cash flows, pose lower default risk, and contribute additional
value through cross-selling opportunities and customer referrals (CRV), thereby demonstrating
higher potential for enduring customer value.
Incorporating maturity preference as a key variable—or even as a risk-weighted input—into
CLV prediction models allows for more precise calibration of risk expectations and significantly
enhances both the accuracy and robustness of long-term value estimation.
Understanding maturity preference enables financial institutions to move from passive
accommodation to proactive guidance. For high-potential customers demonstrating long-term
preferences, institutions can design and recommend product bundles that align with those
preferences, even offering more flexible repayment options to reinforce loyalty through value
creation. For customers favoring short-term arrangements, targeted communication strategies
aimed at improving financial literacy and loyalty may be more appropriate. This preference-
informed differentiation forms the core mechanism for striking a dynamic balance between value
creation for the customer and value extraction by the firm.
In summary, systematically examining loan maturity preferences and integrating them into
multidimensional segmentation and valuation frameworks represents a necessary step in
translating customer value theory from conceptual abstraction to operational precision. Future
research could explore the use of advanced and interpretable machine learning models—such as
gradient boosting trees combined with SHAP (SHapley Additive exPlanations)—to both predict
maturity preferences with high accuracy and identify their key behavioral and financial drivers.
This would provide a scientifically grounded path toward truly customer-centric risk pricing,
product innovation, and value management in the financial services sector.
In the early stages of applying machine learning to financial prediction, classical models such
as logistic regression, decision trees, and random forests played a foundational role in credit risk
assessment. These models were particularly effective in forecasting default risk at a macro level
7
and provided interpretable, rule-based decision structures.
However, as the focus of prediction has shifted from broad risk estimation to understanding
granular customer preferences—such as loan maturity choices—the limitations of these models
have become increasingly apparent. Their relatively rigid functional forms and limited capacity to
capture complex, nonlinear interactions make them less suitable for preference prediction in
heterogeneous customer bases.
To clarify the relative strengths and weaknesses of commonly used models in this context,
the following table provides a comparative summary of their characteristics:
Table 3-2 Model Introduction
8
ensemble averaging, their majority-vote mechanism may dilute the influence of critical but rare
patterns, and their "black box" nature poses challenges for business interpretability and regulatory
compliance.
To overcome these limitations—particularly in predictive accuracy, nonlinear pattern
recognition, and complex feature interaction modeling—we turn to boosting models based on
gradient boosting algorithms. Empirical research has consistently demonstrated the superior
performance of boosted tree models in financial risk assessment. For example, in real-world credit
card default prediction tasks, XGBoost has achieved accuracy rates exceeding 99%, significantly
outperforming traditional logistic regression and even neural networks (Chen & Guestrin, 2016).
LightGBM, with innovations such as Gradient-based One-Side Sampling (GOSS) and Exclusive
Feature Bundling (EFB), maintains high accuracy while substantially reducing computational cost
and memory usage, making it especially suitable for large-scale financial datasets (Ke et al.,
2017).
Loan maturity preference is a prototypical multi-factor, nonlinear decision problem. Factors
such as income, age, credit history, and debt burden interact in non-additive and often
unpredictable ways. Boosting models are well-suited to this context, as they can automatically
learn high-order feature interactions without the need for manual feature engineering. This
capacity makes them an ideal solution for high-precision modeling of customer preferences,
enabling financial institutions to uncover subtle behavioral drivers and support the development of
truly personalized financial products.
9
Table 4-3 Data Preparation Table
Step Description
1 Import the Loan_Tap dataset and perform an initial data quality check.
Numeric Convert :1. employment length to numeric (e.g., '< 1 year'=0, '10+
years'=10).2. Sequential: grade and subgrade as numeric codes 3. Map home
ownership, verification status, and purpose categories to numeric codes; group
rare purposes as 0.4. Binary: Convert loan status to numeric (Fully Paid=0,
3 Charged Off=1, others as missing)(Figure 4-2)
6 Split data into train/test sets, create the final model-ready dataset,
export to SAS Viya Model Studio, document all transformations, and automate
8 the data preparation pipeline to reduce errors and ensure consistency
10
Figure 4-2 Numeric Conver
4.2. IDEA
Target Variable: term_binary is correctly created: 0 for 36 months (short term), 1 for 60 months
(long term).
This indicates a strong class imbalance, with 36-month loans being about three times more
common than 60-month loans
11
Figure 4-4 Category variable
Category variable: The frequency distribution of both grade and sub_grade is skewed, with most
loans concentrated in the middle categories and fewer loans in the highest and lowest categories.
The grade variable shows the highest frequencies in the middle grades (around 3 and 4), while
both lower and higher grades are less common.
The sub_grade variable displays a similar pattern, peaking near the middle sub_grades and
tapering off at both extremes, indicating a strong central tendency in loan quality distribution after
clarifying.
Continuous variable: The distributions of installment amount, interest rate, number of mortgage
accounts, and number of open accounts are all right-skewed, with most values concentrated at the
lower end of their ranges.
12
Each variable shows a clear central tendency, with frequencies peaking at specific values and
tapering off as values increase, indicating that most borrowers have moderate installment amounts,
interest rates, and account numbers.
There are visible outliers and long tails in all distributions, highlighting the importance of outlier
detection and treatment during data preparation to ensure robust modeling.
which research methods will be used • describing the data including the data source
This project aims to predict whether a customer will select a long-term or short-term loan, using
advanced classification models and focusing on maximizing the F1 score for model evaluation.
The F1 score is chosen as the primary metric because it balances precision and recall, providing a
fair assessment of model performance, especially when both types of classification errors (false
positives and false negatives) are important to minimize
Data Description
Data Source:
13
The dataset is sourced from LoanTap’s lending data, containing customer information, loan terms,
employment details, credit grades, home ownership, verification status, loan purpose, and loan
status. All variables are preprocessed to numeric or binary format, as described in the data
preparation section.
Key Variables:
Predictors: Employment length, grade, sub-grade, home ownership, verification status, loan
purpose, and other demographic or behavioral features. Loan amount is excluded as per project
scope.
14
5. Analysis and Results •
The analysis section is concerned with analyzing the results you have achieved through the data
analytics project. • Results are actual statements of observations, including statistics, tables and
graphs. • You should present both negative and positive results with sufficient details so that
readers can draw their own inferences and construct their own explanations.
Four supervised machine learning algorithms are used for binary classification:
15
especially in cases where the costs of misclassification are asymmetric. The KS statistics are a
widely used metric for evaluating the discriminatory power of classification models and is
particularly valuable in credit scoring and risk modeling.
16
5.2. Decision Tree
Figure 5-8Accuracy
17
The most important variables for classification are (by far the most influential)。
Low and Stable Misclassification Rate:
The misclassification rate is 17.6% across all data partitions, indicating that the tree is well-pruned
and not overfitting. This is supported by the pruning error plot, which shows that the selected
subtree with 36 leaves achieves the lowest validation error without unnecessary complexity
18
5.4. Boosted Model
19
Robust Generalization and Nuanced Feature Importance
The model maintains a low and stable misclassification rate (13%) across TRAIN, VALIDATE,
and TEST partitions, indicating strong generalization and minimal overfitting. Gradient Boosting
also provides more nuanced feature importance scores, with sub_grade as the most influential,
followed by installment, mort_acc, and annual_inc. This allows for finer feature selection and
deeper interpretability compared to other models.
20
Discrimination and Ranking (AUC & Lift)
Gradient Boosting and Random Forest both have high AUC and lift, indicating excellent ability to
rank and separate positive cases from negatives.
Gradient Boosting slightly edges out Random Forest in lift and AUC, making it better for
prioritizing top candidates for action
21
Key Results:
Gradient Boosting is the top performer for both accuracy (87.0%) and F1 score (0.70), with the
highest AUC (0.93) and lift (3.69), and the lowest misclassification rate (13.0%)—making it the
best model for binary classification on this dataset.
Random Forest offers a strong balance of performance (accuracy 85.6%, F1 0.65), discrimination
(AUC 0.91, lift 3.62), and robustness (misclassification 14.4%), with lower risk of overfitting than
single trees.
Logistic Regression and Decision Tree are easier to interpret but less powerful, with lower
accuracy, F1, AUC, and higher misclassification rates
Log(installment) 0.669
22
Result: The model using the raw installment variable achieves a slightly higher F1 score (0.672 vs.
0.669), indicating marginally better balance between precision and recall.
Interpretation: Log transformation does not improveand may even slightly reduce predictive
performance.
Variable Importance Remains Strong Without Log Transformation
In both models, sub_grade is the most important feature.
In the model using the raw installment variable, installment is the second most important
predictor.
Result: The model effectively uses the original (untransformed) variable, showing no loss of
predictive power.
Misclassification Rates and Classification Results Are Nearly Identical
Iteration plots for both models show similar misclassification rates as the number of trees
increases.
Confusion matrices are nearly identical, confirming that the use of the raw variable does not
negatively impact event classification.
23
6. Discussion
24
features, thereby obscuring fine-grained insights into specific customer subsegments.
25
one financial platform achieved an 80% increase in approval efficiency and a 35%
reduction in default rates after implementing the XGBoost model. These outcomes
underscore a critical insight: improvements in predictive performance delivered by
high-capacity models can be directly translated into measurable commercial value and
more resilient risk control mechanisms. Such benefits cannot be replicated by models
that offer interpretability alone but fall short in accuracy.
26
7. Conclusions and Contributions
27
Policy Recommendation
Given the high importance of sub_grade in both loan approval and term assignment, it is crucial
for policymakers and financial institutions to:
Ensure that all customers have fair access to higher credit grades, as these directly impact their
ability to secure favorable loan terms and, by extension, their long-term financial health and
inclusion.
28
References •
All factual material that is not original must be accompanied by a reference to its source. • Please
use the APA citation style. There are also citation management tools such as EndNote or Zotero •
Please contact SMU Libraries (library@smu , etc. .[Link]) if you need assistance with citation
management tools
29