0% found this document useful (0 votes)
11 views31 pages

Predictive Modeling for Loan Default Analysis

Class project about SAS Viya

Uploaded by

liuyunjiu96
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
11 views31 pages

Predictive Modeling for Loan Default Analysis

Class project about SAS Viya

Uploaded by

liuyunjiu96
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

ISSS602 Data Analytics Lab

Predictive Modelling for Loan Default


in LoanTap Using SAS Viya
Content
1. Introduction...............................................................................................3
2. Problem Statement and Objectives...................................................................3
2.1. Problem statement.........................................................................3
2.2. Objective.........................................................................................4
3. Literature Review........................................................................................4
3.1. Early Theories.................................................................................5
3.2. Theoretical Foundations of Financial Customer Segmentation and
Value 6
3.3. Extending the Theory of Enduring Customer Value.........................8
3.4. Applications of Machine Learning....................................................9
4. Data and Methods......................................................................................11
4.1. Data Preparation...........................................................................11
4.2. IDEA..............................................................................................13
4.3. Research Methods.........................................................................14
5. Analysis and Results •................................................................................16
5.1. Logistic Regression.......................................................................16
5.2. Decision Tree.................................................................................18
5.3. Random Forest..............................................................................19
5.4. Boosted Model..............................................................................20
5.5. Model comparison.........................................................................21
5.6. Standardization improvement.......................................................23
6. Discussion...............................................................................................25
6.1. Theoretical Limitations of Traditional Quantitative Models............25
6.2. Performance Breakthroughs of Boosted Tree Models....................26
6.3. The Strategic Necessity of Adopting Boosted Tree Models............26
7. Conclusions and Contributions......................................................................28
References......................................................................................................30
1. Introduction

As 2025 unfolds, financial institutions contend with persistent high interest rates (4.25%-4.50%),
stagflation Ary pressures (3% inflation, 1.4% GDP growth), tighter credit conditions, rising
delinquencies and charge-offs, declining small business lending, compressed net interest margins,
and increased competition from fintechs—all while navigating regulatory changes and economic
uncertainty that demand enhanced analytics, risk management, and modernization of lending
strategies.
Strategic Imperative for Customer Value Segmentation
Lenders must move beyond traditional risk-based models and adopt advanced analytics for
customer value segmentation—especially targeting high-value, long-term loan customers—to
optimize portfolios and drive sustainable growth, as research and industry practice show that
focusing on Customer Lifetime Value (CLV) enables banks to improve profitability, enhance
cross-sell opportunities, strengthen customer relationships, and achieve significant gains in
marketing effectiveness and retention, particularly in a challenging, high-rate, and competitive
lending environment
Environment: financial institutions must balance risk management with maximizing customer
lifetime value by deploying sophisticated segmentation strategies—such as segmenting borrowers
by loan term—to identify and target valuable customer segments, enabling tailored products,
optimized risk-adjusted returns, and stronger long-term relationships
Analytical: By leveraging advanced machine learning models—including logistic regression,
decision trees, random forests, and boosted forests—lenders can robustly identify the factors
driving long-term customer value and behavior, achieving both high predictive accuracy and
interpretability essential for regulatory compliance and effective stakeholder communication
Research Gap and Intended Outcomes
While current literature and industry practice primarily focus on default prediction and risk
management, they often neglect customer value segmentation based on loan term preferences,
creating a research gap that this project addresses by shifting from risk-only models to
comprehensive customer value identification through loan term prediction; the intended outcomes
are actionable insights for product design, risk management, and customer retention, as well as
empirical evidence from comparative machine learning analysis to support optimal, interpretable
customer segmentation strategies that enhance institutional capabilities and meet regulatory
requirements in the competitive digital lending landscape.

1
2. Problem Statement and Objectives

2.1. Problem statement

The problem is the inability of financial institutions to efficiently identify and segment valuable
customers based on their loan term preferences, resulting in suboptimal customer targeting,
resource allocation, and revenue optimization.
Challenges: Distinguishing between customers who prefer long-term versus short-term loan
products, which directly impacts their ability to develop targeted marketing strategies, optimize
product offerings, and maximize customer lifetime value. Current manual assessment processes
are time-consuming, inconsistent, and fail to leverage the full potential of available customer data
for predictive insights

2.2. Objective

To develop and compare multiple machine learning algorithms (Logistic Regression, Decision
Tree, Random Forest, and Boosted Forest) for predicting loan term preferences using customer
financial and demographic characteristics, thereby enabling financial institutions to identify
valuable customers and optimize loan product targeting strategies.
Model Development and Performance Optimization
Develop four distinct machine learning models to classify customers based on loan term
preferences with over 70% accuracy by implementing comprehensive data preprocessing—
including handling missing values, outlier detection, and feature scaling—and conducting rigorous
train-validation-test splits to ensure model generalizability and prevent overfitting.
Comparative Algorithm Analysis
Evaluate and compare the predictive performance of Logistic Regression, Decision Tree,
Random Forest, and Boosted Forest algorithms using standardized metrics (accuracy, precision,
recall, F1-score, and AUC-ROC), analyze their computational efficiency and training time to
assess practical feasibility, and determine the optimal algorithm by balancing predictive
accuracy, interpretability, and operational requirements.
Feature Importance and Variable Analysis
Identify and rank the relative importance of predictor variables—including financial capacity
indicators (annual income, debt-to-income ratio, installment amount, revolving balance), credit
profile metrics (grade, sub-grade, interest rate, revolving utilization, verification status), credit
history factors (public records, bankruptcy records, mortgage accounts, open accounts, total
accounts), and employment/housing stability (employment length, home ownership status, loan
purpose)—across all four modeling approaches; analyze interaction effects between predictors to
uncover synergistic relationships that enhance prediction accuracy; and validate feature
importance findings through cross-algorithm comparison to ensure robust and reliable variable
selection.

2
3. Literature Review

The quantification and management of customer value form the foundation of modern
marketing and financial strategy. Related theories have evolved from early models focused solely
on financial valuation to more sophisticated, multidimensional frameworks for managing
customer relationships.

3.1. Early Theories

Early foundational theories in customer value management explicitly conceptualized


customers as quantifiable corporate assets. Gupta, Lehmann, and Stuart (2004) laid the theoretical
groundwork by introducing the concept of Customer Lifetime Value (CLV), which estimates
customer equity by calculating the discounted sum of a customer’s future profitability. This
approach established a direct linkage between marketing metrics—such as customer retention
rates—and firm financial performance and shareholder value, providing a robust quantitative basis
for marketing decision-making.
However, the limitations of CLV models that rely solely on historical financial data soon
became apparent. Hwang, Jung, and Suh (2004) critiqued traditional lifetime value (LTV)
frameworks for neglecting future customer potential and churn risk. In response, they proposed a
more comprehensive three-dimensional model, incorporating current value, potential value, and
customer loyalty—the latter operationalized through defection probability. This enriched
framework marked a significant advancement by integrating risk into customer valuation, thereby
enabling more targeted and effective marketing strategies.
Building on this foundation, Kumar and Reinartz (2016) elevated customer value theory to a
strategic level. They emphasized that sustainable customer value arises from a dynamic balance
between value creation for customers and value extraction by the firm. Their framework extends
beyond the transactional scope of CLV to include non-transactional, engagement-driven
dimensions of customer value, such as Customer Referral Value (CRV) and Customer Knowledge
Value (CKV), derived from advocacy and knowledge sharing, respectively. This shift in
perspective reflects a broader transition from short-term revenue orientation to long-term
relationship-building grounded in comprehensive customer engagement.
Strategic practices of major financial institutions further validate and extend the application
of these theoretical frameworks. For instance, NatWest Group (2024) reports that the bank is
leveraging advanced data science models to enhance its long-term understanding of customer

3
behavior. Its methodological approach goes beyond traditional segmentation and valuation,
incorporating complex Expected Credit Loss (ECL) models and climate risk scenario analysis.
This reflects an integrated perspective in real-world business environments—one that
simultaneously considers customer value, credit risk, and macroeconomic conditions.
The existing literature adopts a wide range of research paradigms to investigate customer
value, spanning from quantitative modeling to conceptual integration. Each approach brings
distinct analytical priorities, contributing to a more nuanced and multidimensional understanding
of the construction. the following table is provided:
Table 3-1 Research Paradigms

Primary
Study Research Paradigm Description
Paradigm
Developed a mathematical model based on customer acquisition
Financial
rate, retention rate, profit margin, and discount rate to quantify
Valuation
Gupta et total customer equity. Empirically tested the model using public
Model &
al. (2004) financial data from five listed companies, demonstrating a strong
Empirical
correlation between model-based valuations and actual market
Analysis
capitalization.
Proposed a three-dimensional LTV model incorporating current
Theoretical
value, potential value, and customer loyalty. Applied the model to
Hwang et Model
real customer data from a major South Korean wireless telecom
al. (2004) Building &
company, validating its effectiveness in customer segmentation and
Case Study
high-value customer identification.
Conducted a systematic review and synthesis of literature in
Kumar & customer value, customer relationship management, and customer
Conceptual
Reinartz engagement. Constructed an integrative theoretical framework
Review
(2016) through critical analysis and conceptual innovation rather than
empirical testing.
Analyzed corporate reports and strategic documents to examine
how the institution applies data science to customer management
NatWest Corporate
and risk control. The report outlines the internal use of customer
Group Case
segmentation models, Expected Credit Loss (ECL) models, and
(2024) Analysis
climate risk assessments, showcasing the in-depth application of
theory in large financial institutions.
Accordingly, it is evident that the definition, measurement, and application of customer value
remain subjects of ongoing theoretical development and scholarly debate.

3.2. Theoretical Foundations of Financial Customer Segmentation and Value

Against the backdrop of digital transformation in the global financial services industry, the

4
sector is undergoing a fundamental strategic shift—from a product-centric to a customer-centric
orientation. Financial institutions are no longer satisfied with offering standardized credit
products; instead, they increasingly seek to deliver highly personalized services through data-
driven insights. Leading institutions have adopted predictive analytics to enable smarter lending
decisions (Verbraken, Verbeke, & Baesens, 2013). These applications extend beyond conventional
credit risk assessments to encompass refined customer management approaches informed by
behavioral and psychographic profiling (Kumar & Reinartz, 2018).
Accurately predicting customer preferences—such as optimal loan duration—requires a
nuanced understanding of customer heterogeneity. Modern financial customer segmentation has
moved beyond traditional demographic-based models and evolved into multidimensional
frameworks that integrate behavioral patterns, psychological traits, and lifecycle stages. Such
approaches offer a more comprehensive perspective on customer preferences and enable
institutions to design more targeted financial strategies.
Academic research in this domain generally approaches the topic from three major angles:
(1) Behavioral Segmentation
This approach classifies customers based on their actual interactions with financial
institutions. For example, credit card users can be segmented into transactors (who pay their
balances in full each month), revolvers (who pay only the minimum and carry forward balances),
and inactive users (Verbraken, Verbeke, & Baesens, 2013). Such segmentation reflects customers’
financial habits and risk preferences in a direct and observable manner.
(2) Psychographic Segmentation
This paradigm incorporates insights from behavioral economics into customer profiling.
Traits such as loss aversion or present bias significantly influence decision-making processes—for
instance, how individuals trade off short-term high interest against long-term low monthly
payments, thereby shaping their loan term preferences (Kahneman & Tversky, 1979).
(3) Lifecycle Segmentation
This method dynamically segments customers according to their life stages and associated
financial needs. Empirical cases show that age- and segment-based classifications—such as youth,
retail, and affluent clients—can effectively predict product preferences across life events such as
education, marriage, and retirement. These patterns are particularly informative in forecasting
demand for loan products of varying maturities (Verbraken, Verbeke, & Baesens, 2013).
The ultimate goal of customer segmentation is to enable optimal allocation of resources, and
the Customer Lifetime Value (CLV) model provides a scientific and quantitative tool for this
purpose. CLV estimates the total profit a customer is expected to generate over the entire duration
of the business relationship, allowing financial institutions to allocate marketing and service

5
resources more efficiently toward high-value clients. Empirical evidence shows significant
disparities in CLV across different customer segments: while the average CLV of a private
banking client may amount to several tens of thousands of dollars, that of a basic retail account
holder might be only a few hundred (Kumar & Reinartz, 2018).
Traditional CLV models often rely on historical transactional data, which can be limiting in
low-frequency, sparse-data contexts such as banking. To address this, researchers have developed
dynamic models based on stochastic processes, such as Markov chains, which simulate the
probabilities of customers transitioning across segments over time. These models offer more
robust forecasts of future value, especially for products with infrequent purchase cycles
(Verbraken, Verbeke, & Baesens, 2013).
From the perspective of personalization strategy, CLV analysis also serves as a direct
foundation for product customization. For customer groups with high predicted CLV, financial
institutions have stronger incentives to develop and recommend tailored loan offerings that match
specific preferences—such as loan maturity or interest rate structure—thereby enhancing customer
satisfaction and long-term loyalty (Kumar & Reinartz, 2018).
However, current research on customer value has largely overlooked a critical and observable
behavioral dimension: customer preferences for loan maturity. Existing studies tend to adopt a
supply-side perspective, focusing on how financial institutions shorten loan durations to mitigate
risk under asymmetric information. Yet few have investigated, from the demand side, the micro-
level factors driving customers’ deliberate choices of specific repayment horizons. This oversight
has led to an understanding of customers that remains largely superficial—limited to financial and
demographic attributes—without fully capturing the deeper behavioral signals embedded in
customer decisions.
Loan maturity choice is not an isolated financial decision. It can serve as a comprehensive
proxy variable, reflecting an individual’s risk tolerance, time discounting behavior, financial
planning ability, and even broader indicators of financial health. Therefore, incorporating maturity
preference as a distinct analytical dimension not only addresses a meaningful gap in the existing
literature but also deepens and strengthens the conceptual and practical foundations of customer
value theory.

3.3. Extending the Theory of Enduring Customer Value

The theory of Enduring Customer Value proposed by Kumar and Reinartz (2016) emphasizes
a strategic shift from focusing solely on transactional contributions—such as Customer Lifetime
Value (CLV)—to building long-term relationships based on comprehensive customer engagement.

6
In this context, a customer’s loan maturity preference serves as a valuable leading indicator for
anticipating relationship longevity and estimating the customer’s potential enduring value.
When analyzed from a preference-based perspective, the choice of loan maturity may signal
divergent customer trajectories and risk profiles. Short-term preferences often correspond to
transaction-oriented, interest-sensitive customer segments that exhibit lower loyalty and higher
attrition risk. In contrast, long-term preferences may reflect a greater willingness to establish
durable relationships and a higher degree of trust in the financial institution. Such customers are
more likely to generate stable future cash flows, pose lower default risk, and contribute additional
value through cross-selling opportunities and customer referrals (CRV), thereby demonstrating
higher potential for enduring customer value.
Incorporating maturity preference as a key variable—or even as a risk-weighted input—into
CLV prediction models allows for more precise calibration of risk expectations and significantly
enhances both the accuracy and robustness of long-term value estimation.
Understanding maturity preference enables financial institutions to move from passive
accommodation to proactive guidance. For high-potential customers demonstrating long-term
preferences, institutions can design and recommend product bundles that align with those
preferences, even offering more flexible repayment options to reinforce loyalty through value
creation. For customers favoring short-term arrangements, targeted communication strategies
aimed at improving financial literacy and loyalty may be more appropriate. This preference-
informed differentiation forms the core mechanism for striking a dynamic balance between value
creation for the customer and value extraction by the firm.
In summary, systematically examining loan maturity preferences and integrating them into
multidimensional segmentation and valuation frameworks represents a necessary step in
translating customer value theory from conceptual abstraction to operational precision. Future
research could explore the use of advanced and interpretable machine learning models—such as
gradient boosting trees combined with SHAP (SHapley Additive exPlanations)—to both predict
maturity preferences with high accuracy and identify their key behavioral and financial drivers.
This would provide a scientifically grounded path toward truly customer-centric risk pricing,
product innovation, and value management in the financial services sector.

3.4. Applications of Machine Learning

In the early stages of applying machine learning to financial prediction, classical models such
as logistic regression, decision trees, and random forests played a foundational role in credit risk
assessment. These models were particularly effective in forecasting default risk at a macro level

7
and provided interpretable, rule-based decision structures.
However, as the focus of prediction has shifted from broad risk estimation to understanding
granular customer preferences—such as loan maturity choices—the limitations of these models
have become increasingly apparent. Their relatively rigid functional forms and limited capacity to
capture complex, nonlinear interactions make them less suitable for preference prediction in
heterogeneous customer bases.
To clarify the relative strengths and weaknesses of commonly used models in this context,
the following table provides a comparative summary of their characteristics:
Table 3-2 Model Introduction

Model Advantages Limitations and Drawbacks Reference


Relies on the strong assumption of a
Simple model structure;
linear relationship between features
coefficients have clear
Logistic and outcomes; unable to capture Siddiqi
business interpretation;
Regression complex nonlinear patterns common (2017)
highly interpretable and
in financial data, limiting predictive
computationally efficient.
accuracy.
Intuitive model form; Highly sensitive to small variations in
decision paths are data; lacks stability; prone to
Decision Breiman et
transparent and align with overfitting when used alone; struggles
Tree al. (1984)
human reasoning; easy to to account for complex feature
understand. interactions.
As an ensemble method,
builds multiple trees and Offers improved predictive accuracy
averages results, over single models but sacrifices
Random Breiman
significantly reducing interpretability; often regarded as a
Forest (2001)
overfitting risk; robust "black-box," making decision
performance on high- rationales harder to explain.
dimensional data.
Customer preferences for loan maturity do not arise from a simple linear decision process.
Rather, they result from a complex, nonlinear interplay of financial circumstances, risk
perceptions, future expectations, and psychological traits. For instance, income level may not
exhibit a strictly positive or negative correlation with loan term; instead, behavioral turning points
may emerge at certain income thresholds. Similarly, the effects of age and occupation on maturity
preferences are likely intertwined rather than independently linear.
Traditional models often struggle with such complexity. Logistic regression, for example,
relies on the assumption of linear relationships and thus cannot adequately capture these nuanced
nonlinearities. Single decision trees, while interpretable, are prone to overfitting and lack
generalizability to unseen data. Although random forests improve predictive stability through

8
ensemble averaging, their majority-vote mechanism may dilute the influence of critical but rare
patterns, and their "black box" nature poses challenges for business interpretability and regulatory
compliance.
To overcome these limitations—particularly in predictive accuracy, nonlinear pattern
recognition, and complex feature interaction modeling—we turn to boosting models based on
gradient boosting algorithms. Empirical research has consistently demonstrated the superior
performance of boosted tree models in financial risk assessment. For example, in real-world credit
card default prediction tasks, XGBoost has achieved accuracy rates exceeding 99%, significantly
outperforming traditional logistic regression and even neural networks (Chen & Guestrin, 2016).
LightGBM, with innovations such as Gradient-based One-Side Sampling (GOSS) and Exclusive
Feature Bundling (EFB), maintains high accuracy while substantially reducing computational cost
and memory usage, making it especially suitable for large-scale financial datasets (Ke et al.,
2017).
Loan maturity preference is a prototypical multi-factor, nonlinear decision problem. Factors
such as income, age, credit history, and debt burden interact in non-additive and often
unpredictable ways. Boosting models are well-suited to this context, as they can automatically
learn high-order feature interactions without the need for manual feature engineering. This
capacity makes them an ideal solution for high-precision modeling of customer preferences,
enabling financial institutions to uncover subtle behavioral drivers and support the development of
truly personalized financial products.

4. Data and Methods

4.1. Data Preparation

Import the data table Loan_Tap_Data.csv into SAS Viya.


Target Variable:
term_binary is correctly created: 0 for 36 months (short term), 1 for 60 months (long term).
Variable Conversion:
All categorical variables (e.g. emp_length, grade sub_grade home_ownership, verification_status
purpose) are numerically encoded for modeling.

9
Table 4-3 Data Preparation Table

Step Description

1 Import the Loan_Tap dataset and perform an initial data quality check.

Target Variable(Figure 4-1): Convert ‘term’ to numeric (remove 'months');


2 create a binary variable (term_binary: 36=0, 60=1).

Numeric Convert :1. employment length to numeric (e.g., '< 1 year'=0, '10+
years'=10).2. Sequential: grade and subgrade as numeric codes 3. Map home
ownership, verification status, and purpose categories to numeric codes; group
rare purposes as 0.4. Binary: Convert loan status to numeric (Fully Paid=0,
3 Charged Off=1, others as missing)(Figure 4-2)

4 handle missing values for all variables appropriately.

5 IDEA: Check variable distributions and ensure consistency.

6 Split data into train/test sets, create the final model-ready dataset,

7 Identify and treat outliers using the IQR method or capping.

export to SAS Viya Model Studio, document all transformations, and automate
8 the data preparation pipeline to reduce errors and ensure consistency

Figure 4-1 Target Variable Convert

10
Figure 4-2 Numeric Conver

4.2. IDEA

Figure 4-3 Term_binary

Target Variable: term_binary is correctly created: 0 for 36 months (short term), 1 for 60 months
(long term).
This indicates a strong class imbalance, with 36-month loans being about three times more
common than 60-month loans

11
Figure 4-4 Category variable

Category variable: The frequency distribution of both grade and sub_grade is skewed, with most
loans concentrated in the middle categories and fewer loans in the highest and lowest categories.
The grade variable shows the highest frequencies in the middle grades (around 3 and 4), while
both lower and higher grades are less common.
The sub_grade variable displays a similar pattern, peaking near the middle sub_grades and
tapering off at both extremes, indicating a strong central tendency in loan quality distribution after
clarifying.

Figure 4-5 Continuous variable

Continuous variable: The distributions of installment amount, interest rate, number of mortgage
accounts, and number of open accounts are all right-skewed, with most values concentrated at the
lower end of their ranges.

12
Each variable shows a clear central tendency, with frequencies peaking at specific values and
tapering off as values increase, indicating that most borrowers have moderate installment amounts,
interest rates, and account numbers.
There are visible outliers and long tails in all distributions, highlighting the importance of outlier
detection and treatment during data preparation to ensure robust modeling.
which research methods will be used • describing the data including the data source

4.3. Research Methods

This project aims to predict whether a customer will select a long-term or short-term loan, using
advanced classification models and focusing on maximizing the F1 score for model evaluation.
The F1 score is chosen as the primary metric because it balances precision and recall, providing a
fair assessment of model performance, especially when both types of classification errors (false
positives and false negatives) are important to minimize

Table 4-4 Research Methods

Model Strengths Weaknesses

- Assumes linear relationships


- Highly interpretable - Sensitive to multicollinearity
- Efficient for linearly separable data - Limited for non-linear patterns
Logistic - Well-calibrated probabilities - May underperform with complex
Regression - Simple to implement data

- Easy to interpret and visualize - Prone to overfitting


- Handles categorical & numerical data - Sensitive to small data changes
Decision - Models non-linear relationships - Large trees can be hard to interpret
Tree - Robust to outliers - May be biased to dominant classes

- Reduces overfitting vs. single trees - Less interpretable ("black box")


- Handles non-linear relationships - Computationally intensive
Random - Robust to outliers & noise - Limited transparency for individual
Forest - Implicit feature selection predictors

- High predictive accuracy


- Handles complex, non-linear - Least interpretable
relationships - Sensitive to noise/outliers
Boosted - Focuses on correcting errors - Requires careful tuning
Forest - Can optimize for specific metrics - Computationally demanding

Data Description
Data Source:

13
The dataset is sourced from LoanTap’s lending data, containing customer information, loan terms,
employment details, credit grades, home ownership, verification status, loan purpose, and loan
status. All variables are preprocessed to numeric or binary format, as described in the data
preparation section.

Key Variables:

Target Variable: (0 = short-term, 1 = long-term)term_binary

Predictors: Employment length, grade, sub-grade, home ownership, verification status, loan
purpose, and other demographic or behavioral features. Loan amount is excluded as per project
scope.

14
5. Analysis and Results •

The analysis section is concerned with analyzing the results you have achieved through the data
analytics project. • Results are actual statements of observations, including statistics, tables and
graphs. • You should present both negative and positive results with sufficient details so that
readers can draw their own inferences and construct their own explanations.
Four supervised machine learning algorithms are used for binary classification:

5.1. Logistic Regression

Consistent Model Performance Across Data Splits


The logistic regression model demonstrates highly stable and consistent classification results. The
proportion of correctly classified events (both 0 and 1) remains stable in all three partitions. This
indicates strong generalization capability and suggests the model is neither overfitting nor
underfitting the data. Consistency across data splits is a key indicator of a reliable model that is
likely to perform well on new, unseen data.

Figure 5-6 KS Cutoff

Optimal Cutoff Value (KS Cutoff) and Improved Event Detection


The ROC analysis identifies an optimal cutoff value of 0.24 (Figure 5-6), where the difference
between sensitivity (true positive rate) and 1-specificity (false positive rate) is maximized.
Selecting this threshold, rather than the default 0.5, allows the model to achieve a better balance
between capturing true positives and minimizing false positives, thus improving event detection,

15
especially in cases where the costs of misclassification are asymmetric. The KS statistics are a
widely used metric for evaluating the discriminatory power of classification models and is
particularly valuable in credit scoring and risk modeling.

Figure 5-7 Accuracy

High Accuracy, Balanced Classification, and Strong Lift/Gain


At the default cutoff of 0.5, the model achieves an accuracy of 81.8% across all data
partitions(Figure 5-7), confirming that over four out of five predictions are correct. This accuracy
is complemented by balanced classification rates for both event classes, as visualized in the
percentage plot. Additionally, the model shows a lift of 3.52–3.54 in the top 5% quantile and a
gain of 2.2 in the top 10% quantile across all partitions. This means the model is more than three
times as effective as random selection at identifying positive events in the highest-probability
group, which is highly valuable for targeted interventions such as marketing or risk mitigation

16
5.2. Decision Tree

Figure 5-8Accuracy

High and Consistent Accuracy Across Data Splits:


The decision tree achieved an accuracy of 82.4% at the standard cutoff of 0.5, demonstrating
robust and consistent predictive performance across all data subsets

Figure 5-9 F1-score

Optimal Cutoff Improves Sensitivity and key indicator:


The KS (Kolmogorov-Smirnov) cutoff for the VALIDATE partition is 0.22, where the sensitivity
reaches 0.8 and the misclassification rate for events is reduced, indicating that adjusting the cutoff
can significantly improve the model’s ability to identify true positives.

17
The most important variables for classification are (by far the most influential)。
Low and Stable Misclassification Rate:
The misclassification rate is 17.6% across all data partitions, indicating that the tree is well-pruned
and not overfitting. This is supported by the pruning error plot, which shows that the selected
subtree with 36 leaves achieves the lowest validation error without unnecessary complexity

5.3. Random Forest

Figure 5-10 F1 score

Highest and Most Consistent Predictive Performance


The Random Forest model achieves the highest accuracy (85.6%) and F1 score (0.653) on the
TEST partition, outperforming both logistic regression and decision tree models. This superior
performance is consistent across TRAIN, VALIDATE, and TEST datasets, indicating strong
generalization and minimal overfitting due to the ensemble approach.
Superior Lift and Concentration in Top Quantiles
In the top 5% of predicted probabilities, Random Forest attains a lift of 3.92 and a response rate of
92.9% in the TEST partition, exceeding the performance of the other models in identifying
positive events among the highest-risk group. This makes it particularly effective for targeted
interventions where capturing the most likely positives is critical.
Broader Feature Utilization and Flexible Cutoff Selection
While sub_grade remains the most important predictor, Random Forest leverages a broader set of
features—including installment, grade, int_rate, and verification_status—enabling it to model
complex, nonlinear relationships. The model’s optimal cutoff (KS = 0.29) further enhances
sensitivity (82.9% in VALIDATE), allowing for data-driven threshold adjustments to meet specific
business objectives.

18
5.4. Boosted Model

Figure 5-11 F1 score

Best-In-Class Predictive Accuracy and F1 Score


The Gradient Boosting model achieves the highest accuracy (87%) and F1 score (0.698)Figure 5-
11 on the TEST partition, outperforming logistic regression, decision tree, and random forest
models. This demonstrates not only excellent overall prediction but also a strong balance between
precision and recall, especially for the positive class.

Figure 5-12 Lift

Exceptional Lift and Targeting in Top Quantiles


In the top 5% quantile, Gradient Boosting delivers a lift of 3.95 and a response rate of 93.8% in
the TEST partition. This means the model is nearly four times more effective than random
selection at identifying positive events in the highest-risk group, slightly surpassing the random
forest model. This makes Gradient Boosting especially valuable for applications focused on
targeting high-value cases.

19
Robust Generalization and Nuanced Feature Importance
The model maintains a low and stable misclassification rate (13%) across TRAIN, VALIDATE,
and TEST partitions, indicating strong generalization and minimal overfitting. Gradient Boosting
also provides more nuanced feature importance scores, with sub_grade as the most influential,
followed by installment, mort_acc, and annual_inc. This allows for finer feature selection and
deeper interpretability compared to other models.

5.5. Model comparison

Predictive Performance (Accuracy & F1 Score)


Gradient Boosting achieves the highest accuracy and F1 score, confirming its superior ability to
balance precision and recall, especially for positive cases. This aligns with published research,
where Gradient Boosted Trees often outperform other models on both metrics.
Random Forest also performs strongly, notably better than single-tree models, due to its ensemble
approach and reduction in overfitting.
Logistic Regression and Decision Tree are more interpretable but lag in predictive power

20
Discrimination and Ranking (AUC & Lift)
Gradient Boosting and Random Forest both have high AUC and lift, indicating excellent ability to
rank and separate positive cases from negatives.
Gradient Boosting slightly edges out Random Forest in lift and AUC, making it better for
prioritizing top candidates for action

Robustness and Generalization (Misclassification Rate & Stability)


Gradient Boosting and Random Forest have the lowest misclassification rates, reflecting better
generalization to unseen data and less overfitting.
Logistic Regression and Decision Tree are more prone to underfitting or overfitting, as shown by
higher misclassification rates and less stable performance

21
Key Results:
Gradient Boosting is the top performer for both accuracy (87.0%) and F1 score (0.70), with the
highest AUC (0.93) and lift (3.69), and the lowest misclassification rate (13.0%)—making it the
best model for binary classification on this dataset.
Random Forest offers a strong balance of performance (accuracy 85.6%, F1 0.65), discrimination
(AUC 0.91, lift 3.62), and robustness (misclassification 14.4%), with lower risk of overfitting than
single trees.
Logistic Regression and Decision Tree are easier to interpret but less powerful, with lower
accuracy, F1, AUC, and higher misclassification rates

5.6. Standardization improvement

Predictive Performance is Equal or Slightly Better Without Log Transformation

Model Version F1 Score

Raw installment 0.672

Log(installment) 0.669

22
Result: The model using the raw installment variable achieves a slightly higher F1 score (0.672 vs.
0.669), indicating marginally better balance between precision and recall.
Interpretation: Log transformation does not improveand may even slightly reduce predictive
performance.
Variable Importance Remains Strong Without Log Transformation
In both models, sub_grade is the most important feature.
In the model using the raw installment variable, installment is the second most important
predictor.
Result: The model effectively uses the original (untransformed) variable, showing no loss of
predictive power.
Misclassification Rates and Classification Results Are Nearly Identical
Iteration plots for both models show similar misclassification rates as the number of trees
increases.
Confusion matrices are nearly identical, confirming that the use of the raw variable does not
negatively impact event classification.

23
6. Discussion

6.1. Theoretical Limitations of Traditional Quantitative Models

The academic literature has systematically identified the inherent limitations of


traditional quantitative models in addressing the complexity of modern financial data.
These limitations provide a solid theoretical basis for this study’s adoption of more
advanced modeling approaches.
As noted by Siddiqi (2017), generalized linear models—particularly logistic
regression—have long served as the foundation of traditional credit scoring systems.
Their strength lies in the interpretability of model coefficients, which allows for
intuitive business insights. However, their fundamental weakness is the strict
assumption of linearity between independent variables and outcomes. This rigid
structure prevents them from capturing the nonlinear patterns frequently observed in
customer financial behavior. For example, the relationship between customer income
and loan maturity preference may exhibit a “turning point” rather than a
straightforward correlation: low- to middle-income borrowers may prefer longer loan
terms to reduce monthly installments, whereas high-income individuals might opt for
shorter terms due to greater repayment capacity. Logistic regression is ill-equipped to
model such non-monotonic relationships, thereby imposing a hard ceiling on its
predictive performance.
Besides, although decision trees offer transparent, rule-based decision paths that
align with human reasoning (Breiman et al., 1984), they suffer from high variance.
Even small changes in the training data can drastically alter the model structure,
making decision trees unstable and prone to overfitting. As a result, their ability to
generalize to unseen data is often severely limited.
To overcome the deficiencies of single decision trees, ensemble methods such as
random forests were developed (Breiman, 2001). By constructing and aggregating
multiple decision trees through a majority-voting mechanism, random forests
significantly improve robustness and predictive accuracy. However, this improvement
comes at the cost of model transparency. Random forests are frequently criticized as
“black-box” models, with decision logic that is difficult to interpret. More
importantly, their averaging mechanism may suppress the impact of rare but critical

24
features, thereby obscuring fine-grained insights into specific customer subsegments.

6.2. Performance Breakthroughs of Boosted Tree Models

In stark contrast to the limitations of traditional models, the empirical findings of


this study align closely with a growing body of academic literature and industry
practice: boosted tree models such as XGBoost and LightGBM have achieved a
qualitative breakthrough in predictive performance, effectively overcoming the
inherent shortcomings of earlier approaches.
Our study provides strong empirical support for this superiority. In a credit card
churn prediction task, the ensemble tree model achieved an Area Under the Curve
(AUC) of 0.951—significantly outperforming logistic regression, which yielded an
AUC of 0.885. In specific financial datasets, the accuracy of the XGBoost model
remained consistently around 95%. Moreover, as highlighted by Ke et al. (2017),
LightGBM not only demonstrated excellent AUC performance (reaching 78.9%) but
also offered substantial advantages in training speed and computational efficiency.
These characteristics make it especially well-suited for financial applications, where
both predictive accuracy and rapid model iteration are critical.
The fundamental reason behind the superior performance of boosted tree models
lies in their ability to automatically learn and model nonlinear relationships and high-
order feature interactions. Unlike linear models, which rely heavily on manual feature
engineering, boosting algorithms operate sequentially—each new tree is constructed
to correct the residual errors of the previous ensemble (Chen & Guestrin, 2016). This
iterative optimization mechanism allows the model to approximate complex real-
world data distributions more closely and to capture the subtle patterns that drive
customer preferences and risk behaviors—patterns that traditional models often fail to
detect.

6.3. The Strategic Necessity of Adopting Boosted Tree Models

In high-stakes decision-making contexts such as financial risk management,


credit scoring, and customer value assessment, predictive accuracy is not merely a
technical metric—it is a core business imperative that directly affects institutional
viability and growth.
Consequently, selecting a model with superior predictive power is not a matter of
preference but a business-driven necessity. As demonstrated in an industry case study,

25
one financial platform achieved an 80% increase in approval efficiency and a 35%
reduction in default rates after implementing the XGBoost model. These outcomes
underscore a critical insight: improvements in predictive performance delivered by
high-capacity models can be directly translated into measurable commercial value and
more resilient risk control mechanisms. Such benefits cannot be replicated by models
that offer interpretability alone but fall short in accuracy.

26
7. Conclusions and Contributions

Improving Processes in Financial Analysis


This project leverages advanced machine learning(Gradient Boosting models)to enhance
financial analysis by accurately predicting customer loan term preferences. Unlike traditional risk-
focused models, this approach enables financial institutions to segment customers based on their
likely loan term choices (short-term vs. long-term), supporting more targeted product offerings,
optimized marketing strategies, and improved customer retention. The empirical results show that
Gradient Boosting outperforms traditional models (logistic regression, decision trees, random
forests) in both predictive accuracy (87%) and F1 score (0.70), making it a best-in-class solution
for binary classification tasks in lending.
Informing Policy Objectives
A key finding is the central role of the sub_grade variable—not only in loan approval decisions
but also in predicting loan term preference. This dual importance highlights the need for fair and
transparent access to higher credit grades. If sub_grade is disproportionately influencing both
approval and term assignment, there is a risk of systemic bias that could limit fair access to long-
term credit products for certain customer groups. Policy recommendations should therefore focus
on:
Ensuring that credit grading criteria are transparent and free from bias.
Regularly auditing grading models for disparate impact.
Promoting fair access to higher credit grades, which are closely tied to favorable loan terms and
broader financial inclusion.
Strengthening Theory: Customer Value and Segmentation
From a theoretical perspective, this project advances the literature by integrating customer loan
term preference into customer value segmentation frameworks. Traditional models often overlook
behavioral signals like loan maturity choice, focusing instead on demographic or financial
variables. By demonstrating that loan term preference is a powerful proxy for customer lifetime
value, risk tolerance, and relationship longevity, this research deepens and operationalizes the
theory of Enduring Customer Value.
The use of Gradient Boosting models further strengthens the theoretical foundation by capturing
complex, nonlinear interactions among features—something earlier models could not achieve.
This aligns with contemporary theories that emphasize the multidimensional and dynamic nature
of customer value, moving beyond static, transaction-based metrics to include behavioral and
engagement-driven dimensions.
Unique Perspective and Practical Implications
The project’s machine learning approach provides a unique, data-driven perspective on how loan
term preferences are shaped by a mix of financial, demographic, and behavioral factors.
The findings support more personalized and equitable lending practices, as well as more precise
risk pricing and product design.
By identifying sub_grade as a key driver of both approval and term, the research underscores the
importance of fair grading systems for customer analysis and long-term value creation.

27
Policy Recommendation
Given the high importance of sub_grade in both loan approval and term assignment, it is crucial
for policymakers and financial institutions to:
Ensure that all customers have fair access to higher credit grades, as these directly impact their
ability to secure favorable loan terms and, by extension, their long-term financial health and
inclusion.

28
References •
All factual material that is not original must be accompanied by a reference to its source. • Please
use the APA citation style. There are also citation management tools such as EndNote or Zotero •
Please contact SMU Libraries (library@smu , etc. .[Link]) if you need assistance with citation
management tools

29

You might also like