0% found this document useful (0 votes)
4 views8 pages

Enhancing Credit Risk Reports Generation Using LLMs

This document presents a novel approach called Labeled Guide Prompting (LGP) to enhance the generation of credit risk reports using Large Language Models (LLMs), specifically GPT-4. The study demonstrates that LGP significantly improves the insightfulness and quality of reports, leading to a preference for LLM-generated reports over those created by human analysts in 60-90% of cases. By integrating Bayesian networks and promoting abductive reasoning, this method aims to streamline credit risk analysis in the finance industry, offering valuable insights for AI integration in credit risk management.

Uploaded by

fenisha.glsbca21
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views8 pages

Enhancing Credit Risk Reports Generation Using LLMs

This document presents a novel approach called Labeled Guide Prompting (LGP) to enhance the generation of credit risk reports using Large Language Models (LLMs), specifically GPT-4. The study demonstrates that LGP significantly improves the insightfulness and quality of reports, leading to a preference for LLM-generated reports over those created by human analysts in 60-90% of cases. By integrating Bayesian networks and promoting abductive reasoning, this method aims to streamline credit risk analysis in the finance industry, offering valuable insights for AI integration in credit risk management.

Uploaded by

fenisha.glsbca21
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Enhancing Credit Risk Reports Generation using LLMs: An

Integration of Bayesian Networks and Labeled Guide Prompting


Ana Clara Teixeira∗ Vaishali Marar∗ Hamed Yazdanpanah
[Link]@[Link] [Link]@[Link] [Link]@[Link]
Traive Inc. Traive Inc. Traive Inc.
Brookline, Massachusetts, USA Brookline, Massachusetts, USA Brookline, Massachusetts, USA

Aline Oliveira Mohammad M. Ghassemi


[Link]@[Link] [Link]@[Link]
Traive Inc. Traive Inc.
Brookline, Massachusetts, USA Brookline, Massachusetts, USA
ABSTRACT Guide Prompting. In Proceedings of 4th ACM International Conference on AI
Credit risk analysis is a process that involves a wide range of com- in Finance (ICAIF ’23). ACM, New York, NY, USA, 8 pages. [Link]
[Link]
plex cognitive abilities. Automating the credit risk analysis process
using Large Language Models can bring transformative changes
to the finance industry, but not without appropriate measures to 1 INTRODUCTION
ensure trustworthy responses. In this work, we propose a novel Credit risk analysis is a complex process that involves a wide range
prompt-engineering method that enhances the ability of Large of abilities, including contextual understanding, logical reasoning,
Language Models to generate reliable credit risk reports - Labeled the application of domain-specific knowledge, and implicit and
Guide Prompting (LGP). LGP consists of: (1) providing annotated causal reasoning. Evaluating the performance of Large Language
few-shot examples to the LLM that denote sets of tokens in an Models (LLMs) in credit risk assessment provides crucial insights
exemplary prompt that are of greater importance when generating into their practical utility in real-world scenarios. This work intro-
sets of tokens in the exemplary response and (2) providing text in duces the application of LLMs for generating comprehensive credit
the prompt that describes the direction, pathways and interactions risk reports - a critical task in financial decision-making. More
between variables from a Bayesian network used for credit risk specifically, we propose a novel prompt-engineering approach de-
assessment, thus promoting abductive reasoning. Using data from signed to enhance the quality and fidelity of credit risk assessments;
100 credit applications, we demonstrate that LGP enables LLMs we compare the credit risk assessments of the LLM against human
to generate credit risk reports that are preferred by human credit analysts through a user-centered, human-based evaluation, demon-
analysts (in 60-90% of cases) over alternative credit risk reports strating the proposed procedure’s efficacy in dealing with the credit
created by their peers in a blind review. Additionally, we found risk assessment task.
a statistically significant improvement (p-value < 10 −10 ) in the Automating the credit risk analysis process can bring transfor-
insightfulness of the responses generated using LGP when com- mative changes to the finance industry. By leveraging LLMs, we
pared to identical prompts without LGP components. We conclude can streamline processing vast amounts of data, enabling real-time
that Labeled Guide Prompting can enhance LLM performance in analysis, surpassing traditional methods. Furthermore, automation
complex problem-solving tasks, achieving a level of competency can also contribute to more consistent and objective analyses; while
comparable to or exceeding human experts. human analysts might be influenced by biases or varying expertise
levels, using LLMs for credit risk assessment can maintain a stan-
KEYWORDS dard level of analysis based on its prompt engineering infrastructure.
GPT-4, prompt engineering, credit risk report, Bayesian network, Thus, the importance of automated credit risk assessment is due to
labeled guide prompting the growing influence of automation and artificial intelligence in
ACM Reference Format: the finance sector. By evaluating the capabilities of LLMs for credit
Ana Clara Teixeira, Vaishali Marar, Hamed Yazdanpanah, Aline Oliveira, risk analysis, we offer valuable insights to inform AI integration
and Mohammad M. Ghassemi. 2018. Enhancing Credit Risk Reports Gen- in credit risk management. Considering the enormous volume of
eration using LLMs: An Integration of Bayesian Networks and Labeled credit decisions made regularly, even marginal enhancements in
LLM performance can translate into significant efficiency gains and
Permission to make digital or hard copies of all or part of this work for personal or
classroom use is granted without fee provided that copies are not made or distributed better decision accuracy.
for profit or commercial advantage and that copies bear this notice and the full citation The task of using GPT-4 for credit risk analysis is highly com-
on the first page. Copyrights for components of this work owned by others than ACM plex due to several factors: (1) Credit risk analysis is inherently
must be honored. Abstracting with credit is permitted. To copy otherwise, or republish,
to post on servers or to redistribute to lists, requires prior specific permission and/or a multifaceted, requiring the consideration of numerous variables
fee. Request permissions from permissions@[Link]. ranging from personal credit history and current financial status, to
ICAIF ’23, November 27–29, 2023, New York, NY broader economic conditions and industry-specific factors; (2) the
© 2018 Association for Computing Machinery.
ACM ISBN 978-1-4503-XXXX-X/18/06. . . $15.00 dynamic nature of financial data, which is characterized by contin-
[Link] uous change and occasional volatility. This dynamic nature makes
ICAIF ’23, November 27–29, 2023, New York, NY Ana Clara Teixeira, Vaishali Marar, Hamed Yazdanpanah, Aline Oliveira, and Mohammad M. Ghassemi

interactions among variables more relevant and leads to evolving analysis [19]. Diverging from this, our work uniquely applies GPT-4
relationships between features and credit risk, and (3) there are to individualized credit risk assessments, focusing on the evaluation
underlying difficulties in utilizing a language model like GPT-4 of specific borrowers.
for tasks traditionally performed by credit analysts, professionals
with a deep understanding of financial dynamics and risk. These 1.1.2 Prompt engineering. Prompt engineering and few-shot learn-
issues require bridging the gap between generalized data analysis, ing have become crucial components of effective utilization of LLMs,
a strength of GPT-4, and credit analysts’ specialized knowledge and guiding these models towards more desirable and task-specific out-
intuition, adding another layer of complexity. Generating compre- puts [3, 18]. Prompt engineering involves crafting carefully struc-
hensive credit risk reports tests the limits of current LLMs such tured inputs that elucidate the context and desired outcome of a
as GPT-4, combining domain knowledge, contextual understand- task for the model, proving instrumental in narrowing down the
ing, and multiple types of reasoning. While some elements may model’s broad knowledge to a specific task [17]. On the other hand,
be solvable using existing language models, the comprehensive few-shot learning involves presenting the LLM with a small number
nature of the task is more demanding. The effectiveness of GPT-4 in of examples of a task, thereby assisting the model in identifying
generating credit risk reports depends on its capacity to grasp and and applying the correct pattern for the given task. Our work builds
incorporate this nuanced knowledge and intuition, which requires on these methods to develop a prompt engineering procedure for
a sophisticated prompt engineering strategy. credit risk assessment with GPT-4.
One of the main limitations in the utilization of GPT and cur- 1.1.3 Task definition and benchmarks. Formulating tasks for LLMs
rent prompt engineering strategies is the unpredictable nature of and evaluating their performance are two fundamental steps in
responses when handling unseen data, anomalies, or changing re- applying these models to real-world scenarios. Task formulation
quirements. This unpredictability is a significant concern in finan- involves defining a clear objective, often expressed as a question or
cial decision-making, where consistent quality of response, insight, statement for the model to complete, expand, or react [10].
and analysis is essential. Evaluation, meanwhile, usually entails matching the model’s
In this paper, we propose a unique prompt engineering strategy output against a standard or specific criteria [7]. Widely accepted
that standardizes the content and quality of GPT-4’s output, making benchmarks such as GLUE [15], SuperGLUE [16] and BIG-bench [11,
it more predictable and insightful, even when handling complex, 12] serve to evaluate the model’s competency across various areas.
dynamic tasks like credit risk analysis. The unique strength of this For our study, we design a task for GPT-4 focused on credit risk
approach lies in its potential to act as a ’missing piece’ in prompt en- assessment and measure its performance using well-defined criteria
gineering; it changes how we interact with LLMs, by ensuring that specific to credit risk evaluation.
the quality of insights and analyses remains consistent, even when
dealing with complex, dynamic tasks like credit risk analysis. Mov- 1.1.4 Abductive reasoning with LLM. Abductive reasoning involves
ing beyond traditional applications of LLMs, our research exposes generating the most plausible explanation for a given set of obser-
these models to the financial industry’s requirements, successfully vations and aligns well with the prediction-focused nature of many
meeting the pragmatic needs of credit analysis. AI applications. By enabling LLMs to not only predict but also to
generate plausible explanations, we can enhance the interpretability
of these models [6, 9].
1.1 Related Work In the context of our work, leveraging abductive reasoning allows
1.1.1 LLM milestones. In recent years, language model develop- the model to generate more informative and nuanced credit reports.
ment has seen substantial advancements in architecture, training For example, it can create insights into how different financial indi-
methods, and real-world applications. The Transformer model pi- cators might interact to influence an individual’s credit risk score.
oneered the use of self-attention mechanisms, leading to signif- This makes the assessment process more transparent and under-
icant improvements in many natural language processing (NLP) standable, a significant advantage in in fields where interpretability
tasks [14]. Then, BERT revolutionized language understanding by is crucial.
training bidirectional transformers, allowing models to understand We propose a novel prompt engineering technique, Labeled Guide
the context of a word based on all of its surroundings (left and right Prompting, designed to aid the LLM in responding to multiple com-
of the word) [4]. plex dimensions of a problem such as the "what?", "why?" and
The advent of GPT-3 introduced scaling laws, which confirms "how?". This method ensures that the response requirements of
that as models increase in size, they continue to improve in perfor- a task, however complex are strictly followed by GPT, as the la-
mance [1, 2, 13]. It also showcased impressive few-shot learning beled guide assists the LLM in answering all concepts and putting
capabilities, enabling the model to generalize from a small num- equal thought into each by returning a standard response given a
ber of examples. Our work builds upon the GPT-4 model, which prompt. By combining this method with Bayesian networks rep-
extends these principles and represents the latest milestone in this resentation, we promotes abductive reasoning in Large Language
progression. Models (LLMs). This approach makes the model to hypothesize or
LLMs have demonstrated competitive performance in machine generate the "best guesses" that explain the observed phenomena,
translation [5], summarization [21], dialogue generation [20], and according to the network structure.
writing human-like essays [8]. BloombergGPT is an example of Explicit labeling of output elements ensures control over re-
a domain-specialized LLM, exhibiting superiority to its generalist sponses, facilitating the production of a specific, well-structured
counterparts in broad financial tasks such as market sentiment output. Consequently, this technique yields logically coherent and
Enhancing Credit Risk Reports Generation using LLMs: An Integration of Bayesian Networks and Labeled Guide Prompting ICAIF ’23, November 27–29, 2023, New York, NY

Standard Prompting Labeled Guide Prompting

Prompt Prompt

For this variable X explain its


(1) direction,
For this variable X explain its direction, pathways and interactions in
(2) pathways and
regards to the Credit Risk Score.
(3) interactions
in regards to the Credit Risk Score.
Few shot learning instance: The variable A is at a low level
In parentheses, the label of the item adressed.
and negatively impacts the Credit Risk Score. It directly impacts the score
and through variables B and C. This negative impact is exacerbated by the
Few shot learning instance: The variable A is at a low level
value of variable D.
and negatively impacts the Credit Risk Score(1). It directly impacts the
score and through variables B and C(2). This negative impact is
exacerbated by the value of variable D(3).

Response Response

The high value of variable X positively impacts the Credit Risk Score (1),
The high value of variable X positively impacts the Credit Risk Score, as
directly and by increasing variable Z (2). However, this impact is mitigated by
it increases variable Z.
the value of variable W (3).

Superficial Insightful
Figure 1: Example of how Labeled Guide Prompting leads to more insightful responses, by identifying each sub-task as a
particular item to be addressed.

insightful responses, enhancing LLM’s efficacy in complex problem- Traditionally crafted by human analysts, our study validates
solving [Link] labeled guide aids the LLM in guaranteeing differ- the capacity of LLMs to effectively replicate and even enhance
entiation between sub-tasks, therefore encouraging specificity and credit risk assessment. In an experimental setup involving 100
a more insightful response given that there is clear separation and credit applications, independent evaluators showed a statistically
purpose for each dimension of the problem at hand. The labeled significant preference for GPT-4 generated reports. This validation
guide serves as a rubric for the LLM, requiring the LLM to review of our method confirms the ability of LLMs to effectively automate
the quality of response, format and content as the sub-tasks of the the generation of credit risk reports.
task are being addressed.
In our work, we apply Labeled Guide Prompting to guide GPT-4 1.2.2 Integration of Bayesian Networks. : Our approach integrates
in generating structured and insightful credit risk reports. We found Bayesian networks to guide the LLM’s reasoning tasks. The use of
that this technique effectively increased the frequency and quality this graphical model assists the LLM in understanding the influence
of interactions in reports, thus enhancing their insightfulness and of each variable on the target, aligning with the network structure,
usefulness in the context of credit risk assessment. and producing plausible scenarios, thereby promoting abductive
In this study, we extend the application of LLMs, specifically GPT- reasoning.
4, to the domain of credit risk assessment in agricultural finance,
presenting a novel approach for leveraging LLMs in the financial 1.2.3 Novel Prompt Engineering Technique. : Our technique, La-
sector. beled Guide Prompting, was effective in ensuring that the LLM
accurately addresses the specific aspects of credit risk analysis.
This not only improved the quality of the generated reports but
1.2 Main contributions also demonstrated the broad potential of this technique in other
Below we summarize our main contributions. These findings bear sophisticated tasks. The LLM, post prompt-engineering, showed
significant implications for the future integration of LLMs in credit a statistically significant improvement in the insightfulness of the
risk assessment and other financial risk evaluation tasks. The suc- reports, as measured by well-defined metrics. This underscores the
cessful application of our approach establishes a solid foundation potential of our approach to enhance the performance of LLMs in
for further research and practical implementations in this field. specialized tasks.
1.2.1 LLMs for Credit Risk Reports. : We introduce an innovative
application of Large Language Models (LLMs), specifically GPT- 2 METHOD
4, for the generation of high-quality credit risk reports. We have The primary objective of this research is to apply an LLM, specifi-
developed a principled procedure for prompt engineering tailored to cally GPT-4, to generate credit risk reports that surpass the quality
the specifications of credit risk reports. This innovation significantly of traditional reports written by human analysts. This strategy in-
expands the repertoire of LLM applicability, demonstrating their volves a novel prompt engineering procedure for GPT-4, designed to
potential in tasks requiring specialized knowledge and precision. adhere to the content specifications of the reports. The specific task
ICAIF ’23, November 27–29, 2023, New York, NY Ana Clara Teixeira, Vaishali Marar, Hamed Yazdanpanah, Aline Oliveira, and Mohammad M. Ghassemi

involves generating comprehensive credit risk reports that quantify (1) Question Text: “Which report was more helpful for you to
and explain credit risk in terms of the probability of delinquency assess the credit risk?”
and corresponding score, based on the borrower’s information de- Response Form: choose “A", “B”, or “Both are equally helpful”
rived from the credit application. The overarching goal is to provide
solid support for credit decisions, which means either approval or (2) “Question Text: Do any of the reports have any information
rejection. Accordingly, the report should provide a thorough expla- that is not true or does not make sense?”
nation of the impact of all risk factors to ensure completeness, and Response Form: choose “A”, “B”, “Both”, or “None”
it should offer a decisive assessment of the borrower’s risk profile
to guarantee insightfulness. (3) Question Text: “Why do you prefer the chosen report?”
To assess the quality and effectiveness of these AI-generated Response Form: free-text
credit risk reports, an experiment involving 100 credit applications
was conducted. Each application was evaluated through two sep- The design of this questionnaire ensured an objective evaluation of
arate reports - one generated by a team of human credit analysts, the reports, while the open-ended question provided insight into
and the other produced by GPT-4 using the proposed method in the evaluators’ reasoning behind their choices.
this paper. These reports were designed to provide comprehensive
explanations of the probability of delinquency, taking into account 2.2 Integrating Bayesian Networks into the LLM
all the input information provided. Subsequently, an independent In generating credit risk reports, we leverage a probabilistic graphi-
team of credit analysts blindly reviewed the reports and chose the cal model, specifically a Bayesian network, to estimate the probabil-
most helpful one for making a credit decision. This experiment ity of delinquency. Our choice for a Bayesian network is motivated
enabled a human-based and user-centered evaluation of the GPT-4- by two primary considerations. Firstly, Bayesian networks provide
generated reports, ensuring a practical assessment of the efficacy robust predictive performance. Secondly, their network structure
of the proposed approach. and computation of a joint probability enhance the interpretability
of the model output, enabling us to identify pathways of influence
2.1 Data and experimental design and interactions among variables.
2.1.1 Data source. The credit application data for this research The Bayesian network structure offers a framework conducive to
were obtained from a financial institution’s clients, a representa- engaging the LLM with complex tasks inherent in credit risk reports.
tive set of lenders within Brazil’s agricultural supply chain. These This structure facilitates reasoning across multiple dimensions, as
applications, filled out by farmer borrowers, contain critical in- follows
formation necessary for a thorough credit risk assessment in an
(1) Common-Sense Reasoning: Informative variable names
agricultural context. This risk assessment task is handled by the in-
and a sensible structure dissuade the LLM from generating
stitution: the institution processes the credit application and returns
nonsensical explanations.
a probability of delinquency alongside a corresponding credit score
(2) Logical-Deductive Reasoning: The model proposes a gen-
for each application. The human-generated (HG) reports based
eral rule, prompting the LLM to deduce specific cases that
on these assessments were then created by the institution’s credit
align with this rule.
analysts, professionals with extensive experience in agricultural
(3) Logical-Abductive Reasoning: Factorization of the joint
finance credit risk evaluation.
distribution according to the network helps identify plausi-
2.1.2 Comparison of human and LLM generated reports. To assess ble scenarios - both pathways and interactions - given the
the utility and quality of the credit risk reports, an independent information at hand.
evaluation was conducted by a team of seven credit analysts exter- (4) Implicit Reasoning: The network paths offer a natural
nal to the institution. These analysts, industry professionals from means of decomposing relationships into intermediate steps.
Brazil’s agricultural finance sector, were blind to the sources of (5) Causal Reasoning: The structure of Bayesian networks
the reports and unaware that one was LLM-generated (LLM-G). enables the identification and interpretation of potential
This blind-review approach ensured an unbiased comparison of the causal relationships.
reports based on their practical utility in credit risk assessment.
In deploying this approach, it is essential to ensure that the rela-
The evaluators, who had no contact with each other, were not
tionships between each variable and the target are intuitive. We rely
aware that the comparison’s purpose was to compare HG and LLM-
on the LLM’s capacity to understand and predict the relationship
G reports. They simply believed they were comparing two different
between each variable and the target, predicated on the variable’s
reports to determine which would be more useful in making a credit
name. We achieve this through the setting of appropriate priors
decision. Six evaluators reviewed reports in Portuguese, while one
and an analysis of how the target estimate fluctuates in response
proficient in English reviewed reports in that language. They had
to changes in the values of the features.
three days to assess the assigned cases.
The evaluation process consisted of a pair of credit reports for
each credit application and a questionnaire for evaluators to com- 2.3 Generating Credit Risk Reports
plete. The questionnaire consisted of three questions. In the ques- The specific task of generating credit risk reports aims to assess
tionaire, ‘A’ corresponded to the report generated by human (HG) the LLM’s ability to produce a proficient credit risk report. This
and ‘B’ corresponded to the LLM-generated (LLM-G) report: specific task challenges the LLM in the following areas:
Enhancing Credit Risk Reports Generation using LLMs: An Integration of Bayesian Networks and Labeled Guide Prompting ICAIF ’23, November 27–29, 2023, New York, NY

(1) Contextual Question-Answering: The LLM should accu- • Model Output: This is the computed credit risk score, which
rately interpret the comprehensive instruction, which in- forms the main subject of the report. The LLM is expected to
cludes explaining the credit risk score based on the Bayesian elaborate on it extensively, given the inputs and the structure
network’s structure, input, and output. of the Bayesian network.
(2) Domain-Specific Tasks: The LLM needs to demonstrate • Items to be Addressed: These detail the elements that the
proficiency in agricultural finance, credit risk, and machine LLM should emphasize in the report. These aspects stem
learning to effectively produce a relevant report. from the Bayesian network’s learning and inference process,
(3) Decomposition: The task of explaining the impact of each thereby enabling the LLM to carry out reasoning tasks with
feature on the target should be broken down into simpler greater accuracy. The items to be addressed are:
steps by the LLM, aligning with the Bayesian network struc- (1) The direction of the impact of each model input on the
ture. output (does this model input increase or decrease the
(4) Common-Sense Reasoning: The LLM must comprehend output?)
the phenomenon that the Bayesian network describes and (2) The network pathway from input to output (how does this
create plausible scenarios given the feature values to justify model input impact the output, considering the network
the estimated target. structure?)
(5) Logical Reasoning: Beyond generating logically coherent (3) Input interactions affecting the output (how does this
statements, the LLM should: model input interact with neighboring inputs to impact
• Deduce the impact of the features on the target in line the output?)
with the Bayesian network structure, and • Few-shot Learning Instances: These are samples of com-
• Engage in abductive reasoning to provide the most plausi- pleted tasks that the LLM can learn from. By generalizing
ble explanation of how the inputs yield the model’s output. from these instances to the present task, the LLM can under-
(6) Implicit Reasoning: The LLM is charged with illuminat- stand the structure, style, and content of an effective credit
ing steps implicit in the network structure, explaining how risk report.
a feature impacts the target through various pathways or
interactions. Table 1: How to represent the Bayesian network structure as
(7) Causal Reasoning: The LLM should identify potential causal a prompt component.
relationships between the features and the target variable as
suggested by the Bayesian network structure.
Network Structure
The primary criterion of our evaluation is human-based and "The Bayesian network structure is as follows:
user-centric, where credit analysts judge the utility of the reports. - Variable 1 is connected to Variables 2, 3, 4, and 5 (the output).
Alongside this, we use the blind review’s results to objectively - Variable 2 is connected to Variables 3 and 5 (the output).
assess the LLM’s performance on the general tasks. This approach - Variable 3 is connected to Variables 4 and 5 (the output).
ensures a robust, comprehensive assessment of the LLM’s capability - Variable 4 is connected to Variables 6 and 5 (the output).
to generate credit risk reports, striking a balance between subjective - Variable 7 is connected to Variable 4.
user experience and objective task performance measures. - Variable 8 is connected to Variables 9 and 5 (the output).
- Variable 9 is connected to Variable 5 (the output)."

2.4 Prompt Components


The prompt engineering design used in this research is constructed 2.5 Labeled Guide Prompting
around several key components. These components guide the LLM
in generating comprehensive and insightful credit risk reports. To ensure that the LLM retrieves pertinent information and reasons
These components are listed below. in a prescribed manner, we introduce a novel prompt engineering
technique called Labeled Guide Prompting. This method works
• Role and Instruction: The role situates the LLM as a data by splitting the task into sub-tasks, each with a unique label. The
scientist expert in agricultural finance. This positioning pro- LLM is given response requirements that define these labels and
vides the context and performance expectations for the LLM. is instructed to address each label. This instruction is further cor-
The instruction propels the LLM to create an in-depth credit roborated by the few shot learning example, which demonstrates
risk report that explains the probability of delinquency and a perfect sample wherein each sentence is labeled in reference to
the corresponding credit risk score, by taking into account which sub-task it is addressing. As a result of this method, the LLM
both the provided data inputs and the underlying network dedicates more sentences to each proposed labeled sub-task . It
structure. views each label as a separate concept leading to a more detailed
• Network Structure: This component denotes the structure response with little to no overlap between the information and
of the Bayesian network, a graphical illustration of the causal insights used in one label to the next.
relationships among different variables. Table 1 shows how To evaluate the effectiveness of our proposed technique, we con-
it equips the LLM with a roadmap to comprehend and reason ducted a series of experiments with 100 credit applications, compar-
about the dependencies between various factors, and how ing the LLM’s performance with and without the implementation
they contribute to the credit risk. of this method. We used several metrics for the evaluation:
ICAIF ’23, November 27–29, 2023, New York, NY Ana Clara Teixeira, Vaishali Marar, Hamed Yazdanpanah, Aline Oliveira, and Mohammad M. Ghassemi

(1) Number of complete responses: a complete response has Table 2: Responses to Question 1 - “Which report was more
the three items addressed for each variable. helpful for you to assess the credit risk?” Preferences ex-
(2) Number of occurrences of the targeted items in each pressed by evaluators, as counts and percentages of total
response. responses. The p-values are from chi-squared tests for equal
likelihood across categories. The first row for each language
In the response, each feature is addressed in a separate paragraph, allows for the choice of both reports, while the second
which details its impact on the target. This explanation includes row indicates exclusive preference for either report human-
interactions whenever it references any parent nodes. For instance, generated (HG) or LLM-generated (LLM-G).
if "Crop Yields" is linked to "Profitability," which in turn is linked Language HG LLM-G Both Total p-value
to "Credit Performance," an interaction for "Profitability" might
include a reference to "Crop Yields." English 6 55 57 118 5.095 × 10 −10
Additionally, we introduce insightfulness measures to evaluate (5.1%) (46.6%) (48.3%)
the depth of the LLM’s reasoning: -
English 6 55 61 2.496 × 10 −10
(9.8%) (90.2%)
(1) Insightfulness of Interactions: For each feature’s para-
graph, we first compute the ratio of the number of mentioned Portuguese 29 43 25 97 0.061
(29.9%) (44.3%) (25.8%)
parent nodes to the total number of parent nodes. We then
-
average this ratio across all features that show interactions. Portuguese 29 43 72 0.098
(40%) (60%)
(2) Insightfulness of the Response: For each feature’s para-
graph, we first compute the ratio of the number of mentioned
parent nodes to the total number of parent nodes. We then
Table 3: Responses to Question 2 - “Do any of the reports have
average this ratio over all features that are explained in the
any information that is not true or does not make sense?”
answer.
Responses as counts and percentages of total responses. The
By quantifying the insightfulness of interactions and responses, first row for each language includes all responses, including
we can assess the LLM’s capability to integrate and reason with mul- those who found no issues in any or both of the reports. The
tiple variables simultaneously. This in-depth analysis enables us to second row includes only the responses of those who found
optimize the performance of the LLM in generating comprehensive issues in precisely one report.
credit risk reports. Language None HG LLM-G Both Total
The Label Guide Prompting technique works in synergy with the
Bayesian network, as they facilitate an environment where the LLM English 84 9 3 22 118
(71.2%) (7.6%) (2.5%) (18.6%)
can engage in abductive reasoning – the process of considering
- -
plausible scenarios and outcomes based on available data. Specifi- English 9 3 12
(75%) (25%)
cally, the Bayesian network representation provides a framework
that allows the LLM to explore various pathways and interactions
Portuguese 71 10 10 6 97
embedded within it, while Label Guide ensures that these intricate (73.2%) (10.3%) (10.3%) (6.2%)
relationships are adressed by the LLM in its output. Consequently, - -
Portuguese 10 10 20
the LLM becomes more proficient in managing complex tasks that (50%) (50%)
require a deep understanding of multiple factors.

3 RESULTS
3.1 Blind Review
40 Table 2 presents the results of Question 1: "Which report was more
35 helpful for you to assess the credit risk?". They indicate the evalu-
ators’ preference for the report generated by GPT-4 (LLM-G) and
30
"Both" (indicating a positive reception to LLM-G) across both Eng-
25
Count

lish and Portuguese. The preference for LLM-G is noticeable in


20 the English version, with 90. 2% favoring it when disregarding
15 the "Both" option. In the Portuguese version, evaluators still show
10 a significant preference, with 60% favoring LLM-G in the same
5 conditions.
0 The p-values presented in Table 2 are derived from a chi-squared
0.0 0.2 0.4 0.6 0.8 1.0
Entropy goodness-of-fit test. In the case of Portuguese evaluators, the p-
value, when we assess preferences for HG versus LLM-G or "Both",
Figure 2: Histogram of the entropy of binary distributions is approximately 6.841 × 10 −10 . Hence, the evaluators in Portuguese
of preferences for human-generated and LLM-generated re- also show a significant preference for LLM-G, as we assume "Both"
ports among multiple evaluations of credit applications. is favorable to LLM-G due to the benefits of automation.
Enhancing Credit Risk Reports Generation using LLMs: An Integration of Bayesian Networks and Labeled Guide Prompting ICAIF ’23, November 27–29, 2023, New York, NY

Table 4: Responses to Question 3 - “Why do you prefer the Table 5: Impact of Labeled Guide Prompting. The p-values
chosen report?” We illustrate the distribution of Semantic are from the Wilcoxon signed-rank test.
Domains. Average Occurrences Non Labeled Labeled p-value
Semantic Domain HG LLM-G Both Total
1 - impact 7.55 8.57 1.557 × 10 −6
Agricultural Finance 12 8 8 15 2 - pathways 4.97 8.46 2.04 × 10 −14
Information 6 5 1 9 3 - interactions 4.57 8.42 3.221 × 10 −16
Credit Risk 0 3 1 3
0.5 Labeled
0.4 Unlabeled

Insightfulness of
Interactions
0.3
In the blind review process for credit applications, each applica-
tion received between 1 to 5 evaluations. Among 72 applications 0.2
evaluated more than once, 41 were evaluated twice, 29 once, 23 three 0.1
times, 5 four times, and 3 five times. By combining the categories 0.0
0 20 40 60 80 100
LLM-G and "Both" into LLM-G exclusively for this analysis, each Samples
credit application evaluated more than once generated a binary
distribution of preferences (HG versus LLM-G). Figure 3: Comparison of Insightfulness of Interactions for
To assess the agreement among reviewers, we calculated the each report, with and without LGP. The corresponding p-
entropy for each of these 72 binary distributions. Entropy, in this value from the Wilcoxon signed-rank test was 3.214 × 10 −11
context, serves as a measure of diversity in preference, with 0
representing unanimous agreement (all reviewers selected the same words associated with ’understanding’ and ’analysis’, showing the
category), and 1 denoting total disagreement (votes for categories respective strengths of each preferred report.
HG and LLM-G were evenly distributed).
Figure 2 depicts the histogram of entropies and shows that 41 3.2 Performance of Labeled Guide Prompting
out of the 72 applications evaluated multiple times had unanimous Table 5 shows the impact of Labeled Guide Prompting (LGP) in
decisions, as evidenced by entropy of 0, indicating a high level the responses. The average occurrence of labels 1, 2, and 3 all rose
of consensus among reviewers. Conversely, we observed that the significantly, with p-values derived from a Wilcoxon signed-rank
entropy was 1 for 17 applications, highlighting a uniformly split test indicating that these increases were statistically significant.
decision among evaluators. Thus, while a significant degree of There was a noteworthy increase in complete responses, from 2 in
consensus was noted among evaluators overall, 23.6% of the 72 ap- the non-labeled scenario to 56 when LGP was employed.
plications had disagreement on the reviewers’ choices. This graph Figure 3 depicts the insightfulness of interactions for each report,
also shows that fewer applications had entropy values falling be- whereas Figure 4 illustrates the insightfulness of each response. A
tween 0 and 1. This signifies a split preference among reviewers, noticeable increase in overall insightfulness, measured by both met-
but not an equal distribution between the two categories. rics, is clearly demonstrated in these figures. The implementation
Table 3 presents the results of Question 2: "Do any of the reports of the proposed technique effectively enhances the insightfulness
have any information that is not true or does not make sense?" The of interactions and responses.
results of whether the evaluators found any information that was The insightfulness of interactions is related to the number of
untrue or did not make sense show that LLM-G has no more errors parent nodes mentioned in each feature paragraph (see 2.5). The
than HG. Also, these results show that the translation tool does not insightfulness of the response is influenced by both the quantity of
increase the overall occurrence of errors. interactions and their individual insightfulness.
When the responses identifying errors in both reports HG and This observation explains why some samples in Figure 3 have
LLM-G are considered, and the counts for "Both" are aggregated higher unlabeled values than labeled ones. These particular cases
to the totals of HG and LLM-G, a chi-squared goodness-of-fit test are instances where only one or a few variables exhibit interac-
yields p-values of 1.095 × 10 −10 and 1.312 × 10 −13 for English and tions. Consequently, when averaged over features with interactions,
Portuguese responses respectively. the insightfulness of interactions of these features predominates,
Table 4 shows the top 30 words in three semantic domains: agri- regardless of how many interactions occur in the response.
cultural finance, information, and credit risk, across the HG and In contrast, in Figure 4, when the averaging process extends
LLM-G reports. The Human report emphasizes agricultural finance over all features explained in the response, it also accommodates
terms such as ’crop yield’ and ’season’, showcasing its ability to the increased number of occurrences of interactions, not just their
focus on the most pertinent information for each credit application, insightfulness. Therefore, a more comprehensive view is provided,
as highlighted by reviewers that preferred this type of report. integrating both the frequency and insightfulness of interactions.
In contrast, the LLM-G report often uses credit risk terms such The introduction of more interactions fosters abductive reason-
as ’credit score’ and ’decision’, showing its strength in explaining ing, as it integrates information about the parent nodes, thus enrich-
credit risk scores in relation to all factors. In the ’information’ ing the relationship observed between the feature and the target.
category, both reports share a similar count of common terms, yet To illustrate this, consider the following examples of interactions
the specific words they use vary. For instance, the LLM-G report provided by GPT-4 using our methods, where the model performs
leans toward ’detail’ and ’specify’, while the Human report favors abductive reasoning.
ICAIF ’23, November 27–29, 2023, New York, NY Ana Clara Teixeira, Vaishali Marar, Hamed Yazdanpanah, Aline Oliveira, and Mohammad M. Ghassemi

0.30 Labeled integrating the Bayesian network feature importance directly into
0.25 Unlabeled
Insightfulness of
the Response

the LLM to offer case-specific insights presents an avenue for future


0.20
work. The continual refinement of these elements furthers the
0.15
proficiency of LLMs in intricate domain-specific tasks, contributing
0.10
0.05
towards the creation of more precise and reliable AI systems.
0.00
0 20 40 60 80 100 REFERENCES
Samples [1] Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan
Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, et al. 2023. A multitask,
Figure 4: Comparison of Insightfulness of the Response for multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and
each report, with and without LGP. The corresponding p- interactivity. arXiv preprint arXiv:2302.04023 (2023).
[2] Tom B. Brown et al. 2020. Language models are few-shot learners. CoRR
value from the Wilcoxon signed-rank test was 1.087 × 10 −17 abs/2005.14165 (2020). arXiv:2005.14165 [Link]
[3] Zhiyu Chen, Harini Eavani, Yinyin Liu, and William Yang Wang. 2019. Few-
shot NLG with pre-trained language model. CoRR abs/1904.09521 (2019).
Crop Yields (high): Considering that the farmer has a high yield arXiv:1904.09521 [Link]
and is cultivating rice (3), this directly enhances the profitability [4] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT:
per hectare (2) and improves the Credit Risk Score (1). Pre-training of deep bidirectional transformers for language understanding. CoRR
abs/1810.04805 (2018). arXiv:1810.04805 [Link]
State (Rio Grande do Sul (RS) - Brazil) and Cultivated Crop (rice): [5] Amr Hendy et al. 2023. How good are GPT models at machine translation? A
These inputs interact with the crop yields and profitability per comprehensive evaluation. arXiv:2302.09210
[6] Seungone Kim. 2022. Can Language Models perform Abductive Commonsense
hectare influencing the Credit Risk Score (3). Cultivating a high- Reasoning? arXiv:2207.05155
yield crop like rice in a state like RS with suitable conditions can [7] Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michi-
enhance profitability and hence, the credit risk score (2). hiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al.
2022. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110
Short Term Debt to Total Planted Area (low): Coupled with the (2022).
low levels of long-term borrowing (3), the lower score for this vari- [8] Zeyan Liu, Zijun Yao, Fengjun Li, and Bo Luo. 2023. Check me if you can: Detect-
able also reflects good cash flow management, thereby improving ing ChatGPT-generated academic writing using CheckGPT. arXiv:2306.05524
[9] Remo Pareschi. 2023. Abductive reasoning with the GPT-4 language model:
the overall Credit Risk Score (1). Case studies from criminal investigation, medical practice, scientific research.
arXiv:2307.10250
[10] Ananya B. Sai, Akash Kumar Mohankumar, and Mitesh M. Khapra. 2020. A
4 DISCUSSION survey of evaluation metrics used for NLG systems. CoRR abs/2008.12009 (2020).
Our results indicate GPT-4 generated reports are as useful as tradi- arXiv:2008.12009 [Link]
[11] Aarohi Srivastava et al. 2023. Beyond the imitation game: Quantifying and
tional credit analyst reports for credit risk assessment. The selection extrapolating the capabilities of language models. arXiv:2206.04615
of the "Both" option signals approval of report LLM-G, because of [12] Mirac Suzgun et al. 2022. Challenging BIG-Bench Tasks and Whether Chain-of-
the inherent efficiency and scalability benefits of automation. Thought Can Solve Them. arXiv:2210.09261
[13] Eva AM Van Dis, Johan Bollen, Willem Zuidema, Robert van Rooij, and Claudi L
The blind review process analysis indicates inter-rater reliability, Bockting. 2023. ChatGPT: five priorities for research. Nature 614, 7947 (2023),
as evidenced by the majority of low entropy values, indicating 224–226.
[14] Ashish Vaswani et al. 2017. Attention is all you need. CoRR abs/1706.03762 (2017).
consistent, non-arbitrary reviewer decisions. arXiv:1706.03762 [Link]
The significant deviation from an equal distribution across non- [15] Alex Wang et al. 2018. GLUE: A multi-task benchmark and analysis platform for
sensical content categories demonstrates that errors are statistically natural language understanding. CoRR abs/1804.07461 (2018). arXiv:1804.07461
[Link]
less frequent in the reports, especially in report LLM-G, suggesting [16] Alex Wang et al. 2019. SuperGLUE: A stickier benchmark for general-purpose
GPT-4 did not introduce harmful hallucinations. language understanding Systems. CoRR abs/1905.00537 (2019). arXiv:1905.00537
The success of our approach lies in the combination of Bayesian [Link]
[17] Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed H. Chi, Quoc
networks and the Label Guide Prompting technique. The former Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in
provides a robust framework for GPT-4 to explore complex credit large language models. CoRR abs/2201.11903 (2022). arXiv:2201.11903 https:
//[Link]/abs/2201.11903
risk factors, contributing to report LLM-G’s comprehensiveness. [18] Jules White et al. 2023. A prompt pattern catalog to enhance prompt engineering
The latter ensures more detailed responses and facilitates the un- with ChatGPT. arXiv:2302.11382
derstanding of each sub-task’s unique characteristics. [19] Shijie Wu et al. 2023. BloombergGPT: A large language model for finance.
arXiv:2303.17564
Despite the success, GPT-4’s tendency to give equal importance [20] Yankai Zeng, Abhiramon Rajasekharan, Parth Padalkar, Kinjal Basu, Joaquín
to all features was seen as a shortcoming. Future work could aim Arias, and Gopal Gupta. 2023. Automated interactive domain-specific conversa-
to refine GPT-4’s summarization capabilities to better reflect the tional agents that understand human dialogs. arXiv:2303.08941
[21] Tianyi Zhang, Faisal Ladhak, Esin Durmus, Percy Liang, Kathleen McKeown, and
Bayesian network feature importance. Tatsunori B. Hashimoto. 2023. Benchmarking large language models for news
summarization. arXiv:2301.13848
5 CONCLUSION Received 20 February 2007; revised dd mm yyyy; accepted dd mm yyyy
In summary, the combination of Bayesian networks and Labeled
Guide Prompting can enhance GPT-4’s performance in complex
problem-solving tasks, achieving a level of competency comparable
to human experts. Despite these advancements, several research
challenges remain. Specifically, there is a need to develop critical
summarization methods that allow GPT-4 to pinpoint and succinctly
communicate the most important aspects of a task. Furthermore,

You might also like