Ppt 1: introduction to artificial
intelligence
Tags exam
Exam: you only need to know what is in the pdf (presentations). Open-book or not
—> he will think about it.
So for example with credit card fraud you can use classification (group them in
yes / no fraud) or with regression you can predict the stock price.
Ppt 1: introduction to artificial intelligence 1
Ppt 1: introduction to artificial intelligence 2
Neural networks
Ppt 1: introduction to artificial intelligence 3
What is an AI system?
AI - some degree of autonomy to achieve specific goals
Generative AI - Q&A
what impacts the return on investment (ROI) for deploying generative AI in
financial functions?
complexity of the AI model
quality of the data
specific use cases
What considerations should CFOs keep in mind when investing in generative
AI technology?
Initial investment costs
Ppt 1: introduction to artificial intelligence 4
Scalability of the AI solution
Data management requirements
Integration with existing systems
Potential risks
Ethical implications of AI usage
Business use case
Expected outcomes
How does generative AI impact financial reporting and compliance?
Automate data aggregation and reports
Ensure accuracy
Reduce the time needed for financial close
What skills and capabilities should finance teams develop to effectively use
generative AI?
Data analytics
Understand AI algorithms
Model interpretation
Data governance
Ethical AI use
Collaborative decision-making
How should CFOs approach the integration of generative AI with existing ERP
and financial systems?
Assessment of current systems and data compatibility
Collaborate with IT teams
how does it influence businesses:
Ppt 1: introduction to artificial intelligence 5
Generative artificial intelligence
what is it, key components:
Data: text-based, video-based, internet, books, news, private data
trained on large datasets that include examples of the content they are
expected to generate. For instance, a generative text model might be
trained on a large corpus of written texts.
Models: Generative Adversial Networks (GANs), Variational Autoencoders
(VAEs), Transformer Models (e.g. GPT4).
Applications: text generation, image creation, music composition, video
production and game development
Challenges and ethical considerations: quality and accuracy, biases and
fairness, ethical use.
Ppt 1: introduction to artificial intelligence 6
you don’t know whether the answer is correct and how much is correct.
biased input gives biased output
Recent developments: RAG (Retrieval Augmented Generation), LLM Agents
A category of artificial intelligence systems designed to create new content, such
as text, images, music, or even videos, by learning patterns from existing data.
These AI models generate new data that is similar to the data they were trained
on, but not identical.
AI is getting to expensive?
Supply shortage has subsided: in 2023 there was a shortage in GPUs. Now
not so much anymore
GPU stockpiles are growing
OpenAI still has the lion’s share of AI revenue: most people use openAI such
as chatGPT. All companies will need to deliver significant value for consumers
to continue opening their wallets.
A 500 billion dollar hole: even by predicting profits, you still come up short of
500 billion dollars for the companies investing in AI models.
It’s not over the B100 is coming: a better chip is coming so likely a shortage
will follow.
Building a railroad you know it’s going to be used for the next years, no one is
placing one next to yours. with chip buildings, you don’t know whether new
technology has emerged by the time your ready to start producing.
GPU capital expenditure is like building railroads?
Lack of pricing power
Intrinsic value associated with the infrastructure you are building
For GPU data centers, there is much less pricing power. GPU computing is
increasingly turning a commodity, metered per hour.
Ppt 1: introduction to artificial intelligence 7
Without a monopoly or oligopoly, high fixed cost + low marginal cost
business almost always see prices competed down to marginal cost.
Investment incineration
Even in the case of railroads - and in the case of many new technologies -
speculative investment frenzies often lead to high rates of capital
incineration.
A lot of people lose a lot of money during speculative technology waves.
It’s hard to pick winners, but much easier to pick losers.
Depreciation
Better next-generation chips
More rapid depreciation
This parallel doesn’t exist for physical infrastructure, which does not follow
any ‘Moore’s Law’ curve, such that cost vs. performance keeps improving.
Winners vs. losers
There are always winners during periods of excess infrastructure building
AI is likely to be the next transformative technology wave.
Declining prices for GPU computing is actually good for long-term
innovation and good for startups
A huge amount of economic value is going to be created by AI
not part of the exam generative AI … investmetn and productive
—> misschien wel zelf even naar kijken voor de zekerheid.
Generative AI - the business perspective
Product innovation and customization
Product development: new designs, content and features
Customization: personalized products and services.
Ppt 1: introduction to artificial intelligence 8
Operational efficiency
Automation: automate repetitive and creative tasks
Cost reduction: reduce labor costs and improve operational efficiency
Marketing and customer engagement
Content generation
customer interaction: chatbots and virtual assistants.
Intellectual property and brand differentiation
Innovation: unique AI-generated content
Intellectual property: AI generated creations can be patented.
What else?
Generative AI - the economic perspective - interactive session
Market dynamics
new markets and industries: AI-generated products. AI tools and services
Disruption: new business models and shifts of value chains
Productivity gains
Labour productivity: automation of creative and analytical tasks
Economic output: increased productivity and efficiency
Job market impact
Job displacement
Job creation
Investment and capital allocation
Venture capital
Resource allocation: towards AI infrastructure, tools and talent.
Generative AI - strategic considerations
Ppt 1: introduction to artificial intelligence 9
Ethical and regulatory challenges
Data and privacy
Data dependency
privacy concerns
Competitive landscape
first-mover advantage
collaboration and competition
Generative AI - the problems:
1. Bias in, bias out
a. generative AI tools reproduce content as biased as the data they were
trained on.
2. Black box
a. Generative AI decisions are opaque and unexplainable they hinder
accountability, trust and potentially lead to unjust outcomes.
3. Expensive
a. Ex-CEO of OpenAI, Sam Altman, confirmed GPT-4 cost more than 100
million to train.
4. Mindless parroting
a. Generative AI’s output is tightly bound to the caliber and volume of its
training data. Its output can only be as good as its training input.
5. Alignment with human values
a. Generative AI lacks the capacity to model the consequences or ethical
implications of its decisions.
6. power hungry
a. ChatGPT’s daily queries are estimated to cost the equivalent of powering
33,000 US households.
Ppt 1: introduction to artificial intelligence 10
7. hallucinations
a. Generative AI has the tendency to confidently spew inaccurate information
or simply make up facts.
8. copyright & IP infringement
a. Several Gen AI models appropriated copyrighted material and intellectual
property with no consent, credit, or compensation.
9. static
a. Generative AI models cannot update their knowledge in real-time or
generate new ideas which may lead to misinformation. .
What is AI
There are three types of AI:
Artificial Narrow Intelligence (ANI)
ANI describes AIs that are good at particular task at a level equal or better
than a human being (Siri, Alexa)
Artificial General Intelligence (AGI)
AGI is an AI that can perform any task that a human being can. This is what
most of us think of when we think of AI (J.A.R.V.I.S from Marvel)
Artificial Super Intelligence (ASI)
This is an intelligence that surpasses anything that humans can do (only in
sci-fi).
Can current state-of-the-art AI achieve thinking machines?
Image Classification
Object detection
Visual reasoning
Ppt 1: introduction to artificial intelligence 11
English language understanding
Question answering
You can also type AI in a few different ways:
Reactive AI
Good for simple classification and pattern recognition tasks
Great for scenarios where all parameters are known; can beat humans
because it can make calculations much faster.
Incapable of dealing with scenarios including imperfect information or
requiring historical understanding.
Limited memory
can handle complex classification tasks
able to use historical data to make predictions
capable of complex tasks such as self-driving cars, but still vulnerable to
outliers or adversarial examples.
This is the current state of AI, and some say we have hit a wall.
Theory of mind
Able to understand human motives and reasoning. Can deliver personal
experience to everyone based on their motives and needs.
Able to learn with fewer examples because it understands motive and
intent.
Considered the next milestone for AI’s evolution.
Self-aware
human-level intelligence that can bypass our intelligence too.
Dartmouth Summer Research Project on Artificial Intelligence - summer 1956
Ppt 1: introduction to artificial intelligence 12
Artificial Intelligence definitions:
Artificial intelligence (AI), in its broadest sense, is intelligence exhibited by
machines, particularly computer systems. It is a field of research in computer
science that develops and studies methods and software that enable
machines to perceive their environment and use learning and intelligence to
take actions that maximize their chances of achieving defined goals. Such
machines may be called AIs.
Artificial intelligence refers to systems that display intelligent behaviour by
analysing their environment and taking actions – with some degree of
autonomy – to achieve specific goals.
AI-based systems can be purely software-based, acting in the virtual world
(e.g. voice assistants, image analysis software, search engines, speech and
face recognition systems) or AI can be embedded in hardware devices (e.g.
advanced robots, autonomous cars, drones or Internet of Things applications).
Machine Learning
Machine learning —> ‘A method of designing a sequence of actions to solve a
problem that optimises automatically through experience and with limited or no
human intervention’
Categories of machine learning:
supervised machine learning (classification)
unsupervised machine learning (clustering)
reinforcement learning: teaching a dog how to sit. Sometimes you immediately
get the reward and sometimes it takes a bit of time (chess).
deep learning
What is AI?
Data
a day in data
Ppt 1: introduction to artificial intelligence 13
The Mathematics
machine learning
neural networks
numerical optimizations
Computing power
10^21 FLOPS (floating point operations per second) globally available
80 x 10^12, is the global GDP
Cost of 1 GFLOP
1945: 1800 trillion USD
2000: 1500 USD
2020: 0.04 USD
Moore’s Law
The number of transistors that can be packed into a given unit of space will
double about every to years.
Moore’s Law has been a driving force of technological and social change,
productivity, and economic growth.
The Mathematics
You can get stuck in a local maximum/minimum —> we can’t solve this yet.
slide of neural network with the mathematics (his fav. slide of the day)
Ppt 1: introduction to artificial intelligence 14
x1 is your outputs, and you take weights, you take a weighted average then the
decisions, you buy when it’s above and you sell when it’s below.
Will you ever get an extreme answer? —> no! Because it’s based on averages. By
taking averages of average (by trying out different weights and different methods)
you will never get extremes.
Neural Networks and the universal approximation theorem
usually one questions about this slide in the exam.
Neural networks can approximate (almost) arbitrary mathematical functions
Cybenko (1989) states that any continuous mathematical function on a compact
domain can be approximated with any precision by an appropriate neural network
with sufficient width and depth.
Neural networks are the most powerful function we have ever had.
Neural Networks - The consequences
Everyone we use as an input is a form of historical data? So how can we predict
the future?
Ppt 1: introduction to artificial intelligence 15
Ppt 2: Data Driven Business
models
Tags exam
2.2: data-driven business models in
finance.
First we explore three key theories that underpin data-driven business models in
finance: Information Asymmetry, Firm Structure (as explained by the Theory of the
Firm and Transaction Cost Economics), and Network Effects.
Information Asymmetry: this says that one party often has better information
than the other during transactions, leading to potential market failure. Data-
driven models can help to reduce this asymmetry. (Google)
Firm structure: both of the theories mentioned above explain why firms exist
and their structures. Data-driven business models can use data to reduce
transaction costs and potentially alter firm structures. (uber?)
Network effects: is highly relevant to digital and data-driven business models.
It suggests that the value of a product or service increases as more people
use it, a characteristic common to many data-driven businesses. (meta)
EXAMPLES!
So let’s see all of these theories in the context of fintech startups:
Information asymmetry:
Traditionally large financial institutions had access to more information and
analytics capabilities than individual investors.
Fintech startups, by using big data and machine learning, have been able
to provide sophisticated financial information to individual investors,
Ppt 2: Data Driven Business models 1
reducing the information asymmetry.
Firm structure:
Traditional financial institutions are often large, hierarchical organisations.
Many fintech startups, on the other hand, are small, agile teams that use
data-driven approaches to disrupt traditional finance.
Network effects:
Fintech platforms often benefit from network effects: as more users join
the platform, more data is generated, which improves the platform’s
services and attracts more users.
This creates a positive feedback loop that can enable rapid growth.
The impact of Robo-advisors on the financial service industry
What are robo-advisors: digital platforms that provide automated, algorithm-driven
financial planning services with minimal human intervention. These platforms use
large amounts of data and sophisticated algorithms to provide personalized
investment advice.
Use the theories again to analyze this:
Information asymmetry: robo-advisors reduce information asymmetry by
making financial advice more accessible and transparent.
Disruptive innovation theory: robo-advisors represent a disruptive innovation
that challenges traditional financial advisors.
Network effect: robo-advisors can benefit from network effects as more users
join the platform and contribute data.
Information Assymetry: Enhancing market efficiency - The Use of Big data in
credit scoring
Traditional credit scoring models rely on a limited set of variables and might
exclude potential borrowers who lack a credit history. Big data technologies
Ppt 2: Data Driven Business models 2
enable the collection and analysis of a wider range of data, providing a more
comprehensive view of a borrower’s creditworthiness.
What are the implications of this:
This use of big data can reduce information asymmetry between lenders and
borrowers, leading to more accurate credit decisions and greater financial
inclusion.
It can also enhance market efficiency by reducing the risk of default and
enabling lenders to offer more competitive interest rates.
Transaction cost economics: Data’s influence: Blockchain technology in supply
chain management
Blockchain technology can create a decentralized, transparent, and immutable
ledger of transactions, which can be used to track and verify goods in a supply
chain. This technology can reduce the need for intermediaries and lower
transaction cost.
Implications:
This can lead to changes in the firm structure, as companies can streamline
operations, reduce the need for certain roles, and increase efficiency.
It can also reduce information asymmetry and increase trust among parties in
the supply chain.
Network effects: The rise of peer-to-peer lending platforms
Peer-to-peer (P2P) lending platforms connect borrowers and lenders directly,
bypassing traditional financial institutions. These platforms use data to assess
credit risk and determine interest rates.
Implications:
P2P lending platforms benefit from network effects: the more users they
attract, the more data they can collect to improve their services, which in turn
attract more users.
Ppt 2: Data Driven Business models 3
This can disrupt traditional lending models and create new opportunities and
challenges in the financial industry.
Let’s do a more deep dive on the three theories mentioned above, starting with
information asymmetry:
Information asymmetry arises when one party in a financial transaction possesses
more or better information than the other. This discrepancy can lead to market
failures, higher risk premiums, and reduced liquidity. Recognizing and mitigating
these imbalances is a major focus in credit markets, insurance, and equity
investments.
It has profound implications:
Adverse selection: higher-risk borrowers may dominate lending pools.
Market inefficiency: investors might misprice securities if they lack crucial
data.
Increased monitoring costs: Lenders must expend resources on due diligence.
Theory:
Akerlof (1970): Lemons problem shows how poor information can degrade
market quality.
Signaling (Spence, 1973): Borrowers or firms provide credentials to convey
their quality.
Screening (Stiglitz): Lenders or underwriters devise mechanisms to extract
hidden information.
Examples:
Credit scoring: Banks use FICO or internal algorithms to reduce uncertainty.
Insurance underwriting: actuarial models account for hidden risk factors.
P2P lending platforms: Detailed borrower data mitigates unseen default risks.
Now let’s focus on the firm structure:
Ppt 2: Data Driven Business models 4
Firm Structure addresses how organizations arrange their internal and external
transactions. The boundaries of the firm, along with governance choices, are
often explained through Transaction Cost Economics (TCE) and the broader
theory of the Firm. In a data-driven context, firms may restructure to optimize
analytics capabilities.
Some concepts:
Transaction Cost Reduction: Data sharing within a firm may be cheaper than
relying on external markets.
Control and Coordination: centralizing analytics can unify data standards and
algorithms.
Flexibility: some organizations adopt hybrid models for specialized functions
(e.g., outsourced AI).
Theory:
Coase (1937): Firms exist to lower transaction costs that arise in open
markets.
Williamson (1979): the governance of contractual relations.
Theory of the Firm: explain vertical integration, outsourcing decisions and
how data-driven assets shifts boundaries.
Examples:
Centralized Analytics teams: banks consolidate data expertise to standardize
risk models.
Vertical integration: payment firms acquiring data providers to reduce
dependency on third parties.
Cloud partnership: outsourcing storage and computation to cloud platforms for
scalability.
Network effects - Deep Dive
Network effects occur when a product or service gains additional value as more
people use it. In finance, platforms like payment networks, crowdfunding sites, or
Ppt 2: Data Driven Business models 5
social trading applications benefit from direct or indirect network externalities that
can foster rapid growth or create ‘winner-takes-most’ scenarios.
Motivation:
User adoption: More participants increase liquidity or funding availability
Positive Feedback Loops: growth drives additional data generation, improving
analytics.
Switching Costs: platforms with large user bases can lock in consumers and
merchants.
Theory:
Direct Network Effects: the value to each user grows with every new
participant (e.g. social trading)
Indirect Network Effects: complementary products or services enhance
platform appeal (e.g., credit card rewards).
Two-Sided Markets: platforms act as intermediaries between distinct user
groups (e.g., merchants and consumers).
Examples:
Payment Networks: Visa, mastercard benefit from broad acceptance, fueling
more usage.
Crowdfunding: Kickstarter’s large user community attracts high-quality
projects and vice versa.
Cryptocurrency exchanges: larger exchanges offer deeper liquidity, attracting
additional traders.
Then, lets bring all of these theories together: Information asymmetry, firm
structure, and network effects each offer unique lenses for examining data-driven
finance. However, they often intersect:
Information asymmetry:
reduced via extensive data sharing within the platform
Ppt 2: Data Driven Business models 6
signaling and screening become more effective when analytics are integrated
at scale.
firms use real-time monitoring to detect anomalies or moral hazard.
Firm Structure
Data-driven insights may encourage vertical integration or strategic
partnerships
TCE suggests lowering transaction costs by internalizing core analytics.
Organizational design can shift rapidly to capture emerging data opportunities.
Network effects:
platform benefits from growing users bases that generate more data.
positive feedback loops accelerate scale, potentially creating dominant market
players.
policy questions arise around fair competition and platform neutrality.
Resource based view (RBV): Motivation
The Resource-Based View (RBV) is a strategic management framework
emphasizing unique, hard-to-imitate resources as the main drivers of sustainable
competitive advantage. These resources can be tangible or intangible, and they
include proprietary data, skilled personnel, and brand reputation. Within finance,
RBV helps explain why certain institutions outperform peers by exploiting
distinctive assets and capabilities.
Context & relevance
Global competition: Firms face intense pressure to differentiate themselves
through superior resources.
Strategic Assets: Patented technology, data analytics platform, or specialized
teams can create lasting advantages.
Sustainability: Resources that are valuable, rare, and inimitable generate
defensible market positions.
Ppt 2: Data Driven Business models 7
Key drivers of competitive Advantage:
VRIO Framework: Resources must be Valuable, Rare, Inimitable, and
Organized to capture value.
Long-Term Returns: Building and maintaining such resources can yield
above-average profitability.
Internal Development: History and path dependence show resources
accumulate over time.
Implications in Finance:
Risk Management: Proprietary risk models can significantly improve lending
decisions.
Asset Management: Unique analytics or research capabilities may lead to
consistent alpha.
FinTech Innovation: Specialized startups use data-drive IP to challenge
incumbents.
Theory:
Foundational concepts: The RVB is strongly associated with the work of Barney
(1991), who argued that resources must fulfill VRIO criteria (Valuable, Rare,
Inimitable, and Organized) to lead to sustained competitive advantage. Tangible
assets can be replicated more easily than intangible resources, such as
reputational capital or organizational culture. The firm’s historical path and prior
decisions shape how resources develop, leading to firm-specific capabilities.
VRIO in Detail:
Valuable: Contributes to efficiency or effectiveness.
Rare: Not widely possessed by competitors.
Inimitable: Difficult or costly to replicate.
Organized: Firm structure must align to exploit the resource.
Tangible vs. Intangible:
Tangible Resources: Physical assets like servers, buildings, or capital.
Ppt 2: Data Driven Business models 8
Intangible Resources: Culture, brand, data analytics expertise, or trade
secrets.
Defense: Intangibles often provide stronger barriers to imitation.
Path Dependence:
Historical Trajectory: Past investments and routines shape current resource
sets.
Lock-In Effects: Firms may become entrenched, reinforcing unique
competencies.
Strategic Lockout: Competitors face higher costs or hurdles to catch up.
Practical Applications: Across the financial sector, institutions exploit key
resources to differentiate themselves. Whether it is a major bank refining
proprietary risk models or a hedge fund cultivating a specialized research team,
RBV helps explain why some firms consistently outperform.
Major Banks:
Brand Reputation: Long history can bolster trust, reducing customer
acquisition costs.
Large Datasets: Legacy relationships generate proprietary data for advanced
analytics.
Capital Scale: Enables significant tech investments, reinforcing competitive
barriers.
Hedge Funds:
Quant Teams: Skilled personnel design unique trading algorithms.
Proprietary Models: Combine financial theory with advanced mathematics for
consistent alpha generation.
High Switching Costs: Competitors cannot easily replicate the fund’s internal
knowledge base.
FinTech Startups:
Agility & Culture: Small, dynamic teams cultivate rapid innovation cycles.
Ppt 2: Data Driven Business models 9
Tech-Based Resources: Cloud platforms, specialized APIs, or unique user
interface designs.
Path Dependency: Early tech choices can evolve into a distinctive competitive
edge if scaled effectively.
Extending RBV: While RBV remains a core strategic management theory, it
evolves alongside new research on dynamic markets, digital transformation, and
knowledge diffusion. Scholars integrate RBV with constructs like dynamic
capabilities and ecosystem-based models to better reflect modern competitive
environments.
Dynamic Environments:
Accelerated Change: Continuous resource renewal can be necessary for
high-tech finance sectors.
Real-Time Data: Access to up-to-date market or consumer info can lead to
ephemeral but impactful advantages.
Disruptive Innovation: New entrants armed with novel resources may
challenge incumbents.
Integration with Other Theories:
TCE Overlaps: Resource decisions can reflect transaction cost minimization.
KBV Links: Knowledge development acts as a specialized intangible resource.
Dynamic Capabilities: Emphasizes reconfiguration and strategic shifts under
uncertainty.
Future Directions:
Data Governance: Firms need to manage and protect information resources
effectively.
AI Integration: Automated tools can amplify or erode resource advantages
depending on adoption speed.
Industry Convergence: Cross-sector collaborations highlight novel resource
combinations.
Ppt 2: Data Driven Business models 10
Knowledge-Based View (KBV): Motivation
Overview: The Knowledge-Based View (KBV) emphasizes knowledge as the
principal resource driving organizational performance and competitive advantage.
Distinct from RBV’s broader resource categories, KBV focuses on how knowledge
is created, shared, and applied within and across firm boundaries. In the finance
sector, this perspective illuminates how specialized expertise, collaborative
learning, and continuous innovation can yield better decisions, advanced
products, and overall resilience.
Rationale:
Complex Decision-Making: Financial products often require a high level of
specialized knowledge.
Rapid Innovation: Knowledge-rich processes underpin frequent new service
launches and refinement.
Globalized Markets: Competition across borders demands continuous
learning and adaptation.
Strategic Significance:
Learning Routines: Systematic methods for capturing and
reusing insights foster agility.
Knowledge Spillovers: Cross-functional teams boost creativity and integration
of diverse perspectives.
Human Capital: Skilled analysts, researchers, and data scientists form a key
knowledge base.
Relevance in Finance:
Risk Analysis: Continual updates to regulatory, market, and consumer data
improve risk models.
Investment Research: KBV clarifies how proprietary insights generate above-
average returns.
Ppt 2: Data Driven Business models 11
Collaborative Ecosystems: Partnerships and networks expedite knowledge
exchange (e.g., FinTech alliances).
Foundational Concepts: KBV posits that a firm’s primary source of competitive
advantage lies in creating, storing, and applying knowledge. Grant (1996) argued
that knowledge integration across individuals and teams enhances organizational
capabilities. Tacit knowledge—rooted in personal experience or complex routines
—often resists codification, adding barriers to imitation.
Tacit vs. Explicit Knowledge
Tacit: Personal, experience-based, difficult to transfer (e.g., trader’s intuition).
Explicit: Codified in manuals, databases, or documents (e.g.,
standard operating procedures).
Knowledge Lock-In: Tacit knowledge can become a key differentiator if well
integrated.
Knowledge Integration
Routines and Processes: Formal mechanisms that encourage sharing across
departments.
Cross-Functional Collaboration: Joint problem-solving draws on
multiple expertise sets.
Absorptive Capacity: Ability to acquire and apply external
knowledge effectively.
Learning Curves
Experience Accumulation: Repetition refines tacit understanding, enhancing
performance.
Organizational Memory: Knowledge repositories preserve lessons from past
successes or failures.
Competitive Shield: Longstanding learning curves hinder rivals from quickly
duplicating expertise.
Ppt 2: Data Driven Business models 12
Practical Application: In finance, organizations continuously generate insights
from data, regulations, and market behaviors. KBV explains how firms transform
diverse forms of knowledge into strategic outcomes, whether in consumer
lending, investment banking, or insurance underwriting.
Lending & Credit
Credit Scoring Expertise: Specialized teams interpret credit reports,
transaction histories, and demographic data.
Risk Models Update: Continuous improvement of underwriting guidelines
based on learned outcomes.
Tacit Insights: Seasoned underwriters incorporate nuances not found in
purely quantitative models.
Capital Markets
Equity Research: Analysts synthesize industry data and company insights for
investment recommendations.
Trader Intuition: Seasoned professionals use experience to recognize market
anomalies early.
Knowledge Sharing Platforms: Intranets and specialized databases
disseminate firm-wide updates.
Insurance & Actuarial Science
Claims Analytics: Deep historical records inform premium pricing and risk
categories.
Actuarial Judgment: Merges statistical models with professional expertise on
uncertainty factors.
Continuous Learning Cycles: Feedback from claim outcomes refines
underwriting guidelines over time.
Ppt 2: Data Driven Business models 13
Ppt 3: Artificial intelligence
presentation
Tags exam
Episode I: Concepts of Large Language Modelling
This section lays the groundwork for understanding how LLMs function.
1. Introduction:
The talk aims to explain how ChatGPT works, referencing works by Stephen
Wolfram and Andrej Karpathy.
The Transformer architecture, introduced by Google in 2017 ("Attention Is All
You Need"), revolutionized sequence transduction by relying solely on
attention mechanisms, replacing recurrent and convolutional neural networks.
Before Transformers, Natural Language Processing (NLP) models relied
heavily on supervised learning with manually labeled data. This limited their
use on datasets that were not well-annotated and made the training of Large
Language Models (LLMs) prohibitively expensive and time-consuming.
OpenAI's 2018 introduction of Generative Pre-trained Transformers (GPT)
involved unsupervised pre-training followed by supervised fine-tuning.
LLM landscape:
ChatGPT, Claude-3, Gemini, Bard, and the most powerful LLMs are
proprietary (model architecture and parameters aren’t disclosed) — they
are only accessible through limited APIs (if at all).
Proprietary: GPT-04, Claude 4, Bard, Gemini, Grok 2
often lead in terms of performance
Open: Meta/Microsoft, deepseek, Grok 3(?)
LLMs as Kernel Process of an Operating System
Ppt 3: Artificial intelligence presentation 1
In a few years it can: read and generate text, more knowledge than any
human, browse the internet, use existing software infrastructure
(calculator, Python, mouse, keyboard), see and generate images and
videos, think for a long time using system 2, self-improve in domains that
offer a reward function, customized and fine tuned, communicate with
other LLMs.
2. Predicting the next word in a sequence:
LLMs like ChatGPT generate text by predicting the next word (token) in a
sequence, producing a ranked list of possible continuations with probabilities.
The distribution of these probabilities follows a power-law decay. (n-1)
If we always pick the highest-ranked word, we’ll get a flat essay (zero
temperature case) and what comes out can get confusing and repetitive, but if
at random we pick lower-ranked words, we get a more interesting essay.
In analogy to exponential distribution sfrom statistical physics we define a
‘temperature’ parameter that determines how often lower-ranked words will be
used.
The probabilities are derived from statistical patterns observed in vast
amounts of text data.
While n-gram probabilities (sequences of n words) could theoretically capture
language statistics, the sheer number of possibilities makes direct calculation
infeasible.
Ppt 3: Artificial intelligence presentation 2
LLMs create models to estimate these probabilities, analogous to how neural
nets recognize images of digits by learning underlying patterns.
3. Neural Nets:
Neural networks are composed of interconnected "neurons" - usually
arranged in layers - that evaluate simple numerical functions.
Weights and biases within the network are learned through a "training"
process.
Each neural net represents an overall mathematical function, albeit a complex
one.
Machine learning is used to find the optimal weights for a given task.
Increasing the size and complexity of the network generally improves
accuracy.
The layers of a neural net often learn hierarchical features of data, such as
edges in images.
Training involves feeding the network examples and adjusting weights to
minimize a "loss function" that measures the difference between the
network's output and the desired output.
The training process often follows the gradient of the "weight landscape."
Interestingly, very large neural nets (with billions of weights) can sometimes
be easier to train than smaller ones, potentially due to avoiding local minima.
Neural network architectures can often be applied across different tasks
without significant customization.
Training can be supervised (with labeled data) or unsupervised (without
explicit labels). Data augmentation techniques can expand training datasets.
4. The Concept of Embeddings:
Embeddings represent words (or other data) as arrays of numbers, where
semantically similar items are located closer together in the embedding space.
These embeddings are learned by training models on large amounts of text,
observing the "environments" in which different words appear.
Ppt 3: Artificial intelligence presentation 3
For example, "alligator" and "crocodile" would have close embeddings.
In ChatGPT, text is broken into tokens, and each token is assigned a numerical
embedding.
The dimensionality of these embedding vectors is typically large.
The state of a neural network before the final output layer can serve as a good
representation of features important for the input data, forming feature
embeddings.
5. Inside ChatGPT:
ChatGPT utilizes a Transformer neural network architecture with billions of
parameters (e.g., 175 billion in an older version mentioned).
Unlike recurrent or convolutional networks, Transformers use attention
mechanisms to weigh the importance of different preceding words when
processing a sequence.
The contribution of each word in the input sequence is considered differently
by the network.
Feature vectors are processed through multiple "Attention Blocks."
The number of parameters in LLMs is substantial, arising from embeddings,
multi-head self-attention, feed-forward networks within transformer blocks,
and the output layer. For instance, GPT-3 had approximately 174.5 billion
parameters.
The final embedding from the network is used to calculate the probabilities of
the next token.
The presentation emphasizes that the inner workings of these billions of
parameters are largely inscrutable. "=> think of LLMs as mostly inscrutable
artifacts, develop correspondingly sophisticated evaluations."
6. Training Large Language Models:
Training LLMs is likened to a "lossy compression of data (compression ratio
~100) collected from the internet, maintaining essentially the 'gestalt'."
Approximately 10TB of text might be compressed into a ~140GB file of
parameters.
Ppt 3: Artificial intelligence presentation 4
The process is computationally intensive, requiring thousands of GPUs for
extended periods (e.g., 6000 GPUs for 12 days, costing around $2 million and
involving ~1e24 FLOPS).
A typical training process involves two stages:
Pre-training: Unsupervised learning on a massive dataset to create a base
model capable of generating internet-style documents.
Fine-tuning: Supervised learning on a smaller, high-quality dataset of
question-answer pairs to create an "Assistant Model" that can respond to
questions in a helpful, truthful, and harmless manner. This involves manually
collected and labeled data. "Just swap the dataset, then continue training."
Increasingly, labeling involves human-machine collaboration, where LLMs can
assist in generating and evaluating training data.
An optional Stage 3: Reinforcement Learning from Human Feedback (RLHF)
can further fine-tune the model based on comparisons of generated answers,
as it's often easier to judge quality than to generate it.
7. LLM Security:
LLMs introduce new security and privacy challenges, including: Jailbreaking,
Prompt injection, Backdoors & data poisoning, Adversarial inputs, Insecure
output handling, Data extraction & privacy, Data reconstruction, Denial of
service, Escalation, Watermarking & evasion, and Model theft.
The OWASP 2025 Top 10 list for LLMs and GenAI highlights these risks.
Examples like "Jailbreak" prompts demonstrate vulnerabilities.
Ppt 3: Artificial intelligence presentation 5
Ppt 4: Credit Risk Economic
Capital
Tags exam
How to withstand a severe crisis?
Economic capital should cover the Unexpected loss, which is defined as the
difference between the Expected Loss and the Value at Risk
It is straight-forward to calculate Expected Losses (EL) using the probability
of default, the Loss Given Default and the Exposure at Default. Alternatively,
this can be calculated as the mean of the loss distribution.
The Value at Risk (VaR) is more complex to estimate and requires to derive the
loss distribution, taking into account joint defaults and rating migrations to
estimate the tail end of the loss distribution. The VaR corresponds to a certain
percentile of the loss distribution, which is usually chosen to be 99.9% (i.e., 1
in 1000 years)
Goal: create a loss distribution of joint default and risk rating migration events of
the portfolio exposures by modelling their correlated behavior to estimate EL and
UL coherently.
What is economic loss?
Ppt 4: Credit Risk Economic Capital 1
The economic loss is defined as the difference in net present value due to a
change in an obligor’s creditworthiness. Default risk only reflects losses due to
default events. Migration risk includes losses (profits) from migrations to other
performing ratings.
What drives the loss distribution?
Correlated default/migration events drive the unexpected loss. Without such
correlation, in an infinitely large portfolio, each year the loss would be equal to
the expected loss and the unexpected loss would be zero.
Higher correlation leads to more joint defaults/migrations and therefore higher
unexpected losses, even if the credit quality (average default rate) remains
unchanged. This is a result of the fatter tails of the loss distribution that in turn
lead to a higher Value at Risk (99.9% quantile of the loss distribution).
The effects of correlation are also observed when looking at default rate time
series
Portfolios with high levels of correlations show high volatility of default rates
which results into spikes in the time series and ‘fat tail’ patterns in the default
rate distribution.
Below an example of 50 observation moments (both 50k obligors, PD = 1%) -
low correlation = 5% and high correlation = 30%.
Correlations can be estimated by historical default rate timeseries (amongst
others)
Apart from correlations, high concentration of exposure towards few
customers increase the unexpected loss.
High single-name concentration leads to fatter tails of the loss distribution.
This in turn leads to higher unexpected losses and higher EC.
Examples on slide
Ppt 4: Credit Risk Economic Capital 2
Ppt 5: Credit Risk Model
Implementation
Tags exam
PD = probability of default
EAD = expose at default
LGD = Loss given default
Different types of risk a bank has
Credit risk
The situation that will arise for the lender when the borrower/obligor fails to
pay them back the amount they owe.
Ppt 5: Credit Risk Model Implementation 1
In simplified terms, the banking systems run on two principles. The first being
the customers using the banking systems to deposit their savings and then the
bank pays them an interest to do this which makes it favorable for customers
to using savings accounts in banks. Then, the bank uses a certain percentage
of these deposits to lend loans/credits to the people/entities and charges
interests to make this lending profitable.
In order to remain profitable, banks need to know what is happening with their
money. This is a big task to do for thousands or millions of customers
So they build models, PG, LGD, and EAD
And what is needed to build these models?
What is a model?
Model = Data + Algorithm
They predict reality, but reality is often different.
‘All models are wrong, but some are useful’.
Why do we need them?
To aid banks in quantifying, aggregating and managing risk across
geographical and product lines.
The outputs of these models also play increasingly important roles in banks’
risk management and performance measurement processing, including
performance-based compensation, customer profitability analysis, risk-based
pricing and, active portfolio management and capital structure decisions.
To result in better internal risk management
To be used in the supervisory oversight of banking organisations.
Credit risk data
Ppt 5: Credit Risk Model Implementation 2
Data Categories
Regimes
Regime What?
Standardized Approach (SA) PD, EAD< and LGD prescribed by regulator
Ppt 5: Credit Risk Model Implementation 3
Internal Ratings Based: Foundation -IRB PD is internal model, EAD and LGD prescribed by
(F-IRB) regulator
Internal Ratings Based: Advanced-IRB
PD, EAD, LGD internal
(A-IRB)
Ppt 5: Credit Risk Model Implementation 4
Ppt 6: Interest Rate Risk in the
Banking Book Models
Tags exam
Models are critical for the future
banks rely more and more on quantitative analysis & models in most aspects
of financial decision making.
banks routinely use models for a broad range of activities, including:
valuing & hedging financial products / portfolios (e.g., options (mortgage)
loans, savings).
Measuring (remaining) risks within the business (e.g., credit, market &
operational risks)
Calculating regulatory and economic capital to hold to remain solvent
Stress testing (e.g., solvency & liquidity stress testing)
Loan/credit approvals
Wealth management for customers (e.g., creating an optimal investment
portfolio)
Making tailor-made customer offers (using data analytics/ML, taking
privacy/ethics into account)
Detecting fraud & money laundering activities via transaction monitoring
(FEC).
The fact that banks rely more and more on quantitative analysis & models is driven
by several factors:
1. Increasing regulations
a. An explosion of new regulations following the global financial crisis &
increased regulatory scrutiny.
Ppt 6: Interest Rate Risk in the Banking Book Models 1
2. Technological advances
a. Technological advances, including increasing availability of data & storage
capacity, computing power, and new techniques to analyse this (big) data
(e.g., machine learning). This also leads to new risks that need to be
measured & managed.
3. Digital ambitions
a. Moreover, in line with ING’s innovation tradition, ING’s making the
difference strategy includes a.o. increasing the pace of (digital) innovation
to serve changing customer needs, and become the next generation
digital bank in which data-driven, quantitative decision making is key.
Models play an important role herein.
The model risk management
What is a model?
“A quantitative method, system or approach that applies statistical, economic,
financial or mathematical theories, techniques & assumptions in order to process
input data into quantitative estimates.”
(the inputs may be (partially) qualitative or based on expert judgment).
A model consists of 3 components:
1. an input component: data & assumptions
2. a processing component: transform inputs into estimates (using statistics,
economics, mathematics)
3. an output/reporting component: which forecasts and estimates and translates
this into business info.
What is a model risk?
Are simplified descriptions of reality, so they are not perfect —> model risk.
Model risk is the potential for adverse consequences from decisions based on
incorrect or misused model outputs.
Ppt 6: Interest Rate Risk in the Banking Book Models 2
—> financial losses, poor business & strategic decision making, or damaging a
bank’s reputation.
Occurs primarily for 2 reasons:
1. the model may have fundamental errors and may produce inaccurate outputs
(in light of its design objectives & intended use).
2. incorrect or inappropriate use of the model (i.e. when its actual use is not in
line with its intended use).
Model risk: Root cause of the global financial crisis (2007-2009)
During 2003-07, US banks started bundling subprime mortgages to create
derivative assets, such as CDOs/CMOs (securitization)
They sold these products to other financial institutions worldwide, thereby
transferring the credit risk to all parts of the global financial system.
When the FED increased the interest rates significantly, the monthly payments
of subprime borrowers increased drastically
Subprime borrowers started to default.
Models did not account for tail dependence. The probability that a large
number of subprime borrowers would default at the same time was completely
underestimated.
Hence, one of the root casues of the global crisis was model risk, in particular
related to the valuation of these CDO-type of assets.
Given recent fines, the role of MV & MoRM (Model Risk Management) cannot be
overstated
J.P. Morgan: 2012, trading losses of 6 billion & fine of 1 billion.
Why? Due to a.o. flaws in a new VaR model created by someone without
experience and with no support.
Mizuho Capital Markets: 2018, fine of 900 million.
Ppt 6: Interest Rate Risk in the Banking Book Models 3
Why? For deficiencies related to a.o. using inadequate processes to
assess the risks of its uncleared swaps, and backtest, benchmark &
validate its margin model.
Aegon: 2019: fine of 100 million
Why? For misleading investors, the SEC said that they had sold
investments that were supposedly based on quantitative models, but
which did not work as intended.
Due to the financial crisis & the increasing use of models in all aspects of banking,
regulators worldwide have increasingly been shifting their attention to models &
active model risk management by banks in recent years.
Spatial Finance: Leveraging Geospatial Data for Financial Decision Making
Climate change —> new set of risks?
Especially important for mortgages as if the house disappears and someone
defaults the bank cannot take the asset back.
We can use satellite data for this.
You have 3 different kinds of satellites:
GEO
MEO
LEO
Satellite industry: also shows exponential growth over the years. Partly due to
decreasing costs.
They are also some public datasets showing this satellite data.
How is these data transmitted to us: electromagnetic spectrum.
How do we work with spatial Finance?
Ppt 6: Interest Rate Risk in the Banking Book Models 4
1. Industry risk assessment
2. Climate patterns detection
3. Spotting opportunities
For example: deforestation risk assessment.
Ppt 6: Interest Rate Risk in the Banking Book Models 5
Ppt 8: Trust in Algorithms
Tags exam
Trust in Algorithms: How reliable are their
predictions?
Executive summary:
Every ML model prediction has an uncertainty
Many sources of uncertainty around
Not same as ‘probability of predicted label’
Proper uncertainty estimate normally NOT provided by a trained model
Without uncertainty, ML model predictions can be (highly) unreliable or
meaningless
Max Baak’s View on Explainable AI
ML model predictions should be statistically rigid and sound
Every individual ML model prediction should ideally come with a (correct)
uncertainty estimate
Regression, classification, LMs, etc.
His research interests include developing statistical techniques and best
practices to achieve this goal
Types of uncertainty
Ppt 8: Trust in Algorithms 1
Aleatoric: uncertainty due to intrinsic randomness (goes down with more data)
Also known as: statistical uncertainty
Epistemic: uncertainty due to lack of knowledge
Also known as: systematic uncertainty
Many different sources of uncertainty!
labelling noise, dropout, …
Ppt 8: Trust in Algorithms 2
Warning:
Result from model.predict_proba() is an approximation of a probability
Cannot be trusted out-of-the-box, often unreliable!
E.g. does not warn if data point is out of distribution
Prediction may be high on probability, yet low on confidence
Solution: Anomaly detection
Use anomaly detection before applying model prediction.
if anomaly found: skip model prediction
E.g. look at the similarity (= distance) between the point you want to predict
and the training data
average distances to a set of k nearest neighbours from predicted class
and to all other classes.
does not work for categorical features
Ppt 8: Trust in Algorithms 3
Dataset shift in ML
Application data often looks different from training data
= dataset shift
x = variables, y = target / class
Covariate shift: shift in the independent variables (p(x)). p (y|x) is unchanged.
Prior probability shift: shift in the target variable (the class, p(y)). p(x|y) is
unchanged
Concept shift: shift in the relationship between the independent and target
variables (i.e. p(x|y)).
Transaction monitoring: Name matching
Why? —> to join datasets
Match (high-risk) names to international watch lists
Match external bank accounts to ING accounts
Look at names on transactions
We focus on Dutch company names.
1. Differentiate between personal and company name.
2. Match company name to ground truth\
Adapting name matching to different name sets.
1. positive name: the name-to-match belongs to a name in the ground truth
2. negative name: the name should NOT match to the ground truth
Existing model is giving a score based on assumed ratio of positive/negative
names
In reality we don’t know the negative fraction!
the correct value may be very big
Ppt 8: Trust in Algorithms 4
We would like our model to give a calibrated probability that a name is a match
or not.
ING vs non-ING datasets behave differently
The distributions are quite different between ING and non-ING names
both negative and positive name-pairs behave differently
out-of-the-box name-matching is uncalibrated
Can one correct for these two types of dataset shift?
Another form of dataset shift:
the (linear) model does not extrapolate well
By weighing the training data, the (linear) model extrapolates better.
Uncertainties on ML model predictions
Ppt 8: Trust in Algorithms 5
(methods and techniques for assessing the uncertainties on ML model predictions
(both systematic and statistical)).
quantifying uncertainty on ML predictions is difficult for many types of ML
algorithms
Actually doable for statistical uncertainty with linear models.
Titanic survival rate
Titanic dataset: model the passenger life for survival rate
Band: statistical uncertainty on the survival rate estimate
Calculated using error propagation on a logistic regression model
From just looking at picture 1, you would say that a high fare rate would lead to a
higher survival rate, but looking at picture 2, you might not be so certain anymore.
Example of systematic uncertainty
Model validation: predicted vs observed probability
Ppt 8: Trust in Algorithms 6
Highly encouraged: slice and dice the (test) data, and show predicted vs
observed probability
(example from an ING project)
NB difference between predicted and observed fractions!
Probability calibration
Classifiers are typically not well calibrated
Scores are only approximate probabilities.
Use observed vs. predicted probability curves to cross-check calibration
Also good for model validation!
To recalibrate: if possible, use isotonic regression to fit the reliability curve.
This remains a hack: it applies an average correction. But works pretty well
in practice.
Model performance monitoring
( keeping one’s models up-to-date over time, under changing conditions)
Popmon - population shift monitoring made easier.
To monitor the stability of a pandas or spark dataset
Automatically detect changes over time from trends, shifts, peaks, outliers,
anomalies, correlations, etc.
support numerical, ordinal, categorical features
Alerting based on static or dynamic business rules.
Why?
When data changes, are ML predictions still reliable?
are our ML models in production monitored carefully enough?
No good open-source solution available…
Ppt 8: Trust in Algorithms 7
Past experience at CERN in doing this right
Precision-Recall curve confidence intervals
Test-set sampling uncertainties on the Precision-Recall curve
By sampling uncertainty on recall and precision.
evaluate and plot the related uncertainty band of the PR (or ROC) curve
Where to set your threshold?
Explainable AI
uncertainty affects decision making
XAI: not only ‘how does it work?’, but also ‘how well does it work?’.
How reliable are your (ML) model predictions ?
—> every prediction should come with an uncertainty
Bad practices in data science.
Quoting robust uncertainties on machine learning (ML) model metrics,
typically not done in the field of data science.
Even though these are essential for the proper interpretation and comparison
of ML models.
Metrics, such as f1-score, precision, recall, etc.
Many possible sources of uncertainty.
Example: precision and recall
Provided: a trained, binary classifier.
for example, fraud detection. Is fraudulent? Yes or no.
Ppt 8: Trust in Algorithms 8
Classifier has a discrimination threshold. Typically fixed by business
requirements.
Have measured recall-precision values on confusion matrix of the test set.
Recall: fraction of true fraud cases identified as fraud
Precision: fraction of all correctly-classified fraud cases.
Executive summary:
Uncertainty affects decision making!
Not only consider ‘what is prediction?’, but also ‘how reliable is the
prediction?’.
How reliable are your (ML) model predictions?
—> every prediction should come with a validity and an uncertainty.
Ppt 8: Trust in Algorithms 9
Ppt 9: Two GenAI projects
Tags exam
Environmental, Social & Governance (ESG)
ESG: "Energy consumption, climate, availability of raw materials, health, safety
and good corporate governance are taken into account in the selection and
management of investments in companies.”
Whole sale banking:
Branch that caters specifically to large companies.
Large Language Model:
Large Language Model. Transformer-based next-token prediction model
Also called a ‘GenAI’ model, for generative AI.
Like ChatGPT from OpenAI.
Why reliable ESG data is relevant
ESG data gives insight into current and future alignement towards ING’s
NetZero 2050 promise; and enables steering our portfolio towards it.
—> the bank need to report on ESG data of its client portfolio.
Our projects provide dashboards for sectors and displays how the sector’s
emission intensities are compared to the target scenario and the market
average
To measure is to know!
Big companies publish their ESG data in annual ‘sustainability reports’.
Example sustainability report
Annual / sustainability report
Ppt 9: Two GenAI projects 1
more than 100 pages
No standard format, unstructured.
Information in text / tables / graphs, not always present
More challenging….
Data Collected
There are 7 main sections of form fields for collecting CO2 transition plan data:
1. Reporting period
2. GHG emissions - Emission values from a company in tonnes CO2.
a. Broken down - Scope 1, Scope 2, Scope 3, further breakdowns.
3. GHG Emission intensities - Emission values normalized by a denominator e.g.
tonnes CO2 / number of employees.
4. Governance - Assurance, Audit, Strategy and persons responsible for the
transition plan
5. Targets - A target set to reduce GHG emission by a certain percentage by a
set date
6. Actions - Description of what the company will do to reduce their emissions.
7. EU taxonomy - the size of the companies’ activities contributing to climate
mitigation under EU taxonomy.
Ppt 9: Two GenAI projects 2