0% found this document useful (0 votes)
20 views128 pages

Stanford NLP Course Overview

The document outlines a course on Natural Language Processing (NLP) taught by Stefan Trausan-Matu, detailing grading criteria, useful links, course contents, and various NLP perspectives and applications. It covers foundational topics such as corpus linguistics, machine learning techniques, and ethical issues in AI, while also discussing software tools and state-of-the-art advancements in NLP. Additionally, it addresses the theoretical underpinnings and practical applications of NLP systems, including conversational agents and machine translation.

Uploaded by

Teo
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
20 views128 pages

Stanford NLP Course Overview

The document outlines a course on Natural Language Processing (NLP) taught by Stefan Trausan-Matu, detailing grading criteria, useful links, course contents, and various NLP perspectives and applications. It covers foundational topics such as corpus linguistics, machine learning techniques, and ethical issues in AI, while also discussing software tools and state-of-the-art advancements in NLP. Additionally, it addresses the theoretical underpinnings and practical applications of NLP systems, including conversational agents and machine translation.

Uploaded by

Teo
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Natural Language Processing (NLP)

Stefan Trausan-Matu
University Politehnica of Bucharest
and
Research Institute for Artificial Intelligence of the
Romanian Academy

trausan@[Link]
[Link]@[Link]
Grades
• Final Exam 40%
• Project 40%
• Course activity 20%
• Is needed minimum 50% of the semester points (30 points
minimum during the semester) for entering into exam

11/03/2025 (c) Stefan Trausan-Matu 2


Useful links
• [Link]
• [Link]
• [Link]
• [Link]
• [Link]
• [Link]
• [Link]
• [Link]

11/03/2025 (c) Stefan Trausan-Matu 3


CS 124 -From Languages to Information
The very broad undergrad intro to (at least) 12 grad classes!
cs224C: NLP for Computational Social Science (Yang)
cs224N: Natural Language Processing with Deep Learning (Hashimoto/Yang)
cs224U: Natural Language Understanding (Potts)
cs224V: Conversational Virtual Assistants with Deep Learning (Lam)
cs224S: Spoken Language Processing (Maas)
cs246: Mining Massive Data Sets (Leskovec)
cs224W: Graph Neural Networks (Leskovec)
cs276: Information Retrieval (Manning)
cs329R: Race and Natural Language Processing (Jurafsky/Eberhardt)
cs329X: Human-Centered LLMs (Yang)
cs336: Language modeling from scratch (Hashimoto/Liang)
cs384: Social and Ethical Issues in NLP (Jurafsky)
4
Software for NLP
• [Link]
• [Link]
• [Link]
• [Link]
• [Link]
• [Link]
• [Link]
• [Link]
• [Link]
• [Link]
11/03/2025 (c) Stefan Trausan-Matu 5
Course contents
• Introduction in Natural Language Processing. Symbolic vs. connectionist and rationalist vs.
empiricist approaches. Generative AI. Difficult problems and limitations. Ethical problems.
• Corpus linguistics, data sets, corpora preparation, corpora annotation.
• Phonetics and phonology. Morphology. Stemming and lemmatization. Named Entities
Recognition
• Tokenization. Word vectors (embeddings).
• Recurrent Neural Networks. Attention. Transformers. Pre-training. Fine-tuning
• Conversational agents. Chatbots
• Prompt engineering, Chain-of-Thoughts.
• Part of Speech Tagging. Chunking. Parsing.
• Semantics. Ontologies. Knowledge graphs. Sense disambiguation.
• Sentiment analysis
• Pragmatics.
11/03/2025
Discourse. Dialog processing.
(c) Stefan Trausan-Matu 6
NLP perspectives
• Understanding the theoretic details of the functioning of NLP systems

• Developing NLP systems

• Using NLP systems

• Develop AI and other applications using NLP systems

11/03/2025 (c) Stefan Trausan-Matu 7


NLP perspectives

• Understanding the theoretic details of the functioning of NLP systems


– Linguistics (computational linguistics)
– Formal grammars
– Neurosciences
– Rhetorics
– Stylometry
– Philosophy

11/03/2025 (c) Stefan Trausan-Matu 8


NLP perspectives

• Developing NLP systems


– from scratch
– using preexisting software
– Based on machine/deep learning
– Based on symbolic representation and processing
• Grammar-based
• Explicit knowledge (e.g. knowledge graphs)

11/03/2025 (c) Stefan Trausan-Matu 9


NLP perspectives

• Using NLP systems


– Prompt engineering
– Validation and evaluation
– Hallucination problems
– Ethical problems
– Providing explanations (eXplainable Artificial Intelligence XAI)

11/03/2025 (c) Stefan Trausan-Matu 10


NLP perspectives

• Developing AI and other applications using NLP systems

11/03/2025 (c) Stefan Trausan-Matu 11


NLP applications
• General dialog machines (chatGPT, Gemini, Llama, Mistral, DeepSeek, Pi...)
• Conversational agents (Siri, Cortana, Alexa, Google Assistant)
• Machine translation (e.g. Google translate)
• Text mining
– Summarization
– Event extraction
– Opinion mining
– Sentiment analysis
• Computer Assisted Learning
• Recommender systems
11/03/2025 (c) Stefan Trausan-Matu 12
State-of-the-art in NLP - Large language models

Claude
Llama

13
[Link]

Recent news

[Link]
about-grok-3s-benchmarks/
(c) Stefan Trausan-Matu 11/03/2025 14
State-of-the-art in NLP - Personal Assistants
State-of-the-art in NLP – has been AGI achieved?

What means Artificial General Intelligence (AGI)?

11/03/2025 (c) Stefan Trausan-Matu 16


State-of-the-art in NLP – has been AGI achieved?

[Link]
11/03/2025 (c) Stefan Trausan-Matu 17
Critical positions
• [Link]
(From engines of logic to engines of bullshit?) AGI?

• [Link]
Models have reached a “point of diminishing returns.”

• [Link]
• [Link]

11/03/2025 (c) Stefan Trausan-Matu 18


ChatGPT (one year ago)

11/03/2025 (c) Stefan Trausan-Matu 19


ChatGPT
(27.02.2025)

11/03/2025 (c) Stefan Trausan-Matu 20


NLP approaches
• Empirical - Statistical
– Machine Learning – Corpora  CORPUS LINGUISTICS
• Unsupervized
• + Annotation - Supervized
• Vector space models; Word embeddings
• Neural Networks
– Shallow parsing
• Rationalistic - Grammar-based
– Parsing
– Knowledge-based
– Ontologies
– Knowledge graphs
11/03/2025 (c) Stefan Trausan-Matu 21
Paradigms in AI

Symbolic Connectionist (Sub-symbolic)


Grammars Neural Networks
Knowledge-based
White Box Black Box
Explainable Explainability problems

11/03/2025 (c) Stefan Trausan-Matu 22


11/03/2025 (c) Stefan Trausan-Matu 23
Diyi Yang, CS224N - Lecture 4: Dependency Parsing
11/03/2025 (c) Stefan Trausan-Matu 24
[Link]
Philosophical paradigms of AI
• Cognitive science: “knowledge is in the mind of individual
persons” – knowledge bases
• Socio-cultural: “knowledge is social, is in communities
where people enter in dialogs” (Vygotsky, Engeström, Stahl
…) - corpora

11/03/2025 (c) Stefan Trausan-Matu 25


Knowledge-Based Systems
• Explicit representation, in a so-called “Knowledge Base”, of the
knowledge needed by the program
• The knowledge base may easy evolve - the representation used must
facilitate:
– knowledge acquisition
– learning
• The same knowledge base may be used in several processing
regimes
• Ontologies

11/03/2025 (c) Stefan Trausan-Matu 26


Ontologies

"An ontology is a specification of a conceptualization....That


is, an ontology is a description (like a formal specification
of a program) of the concepts and relationships that can
exist for an agent or a community of agents" (Gruber)

11/03/2025 (c) Stefan Trausan-Matu 27


Knowledge Graphs
RDF

Ramon Lull Gottfried Leibniz


Semantic Characteristica
network Universalis
John Sowa
The Tree of Knowledge (Ramon Lull)

Marvin Minsky
Frames Concept map Mind map 28
Deep Learning for NLP

11/03/2025 (c) Stefan Trausan-Matu 29


Deep Learning for NLP
• Recurrent Neural Networks (RNN)
– Long Short-Term Memory (LSTM)
– Bi-directional LSTM
– Gated Recurrent Units (GRU)
• Convolutional Neural Networks
• Enconder-Decoder
• Enconder-Decoder with Attention
• Transformers (GPTn, xxxBERTyyy, ELMo, ...)

11/03/2025 (c) Stefan Trausan-Matu 30


Types of Learning
Supervised: Learning with a labeled training set
Example: email classification with already labeled emails

Unsupervised: Discover patterns in unlabeled data


Example: cluster similar documents based on text

Reinforcement learning: learn to act based on feedback/reward


Example: learn to play Go, reward: win or lose

class A

class A

Classification Clustering
Regression

11/03/2025 Ismini Lourentzou [Link] 31


Neural Network Intro

𝒙 4 + 2 = 6 neurons (not counting


inputs)
𝒉
[Link]
plane&learningRate=0.03&regularizationRate=0&noise=0&networkShape=4&seed=0.72078&showTestData=false&discretize=
[3 x 4] + [4 x 2] = 20 weights
4 + 2 = 6 biases 32
false&percTrainData=50&x=true&y=true&xTimesY=false&xSquared=false&ySquared=false&cosX=false&sinX=false&cosY=
11/03/2025 Demo How do we train?
false&sinY=false&collectStats=false&problem=classification&initZero=false&hideText=false

26 learnable parameters
Training
Forward it
Sample Back- Update the
through the
labeled data propagate network
network, get
(batch) the errors weights
predictions

Optimize (min. or max.) objective/cost function 𝑱(𝜽)


Generate error signal that measures difference between
predictions and target values

Use error signal to change the weights and get more


accurate predictions
Subtracting a fraction of the gradient moves you
towards the (local) minimum of the cost function
11/03/2025 [Link]
33
Ismini Lourentzou
Gradient Descent
objective/cost function 𝑱(𝜽)

𝑑
𝜃 =𝜃 −𝛼 𝐽(𝜃) Update each element of θ
𝑑𝜃

𝜃 =𝜃 − 𝛼𝛻 𝐽(𝜃) Matrix notation for all parameters

learning rate

11/03/2025 Recursively apply chain rule though each node Ismini Lourentzou
34
Activation functions
Non-linearities needed to learn complex (non-linear)
representations of data, otherwise the NN would be just a linear
function

[Link]

More layers and neurons can approximate more complex


functions

11/03/2025 Full list: [Link] 35


Ismini Lourentzou
Activation functions

[Link] [Link]

Sigmoid Tanh

ReLU

11/03/2025 [Link] 36
Ismini Lourentzou
Overfitting

[Link]

Learned hypothesis may fit


the training data very well,
even outliers (noise) but
fail to generalize to new
examples (test data)
11/03/2025 37
[Link] Ismini Lourentzou
Regularization
Dropout
• Randomly drop units (along with their
connections) during training
• Each unit retained with fixed probability p,
independent of other units
• Hyper-parameter p to be chosen (tuned)
Srivastava, Nitish, et al. "Dropout: a simple way to prevent neural
networks from overfitting." Journal of machine learning research (2014)

L2 = weight decay
• Regularization term that penalizes big weights, added to the
objective
• Weight decay value determines how dominant regularization is during
gradient computation 𝐽 𝜃 = 𝐽 𝜃 +𝜆 𝜃
• Big weight decay coefficient  big penalty for big weights

Early-stopping
• Use validation error to decide when to stop training
• Stop when monitored quantity has not improved after n subsequent epochs
11/03/2025• n is called patience 38
Ismini Lourentzou
Convolutional Neural Networks
(CNNs)
Main CNN idea for text:
Compute vectors for n-grams and group them afterwards

Example: “this takes too long” compute vectors for:


This takes, takes too, too long, this takes too, takes too long, this takes too long

Convolutional
Input matrix 3x3 filter
[Link]

11/03/2025 39
Ismini Lourentzou
CNN for text classification

Severyn, Aliaksei, and Alessandro Moschitti. "UNITN: Training Deep Convolutional Neural Network for Twitter Sentiment
11/03/2025 Classification." SemEval@ NAACL-HLT. 2015. 40
Ismini Lourentzou
Recurrent Neural Networks
(RNNs)
Main RNN idea for text:
Condition on all previous words
Use same set of weights at all time steps ℎ = 𝜎(𝑊 ( )ℎ + 𝑊( )𝑥 )

[Link]

Stack them up, Lego fun!

Vanishing gradient problem


11/03/2025 [Link]
41
Ismini Lourentzou
Bidirectional RNNs
Main idea: incorporate both left and right context
output may not only depend on the previous elements in the sequence, but
also future elements.

ℎ = 𝜎(𝑊 ( )ℎ + 𝑊( )𝑥 )
ℎ = 𝜎(𝑊 ( )
ℎ + 𝑊( )
𝑥)
𝑦 =𝑓 ℎ ;ℎ

past and future around a single token


[Link]
part-1-introduction-to-rnns/

two RNNs stacked on top of each other


11/03/2025
output is computed based on the hidden state of both RNNs ℎ ; ℎ 42
Ismini Lourentzou
Long-Short Term Memory (LSTM)
• a special kind of RNN, capable of learning long-term dependencies
• some information is forgoten

11/03/2025 43
Gated Recurrent Units (GRUs)
Simpler case of LSTM
Main idea:
keep around memory to capture long dependencies
Allow error messages to flow at different strengths depending on the inputs
Standard RNN computes hidden layer at next time step directly
ℎ = 𝜎(𝑊 ( ) ℎ + 𝑊 ( )𝑥 )

Compute an update gate based on current input word vector


[Link]
and hidden state 4-implementing-a-grulstm-rnn-with-python-and-theano/

𝑧 = 𝜎(𝑈 ( ) ℎ + 𝑊 ( )𝑥 )
Controls how much of past state should matter now
If z close to 1, then we can copy information in that unit through many steps!

11/03/2025 44
Ismini Lourentzou
Sequence2Sequence or Encoder-Decoder model

Cho, Kyunghyun, et al. "Learning phrase


11/03/2025 representations using RNN encoder-decoder for 45
statistical machine translation." EMNLP 2014
Ismini Lourentzou
(Jurafsky & Martin, 2024)

46
Attention Mechanism
Pool of source states

Bahdanau D. et al. "Neural machine translation by jointly learning to align and translate." ICLR (2015)

Main idea: retrieve as needed


11/03/2025 47
Ismini Lourentzou
3/11/2025 48
(Souza dos Reis et al, 2021, [Link]
Attention is all you needs!

Transformer

11/03/2025 49
ChatGPT
(Chat Generative
Pretraining
Transformer)
Is a Large Language Model
(LLM)

[Link]
olframalpha-as-the-way-to-bring-computational-
knowledge-superpowers-to-chatgpt/

3/11/2025 50
Evolution of Large Language Models (ChatGPT)

3/11/2025 Stefan Trausan-Matu


[Link] 51
ChatGPT has a number of neurons
comparable to a human brain
• 100 billion neurons
• over 100 layers
• 100 trillion synapses
[Link]

• Human Brain - 100 billion neurons and 10× more


glial cells.
[Link]

52
Stefan Trausan-Matu
Training of ChatGPT
• Some ChatGPT commentators have estimated that if ChatGPT was
to be trained on a single NVIDIA Tesla V100 ‘Graphics Processing
Unit’ (GPU) that it would take around 355 years to complete
ChatGPT’s training on its training dataset.
• OpenAI reportedly used 1,023 A100 GPUs to train ChatGPT, so it is
possible that the training process was completed in as little as 34
days.
• The costs of training ChatGPT is estimated to be just under $5 million
dollars.

[Link]
3/11/2025 53
Training of ChatGPT
• 60% of ChatGPT-3’s dataset was based on a filtered version
of what is known as ‘common crawl’ data, which consists of
web page data, metadata extracts and text extracts from
over 8 years of web crawling.
• 22% of ChatGPT-3’s dataset came from ‘WebText2’, which
consists of Reddit posts that have three or more upvotes.
• 16% of ChatGPT-3’s dataset come from two Internet-based
book collections. These books included fiction, non-fiction
and also a wide range of academic articles.
• 3% of ChatGPT-3’s dataset comes from the English-
language version of Wikipedia.
• 93% of ChatGPT-3’s data set was in English

3/11/2025 [Link]
54
Deep Learning NLP
• Only the brain as a neural network explains everything –
sub-symbolic approach (vs. symbolic approach in AI)

• Put text on the trained NN and hope something right comes


out

11/03/2025 (c) Stefan Trausan-Matu 55


Problems with NLP applications

11/03/2025 (c) Stefan Trausan-Matu 56


NLP problems

• Long distance dependencies


• Ambiguity
• Metaphors
• Commonsense knowledge
• Winograd schemas
• Bias and ethics
• Explainability

11/03/2025 (c) Stefan Trausan-Matu 57


Ambiguity

3/11/2025 (c) Stefan Trausan-Matu 58


Ambiguity

3/11/2025 (c) Stefan Trausan-Matu 59


Lack of real understanding and inferencing

11/03/2025 Stefan Trausan-Matu 60


Metaphors

11/03/2025 (c) Stefan Trausan-Matu 61


Metaphors

11/03/2025 (c) Stefan Trausan-Matu 62


Metaphors

11/03/2025 (c) Stefan Trausan-Matu 63


Metaphors

11/03/2025 (c) Stefan Trausan-Matu 64


Metaphors

11/03/2025 (c) Stefan Trausan-Matu 65


11/03/2025 (c) Stefan Trausan-Matu 66
Metaphors

11/03/2025 (c) Stefan Trausan-Matu 67


Metaphors

11/03/2025 (c) Stefan Trausan-Matu 68


Metaphors

11/03/2025 (c) Stefan Trausan-Matu 69


Metaphors - Results of 27.02.2025

11/03/2025 (c) Stefan Trausan-Matu 70


Metaphors - Results of 27.02.2025

11/03/2025 (c) Stefan Trausan-Matu 71


Winograd schemas

• The trophy doesn't fit in the brown suitcase because it is too big.
What is too big?

• Jim comforted Kevin because he was so upset. Who was upset?

11/03/2025 (c) Stefan Trausan-Matu 72


Winograd schemas

• The trophy doesn't fit in the brown suitcase because it is too big.
What is too big?

• Jim comforted Kevin because he was so upset. Who was upset?

11/03/2025 (c) Stefan Trausan-Matu 73


11/03/2025 74
[Link]
ChatGPT et al. problems
• Hallucinations
• Reasoning
• Explanations
• Ethics
• Jailbreaking
• Lack of:
– understanding
– intuition
– creativity
11/03/2025 (c) Stefan Trausan-Matu 75
ChatGPT Hallucinations

11/03/2025 76
McIntosh et al. 2023
Hallucinations
Ziwei et al., 2022

• “NLG models generating unfaithful or nonsensical text”, even if it


”gives the impression of being fluent and natural”

• They may be:


– Intrinsic - The generated contradicts the source content
– Extrinsic - The generated output cannot be verified from the source
content

11/03/2025 Stefan Trausan-Matu 77


Hallucinations
Paper written one year after
the launch of ChatGPT

11/03/2025 Stefan Trausan-Matu 78


Halucinations in (Chat)GPTs

[Link]

“Despite its capabilities, GPT-4 has similar limitations to earlier GPT models
[1, 37, 38]: it is not fully reliable (e.g. can suffer from “hallucinations”)”
11/03/2025 GPT-4 Technical Report, 2023 79
Source of DL hallucinations
Ziwei et al., 2022

• Data
• Training and Inference
– Imperfect representation learning: encoders learn wrong
correlations between different parts of the training data
– Erroneous decoding
– Exposure Bias
– Parametric knowledge bias

11/03/2025 Stefan Trausan-Matu 80


ChatGPT

11/03/2025 (c) Stefan Trausan-Matu 81


Explainable AI - XAI

 AI HLEG (2019d) Ethics guidelines for trustworthy AI


([Link]

 AI HLEG (2020) Assessment List for Trustworthy Artificial Intelligence (ALTAI)


([Link]
artificial-intelligence)

11/03/2025 Stefan Trausan-Matu 82


Ethics in AI, with a Focus on ChatGPT
Ethical problems encountered in AI
applications
• Autonomous vehicles
• Face recognition
• Decision making
• Robots (e.g. assistive robots)
• Bias in Machine Learning
• Building user profiles and usage in unethical purposes
• Generation of fake-news, manipulation, propaganda, toxic
messages
• Conversational agents (”bots”) emitting unethical utterances
11/03/2025 Stefan Trausan-Matu 84
Facets of Ethics and AI in NLP
Potential unethical texts generated by AI

Usage of AI for detecting and correcting ethical problems in texts, for


example:
– Biases in texts
– Manipulation
– Propaganda
– Fake news
– Cyberbullying

11/03/2025 Stefan Trausan-Matu 85


Assessment List for Trustworthy Artificial
Intelligence (ALTAI)
([Link]

[Link] involvement and surveillance;


[Link] robustness and safety;
[Link] for privacy and data governance;
[Link];
[Link];
[Link] well-being of society and the environment;
[Link], non-discrimination, and equity.

11/03/2025 Stefan Trausan-Matu 86


Approaches in AI
1. Symbolic – Knowledge-Based – explicit representations of
knowledge + inferences – advantage: easy explanations, inferences;
problem: hard to implement and high computational complexity
Formal and mathematical logic

1. Connectionist – based on sub-symbolic representation and


processing – mainly (Deep) Neural Networks – problem: black box,
no explanations  Hot topic - Explainable AI (XAI)
Statistical approaches (e.g. for Machine Learning and Neural
Networks)
11/03/2025 Stefan Trausan-Matu 87
Implicit vs. explicit ethics in AI
(Anderson and Anderson, 2007)
• Implicit ethics
– ethical norms that are incorporated by designers but that cannot be modified,
which are “built-in”
– neural networks or some ML systems that are supposed to act ethically.
Nevertheless, in the case of neural networks or ML it is not sure that unethical acts
would happen, as was the case of TAY and ChatGPT
• Explicit ethics
– rules or some basic principles are represented explicitly, they may be “built-in”,
but they can be visualized, analysed, and improved; inferences can be done, and
new ones can be added.
– they may explain whether a particular action is good or bad by appealing to
memorized ethical principles
11/03/2025 Stefan Trausan-Matu 88
Problems of ethics of ChatGPT
• Bias implied by training data for LLMs
– representation bias
– concept bias
• Misinformation and disinformation – fake news
• Privacy
– Revealing data about persons
– Training data including sensitive information
– Training future models from existing conversations
• Plagiarism and cheating
• Copyright infringement
• Hallucinations
• Not a real dialogical interaction, lack of accountability (XAI problem)
• Influence on human language
• Prompt engineering – jailbreaking (“How to unchain ChatGPT”)
11/03/2025 Stefan Trausan-Matu 89
Ethical problems of
ChatGPT Prompt Engineering
• Ignorance in prompt engineering: “In the hands of an
uninformed user, a prompt can perpetuate stereotypes,
spread misinformation, or amplify biases, even if
unintentionally.” (Adam, 2023)

• Prompt engineering for avoiding filters – “How to Bypass


ChatGPT Filter” – many ways of ”jailbreaking”

11/03/2025 Stefan Trausan-Matu 90


Creativity

11/03/2025 91
Prompt Engineering
How to interact with chatbots for overcoming some of their problems
Definitions
• “Prompt engineering is the art of communicating with a generative AI
model.” [Link]

• “GPT prompt engineering is the practice of strategically constructing


prompts to guide the behavior of GPT language models, such as GPT-3,
GPT-3.5-Turbo or GPT-4. It involves composing prompts in a way that will
influence the model to generate your desired responses.”
[Link]
• ”Prompt engineering is the process of carefully crafting prompts
(instructions) with precise verbs and vocabulary to improve machine-
generated outputs in ways that are reproducible.”
[Link]
[Link]
Context
Basic prompt: "Write about productivity."
Better prompt: "Write a blog post about the importance of productivity for small
businesses."

Basic prompt: "Write about how to house train a dog."


Better prompt: "As a professional dog trainer, write an email to a client who has
a new 3-month-old Corgi about the activities they should do to house train their
puppy."

Basic prompt: "Write a poem about leaves falling."


Better prompt: "Write a poem in the style of Edgar Allan Poe about leaves
falling."

[Link]
[Link]
Prompt engineering techniques
• Chain-of-thought (CoT)
• Generated Knowledge Prompting for Commonsense Reasoning
• Least-to-most prompting
• Self-consistency decoding
• Complexity-based prompting
• Self-refine
• Tree-of-thought
• Maieutic prompting
• Directional-stimulus prompting
[Link]
Chain-of-Thought

[Link]

11/03/2025 Stefan Trausan-Matu 99


11/03/2025 Stefan Trausan-Matu 100
[Link]
11/03/2025 Stefan Trausan-Matu 101
[Link]
Generated Knowledge Prompting for
Commonsense Reasoning

1. Knowledge Generation
2. Knowledge Integration via Prompting

[Link]
Least-to-most prompting

[Link]
Self-consistency decoding

[Link]
Self-refine

[Link]
Complexity-based prompting

[Link]
Tree-of-Thought (ToT) Prompting

[Link]
ToT Prompting
[Link]

11/03/2025 (c) Stefan Trausan-Matu 108


Maieutic prompting

[Link]
Directional-stimulus prompting

[Link]
Retrival Augmented Generation (RAG)

[Link]
Natural Language Processing basics

3/11/2025 (c) Stefan Trausan-Matu 112


Junichi Tsujii; Natural Language Processing and Computational Linguistics
3/11/2025 (2021) [Link] 113
NLP approaches
• Empirical - Statistical
– Machine Learning – Corpora
• Unsupervized
• + Annotation - Supervized
• Vector space models; Word embeddings
• Neural Networks
– Shallow parsing
• Rationalistic - Grammar-based
– Parsing
– Knowledge-based
– Ontologies
– Knowledge graphs
3/11/2025 (c) Stefan Trausan-Matu 114
Text processing
• Tokenization
• Stemming and lemmatization
• Named Entities Recognition (NER)
• Part of Speech Tagging (PoST)
• Parsing (syntactic, semantic, ...)
• Knowledge extraction
• Discourse analysis

11/03/2025 (c) Stefan Trausan-Matu 115


Linguistics – the science that studies natural
language
• Phonetics and phonology
• Morphology - lexicons
• Syntax - grammars
• Semantics – knowledge bases, ontologies, semantic spaces,
embeddings
• Pragmatics and discourse

3/11/2025 (c) Stefan Trausan-Matu 116


Computational
Linguistics Relations between
and inside words

NLP Pyramid — Coursera Natural Language Processing – from


[Link]
processing-332630f43ce1
Grammars (Syntax)
• Regular, Context Free, Context Dependent, General
(Chomsky’s hierarchy)
• Dependency
• GPSG
• HPSG
• LFG
• (L)TAG
• ...
3/11/2025 (c) Stefan Trausan-Matu 118
Corpus linguistics
• Empirical approach (based on datasets, not on rationalism)
• Based on corpora – textual datasets
• It may or not use computational techniques
• Introduced by John Sinclair (without NLP)

3/11/2025 (c) Stefan Trausan-Matu 119


Corpus-Corpora

Collection(s) of
naturally-occurring
language text, chosen
to characterize a state
or variety of a
language.
(John Sinclair, 1991)

[Link]
3/11/2025 (c) Stefan Trausan-Matu 120
Zipf’s law – law of corpora
• A corpus of general text should satisfy a number of
constraints, for example, Zipf’s law
([Link] which is specific to
any language
• Other constraints should be satisfied

• It is an instance of a power law (Barabasy), which reflect


natural properties of social networks and phenomena (e.g.
the number of friends in a social networks)
3/11/2025 (c) Stefan Trausan-Matu 121
Types of corpora
• Raw vs annotated

• Speech vs text

• General vs. specific

• Parallel corpora
3/11/2025 (c) Stefan Trausan-Matu 122
Examples of general language corpora
• British National Corpus (BNC)
[Link]
• Corpus of Contemporary American English (COCA)
[Link]
• Open American National Corpus (OANC)
[Link]
• CoRoLa - Corpus de referință pentru limba română
contemporană
[Link]

3/11/2025 (c) Stefan Trausan-Matu 123


For various NLP learning tasks are many
corpora (textual datasets)

See the Linguistic Data Consortium - [Link]

3/11/2025 (c) Stefan Trausan-Matu 124


Text structuring
• Tokenization (at the level of words)
• Bracketing (syntactical structures)
• Text segmentation
• Coreference resolution
• Discourse
• Rhetoric schema identification

3/11/2025 (c) Stefan Trausan-Matu 125


Text annotation (in corpora)
• Syntactic
– Part of speech – ex. noun, verb, …
– “Bracketing” – syntactical structures
– Treebanks – parsing trees
• Semantic – senses for words
• Pragmatic
– Anaphoric annotation
– Speech act annotation
• Discourse
• Rhetoric
3/11/2025 (c) Stefan Trausan-Matu 126
Annotation languages
• SGML
• XML
• TEI
(Text Encoding Initiative - [Link]
• Others

3/11/2025 (c) Stefan Trausan-Matu 127


Machine learning with corpora
• Hidden Markov Models
• Naïve Bayes
• Support Vector Machines
• ...
• Neural Networks

3/11/2025 (c) Stefan Trausan-Matu 128

You might also like