MODULE 4
SEMANTIC ANALYSIS &
DISCOURSE
GAYANA M N
ASST PROFESSOR,DEPT OF ICBS
ST JOSEPH ENGINEERING COLLEGE
VAMANJOOR, MANGALORE
Module 4 Topics
Semantic Analysis Computational Discourse
• The representation of meaning • Introduction
• Syntax driven semantic analysis • Discourse Segmentation
• Word Senses • Text Coherence Relations
• Relations between senses, • Reference Resolution
• WordNet: A Database of Lexical • Anaphora resolution
Relations
• Word Sense Disambiguation
Gayana M N Module 4: Semantic Analysis & Computational Discourse
2
Semantic Analysis
• Semantic interpretation is to take natural language sentences or
utterances and map them onto some representation of meaning.
• Semantics can be divided into two parts:
• The study of the meaning of individual words(lexical semantics)
• The study of how individual words combine to give meaning to a sentence (or
larger units)
Gayana M N Module 4: Semantic Analysis & Computational Discourse 3
Semantic Analysis
• Components of Semantics:
• Word meaning
• The relationship that exists between words
• The domain
• Word Order
• Syntactic Structure
• The underlying context
• Real world knowledge
………………….. etc
Gayana M N Module 4: Semantic Analysis & Computational Discourse 4
Semantic Analysis
• Semantic Analysis :
• Begins with Lexical Semantics
• To handle compositional semantics-Model Theoretic Semantics
• Takes into consideration logical words such as and, or, all, some etc.
• Efficiently determines how the meaning of a sentence is determined from the meanings
of its parts.
• In this theory, the truth of a sentence does not mean that sentence is actually true. It
simply means the sentence is true in the world being modelled.
• Usage of Structural Semantics
Gayana M N Module 4: Semantic Analysis & Computational Discourse 5
Module 4 Topics
Semantic Analysis Computational Discourse
• The representation of meaning • Introduction
• Syntax driven semantic analysis • Discourse Segmentation
• Word Senses • Text Coherence Relations
• Relations between senses, • Reference Resolution
• WordNet: A Database of Lexical • Anaphora resolution
Relations
• Word Sense Disambiguation
Gayana M N Module 4: Semantic Analysis & Computational Discourse
6
Meaning Representation
• A meaning representation can be understood as a bridge between
subtle linguistic nuances and our common-sense non-linguistic
knowledge about the world.
• It can be seen as a formal structure capturing the meaning of
linguistic input.
• We assume that any given linguistic structure has some
stuff/information that can be used to express the state of the world.
Gayana M N Module 4: Semantic Analysis & Computational Discourse 7
Meaning Representation
• Commonly used meaning representation languages:
1. Logic based representation – Propositional logic , first order
predicate logic
2. Network based representation-Semantic networks, conceptual
graphs
3. Structured Representations- Frames, Scripts
Gayana M N Module 4: Semantic Analysis & Computational Discourse 8
Meaning Representation
Logic based representation – Propositional logic , first order predicate
logic(FOPL)
• Propositional logic uses inference rules to perform logical proofs or
deductions. It uses predicates, variables, functions, constants,
quantifiers and logical connectives.
• FOPL uses two variable quantifiers –existential and universal
• Ex: “All children like Apple”
Can be represented in FOPL as
(V)(children(x) →like(apple, x)
Gayana M N Module 4: Semantic Analysis & Computational Discourse 9
Meaning Representation
Network based representation-Semantic networks, conceptual graphs
• Representing knowledge as graph with labelled nodes and arcs
• Eg: “Suha eats apple” can be represented as
eats
Suha Apple
Gayana M N Module 4: Semantic Analysis & Computational Discourse 10
Gayana M N Module 4: Semantic Analysis & Computational Discourse 11
Characteristics of Meaning Representation Languages
1. Verifiability
• We must be able to determine the truth of our representation.
• Truth is determined by comparing the meaning representation of an
input with the repository of facts existing in the domain (i.e,
knowledge base )
• Eg: consider the question “Does Kingfisher serve Hyderabad?”
the proposition is Serves(Kingfisher, Hyderabad)
If the system finds a matching proposition, return yes. Returns
No, if it has reason to believe that its knowledge base lacks
information.
Gayana M N Module 4: Semantic Analysis & Computational Discourse 12
Characteristics of Meaning Representation Languages
2. Unambiguous
• It should support representations that have only one possible
interpretation.
• As the reasoning and subsequent action is based upon the semantic
content of inputs, it is important that the final meaning
representation of an input be unambiguous, regardless of ambiguity
in the raw input.
• A meaning representation language should support a certain level of
vagueness.
• Eg “I want to go to a hill station.”
Gayana M N Module 4: Semantic Analysis & Computational Discourse 13
Characteristics of Meaning Representation Languages
3. Canonical Form
• The idea is to have same representation meaning for the input that
means the same.
• Consider : “Does Kingfisher serve Hyderabad?”
“Does Kingfisher offer a flight to Hyderabad?”
“Does Kingfisher have a flight to Hyderabad?”
• The process involves choosing systematic meaning relationships
among word senses and among grammatical constituents.
• i.e. choosing the right sense in context.
Gayana M N Module 4: Semantic Analysis & Computational Discourse 14
Characteristics of Meaning Representation Languages
4. Inference and Variables
• Inference refers to system’s ability to draw valid conclusions based on
the meaning representation of inputs and representations of facts in
its knowledge base.
• Meaning representation should support this type of
derivation(inferencing).
• It should also allow the use of variables.
Gayana M N Module 4: Semantic Analysis & Computational Discourse 15
Characteristics of Meaning Representation Languages
4. Inference and Variables
• Consider the sentence:
“I would like to catch a flight to Hyderabad”
It requires complex matching, which involves the usage of
variables such as
Goes(x, Hyderabad)
– search for a known object in the knowledge base such that,
substituting x by it matches the whole proposition.
Gayana M N Module 4: Semantic Analysis & Computational Discourse 16
Characteristics of Meaning Representation Languages
5. Expressiveness:
• Natural language cover a wide variety of content(or subject matter).
• A meaning representation language must be able to represent the
meaning of this wide range of content.
• i.e it must be equipped with expressive power.
Gayana M N Module 4: Semantic Analysis & Computational Discourse 17
Module 4 Topics
Semantic Analysis Computational Discourse
• The representation of meaning • Introduction
• Syntax driven semantic analysis • Discourse Segmentation
• Word Senses • Text Coherence Relations
• Relations between senses, • Reference Resolution
• WordNet: A Database of Lexical • Anaphora resolution
Relations
• Word Sense Disambiguation
Gayana M N Module 4: Semantic Analysis & Computational Discourse
18
Syntax Driven Semantic Analysis
• It is a computational approach to semantic analysis that utilizes static
knowledge from the lexicon and the grammar.
• The basic notion that drives this approach is the principle of
compositionality, which states that the meaning of the whole can be
composed from the meaning of its parts.
• i.e we can create meaning representations from sentences from the
meanings of their consistent words
Input Semantic Output Meaning
Parser Parse Representation
Tree Analyzer
Fig 5.1 A Simple Approach to syntax driven semantic analysis
Gayana M N Module 4: Semantic Analysis & Computational Discourse
19
How semantic analyzer builds the meaning representation of a sentence?
• Usually, a semantic analyzer produces multiple ambiguous meaning
representation as output.
• Semantic analyzer requires access to domain specific and contextual
knowledge.
• POS tagger, PP attachment mechanism, word sense disambiguation
mechanism to reduce the number of possible meaning
representations produced by the semantic analyzer.
Gayana M N Module 4: Semantic Analysis & Computational Discourse 20
How semantic analyzer builds the meaning representation of a sentence?
• Consider the sentence “President nominates speaker”
• Parse tree
Gayana M N Module 4: Semantic Analysis & Computational Discourse 21
How semantic analyzer builds the meaning representation of a sentence?
• Using the parse tree, a semantic analyzer produces a semantic
representation in following steps:
1. First, it retrieves a meaning representation from the sub tree corresponding to
the verb “nominates”. The meaning representation of the verb acts as a
template for the meaning representation of the whole sentence. It contains
variables that are filled later by noun phrases.
2. It then identifies the meaning representations corresponding to the two noun
phrases.
3. Finally, it binds or associates the meaning representations of noun phrases to
the variables appearing in the meaning representation of the verb, to give the
meaning representation for the sentence as whole.
Gayana M N Module 4: Semantic Analysis & Computational Discourse 22
How semantic analyzer builds the meaning representation of a sentence?
• Using the parse tree, a semantic analyzer produces a semantic
representation in following steps:
1. First, it retrieves a meaning representation from the sub tree corresponding to
the verb “nominates”. The meaning representation of the verb acts as a
template for the meaning representation of the whole sentence. It contains
variables that are filled later by noun phrases.
2. It then identifies the meaning representations corresponding to the two noun
phrases.
3. Finally, it binds or associates the meaning representations of noun phrases to
the variables appearing in the meaning representation of the verb, to give the
meaning representation for the sentence as whole.
Gayana M N Module 4: Semantic Analysis & Computational Discourse 23
How semantic analyzer builds the meaning representation of a sentence?
• Mapping every possible tree to its semantic representation is not a viable
approach.
• We need to identify general mappings which is achieved by augmenting the
lexicon and the grammar rule with semantic attachment, and devising a mapping
between rules of the grammar and the rules of semantic representation. This is
known as rule-to-rule hypothesis.
• The semantic attachments are instructions on mapping components of a rule to a
semantic representation .
• An augmented rule takes the form:
f is a Function of the semantic attachments of A’s constituents.
Gayana M N Module 4: Semantic Analysis & Computational Discourse 24
Mapping Illustration:
• First step is to associate the constants President and Speaker with
constituents(rules) that introduce them into the sentence:
Noun President{President}
Noun speaker {speaker}
• These two augmented rules state that the meaning associated with the sub-trees
produced by the application of these rules consist of the constants President and
speaker.
• These meaning representations are passed on to their parents who contribute to
the final meaning representation as indicated by the dotted arrows.
• The augmented rule for noun phrase is NP → Noun { [Link] }
• This rule states that the meaning representation of the noun phrase is the same
as the meaning representation of its constituents.
Gayana M N Module 4: Semantic Analysis & Computational Discourse 25
Mapping Illustration:
• Next is to specify the semantics of the event underlying the sentence, i.e., the
verb nominates. The nomination event involves a nominator and a nominee.
• The semantics for nominate can be represented by the logical formula:
• This logical formula is used as the semantic attachment of nominate, resulting in
the following augmented rule.
Gayana M N Module 4: Semantic Analysis & Computational Discourse 26
Mapping Illustration:
• The next constituent to be considered is VP, which dominates both nominates
and the speaker. The grammar rule used is
• To achieve the goal, we need to incorporate the meaning of NP into the meaning
of the verb and assign it to the VP. This requires that the variable y be replaced
with the logical term speaker, as the second argument of the nominee role of
nomination event. Hence, VP. sem must tell us two things:
1. Which variable in the verb's semantic attachment is to be replaced by which
argument.
2. How the replacement is to be performed.
It demands revising the verb's semantic attachment.
Gayana M N Module 4: Semantic Analysis & Computational Discourse 27
lambda reduction
• We can use lambda calculus as the 'glue-language' to combine semantic
representations systematically.
• Lambda calculus is an extension of FOPC.
• A lambda expression λxP(x) consists of
• a λ-operator,
• a list of variables (parameters), and
• an FOPC expression in those variables.
• The parameter list in the lambda expression makes available the variable within
the body of logical expressions, for binding to external arguments provided by
the semantics of other constituents. This process of binding is known as lambda
reduction
Gayana M N Module 4: Semantic Analysis & Computational Discourse 28
lambda reduction
• The meaning of a verb phrase nominates speaker is thus:
• We now create the semantic attachment for the S rule.
Gayana M N Module 4: Semantic Analysis & Computational Discourse 29
• idioms and collocates represent sentences in which the meaning does not
depend on the meaning of the constituents or depends on it only partially. This
poses a challenge to the principle of compositionality.
• For example, consider the following sentence:
The old man finally kicked the bucket.
• The phrase kick the bucket has nothing to do with a kick or a bucket.
• To handle these situations, we need to introduce new grammar rules with
semantic attachments that introduce logical terms and predicates, which are
not related to any of the constituents of the rule.
• For example, for the idiom kick the bucket, we can introduce following grammar
rule.
VP→ kicked the bucket {died}
Gayana M N Module 4: Semantic Analysis & Computational Discourse
30
Approaches to semantic analysis:
1. Pipeline approach
2. Integrated semantic approach.
Gayana M N Module 4: Semantic Analysis & Computational Discourse
31
Approaches to semantic analysis:
1. Pipeline approach
2. Integrated semantic approach.
Gayana M N Module 4: Semantic Analysis & Computational Discourse
32
Approaches to semantic analysis:
1. Pipeline approach :
• In this approach, syntactic analysis is performed to give a parse tree.
• Then we walk through the parse tree applying semantic attachments
in a bottom-un fashion.
Gayana M N Module 4: Semantic Analysis & Computational Discourse 33
Approaches to semantic analysis:
2. Integrated semantic approach :
• In this approach, we modify the parse to include operations on
semantic attachments.
• A semantic representation of the input fragments is created as they
are being parsed.
• The advantage of this approach is that it can block a state as soon as
an ill-formed semantic fragment is detected.
Gayana M N Module 4: Semantic Analysis & Computational Discourse 34
Lexical Semantics
- Understanding the Semantics of Words
Quote: “When I use a word, it means just what
I choose it to mean – neither more nor less.” —
Lewis Carroll
Definition: Study of word meaning in language.
Goal: Explore how word meanings are
represented and structured.
Gayana M N Module4:Semantic Analysis and Discourse
1
Key Terms
Term Definition Example
Lexeme Pairing of a word's form "Cat" as a lexeme refers to the concept of a small,
with its meaning domesticated feline animal. The word “cat” is the
form, and the concept of the animal is its meaning.
Lexicon A finite set of lexemes English Lexicon:
Lexemes: Dog, Run, Happy, Technology, Eat,
Beautiful.
Lemma The base or dictionary Wordforms: Swim, Swims, Swam, Swimming
form of a word Lemma: Swim
Wordform Different morphological Lemma :Go
forms of a lemma Wordform: goes, went, gone
Lemmatization Mapping a wordform to Children → Child
its lemma/base form Better-> Good
Running->Run
Gayana M N Module4:Semantic Analysis and Discourse 2
Module 4 Topics
Semantic Analysis Discourse
• The representation of meaning • Introduction
• Syntax driven semantic • Discourse Segmentation
analysis • Text Coherence Relations
• Word Senses • Reference Resolution
• Relations between senses • Anaphora resolution
• WordNet: A Database of
Lexical Relations
• Word Sense Disambiguation
Gayana M N Module4:Semantic Analysis and Discourse 3
Word Senses
• Definition: Different meanings a
word can have.(we say that word has
different senses.)
• A word sense is a discrete
representation of one aspect of the
meaning of a word.
• Example: "Bank" can mean financial
institution (bank1) or riverbank
(bank2).
Gayana M N Module4:Semantic Analysis and Discourse 4
Word Senses - Types
Homonymy: Unrelated meanings
Polysemy: Related meanings
Metonymy: part of the same entity or institution.
Gayana M N Module4:Semantic Analysis and Discourse 5
Word Senses - Types
Homonymy
• Homonymy occurs when a single word has completely different, unrelated
meanings.
• In this case, "bank" has two unrelated meanings:
[Link] institution: "She went to the bank to deposit her paycheck."
[Link] or edge of a river: "They enjoyed a picnic on the bank of the
river."
• Here, "bank" has two entirely different meanings, and there’s no conceptual
link between a financial institution and the side of a river.
Gayana M N Module4:Semantic Analysis and Discourse 6
Word Senses - Types
Polysemy
• Polysemy occurs when a word has multiple related meanings, where each
meaning is connected in some way.
[Link] institution: "The bank offers loans at competitive rates."
[Link] to store or accumulate something: "She has a memory bank filled
with vivid childhood moments."
[Link] of something valuable: "The data bank contains a wealth of
information."
• In polysemy, the meanings of "bank" all relate to the concept of storing or
holding something valuable, whether it’s money, memories, or data.
Gayana M N Module4:Semantic Analysis and Discourse 7
Word Senses - Types
Metonymy
Metonymy involves using a word to refer to something closely associated with
it, often part of the same entity or institution.
1. Bank as an institution: "The bank raised interest rates to combat inflation.”
"Here, "bank" refers not to the physical building or the actual financial reserves,
but rather to the people or policies of the financial institution itself.
2. Bank as its employees: "The bank was supportive of her loan application.”
In this case, "bank" stands in for the people or staff within the institution who helped
with her application.
Gayana M N Module4:Semantic Analysis and Discourse 8
Distinguishing Word Senses
• Different criteria for identifying word senses.
• Zeugma Test: Combining different meanings in one sentence to test clarity.
• Example: Consider the following sentences:
S1: Which of those flights serve breakfast?
S2: Does Midwest Express serve Philadelphia?
S3: Does Midwest Express serve breakfast and Philadelphia?“
• Statement S3(a case of zeugma) indicates there is no sensible way to make a
single sense of serve work for both breakfast and Philadelphia
Gayana M N Module4:Semantic Analysis and Discourse 9
Homophones and Homograph
• Homophones : Are the senses that are linked to lemmas with the same
pronunciation but different spellings.
• Examples:
• Right / Write:
• "You have the right to remain silent."
• "Please write your name on the form.“
• Sea / See:
• "The sea was calm today."
• "I can see the mountains from here."
Gayana M N Module4:Semantic Analysis and Discourse 10
Homophones and Homograph
Homographs:
• These are distinct senses linked to lemmas with the same orthographic form
but different pronunciations.
• words that are spelled the same but have different meanings and sometimes
different pronunciations.
• Example:
• Bow (to bend forward) / Bow (a type of knot or weapon):
• "He took a bow after his performance."
• "She tied the ribbon in a neat bow."
Gayana M N Module4:Semantic Analysis and Discourse 11
Module 4 Topics
Semantic Analysis Discourse
• The representation of meaning • Introduction
• Syntax driven semantic • Discourse Segmentation
analysis • Text Coherence Relations
• Word Senses • Reference Resolution
• Relations between senses • Anaphora resolution
• WordNet: A Database of
Lexical Relations
• Word Sense Disambiguation
Gayana M N Module4:Semantic Analysis and Discourse 12
Relations between senses
Synonymy: Nearly identical meanings (e.g., car/automobile).
Antonymy: Opposites (e.g., big/small, rise/fall).
Hypernymy/Hyponymy: Class inclusion relationships (e.g., animal/dog).
Meronymy :Part-whole relationship.(e.g., wheel/car)
Gayana M N Module4:Semantic Analysis and Discourse 13
Synonymy and Antonymy
• When the meaning of two senses of two different words are identical or
nearly identical, we say the two senses are synonyms.
• i.e, they are substitutable one for the other in any sentence without changing
the truth condition of the sentence. We often say in this case, that the two
words have the same propositional meaning.
• Example:
Big / Large:
“He lives in a big house.”
“He lives in a large house.”
Gayana M N Module4:Semantic Analysis and Discourse 14
Synonym Usage -Example
• Big / Large:
• Here, we cannot substitute large for big:
• Miss Nelson became a kind of big sister to Benjamin.
• Miss Nelson became a kind of large sister to Benjamin.
• The word big has a sense that means being older, or grown up, while large
lacks this sense.
• Thus, it will be convenient to say that some senses of big and large are (nearly)
synonymous while other ones are not.
Gayana M N Module4:Semantic Analysis and Discourse 15
Synonymy and Antonymy
• Antonymy occurs when two words have opposite meanings.
• Examples:
• Hot / Cold:
• “The coffee was too hot to drink.”
• “The ice cream was too cold to hold.”
• Happy / Sad:
• “He was happy to see his friends.”
• “She felt sad when they left.”
• Another groups of antonyms is reversive, which describe some sort of
change or movement in opposite directions, such as rise/fall or up/down.
Gayana M N Module4:Semantic Analysis and Discourse 16
Hyponymy
• Hyponym: One sense is hyponym of another sense if the first
sense is more specific denoting a subclass of the other.
• Car is a hyponym of vehicle.
• Hypernym: One sense is Hypernym of another sense if the first
sense is more Generic denoting a Broad class of the other.
• vehicle is a hypernym of car.
• animal is a hypernym of dog.
• Application: Helps create taxonomies or hierarchical structures.
Gayana M N Module4:Semantic Analysis and Discourse 17
Hyponymy and Taxonomies
• It is unfortunate that the two words (hypernym and hyponym) are
very similar and hence easily confused.
• For this, superordinate is often used instead of hypernym.
Superordinate vehicle fruit furniture mammal
Hyponym car mango chair dog
• Application: Helps create taxonomies or hierarchical structures.
Gayana M N Module4:Semantic Analysis and Discourse 18
Hyponymy and Taxonomies
• The concept of hyponymy is closely related to several other notions that play central roles in
computer science, biology, and anthropology and computer science.
• The term ontology usually refers to a set of distinct objects resulting from an analysis of a
domain, or microworld.
• A taxonomy is a particular arrangement of the elements of an ontology into a tree-like
class inclusion structure.
• Normally, there are a set of well formedness constraints on taxonomies that go beyond their
component class inclusion relations.
• For example, the lexemes hound, mutt, and puppy are all hyponyms of dog, as are golden
retriever and poodle, but it would be odd to construct a taxonomy from all those pairs since
the concepts motivating the relations is different in each case.
• Instead, we normally use the word taxonomy to talk about the hypernymy relation between
poodle and dog.
• By this definition, taxonomy is a subtype of hypernymy.
Gayana M N Module4:Semantic Analysis and Discourse 19
Meronymy and Semantic Fields
• Meronymy: Part-whole relations
• Example: Wheel is Meronym of a car.
• Holonymy: Whole-Part relations
• Example: Car is a Holonym of Wheel.
Gayana M N Module4:Semantic Analysis and Discourse 20
Meronymy and Semantic Fields
• Semantic Fields: It is a model of a more integrated or holistic or
relationship among entire set of Words from a single domain
• Example: reservation, flight, travel, buy, price, cost ,fare , rates, meal, plane
• We could assert individual lexical relations of hyponymy etc. between many
of the words in this list. And these basic senses are drawn based on
background information.
Gayana M N Module4:Semantic Analysis and Discourse 21
Module 4 Topics
Semantic Analysis Discourse
• The representation of meaning • Introduction
• Syntax driven semantic • Discourse Segmentation
analysis • Text Coherence Relations
• Word Senses • Reference Resolution
• Relations between senses • Anaphora resolution
• WordNet: A Database of
Lexical Relations
• Word Sense Disambiguation
Gayana M N Module4:Semantic Analysis and Discourse 22
WordNet: A Resource for Lexical Relations
• The most commonly used resource for English sense relations is the
WordNet lexical database
• WordNet consists of three separate databases, one each for
• Nouns,
• Verbs,
• Adjectives & Adverbs.
• Each database consists of a set of lemmas, each one annotated with a set of
senses.
• The WordNet 3.0 release has 117,097 nouns, 11,488 verbs, 22,141
adjectives, and 4,601 adverbs.
• The average noun has 1.23 senses, and the average verb has 2.16 senses.
Gayana M N Module4:Semantic Analysis and Discourse 23
WordNet: A Resource for Lexical Relations
Gayana M N Module4:Semantic Analysis and Discourse 24
WordNet: A Resource for Lexical Relations
• Each word is provided with the following :
• Gloss – Dictionary Style Definition
• Synset- List of synonyms for the sense
• Usage Examples
• Synonym Set (or Synset): The set of near synonyms for a WordNet sense.
Gayana M N Module4:Semantic Analysis and Discourse 25
WordNet: Relations
Gayana M N Module4:Semantic Analysis and Discourse 26
WordNet Relations
• Examples of Relations:
• Hypernymy/Hyponymy: Chains leading from specific to general terms.
• Synonymy: Synsets grouped by similar meanings.
• Applications in Computational Semantics
• Word Sense Disambiguation: Using context to determine word
meanings.
• Machine Translation and Information Retrieval: Enhancing accuracy
through semantic knowledge.
Gayana M N Module4:Semantic Analysis and Discourse 27
Module 4 Topics
Semantic Analysis Discourse
• The representation of meaning • Introduction
• Syntax driven semantic • Discourse Segmentation
analysis • Text Coherence Relations
• Word Senses • Reference Resolution
• Relations between senses • Anaphora resolution
• WordNet: A Database of
Lexical Relations
• Word Sense Disambiguation
Gayana M N Module4:Semantic Analysis and Discourse 28
Word Sense Disambiguation
• The task of selecting the correct sense for a word is called word sense
disambiguation, or WSD.
• Disambiguating word senses has the potential to improve many natural
language processing tasks such as
• Machine translation
• Question Answering
• Information Retrieval
• Text Classification etc..
Gayana M N Module4:Semantic Analysis and Discourse 29
Word Sense Disambiguation
• WSD algorithm takes as input a word in context along with a fixed inventory of
potential word senses and return the correct word sense for that use.
• Both the nature of the input and the inventory of senses depends on the
task.
• Search Engines: Improve relevance of search results.
• Machine Translation: Select correct word meanings for accurate
translation.
• Speech Recognition: Disambiguate homophones using context.
• Text Summarization: Ensure key information is accurately condensed.
Gayana M N Module4:Semantic Analysis and Discourse 30
Word Sense Disambiguation
Gayana M N Module4:Semantic Analysis and Discourse 31
Variants of WSD task
1. Lexical Sample Task
The Lexical Sample Task focuses on disambiguating a predefined
set of target words. Each of these words has annotated
occurrences with specific senses in a corpus.
2. All word Task
The All-Word Task requires disambiguating all content words (e.g.,
nouns, verbs, adjectives, adverbs) in a given corpus or passage, not
just a predefined set of target words.
Gayana M N Module4:Semantic Analysis and Discourse 32
Variants of WSD task
1. Lexical Sample Task Example
Target Word: "Bank"
The system is provided with multiple sentences, and its goal is to identify the sense of the target word
"bank" in each case.
Sentences:
He went to the bank to deposit money.
The river overflowed its bank after heavy rainfall.
The fisherman set up a tent on the river bank and waited.
Sense Inventory for "Bank":
Bank (Financial Institution): A place where money is kept, saved, or exchanged.
Bank (River Edge): The land alongside a river.
Expected Output:
Sense 1: Bank (Financial Institution)
Sense 2: Bank (River Edge)
Sense 2: Bank (River Edge)
Gayana M N Module4:Semantic Analysis and Discourse 33
Variants of WSD task
2. All-Word Task Example • Bank:
Full Sentence: • Financial Institution.
• The fisherman went to the bank to set up his tent • Edge of a river.
and then withdrew money from the bank. • Tent:
• A portable shelter made of cloth or canvas.
The system is required to disambiguate all content
words in the sentence, not just a predefined target • Withdrew:
word. • Removed money from a bank account.
Words to Disambiguate: • Pulled back or retreated.
1. Fisherman
Expected Output:
2. Bank (occurs twice) 1. Fisherman → Sense 1: A person who catches fish.
3. Tent 2. Bank (1st instance) → Sense 2: Edge of a river.
4. Withdrew 3. Tent → Sense 1: A portable shelter.
4. Bank (2nd instance) → Sense 1: Financial
Sense Inventory: Institution.
• Fisherman: 5. Withdrew → Sense 1: Removed money from a
• A person who catches fish as a hobby or for bank account.
livelihood.
Gayana M N Module4:Semantic Analysis and Discourse 34
Variants of WSD task
Gayana M N Module4:Semantic Analysis and Discourse 35
Approaches to WSD
1. Knowledge-Based Methods
• Use lexical resources like WordNet.
• Techniques: Lesk Algorithm, Graph-Based Approaches.
2. Supervised Learning
• Train models on labeled datasets.
• Features: Context words, POS tags.
• Algorithms: Decision Trees, SVMs.
Gayana M N Module4:Semantic Analysis and Discourse 36
Approaches to WSD
3. Unsupervised Learning
• Clustering techniques.
• Relies on word co-occurrence patterns.
4. Deep Learning Approaches
• Contextual embeddings (e.g., BERT, GPT).
• Sequence models like RNNs, LSTMs.
Gayana M N Module4:Semantic Analysis and Discourse 37
Module 4 Topics
Semantic Analysis Computational Discourse
• The representation of meaning • Introduction
• Syntax driven semantic • Discourse Segmentation
analysis • Text Coherence Relations
• Word Senses • Reference Resolution
• Relations between senses • Anaphora resolution
• WordNet: A Database of
Lexical Relations
• Word Sense Disambiguation
Gayana M N Module4:Semantic Analysis and Discourse 38
Introduction to computational Discourse
Gracie: Oh yeah. . . and then Mr. and Mrs. Jones were having
matrimonial trouble, and my brother was hired to watch Mrs. Jones.
George: Well, I imagine she was a very attractive woman.
Gracie: She was, and my brother watched her day and night for six
months.
George: Well, what happened?
Gracie: She finally got a divorce.
George: Mrs. Jones?
Gracie: No, my brother’s Wife
Gayana M N Module4:Semantic Analysis and Discourse 39
Introduction to computational Discourse
• Language doesn’t normally consist of isolated, unelated sentences, but
instead of collocated, structured, coherent groups of sentences.
• Such a Coherent Structured group of sentences is called as discourse.
• Computational discourse deals with understanding the structure and
meaning of text beyond individual sentences, focusing on context and
relationships between sentences.
• Importance in NLP:
• Enables coherent text understanding.
• Critical for tasks like dialogue systems, summarization, and machine
translation.
Gayana M N Module4:Semantic Analysis and Discourse 40
Introduction to computational Discourse
Key Terms:
• Monologue: communication flows in only one direction that is, from the
speaker to the hearer
• Dialogue: generally, consist of many different types of communicative acts:
asking questions, giving answers, making corrections, and so forth.
Key Components:
• Discourse Structure: How ideas are organized.
• Discourse Semantics: Understanding meaning relationships across text.
Gayana M N Module4:Semantic Analysis and Discourse 41
Introduction to computational Discourse
• Doing automatic disambiguation is a difficult task.
• Example: Consider the sentences
“The Tin Woodman went to the Emerald City to see the Wizard of Oz and
ask for a heart. After he asked for it, the Woodman waited for the
Wizard’s response.”
Deciding he and it using WSD is difficult.
• The goal of deciding what pronouns and other noun phrases refer to is
called coreference resolution.
• Coreference resolution is important for information extraction,
summarization, and for conversational agents.
Gayana M N Module4:Semantic Analysis and Discourse 42
Introduction to computational Discourse
Coherence Relations:
• Identifying the relation between two or more sentences by providing
information such as subordinate and background information etc.
• Example:
• Consider Summarization application:
• Input:
• Output:
Gayana M N Module4:Semantic Analysis and Discourse 43
Introduction to computational Discourse
• Coherence is also a property of a good text, automatically detecting
coherence relations is also useful for tasks that measure text quality, like
automatic essay grading.
• In automatic essay grading, short student essays are assigned a grade by
measuring the internal coherence of the essay as well as comparing its
content to source material and hand-labeled high-quality essays.
• Coherence is also used to evaluate the output quality of natural language
generation systems.
• Determining the discourse structure can help in determining coreference.
Gayana M N Module4:Semantic Analysis and Discourse 44
Coherence
• Coherent discourse refers to text or speech where the ideas or
utterances are logically connected, making it easy to understand
and follow.
• These connections, often termed coherence relations, give
structure and meaning to a discourse.
• EXPLANATION is one such coherence relation, where one
utterance explains the reason or cause for the event or state
mentioned in another.
Gayana M N Module4:Semantic Analysis and Discourse 45
Coherence
1. Explanation Relation
• One utterance provides the cause or reason for the statement in another,
enhancing understanding of the discourse.
• Example:
• "John didn’t come to the party. He was feeling unwell."
• Here, the second sentence explains the reason for John not coming to
the party.
• "The car stopped suddenly because it ran out of fuel."
• "Because it ran out of fuel" explains why "the car stopped suddenly."
Gayana M N Module4:Semantic Analysis and Discourse 46
Coherence
2. Other Coherence Relations
• Coherence relations help build logical connections within a discourse.
• Some common types include:
• Cause and Effect:
• "It rained heavily. As a result, the match was canceled."
• The second sentence gives the effect of the first.
• Elaboration:
• "Alice is a doctor. She specializes in pediatrics."
• The second sentence elaborates on Alice's profession.
• Contrast:
• "Mary likes tea; however, John prefers coffee."
• The second sentence contrasts with the first.
• Temporal Relation:
• "She finished her homework before watching TV."
• The second action follows the first in time.
Gayana M N Module4:Semantic Analysis and Discourse 47
Module 4 Topics
Semantic Analysis Computational Discourse
• The representation of meaning • Introduction
• Syntax driven semantic • Discourse Segmentation
analysis • Text Coherence Relations
• Word Senses • Reference Resolution
• Relations between senses • Anaphora resolution
• WordNet: A Database of
Lexical Relations
• Word Sense Disambiguation
Gayana M N Module4:Semantic Analysis and Discourse 48
Discourse Segmentation
• Discourse Segmentation is the task of dividing a text into smaller,
meaningful units, called segments, based on the discourse
structure.
• Each segment corresponds to a subtopic, paragraph, or discourse
unit, contributing to the text's overall meaning and coherence.
Gayana M N Module4:Semantic Analysis and Discourse 49
Discourse Segmentation - Advantages
• Improves Text Understanding: Helps identify the logical flow of ideas within a
document.
• Supports NLP Applications:
• Text Summarization: Ensures segments are summarized appropriately.
• Information Retrieval: Improves relevance by focusing on specific
segments.
• Dialogue Systems: Allows the system to handle context-sensitive
interactions.
• Facilitates Coherence Analysis: Enables better modeling of relationships
between segments for higher-level text analysis.
Gayana M N Module4:Semantic Analysis and Discourse 50
Discourse Segmentation - Types
Linear Segmentation:
• Divides text into a sequence of non-overlapping segments in a linear order.
• Example: Breaking a scientific article into sections like Introduction,
Methodology, and Conclusion.
Hierarchical Segmentation:
• Organizes segments in a tree-like structure, showing relationships between
them.
• Example: Analyzing nested topics or subtopics in a textbook.
Gayana M N Module4:Semantic Analysis and Discourse 51
Discourse Segmentation – Methods
1. Unsupervised Methods
• No labeled data; relies on inherent text properties like cohesion.
• Example Algorithm: TextTiling (Hearst, 1997)
• Measures lexical cohesion across sentences or paragraphs.
• Steps:
• Tokenize text and remove stopwords.
• Create "pseudo-sentences" of fixed length.
• Compute similarity between adjacent text blocks using cosine similarity.
• Identify segment boundaries, where similarity dips significantly.
• Applications: News story segmentation, summarization.
Gayana M N Module4:Semantic Analysis and Discourse 52
Discourse Segmentation – Methods
2. Supervised Methods
• Uses labeled data with pre-annotated segment boundaries.
• Features:
• Lexical cohesion (word overlap, cosine similarity).
• Discourse Markers: Words or phrases indicating transitions (e.g., "however," "next," "in
conclusion").
• Named entities or specific domain cues.
• Techniques:
• Binary classification to predict segment boundaries.
• Sequence models like Hidden Markov Models (HMMs) or Conditional Random Fields
(CRFs).
• Applications: Speech segmentation, document classification.
Gayana M N Module4:Semantic Analysis and Discourse 53
Unsupervised Discourse Segmentation
• An important class of unsupervised algorithms for the linear discourse
segmentation task rely on the concept of cohesion.
• Lexical cohesion is cohesion indicated by relations between words in two
units such as use of an identical word, s synonym or a hypenym.
• Example:
"A dog barked loudly in the park. The barking startled the joggers. The
animal was looking for its owner, who had wandered off. Eventually, the
owner returned and calmed the pet.“
• Cohesion Elements:
• Repetition: "barked" and "barking".
• Synonyms: "dog" and "animal".
• Pronouns: "its" and "who".
• Substitution: "the pet" replaces "dog".
Gayana M N Module4:Semantic Analysis and Discourse 54
Unsupervised Discourse Segmentation
• Cohesion chain
• A cohesion chain connects different parts of a text by referring back to the same concept or
entity using cohesive devices. These devices help readers understand how ideas are related
across sentences or paragraphs.
• Example:
"Alice visited the zoo last weekend. She was thrilled to see the lions. The big cats roared loudly,
which fascinated her. Alice also enjoyed watching the elephants."
• Cohesion Chains in the Text:
[Link] Chain for Alice (Repetition, Pronoun):
Alice → She → her → Alice
[Link] Chain for Lions (Synonym, Lexical Tie):
Lions → big cats
[Link] Chain for Emotional Reaction (Lexical Tie):
Thrilled → fascinated
Gayana M N Module4:Semantic Analysis and Discourse 55
Unsupervised Discourse Segmentation
• Coherence Vs Cohesion
Coherence Cohesion
Coherence is the logical flow of ideas in a text Cohesion refers to the grammatical and
or discourse. It refers to the overall sense or lexical links that connect sentences and
meaning that connects sentences and phrases, ensuring surface-level connectivity.
paragraphs into a unified whole.
Achieved through meaning, logic, and Achieved through cohesive ties such as
thematic progression. pronouns, conjunctions, and repetition.
"John couldn’t open the door. He searched for "John couldn’t open the door. He realized he
the key in his pocket.“ didn’t have the key.“
The two sentences are logically connected The pronoun "he" and the repeated reference
through the idea of needing a key to open the to "key" create linguistic ties between the
door. sentences.
Gayana M N Module4:Semantic Analysis and Discourse 56
Text Tiling Algorithm
• Unsupervised Discourse segmentation algorithm
• Approach: Cohesion-based.
• Objecive: Identify topic boundaries in a document.
• The algorithm has three steps:
1. tokenization,
2. lexical score determination,
3. boundary identification.
Gayana M N Module4:Semantic Analysis and Discourse 57
Text Tiling Algorithm
1. Tokenization:
• In the tokenization stage,
• each space-delimited word in the input is converted to lower-case,
• words in a stop list of function words are thrown out, and
• the remaining words are morphologically stemmed.
• The stemmed words are grouped into pseudo-sentences of length w = 20
(equal-length pseudo-sentences are used rather than real sentences).
Gayana M N Module4:Semantic Analysis and Discourse 58
Text Tiling Algorithm
2. Lexical Score Determination
• Next is to look at each gap between pseudo-sentences and compute a lexical cohesion
score across that gap.
• The cohesion score is defined as the average similarity of the words in the pseudo-
sentences before gap to the pseudo-sentences after the gap.
• We generally use a block of k = 10 pseudo-sentences on each side of the gap.
• To compute similarity, we create a word vector b from the block before the gap, and a
vector a from the block after the gap, where the vectors are of length N (the total number
of non-stop words in the document) and the ith element of the word vector is the
frequency of the word wi .
• Now we can compute similarity using cosine:
Gayana M N Module4:Semantic Analysis and Discourse 59
Text Tiling Algorithm
The dot product between the first two pseudo sentences is:
2*1 + 1*1 + 2*1 1*1 + 2*1 = 8
Gayana M N Module4:Semantic Analysis and Discourse 60
Text Tiling Algorithm
3. Boundary Identification
• Finally, we compute a depth score for each gap, measuring the depth of the similarity
valley at the gap.
• The depth score is the distance from the peaks on both sides of the valley to the valley.
• Here, it is
• Boundaries are assigned at any valley which is deeper than a cutoff theshold
Gayana M N Module4:Semantic Analysis and Discourse 61
Supervised Discourse Segmentation
• For some kinds of discourse segmentation tasks, it is relatively easy to
acquire boundary-labeled training data.
• For spoken discourse task of segmentation of broadcast news, we first need
to assign boundaries between news stories. This is a simple discourse
segmentation task, and training sets with hand-labeled news story
boundaries exist.
• Similarly, for the task of paragraph segmentation, we can find the labeled
data from the web or other sources.
• In this task, discourse markers or cue words are often used.
• Example: In broadcast news segmentation, important discourse markers
might include a phrase like “good evening, I’m (PERSON)”.
• Methods used are classification algorithm such as SVM, decision Tree etc.
Gayana M N Module4:Semantic Analysis and Discourse 62
Module 4 Topics
Semantic Analysis Computational Discourse
• The representation of meaning • Introduction
• Syntax driven semantic • Discourse Segmentation
analysis • Text Coherence Relations
• Word Senses • Reference Resolution
• Relations between senses • Anaphora resolution
• WordNet: A Database of
Lexical Relations
• Word Sense Disambiguation
Gayana M N Module4:Semantic Analysis and Discourse 63
Text coherence
• Text coherence refers to the quality of a
text that makes it understandable and
logically connected.
• It ensures that sentences and ideas in
the text flow smoothly, making it easier
for readers to comprehend.
Gayana M N Module4:Semantic Analysis and Discourse 64
Hobbs' Coherence Relations
Coherence as Inference:
• Hobbs proposed that coherence in discourse arises from the inferences readers
make to establish connections between sentences or segments.
• These inferences are based on logical, causal, or situational relationships.
Types of Coherence Relations :
• Hobbs identified several types of relations that can be inferred to achieve coherence
in discourse:
• Cause
• Result
• Explanation
• Parallel
• Elaboration
• Occasion…. etc
Gayana M N Module4:Semantic Analysis and Discourse 65
a. Cause
• One segment provides the reason or cause for another.
• Example:
• She was late because she missed the bus.
• Here, the second clause explains the cause of being late.
b. Result
• One segment presents the outcome of an action or event described in another.
• Example:
• She studied hard, so she passed the exam.
• The second clause shows the result of the first.
Gayana M N Module4:Semantic Analysis and Discourse 66
c. Explanation
• A segment provides an explanation or rationale for a preceding statement.
• Example:
• She quit her job. She wasn’t happy with the management.
• The second clause explains the reason for the first.
d. Parallelism
• Segments are connected by their similarity in structure or content.
• Example:
• He likes football. She likes basketball.
• Both sentences share a parallel structure, enhancing coherence.
Gayana M N Module4:Semantic Analysis and Discourse 67
e. Elaboration
• One segment provides additional details or expands on the idea of another.
• Example:
• She bought a car. It’s a red sedan with leather seats.
• The second sentence elaborates on the type of car.
f. Occasion
• Signifies that the second event or situation follows naturally or logically from the first. It often
implies that one event sets the stage or conditions for the next to happen.
• Examples:
• She lit the candle. The room filled with a warm glow.
• The act of lighting the candle sets the occasion for the room being illuminated.
• He finished his dinner. Then, he went for a walk.
• The first event (dinner) provides the occasion for the second (walk).
Gayana M N Module4:Semantic Analysis and Discourse 68
• coherence of an entire discourse can be done by considering the hierarchical structure between
coherence relations.
• Example:
• John went to the bank to deposit his paycheck. (S1)
• He then took a train to Bill’s car dealership. (S2)
• He needed to buy a car. (S3)
• The company he works for now isn’t near any public transportation. (S4)
• He also wanted to talk to Bill about their softball league. (S5)
The discourse structure is shown below: Here, each node represents discourse segment.
Gayana M N Module4:Semantic Analysis and Discourse 69
Rhetorical Structure Theory
• It is a framework in computational linguistics and discourse analysis used to
describe and analyze the organization of texts.
• It focuses on the functional relationships between different parts of a text,
aiming to explain how these parts contribute to the overall coherence and
purpose of the discourse.
• Consider the following text:
She didn’t go to the party because she was tired. Besides, she had an early
meeting the next day.
Nucleus: She didn’t go to the party.
Satellite 1: because she was tired. (Cause-Effect relation)
Satellite 2: Besides, she had an early meeting the next day. (Additional
evidence relation)
Gayana M N Module4:Semantic Analysis and Discourse 70
Rhetorical Structure Theory
1. RST defines a set of rhetorical relations that connect units of text, explaining
how they function together.
• These relations are categorized as either nucleus or satellite:
• Nucleus: The central, essential part of the text that carries the primary
meaning.
• Satellite: The supporting part that elaborates, provides evidence, or
serves another function for the nucleus.
Gayana M N Module4:Semantic Analysis and Discourse 71
Rhetorical Structure Theory
2. Considers Text as a Hierarchical Structure
• Texts are represented as hierarchical structures, often visualized as trees.
• The nodes represent parts of the text (e.g., clauses, sentences, or larger
units).
• The relationships between these parts indicate their rhetorical roles.
3. Coherence arises from the rhetorical relations between units, showing how
each part contributes to the text's overall communicative purpose.
Gayana M N Module4:Semantic Analysis and Discourse 72
Rhetorical Structure Theory
Types of Rhetorical Relations: RST proposes various relations, such as:
[Link]: Adds more detail to the nucleus.
Example: The weather is perfect today. The sky is clear, and the temperature is mild.
[Link]: Provides proof or justification for the nucleus.
Example: She must be home because her car is parked in the driveway.
[Link]-Effect: Indicates a causal relationship between nucleus and satellite.
Example: He missed the bus because he woke up late.
[Link]: Highlights differences between the nucleus and satellite.
Example: Although it was sunny, the beach was empty.
[Link]: Acknowledges a counterargument but supports the nucleus.
Example: Even though the roads were icy, she managed to arrive on time.
[Link]: The satellite summarizes the information in the nucleus.
Example: The team worked tirelessly, staying late every night. In short, their dedication paid off.
Gayana M N Module4:Semantic Analysis and Discourse 73
Automatic Coherence Assignment
• Definition:
• Determining coherence relations between sentences algorithmically.
• Goals:
• Extract discourse trees/graphs.
• Assign relations like Contrast, Cause, Explanation.
• Challenges:
• Ambiguity in cue phrases.
• Implicit relations without explicit markers.
Gayana M N Module4:Semantic Analysis and Discourse 74
Automatic Coherence Assignment
• Cue-Phrase-Based Algorithm
• Steps:
1. Identify cue phrases (e.g., because, although).
2. Segment text into discourse units.
3. Classify relationships between units.
• Examples:
“Because of the low atmospheric pressure, liquid water evaporates
instantly.”
• Relation: Cause
Gayana M N Module4:Semantic Analysis and Discourse 75
Automatic Coherence Assignment
1. Identify cue phrases
• Cue phrase (or discourse marker or cue word) is a word or phrase that
functions to signal discourse structure, especially by linking together discourse
segments.
• Connective cue phrases : Conjunctions, adverbs (Eg: Because)
• Ambiguity in Cue Phrases
• Examples:
• With its distant orbit, Mars exhibits frigid weather conditions. (discourse
marker)
• We can see Mars with a telescope. (sentential use)
• Disambiguation Techniques:
• Capitalization, punctuation rules.
• Using syntactic parsers for complex rules.
Gayana M N Module4:Semantic Analysis and Discourse 76
Automatic Coherence Assignment
2. Segment text into discourse units.
• This step deals with determining the correct coherence relation by segmenting the
text into discourse segments.
• Discourse segments generally correspond to clauses or sentences, although
sometimes they are smaller than clauses.
• A clause or clause-like unit is a more appropriate size for a discourse segment.
• One way to segment these clause-like units is to use hand-written segmentation rules
based on individual cue phrases.
Gayana M N Module4:Semantic Analysis and Discourse 77
Automatic Coherence Assignment
3. Classify relationships between units
• The third step in coherence extraction is to automatically classify the relation between
each pair of neighboring segments. For each discourse marker. Rules can be
written(rule based approach).
• However, many coherence relations are signaled by more implicit cues.
• Implicit Coherence Relations
• Challenge:
• Many relations lack explicit markers (e.g., contrast without but).
• Example:
• I don’t want a truck; I’d prefer a convertible.
• Implicit cues: Negation, lexical relations, syntactic parallelism.
• Solution:
• Use implicit features like negation, parallelism.
Gayana M N Module4:Semantic Analysis and Discourse 78
Automatic Coherence Assignment
Machine Learning for Coherence
• Approach:
• Label large corpora using strong discourse markers (e.g., consequently).
• Train supervised models on these examples.
• Features:
• Word pairs (e.g., truck/convertible).
• Parts of speech, stemmed words, and lexical cues.
• Tools:
• Naive Bayes, feature selection.
Gayana M N Module4:Semantic Analysis and Discourse 79
Module 4 Topics
Semantic Analysis Computational Discourse
• The representation of meaning • Introduction
• Syntax driven semantic • Discourse Segmentation
analysis • Text Coherence Relations
• Word Senses • Reference Resolution
• Relations between senses • Anaphora resolution
• WordNet: A Database of
Lexical Relations
• Word Sense Disambiguation
Gayana M N Module4:Semantic Analysis and Discourse 80
Reference Resolution
• Reference resolution is the task of determining what entities are referred to
by which linguistic expressions.
• Example:
• Sentence: "Alice left her bag on the table, and now she can't find it."
• Task in NLP:
• Resolve "her" to "Alice."
• Resolve "she" to "Alice."
• Resolve "it" to "her bag."
Gayana M N Module4:Semantic Analysis and Discourse 81
Reference Resolution
• A natural language expression used to perform reference is called a referring
expression, and the entity that is referred to is called Referent.
• Example:
Sentence: "Tom saw a bird in the garden. The bird was building a nest."
Referring Expression: "The bird."
Referent: The specific bird in the garden.
• Two referring expressions that are used to refer to the same entity are said to
corefer.
• Example:
• Sentence: "The book was fascinating, and I couldn’t put it down."
• Coreferring Expressions: "The book" and "it."
Gayana M N Module4:Semantic Analysis and Discourse 82
Reference Resolution
Discourse Model
• The discourse model contains representations of the entities that have been
referred to in the discourse and the relationship in which they participate.
• There are two fundamental operations to the discourse model.
• Evoke- When a referent is first mentioned in the discourse, we say that a
representation for it is evoked into the model.
• Access-Upon Subsequent mention, this representation is accessed from
the model.
• two reference resolution tasks:
• coreference resolution
• pronominal anaphora resolution
Gayana M N Module4:Semantic Analysis and Discourse 83
Reference Resolution
Discourse Model
Gayana M N Module4:Semantic Analysis and Discourse 84
Reference Resolution
Discourse Model
Consider the following paragraph:
A coreference resolution algorithm would need to find four coreference chains:
1. { Victoria Chen, Chief Financial Officer of Megabucks Banking Corp since 2004, her, the 37-
year-old, the Denver-based financial-services company’s president, She}
2. { Megabucks Banking Corp, the Denver-based financial-services company, Megabucks }
3. { her pay }
4. { Lotsabucks }
Gayana M N Module4:Semantic Analysis and Discourse 85
Module 4 Topics
Semantic Analysis Computational Discourse
• The representation of meaning • Introduction
• Syntax driven semantic • Discourse Segmentation
analysis • Text Coherence Relations
• Word Senses • Reference Resolution
• Relations between senses • Anaphora resolution
• WordNet: A Database of
Lexical Relations
• Word Sense Disambiguation
Gayana M N Module4:Semantic Analysis and Discourse 86
Anaphora Resolution
• A key task in Natural Language Processing (NLP) that involves identifying the
antecedent of an anaphor in a text.
• An anaphor is a word or phrase (commonly pronouns like he, she, it, they or
other referring expressions like this, that, such) that depends on another part
of the text for its meaning.
• The antecedent is the earlier expression to which the anaphor refers.
• Importance of Anaphora Resolution
• Understanding Context: It enables machines to comprehend the meaning
of sentences in context, which is essential for tasks like summarization,
machine translation, and question-answering.
• Coherence in Text: Helps maintain textual coherence by linking different
parts of the text together.
Gayana M N Module4:Semantic Analysis and Discourse 87
Anaphora Resolution
Examples of Anaphora
[Link]:
• John went to the store. He bought milk.
• He refers to John.
[Link] Descriptions:
• I saw a dog. The dog was barking.
• The dog refers to a dog.
[Link]-Anaphora:
• I like the red car, but I prefer the blue one.
• One refers to car.
[Link] Anaphora:
• She fell off the ladder. This scared everyone.
• This refers to the event of falling off the ladder.
Gayana M N Module4:Semantic Analysis and Discourse 88
Anaphora Resolution
Steps in Anaphora Resolution
[Link] the Anaphor:
Detect the referring expression, like a pronoun or definite description.
[Link] Candidate Antecedents:
Locate possible antecedents in the text, often nouns or noun phrases.
[Link] the Correct Antecedent:
Select the antecedent based on contextual, grammatical, and semantic
factors.
Gayana M N Module4:Semantic Analysis and Discourse 89
Anaphora Resolution
Example:
Text: Sara went to the market. She bought apples.
[Link] the Anaphor:
Anaphor: She
[Link] Candidate Antecedents:
Candidates: Sara
(Implicit) Other possible antecedents introduced earlier in a larger text or context.
[Link] the Correct Antecedent:
She is singular and feminine, matching Sara in the preceding sentence.
Sara is the most recent entity and logically fits as the buyer of apples.
Chosen Antecedent: Sara
Gayana M N Module4:Semantic Analysis and Discourse 90
Gayana M N Module4:Semantic Analysis and Discourse 91