Coreference Resolution in NLP
Coreference Resolution in NLP
Module-V
Mr Gunti Spandan
Assistant Professor
Department of CSE
GITAM School of Technology (GST)
Email: sgunti@[Link]
1. Daniel Jurafsky, James H Martin, “Speech and Language Processing: An introduction to Natural Language
Processing, Computational Linguistics and Speech Recognition”, 2/e, Prentice Hall, 2008.
2. C. Manning, H. Schutze, “Foundations of Statistical Natural Language Processing”, MIT Press. Cambridge, MA,
1999.
3. Jacob Eisenstein, Introduction to Natural Language Processing, MIT Press, 2019.
REFERENCE BOOK:
1. Jalaj Thanaki, Python Natural Language Processing: Explore NLP with machine Learning and deep learning
Techniques, Packt, 2017.
(1) Victoria Chen, CFO of Megabucks Banking, saw her pay jump to $2.3 million, as the 38-year-old became the
company’s president. It is widely known that she came to Megabucks from rival Lotsabucks.
• Two or more referring expressions that are used to refer to the same discourse entity are said to corefer;
thus, Victoria Chen and she corefer in (1).
• A question answering system that uses Wikipedia to answer a question about Marie Curie
must know who she was in the sentence “She was born in Warsaw”.
Discourse Model:
This is used by Natural language understanding systems (and humans) to interpret linguistic expressions.
• It is a mental model (Fig. 21.1) that the understander builds incrementally when interpreting a text.
• It contains representations of the entities referred to in the text, as well as
properties of the entities and relations among them.
Anaphora:
Reference in a text to an entity that has been previously introduced into the discourse is Anaphora and
the referring expression used is said to be an anaphor (or anaphoric).
Singleton:
An entity that has only a single mention in a text (like Lotsabucks in (1)) is called a singleton.
Coreference Resolution:
Coreference resolution is the task of determining whether two mentions corefer, by which we mean they
refer to the same entity in the discourse model (the same discourse entity).
• Coreference resolution thus comprises two tasks (although they are often performed jointly):
(1) identifying the mentions, and (2) clustering them into coreference chains/discourse entities.
• Two mentions corefered if they are associated with the same discourse entity.
• We have to decide which real world entity is associated with this discourse entity.
Ex: the mention Washington might refer to the US state, or the capital city, or
the person George Washington (interpretation of the sentence will be very different for each of these).
Department of CSE, GST CSEN4141: NLP 7
21.1 Coreference Phenomena: Linguistic Background
Referring expressions:
They are used to evoke and access entities in the discourse model, and talk about linguistic features of
the anaphor/antecedent relation (like number/gender agreement, or properties of verb semantics).
Exs:
a. Mrs. Martin was so very kind as to send Mrs. Goddard a beautiful goose.
b. He had gone round one day to bring her some walnuts.
c. I saw this beautiful cauliflower today.
(3) Pronouns: Another form of definite reference is pronominalization, used for entities that are extremely salient in
the discourse.
Exs:
(a) Emma smiled and chatted as cheerfully as she could.
Pronouns can also participate in cataphora, in which they are mentioned before their referents are, as in (b).
(b) Even before she saw it, Dorothy had been thinking about the Emerald City every day.
Here, the pronouns she and it both occur before their referents are introduced.
• Under the relevant reading, her does not refer to some woman in context,
but instead behaves like a variable bound to the quantified expression every dancer.
(a) I just bought a copy of Thoreau’s Walden. I had bought one five years ago.
That one had been very tattered; this one was in much better condition.
this NP is ambiguous; in colloquial spoken English, it can be indefinite, as in (21.6), or definite, as in (21.14).
new NPs:
brand new NPs: these introduce entities that are discourse-new and hearer-new like a fruit or some
walnuts.
unused NPs: these introduce entities that are discourse-new but hearer-old like Hong Kong, Marie Curie,
or
the New York Times.
old NPs:
also called evoked NPs, these introduce entities that already in the discourse model, hence are both
discourse-old and hearer-old, like it in “I went to a new restaurant. It was...”.
Neither the chef nor the dough were in the discourse model based on the first sentence of either example,
but the reader can make a bridging inference that these entities should be added to the discourse model
and associated with the restaurant and the ingredients,
based on world knowledge that restaurants have chefs and dough is the result of mixing flour and liquid.
• Many noun phrases or other nominals are not referring expressions, although they may bear a confusing
superficial resemblance.
Department of CSE, GST CSEN4141: NLP 12
Continued...
Ex: The NP a car in the following example does not create a discourse referent:
(21.20) Janet doesn’t have a car.
and cannot be referred back to by anaphoric it or the car:
(21.21) *It is a Toyota.
(21.22) *The car is red.
The four common types of structures that are not counted as mentions in coreference tasks:
(and hence complicate the task of mention-detection)
(i) Appositives:
• An appositional structure is a noun phrase that appears next to a head noun phrase, describing the head.
(21.23) Victoria Chen, CFO of Megabucks Banking, saw ...
(21.24) United, a unit of UAL, matched the fares.
• Appositional NPs are not referring expressions, instead functioning as supplementary description of the head NP.
(iii) Expletives:
Many uses of pronouns like it in English and corresponding pronouns in other languages are not referential.
Such expletive or pleonastic cases include it is raining, in idioms like hit it off, or in particular syntactic situations
like clefts (i) or extraposition (ii):
(i) It was Emma Goldman who founded Mother Earth.
(ii) It surprised me that there was a herring hanging on her wall.
Second, singular they has become much more common, in which they is used to describe singular individuals,
often useful because they is gender neutral.
Thus a third person pronoun (he, she, they, him, her, them, his, her, their) must have a third person antecedent
(one of the above or any other noun phrase).
In English this occurs only with third-person singular pronouns, which distinguish between male (he, him, his),
female (she, her), and non-personal (it) grammatical genders.
Non-binary pronouns like ze or hir may also occur in more recent texts.
Knowing which gender to associate with a name in text can be complex, and may require world knowledge about
the individual.
Exs:
(i) Maryam has a theorem. She is exciting. (she=Maryam, not the theorem)
(ii) Maryam has a theorem. It is exciting. (it=the theorem, not Maryam)
• Oversimplifying a bit, reflexive pronouns like himself and herself corefer with the subject of the most immediate
clause that contains them (Ex: (i)), whereas non-reflexives cannot corefer with this subject (Ex: (2)).
Exs:
(i) Janet bought herself a bottle of fish sauce. [herself = Janet]
(ii) Janet bought her a bottle of fish sauce. [her ≠ Janet]
(v) Recency:
Entities introduced in recent utterances tend to be more salient than those introduced from utterances further back.
Thus, in the below, the pronoun it is more likely to refer to Jim’s map than the doctor’s map.
Ex: The doctor found an old map in the captain’s chest. Jim found an even older map hidden on the shelf.
It described an island.
(i) Billy Bones went to the bar with Jim Hawkins. He called for a burger. [ he = Billy ]
(ii) Jim Hawkins went to the bar with Billy Bones. He called for a burger. [ he = Jim ]
• There are two possible referents for it, the soup and the bowl.
• The verb eat requires that its direct object denote something edible, and
this constraint can rule out bowl as a possible referent.
Coherent discourse:
Creating a text by taking random sentences each from many different sources and pasting them together,
is not a coherent discourse.
(1) Sentences or clauses in real discourses are related to nearby sentences in systematic ways.
Jane took a train from Paris to Istanbul. She had to attend a conference.
(2) In a coherent discourse, some entities are salient and the discourse focuses on them and doesn’t go back and
forth between multiple entities.
This is called entity-based coherence.
Ex: An incoherent passage (the salient entity seems to wildly swing from
John to Jenny to the piano store to the living room, back to Jenny, then the piano
again).
John wanted to buy a piano for his living room.
Jenny also wanted to buy a piano.
He went to the piano store.
It was nearby.
The living room was on the second floor.
She didn’t find anything she liked.
The piano he bought was hard to get up to that floor.
Department of CSE, GST CSEN4141: NLP 23
Continued...
Entity-based coherence models measure this kind of coherence by tracking salient entities across a discourse.
Nearby coherent sentences are generally about the same topic and use the same or similar vocabulary to discuss
these topics.
This is called lexical cohesion (the sharing of identical or semantically related words in nearby sentences).
Ex: The fact that the words house, chimney, garret, closet, and window—
all of which belong to the same semantic field— appear in the two sentences (given below),
or that they share the identical word shingled, is a cue that the two are tied together as a discourse:
• In addition to the local coherence between adjacent or nearby sentences, discourses also exhibit
global coherence.
Examples:
• Many genres of text are associated with particular conventional discourse structures.
Word sense: A sense (or word sense) is a discrete representation of one aspect of the meaning of a word.
WordNet:
• A large online thesaurus —a database that represents word senses—with versions in many languages.
• It also represents relations between senses.
Ex: There is an IS-A relation between dog and mammal (a dog is a kind of mammal) and
a part-whole relation between engine and car (an engine is a part of a car).
Necessity of WordNet:
Knowing the relation between two senses can play an important role in language understanding.
• Glosses are not a formal meaning representation; they are just written for people.
(1) Dictionaries:
Definitions of right, left, red, and blood from the American Heritage Dictionary:
right adj. located nearer the right hand esp. being on the right when facing the same direction as the observer.
left adj. located nearer to this side of the body than the right.
red n. the color of blood or a ruby.
blood n. the red liquid that circulates in the heart, arteries and veins of animals.
Department of CSE, GST CSEN4141: NLP 28
Continued...
Circularity in these definitions:
Definition of right: makes two direct references to itself
Definition of left: contains an implicit self-reference in the phrase this side of the body,
which presumably means the left side
Definitions of red and blood: refer to each other
For humans, such entries are useful since the user of the dictionary has sufficient grasp of these other terms.
Dictionaries often give example sentences along with glosses, and these can again be used to help build a sense
representation.
(2) Thesauruses:
• Sense relations of this sort (IS-A, or antonymy) are explicitly listed in on-line databases like WordNet.
• Given a sufficiently large database of such relations, many applications are quite capable of performing
sophisticated semantic tasks about word senses (even if they do not really know their right from their left).
A zeugma is a literary term for using one word to modify two other words, in two different ways.
Ex: “She broke his car and his heart.”
When you use one word to link two thoughts, you're using a zeugma.
• The task of selecting the correct sense for a word is called word sense disambiguation, or WSD.
• WSD algorithms take as input a word in context and a fixed inventory of potential word senses and
outputs the correct word sense in context.
Ex: SemCor corpus showing the WordNet sense numbers of the tagged words
(standard WSD notation is used)
• the SemCor-based WSD task is to choose the correct sense from the possible senses in WordNet.
• For fruit, this would mean choosing between the correct answer from
fruit1n the ripened reproductive body of a seed plant,
fruit2n yield (an amount of a product) and
fruit3n the consequence of some effort or action.
(1) Choosing the most frequent sense for each word from the senses in a labeled corpus.
For WordNet, this corresponds to the first sense, since senses in WordNet are generally ordered from
most frequent to least frequent based on their counts in the SemCor sense-tagged corpus.
The most frequent sense baseline can be quite accurate, and is therefore often used as a default,
A word appearing multiple times in a text or discourse often appears with the same sense.
Semantic roles, express the role that arguments of a predicate take in the event, codified in databases like
PropBank and FrameNet.
Semantic role labeling, the task of assigning roles to spans in sentences.
• The roles of the subjects of the verbs break and open are Breaker and Opener respectively.
• These deep roles are specific to each event;
Breaking events have Breakers, Opening events have Openers, and so on.
• This approach performs translation by analysing the source language syntax and transforming it into the target
language’s structure.
• It typically involves parsing the source sentence (using a context-free grammar), mapping its tree structure into
• For example, in English (SVO: Subject-Verb-Object) versus Japanese (SOV: Subject-Object-Verb), syntactic
• The source sentence's meaning is represented using semantic frames or logical forms, which are then used
to generate the target text.
• This is effective for resolving structural differences and can handle idiomatic expressions better.
• For example, the German sentence "Ich esse gern" can be represented semantically as LIKE(I, EAT), to
correctly translate as "I like to eat" in English.
Interlingua Approach
• The source sentence is first converted to Interlingua, which captures its universal semantic meaning, and
then rendered into the target language.
• Its main advantages include scalability across many languages and preservation of deep meaning.
• Challenges include the difficulty of designing truly universal representations, disambiguating meaning, and
handling varying cultural concepts.
Statistical Machine Translation
• This approach leverages probabilistic models trained on large parallel corpora (bilingual datasets).
• The translation model learns from example sentence pairs—no hand-crafted linguistic rules are needed.
• Methods like the noisy channel model (using Bayes’ theorem) are central, and large bilingual resources
(such as the Canadian Hansards corpus with 3 million aligned sentences) are important for accuracy.
Language models: By using large monolingual corpora for the target language, the translation output
becomes more fluent and natural.
Statistical Machine Translation
Strengths:
•Can handle language pairs without deep linguistic understanding.
Weaknesses:
•Key Idea: To find the best English sentence that matches a given French sentence, you use probability
models.
•Main Formula: You want the English e that makes P(e|f) highest, which splits into P(e) (how natural
English is) × P(f|e) (how well the French matches the English).
•Three Components:
• Language Model: Checks if the English is fluent.
• Translation Model: Checks if the translation makes sense.
• Decoder: System that finds the best match.
Text & Word Alignment
Sentence Alignment: Matching whole sentences between two languages. Algorithms use things
like sentence length or punctuation.
Word Alignment: Match up individual words inside matching sentence pairs. This is tricky—what
if one English word matches several French words or vice versa?
Challenges: Includes tricky cases like idioms ("kick the bucket") or words that don’t translate
(articles, short words).
IBM Model 1: Foundation
Simplest statistical translation model (Brown et al., 1993).
Model ignores the order of words—just treats them like a bag of words!
NULL Alignment: Some French words (like “le” or “de”) might not need an English word—so they
align to ‘NULL’.
IBM Model 1: Key Assumptions
Independence: Each translation happens separately.
The Formula:
P(f,a∣e)=ε∏t(f∣e)
Goal: Figure out word translations from lots of sentence pairs where alignments aren’t marked.
Solution: Ask the computer to guess alignments, then improve its guesses over several rounds
(Expectation-Maximization).
count(f|e) = Σ δ(f, e)
•End Result: The model’s guess about translations gets really accurate.
Practical Training Example
Corpus: "the house" → "la maison", "the book" → "le livre"
Real-world applications:
Brand monitoring: Companies analyse social media chatter to understand public perception of their products or
campaigns .
Customer feedback analysis: Businesses automatically sort through vast amounts of customer reviews and
support interactions to identify areas for improvement.
Market research: Financial analysts use sentiment analysis on news articles and financial reports to predict
market trends.
Types:
• Extractive summarization: Identifies and extracts key sentences or phrases directly from the source text to
form the summary [1].
• Abstractive summarization: Generates new sentences to describe the main points, much like a human would
summarize an article, often requiring more advanced neural network models [1].
Real-world applications:
• News aggregation: Providing a quick snippet of a news story so users can decide if they want to read the full
article.
• Research paper analysis: Helping researchers quickly grasp the core findings of academic papers.
• Meeting transcripts: Generating concise summaries of long meeting notes or recordings.
Real-world applications:
• Search engines: Providing direct answers at the top of search results pages for factual queries.
• Virtual assistants: Answering quick factual questions asked of devices like smart speakers or phone assistants.
• Enterprise knowledge management: Allowing employees to quickly find specific data points within a
company's vast internal documentation