Module 1-1
Module 1-1
1
Outline
• Introduction to NLP
• Language modelling
• Word level Analysis: Use of FSA
• Spelling error detection
• POS Tagging
• Syntactic Analysis, parsing
• Semantic Analysis
• Information Retrieval system
• Machine Translation system
• Speech synthesis and recognition
• Other Applications
2
Books
3
Question Answering: IBM’s Watson
WILLIAM WILKINSON’S
“AN ACCOUNT OF THE PRINCIPALITIES OF
WALLACHIA AND MOLDOVIA” Bram Stoker
INSPIRED THIS AUTHOR’S
MOST FAMOUS NOVEL
4
Event: Curriculum meet
Information Extraction Date: Jan-16-2012
Start: 10:00am
End: 11:30am
Subject: curriculum meeting Where: Gates 159
Date: January 15, 2012
To: Dan Jurafsky
5
Information Extraction & Sentiment Analysis
Attributes:
zoom
affordability
size and weight
flash
ease of use
Size and weight
• nice and compact to carry!
✓ • since the camera is small and light, I won't need to carry around those
✓ heavy, bulky professional cameras either!
• the camera feels flimsy, is plastic and very light in weight you have to be
✗ very delicate in the handling of this camera
6
Machine Translation
这 不过 是 一 个 时间 的 问题 .
7
Language Technology
making good progress
Sentiment analysis still really hard
Best roast chicken in Bhubaneswar!
mostly solved Question answering (QA)
The waiter ignored us for 20 minutes.
Q. How effective is ibuprofen in reducing
Spam detection Coreference resolution fever in patients with acute febrile illness?
Let’s go to Agra! ✓
✗
Carter told Mubarak he shouldn’t run again. Paraphrase
Buy Sreeram stock …
Word sense disambiguation XYZ acquired ABC yesterday
(WSD)
I need new batteries for my mouse. ABC has been taken over by XYZ
Part-of-speech (POS) tagging
ADJ ADJ NOUN VERB ADV Summarization
Colorless green ideas sleep furiously. Parsing The Dow Jones is up Economy is
I can see Alcatraz from the window! The S&P500 jumped good
Housing prices rose
Named entity recognition (NER) Machine translation (MT)
PERSON ORG LOC 第13届上海国际电影节开幕… Dialog Where is Citizen Kane playing in SF?
Einstein met with UN officials in Princeton
The 13th Shanghai International Film Festival…
Castro Theatre at 7:30. Do
Information extraction (IE) you want a ticket?
Party
You’re invited to our dinner May 27
party, Friday May 27 at 8:30 add
8
Ambiguity makes NLP hard:
9
Why else is natural language
understanding difficult?
non-standard English segmentation issues idioms
Great job @justinbieber! Were
SOO PROUD of what youve dark horse
the New York-New Haven Railroad
accomplished! U taught us 2 get cold feet
the New York-New Haven Railroad
#neversaynever & you yourself lose face
should never give up either♥ throw in the towel
unfriend Mary and Sue are sisters. Where is A Bug’s Life playing …
Retweet Mary and Sue are mothers. Let It Be was recorded …
… a mutation on the for gene …
10
Introduction
Natural Language:
• Language spoken and written by human for general purpose
communication
• Examples : Odia, Hindi, English, French, Chinese, ………,
etc
11
Overview of the NLP Field
Computer Science
12
Need of Processing Human Languages
• To develop different Human-computer Interactive systems
through natural languages.
13
Applications of NLP
14
Example: The Search Engines
15
Applications of NLP
• Information Retrieval
• Machine Translation
• Text Summarization
• Information Extraction Other Applications:
• Question Answering • Dictionary word suggestion
• Text to Speech • Spelling error/ grammar
• Speech to Text correction
• Chat bots
• Sentiment analysis
• …..
16
Applications of NLP Cont…
17
Applications of NLP Cont…
Text Summarization:
• The goal of automatic summarization is to take an information source,
extract contents from it and provide the most important contents to the
user in a concise form as a summary.
Question Answering:
• Concerns with building systems that can automatically answer
questions posed by users in a natural language.
18
Applications of NLP Cont…
Text to Speech:
• A text-to-speech (TTS) system converts natural language text into
speech
• It has its applications in designing speech synthesizers, screen
readers, language learning apps, etc.
Speech to Text:
• Speech to text (STT) conversion is the process of converting spoken
words into written texts.
• This is also called speech recognition.
• It has its applications in designing text dictation systems, command
and control, audio document transcription, etc
19
Advanced Applications
• Virtual Assistants:
• Apple- Siri
• Amazon- Alexa
• Microsoft- Cortana
20
2. A Brief History of NLP
Application Development
• 1950: Mathematical model of computation (Turing machine)
Turing machines, first described by Alan Turing in Turing 1936–7, are simple
abstract computational devices that help investigate the extent and limitations of
what can be computed. Turing’s ‘automatic machines’, were devised for the
computing of real numbers. Today, they are considered to be one of the foundational
models of computability and (theoretical) computer science
21
History
The rules of how to order words help the language parts make
sense. Sentences often start with a subject, followed by a predicate
(or just a verb in the simplest sentences) and contain an object or a
complement (or both), which shows, for example, what's being acted
upon.
22
History
23
History
24
Issues and Processing Complexities
1. Ambiguity:
• Natural languages are highly ambiguous
• Words in a natural language may have a number of different
meanings.
Examples:
• Bank (River bank/ Financial Institution),
• Bat( cricket bat/ species)
• For many NLP task, the proper sense of each ambiguous word
in a sentence must be determine to interpret correct meaning.
25
Cont…
2. Language Variability:
• There are various ways to express meaning
26
Cont…
3. Difficult to incorporate human cognition over
machine:
27
How can a machine understand these
differences?
Examples:
28
Issues in Processing Indian
Languages
1. Unlike English, Indic scripts have a nonlinear structure.
• Example:
Language Script
English English language
Hindi हिन्दी भाषा
Odia ଓଡ଼ିଆ ଭାଷା
Bengali: বাাংলা ভাষা
29
Cont…
2. English uses SVO (Subject-Verb-Object) format while Indian
languages uses SOV (Subject-Object-Verb) format
Example:
English: pooja plays veena
(S) (V) (O)
30
Cont…
Example:
Usne khaanaa khaya
or
Usne khaaya Khanaa
31
Cont…
4. Have rich set of morphological variants
Hindi:
ghodda, ghodde, ghoddi, ghoddon……
32
Cont…
5. Extensive and productive use of complex predicates
Example:
हिन्दी
शब्द
सम्पर्
ू ण
33
Cont…
6. Ambiguity:
Example:
सोना
34
Language and Grammar
35
2 Main Components of NLP
36
3. Phases of Natural Language
Processing
Example:
• cat (1 morpheme)
• cats (2 morphemes: ‘cat’, ‘s’)
Example:
“she is going to the market” valid sentence
“she are going to the market” invalid sentence
39
Semantic Analysis
• Concerned with creating meaningful representation of linguistic
inputs.
40
Discourse Analysis
• Attempts to interpret structure and meaning of even larger units
• The meaning of any sentence depends upon the meaning of the
sentence just before it.
• Requires Discourse Knowledge
• Knowledge of how the meaning of a sentence is determined
by processing sentences
Example:
She forgot her book.
To understand to whom- ‘she’, ‘her’ refers, processing of
previous sentences are required
41
Pragmatic Analysis
• The meaning of a sentence can’t be always derived based on the
meaning of its words. Multiple interpretations of a sentence can be
possible.
• Syntactic structure and compositional semantics fail to explain
these interpretations.
• Pragmatic analysis deals with the purposeful use of sentences in
situations.
• It requires real world knowledge along with language knowledge
Example:
Do u know who am I?
(different context of use can be
possible)
42
Major Approaches of NLP
• Empiricist approach
Rationalist Approach (Symbolic approach)
• Assumes existence of some language faculty in human brain
• No language faculty
Data driven:
• Assumes the existence of large amount of data and techniques to learn
syntactic pattern
• Requires less human effort
• Performance depends on quantity of data
• Adaptive to noisy data
4. Language Modeling
47
4. Language Modeling
• Language: Primary mean of communication used by human
48
Types of Language Modeling
2 types
Example:
n-gram model
49
Grammar-based Language Modeling
• Uses the grammar of a language to create its model
Example:
S -> NP+VP //S: Sentence, NP: Noun Phrase, VP: Verb Phrase
Neha was going to school.
51
Transformational Grammar
(Chomsky 1957)
• Assumes 2 levels of existence of sentences:
• Surface level
• Deep root level Deep structure:
S
Example:
sub obj
S: pooja plays veena reln
Surface structure:
S S
NP VP NP VP
2. Transformational rule
3. Morphophonemic rule
53
1. Phrase Structure Grammar
• Consists of rules that generate natural language sentences
Example:
54
2. Transformational Rule
Example:
• converting active sentence into passive
Plays-> played
55
Consider the active sentence :
The police will catch the snatcher.
The application of phrase structure rules will assign the structure shown
below:
56
• The transformational rules will convert it into:
The + snatcher + will + be + en + catch + by + police
57
3. Morphophonemic rule
• is the study of word formation – how words are built up from smaller pieces
• Match each sentence representation to a string of phonemes:
Example:Morphology
morpho-phonemic rule will convert:
catch +en →caught
58
Limitations of Grammar-based Language
Models
• A large number of re-writable rules which are language specific
59
Probabilistic Language Models
Example:
n-gram model
• More variables:
P(A,B,C,D) = P(A)P(B|A)P(C|A,B)P(D|A,B,C)
• The Chain Rule in General
P(x1,x2,x3,…,xn) = P(x1)P(x2|x1)P(x3|x1,x2)…P(xn|x1,…,xn-1)
The Chain Rule applied to compute joint
probability of words in sentence
•Simplifying assumption:
Andrei Markov
P(the | its water is so transparent that) » P(the | that)
•Or maybe
70
Simplest case: Unigram model
P(w1w2 … wn ) » Õ P(w i )
i
Some automatically generated sentences from a unigram model
Example:
If total no. of words in a corpus= 1,000,00 and the word ‘the’
appears 69971 times.
using un-igram model,
P(the)= 69971/100000
= 0.69971
= 0.7
72
Bigram model
Example-1:
Training Corpus:
<s> I am Sam </s>
<s> Sam I am </s>
<s> I do not like eggs </s>
74
Example-2
Training set:
The Arabian Knights
These are the fairy tales of the east
The stories of the Arabian Knights are translated in many
languages
Test sentence:
The Arabian Knights are the fairy tales of the east
75
Answer
Training set:
The Arabian Knights
These are the fairy tales of the east
The stories of the Arabian Knights are translated in many languages
= 0.00268
77
Raw bigram counts
• Out of 9222 sentences
Raw bigram probabilities
• Normalize by unigrams:
• Result:
Bigram estimates of sentence probabilities
Example:
The Arabian Knights
P (Knights/ the Arabian)
81
Example
Find probability of the below given test sentence using a tri-gram model.
Training set:
The Arabian Knights
These are the fairy tales of the east
The stories of the Arabian Knights are translated in many languages
Test Sentence:
The Arabian Knights are the fairy tales of the east
Answer:
Trigram probabilities:
P(The/<s1><s2>) = 2/3= 0.666=0.67 P(Arabian/<s2>the)=1/2=0.5
P (Knights/ the Arabian)= 2/2=1.0 P(are/Arabian Knights)= ?
……// find the tri-gram probabilities of other sequences
82
Estimating Sentence Probability
P(The Arabian Knights are the fairy tales of the east)
= P(<s1><s2>The Arabian Knights are the fairy tales of the east)
= ?
// calculate the sentence probability
83
Practice Examples on n-gram Model
84
Example-1
Consider the following frequency matrix and Find the likelihood
estimate of the below sentence using bigram model:
85
Answer
= 0.4
86
Example-2
For the below given training set find bi-gram probability of all terms
in training set and find the probability of the given test sentence .
Training set:
I am Sam
Sam I am
I am not Sam
Test sentence:
I am not Sam
87
Answer
Bigram probabilities:
P(I/<s>)=2/3=0.67
P(am/I)=3/3=1.0
P(sam/am)=1/3=0.33
P(sam/<s>)=1/3=0.33
P(I/Sam)=1/3)=0.33
P(not/am)=1/3=0.33
P(sam/not)=1/1=1.0
88
Homework Questions
89
Example-3
Answer: 0.2211
90
Example-4
Answer: 0.0103
91
Data Sparseness Problem in n-gram Model
• An n-gram that does not occur in the training set is assigned a zero
probability
• Smoothing techniques:
92
Add-one Smoothing
c(wi-1, wi )+1
PAdd-1 (wi | wi-1 ) =
c(wi-1 )+V
Where, V=no. of words in the vocabulary
93
N-gram models