0% found this document useful (0 votes)
11 views25 pages

Introduction to NLP Applications and Techniques

The document provides an introduction to Natural Language Processing (NLP), covering various applications such as Question Answering, Information Extraction, and Sentiment Analysis. It discusses rule-based approaches, including the use of regular expressions and the concepts of precision, recall, and edit distance in evaluating NLP models. Additionally, it highlights the importance of minimizing false positives and negatives in NLP tasks.

Uploaded by

kimbull.jinho
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
11 views25 pages

Introduction to NLP Applications and Techniques

The document provides an introduction to Natural Language Processing (NLP), covering various applications such as Question Answering, Information Extraction, and Sentiment Analysis. It discusses rule-based approaches, including the use of regular expressions and the concepts of precision, recall, and edit distance in evaluating NLP models. Additionally, it highlights the importance of minimizing false positives and negatives in NLP tasks.

Uploaded by

kimbull.jinho
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd

1.

Introduction to NLP
Disclaimer

• Some slides are based on Stanford NLP/IR book (and slides)


• [Link]
• [Link]
• Some slides from Berkeley AI course
• [Link]
• All three courses are available on Youtube (excellent follow-up
materials of this course)

2
Introducing NLP Applications:
Question Answering (QA)
• Won Jeopardy on February 16, 2011!

WILLIAM WILKINSON’S
“AN ACCOUNT OF THE PRINCIPALITIES OF
WALLACHIA AND MOLDOVIA” Bram Stoker
INSPIRED THIS AUTHOR’S
MOST FAMOUS NOVEL

3
Information Extraction (IE)

4
Dialogue (Limited)

• SingleMulti-turn
• ClassificationGeneration

5
Sentiment Analysis
Attributes:
zoom
affordability
size and weight
flash
ease of use
Size and weight
✓ • nice and compact to carry!
• since the camera is small and light, I won't need to carry
✓ around those heavy, bulky professional cameras either!
✗ • the camera feels flimsy, is plastic and very light in weight you
have to be very delicate in the handling of this
6 camera
2. Rule-based Approaches

Classification:
F(text) is {spam,non-spam}
{positive, negative}
Let’s build a rule-based spam filter

• “Viagra” or “Cialis”
• Is this a good rule?

8
Regular expressions
• A formal language for specifying text strings
• How can we search for any of these?
• woodchuck
• woodchucks
• Woodchuck
• Woodchucks
Regular Expressions: ? * + .

Pattern Matches
colou?r Optional color colour
previous char
oo*h! 0 or more of oh! ooh! oooh! ooooh!
previous char
o+h! 1 or more of oh! ooh! oooh! ooooh!
previous char
Stephen C Kleene
baa+ baa baaa baaaa baaaaa
beg.n begin begun begun beg3n Kleene *, Kleene +
Example

• Find me all instances of the word “the” in a text.


the
Misses capitalized examples
[tT]he
Incorrectly returns other or theology
[^a-zA-Z][tT]he[^a-zA-Z]
How do we know if our rule is good?

• The process we just went through was based on fixing


two kinds of errors
• Matching strings that we should not have matched (there,
then, other)
• False positives (Type I)
• Not matching things that we should have matched (The)
• False negatives (Type II)
Precision vs. Recall

• In NLP we are always dealing with these kinds of errors.


• Reducing the error rate for an application often
involves two antagonistic efforts:
• Increasing accuracy or precision (minimizing false positives)
• Increasing coverage or recall (minimizing false negatives).
The 2-by-2 contingency table
correct not correct
selected tp fp
not selected fn tn
Precision and recall
• Precision: % of selected items that are correct
Recall: % of correct items that are selected

correct not correct


selected tp fp
not selected fn tn
A combined measure: F
• A combined measure that assesses the P/R tradeoff is F measure
(weighted harmonic mean):

• The harmonic mean is a very conservative average; see


[Link]

• People usually use balanced F1 measure


• i.e., with  = 1 (that is,  = ½): F = 2PR/(P+R)
3. Word (Lexical) Similarity

Definition
How similar are two strings?

• Spell correction • Computational Biology


• The user typed “graffe” • Align two sequences of nucleotides
Which is closest? AGGCTATCACCTGACCTCCAGGCCGATGCCC
• graf TAGCTATCACGACCGCGGTCGATTTGCCCGAC
• graft • Resulting alignment:
• grail
• giraffe -AGGCTATCACCTGACCTCCAGGCCGA--TGCCC---
TAG-CTATCAC--GACCGC--GGTCGATTTGCCCGAC

• Also for Machine Translation, Information Extraction, Speech Recognition


Edit Distance

• The minimum edit distance between two strings


• Is the minimum number of editing operations
• Insertion
• Deletion
• Substitution
• Needed to transform one into the other
Minimum Edit Distance

• Two strings and their alignment:


Minimum Edit Distance

• If each operation has cost of 1


• Distance between these is 5
• If substitutions cost 2 (Levenshtein)
• Distance between them is 8
How to find the Min Edit Distance?
• Searching for a path (sequence of edits) from the start string to
the final string:
• Initial state: the word we’re transforming
• Operators: insert, delete, substitute
• Goal state: the word we’re trying to get to
• Path cost: what we want to minimize: the number of edits

22
Weighted Edit Distance

• Why would we add weights to the computation?


• Spell Correction: some letters are more likely to be mistyped than others
• Biology: certain kinds of deletions or insertions are more likely than
others
Confusion matrix for spelling errors
Keyboard used contributes significantly to model

You might also like