0% found this document useful (0 votes)
5 views58 pages

N-Grams and Language Modeling Insights

The document discusses N-grams and probabilistic language models, emphasizing their utility in various applications such as speech recognition, machine translation, and text completion. It highlights the importance of assigning probabilities to sequences of words and the limitations of N-gram models in handling long-distance dependencies. Additionally, it notes the sensitivity of N-gram models to the training corpus used.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views58 pages

N-Grams and Language Modeling Insights

The document discusses N-grams and probabilistic language models, emphasizing their utility in various applications such as speech recognition, machine translation, and text completion. It highlights the importance of assigning probabilities to sequences of words and the limitations of N-gram models in handling long-distance dependencies. Additionally, it notes the sensitivity of N-gram models to the training corpus used.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Module 2 Part 3

Word Level Analysis


N-Grams-
14
Language Model
15 N-Grams-Language Model
🠶 Formal grammars (e.g. regular, context free)
give a hard “binary” model of the legal
sentences in a language.
🠶 For N L P, a probabilistic model of a language
that gives a probability that a string is a
member of a language is more useful.
🠶 To specify a correct probability distribution,
the probability of all sentences in a language
must sum to 1.
16 Uses of Language Models
🠶 Speech recognition
🠶 “I ate a cherry” is a more likely sentence than “Eye
eight uh Jerry”
🠶 O C R & Handwriting recognition
🠶 More probable sentences are more likely correct
readings.
🠶 Machine translation
🠶 More likely sentences are probably better translations.
🠶 Generation
🠶 More likely sentences are probably better N L
generations.
🠶 Context sensitive spelling correction
🠶 “Their are problems wit this sentence.”
17
Completion Prediction

🠶 A language model also supports


predicting the completion of a
sentence.
🠶 Please turn off your cell _
🠶 Your program does not
🠶 Predictive text input systems can
guess what you are typing and give
choices on how to complete it.
18 Completion Prediction
🠶 C
🠶 Models that assign a probability to each possible next
word.
🠶 The same models will also serve to assign a probability to
an entire sentence.
🠶 Such a model, for example, could predict that the
following sequence has a much higher probability of
appearing in a text:
all of a sudden I notice three guys standing on the
sidewalk
than does this same set of words in a different order:
on guys all I of notice sidewalk three a sudden standing
the
19 Completion Prediction

🠶 Probabilities are essential in any task in which we have to


identify words in noisy, ambiguous input, like speech
recognition.
🠶 For a speech recognizer to realize that you said I will be
back soonish and not I will be bassoon dish, it helps to
know that back soonish is a much more probable sequence
than bassoon dish.
🠶 For writing tools like spelling correction or grammatical
error correction, we need to find and correct errors in
writing .
🠶 Their are two midterms,
in which There was mistyped as Their,
20 Predicting Probability
🠶 Everything has improve, in which improve should have
been improved.
🠶 This model allowing us to help users by detecting and
correcting these errors.
🠶 Assigning probabilities to sequences of words is also
essential in machine translation.
🠶 he introduced reporters to the main contents of the
statement
🠶 he briefed to reporters the main contents of the statement
🠶 he briefed reporters on the main contents of the statement
🠶 Which is the correct one????
🠶 Models that assign probabilities to sequences of words are
called language models or LMs.
21 Language Model
22 Language Model
23 Probabilistic Language Model
24 N-Gram Models
🠶 Estimate probability of each word given prior context.
🠶 P(phone | Please turn off your cell)
🠶 Number of parameters required grows exponentially with
the number of words of prior context.
🠶 An N-gram model uses only N−1 words of prior context.
🠶 Unigram: P(phone)
🠶 Bigram: P(phone | cell) p(cell | your)
🠶 Trigram: P(phone | your cell) P (cell | off your)
🠶 The Markov assumption is the presumption that the
future behavior of a dynamical system only depends on
its recent history.
🠶 In particular, in a kth-order Markov model, the next
state only depends on the k most recent states, therefore
an N-gram model is a (N−1)-order Markov model
25 Cha in Rule of Probability
26 C ha i n Rule of Probability &
Conditional Probabilities
27 Computing Conditional Probabilities
28 N -Grams
29 N-Gram Model Formulas

C ompiled by : Mrs. Angelin Florence A


30 Estimating Probabilities

C ompiled by : Mrs. Angelin Florence A


31 N- Grams

C ompiled by : Mrs. Angelin Florence A


A Problem for N-Grams:
32
Long Distance Dependencies
🠶 Many times local context does not provide the most
useful predictive clues, which instead are provided
by long-distance dependencies.
🠶 Syntactic dependencies
🠶 “The man next to the large oak tree near the grocery store on
the corner is tall.”
🠶 “The men next to the large oak tree near the grocery store on
the corner are tall.”
🠶 Semantic dependencies
🠶 “The bird next to the large oak tree near the grocery store on
the corner flies rapidly.”
🠶 “The man next to the large oak tree near the grocery store on
the corner talks rapidly.”
🠶 More complex models of language are needed to
handle such dependencies.
33 Probabilistic Language Modeling
34 Estimating Bigram Probabilities
35 Example
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56 N Gram Sensitivity to the Training Corpus
57 N Gram Sensitivity to the Training Corpus

🠶 The longer the context on which we train the model, the more
coherent the sentences
58 N Gram Sensitivity to the Training Corpus

🠶 The longer the context on which we train the model, the more
coherent the sentences
🠶 N= 884,647
🠶 V=29,066
🠶 N-gram Probability matrices are spare.
🠶 There are V 2 = 844,000,000 possible bigrams alone.
🠶 To get an idea of the dependence of a grammar on its training set,
let’s look at an n-gram grammar trained on a completely different
corpus: the Wall Street Journal (WSJ) newspaper.
🠶 Shakespeare and the Wall Street Journal are both English, so we
might expect some overlap between our n-grams for the two genres.
59 N Gram Sensitivity to the Training Corpus
60
61
62
63
64
65
66 G o od Turing / Discounting
67 Good Turing / Discounting
68 Good Turing / Discounting
69

C ompiled by : Mrs. Angelin Florence A


70

You might also like