NLP Unit 1 Notes
NLP Unit 1 Notes
PART A
Finding the structure of words
Human language is a complicated thing. We use it to express our thoughts, and through language,
we receive information and infer its meaning. Linguistic expressions are not unorganized, though.
They show structure of different kinds and complexity and consist of more elementary components
whose co-occurrence in context refines the notions they refer to in isolation and implies further
meaningful relations between them.
Linguists have developed whole disciplines that look at language from different perspectives and at
different levels of detail. The point of morphology, for instance, is to study the variable forms and
functions of words, while syntax is concerned with the arrangement of words into phrases, clauses,
and sentences. Word structure constraints due to pronunciation are described by phonology,
whereas conventions for writing constitute the orthography of a language. The meaning of a
linguistic expression is its semantics, and etymology and lexicology cover especially the evolution
of words and explain the semantic, morphological, and other links among them.
Morphology is an essential part of language processing, and in multilingual settings, it becomes
even more important. The discovery of word structure is morphological parsing.
Various Terminologies
Morphology is to find the structure of the words
Syntax is concerned with the arrangement of words into phrases, clauses, and sentences.
The meaning of a linguistic expression is its semantics.
Discourse deals with the sense of the context.
Pragmatics deals with communicative and social content and its effect on interpretation.
Word structure constraints due to pronunciation are described by phonology.
Whereas conventions for writing constitute the orthography of a language.
Etymology is to derive the source of a word, or the study of the source of specific words.
(e.g., tracing a word back to its Latin roots in one or many languages)
Lexicology covers especially the semantic, morphological, and other links among words.
Derivational Morphology :
Creation of a new word from existing words by changing grammatical category. In linguistic
morphology derivation is the process of forming a new word from an existing word, often by
adding a prefix or suffix, such as un- or -ness.
For example, the word “transformation” contains two derivational morphemes: trans (prefix)
-form (root) -ation (suffix)
Some examples of derivational morphemes are:
Words are defined in most languages as the smallest linguistic units that can form a complete
utterance by themselves. The minimal parts of words that deliver aspects of meaning to them are
called morphemes. Depending on the means of communication, morphemes are spelled out via
graphemes—symbols of writing such as letters or characters—or are realized through phonemes,
the distinctive units of sound in spoken language. It is not always easy to decide and agree on the
precise boundaries discriminating words from morphemes and from phrases.
a. Tokens
Suppose, for a moment, that words in English are delimited only by whitespace and
punctuation. If we confront our assumption with insights from etymology and syntax, we
notice two words here: newspaper and won’t. Being a compound word, newspaper has an
interesting derivational structure. We might wish to describe it in more detail, once there is a
lexicon or some other linguistic evidence on which to build the possible hypotheses about the
origins of the word. In writing, newspaper and the associated concept is distinguished from
the isolated news and paper. In speech, however, the distinction is far from clear, and
identification of words becomes an issue of its own.
For reasons of generality, linguists prefer to analyze won’t as two syntactic words, or tokens,
each of which has its independent role and can be reverted to its normalized form. The
structure of won’t could be parsed as will followed by not. In English, this kind of tokenization
and normalization may apply to just a limited set of cases, but in other languages, these
phenomena have to be treated in a less trivial manner. In the writing systems of Chinese,
Japanese, and Thai, whitespace is not used to separate words.
Nonetheless, the elementary morphological units are viewed as having their own syntactic
status. In such languages, tokenization, also known as word segmentation, is the
fundamental step of morphological analysis and a prerequisite for most language processing
applications.
b. Lexemes
By the term word, we often denote not just the one linguistic form in the given context but
also the concept behind the form and the set of alternative forms that can express it. Such
sets are called lexemes or lexical items, and they constitute the lexicon of a language.
Lexemes can be divided by their behavior into the lexical categories of verbs, nouns,
adjectives, conjunctions, particles, or other parts of speech. The citation form of a lexeme, by
which it is commonly identified, is also called its lemma.
Examples: [Playing-play], [singing-sing], [annoyed-annoy], etc.
When we convert a word into its other forms, such as turning the singular mouse into the
plural mice or mouses, we say we inflect the lexeme. When we transform a lexeme into
another one that is morphologically related, regardless of its lexical category, we say we
derive the lexeme: for instance, the nouns receiver and reception are derived from the verb
to receive.
c. Morphemes
Morphological theories differ on whether and how to associate the properties of word forms
with their structural components. These components are usually called segments or morphs.
The morphs that by themselves represent some aspect of the meaning of a word are called
morphemes of some function.
Human languages employ a variety of devices by which morphs and morphemes are
combined into word forms. The simplest morphological process concatenates morphs one by
one, as in dis-agree-ment-s, where agree is a free lexical morpheme and the other elements
are bound grammatical morphemes contributing some partial meaning to the whole word.
In a more complex scheme, morphs can interact with each other, and their forms may become
subject to additional phonological and orthographic changes denoted as morphophonemic.
The alternative forms of a morpheme are termed allomorphs.
There are two main types: free and bound. Free morphemes can occur alone and
bound morphemes must occur with another morpheme. An example of a free
morpheme is "bad", and an example of a bound morpheme is "ly." It is bound because
although it has meaning, it cannot stand alone.
Example: 1 [boy], 2 [desire-able], 3 [desire-able-ity], 4 [gentle-man-li-ness], etc.
Typology
Morphological typology divides languages into groups by characterizing the prevalent morphological
phenomena in those languages. It aims to capture structural and semantic variation across the
world's languages. It can consider various criteria, and during the history of linguistics, different
classifications have been proposed. Let us outline the typology that is based on quantitative relations
between words, their morphemes, and their features:
Isolating, or analytic, languages include no or relatively few words that would comprise more than
one morpheme (typical members are Chinese, Vietnamese, and Thai; analytic tendencies are
also found in English).
Synthetic languages can combine more morphemes in one word and are further divided into
agglutinative and fusional languages.
Agglutinative languages have morphemes associated with only a single function at a time (as in
Korean, Japanese, Finnish, and Tamil, Telugu, etc.).
Fusional languages are defined by their feature-per-morpheme ratio higher than one (as in Arabic,
Czech, Latin, Sanskrit, German, etc.).
In accordance with the notions about word formation processes mentioned earlier, we can also
discern:
Approaches to morphology
There are three principal approaches to morphology:
Morpheme based morphology:
By irregularity, we mean existence of such forms and structures that are not described
appropriately by a prototypical linguistic model. Some irregularities can be understood by
redesigning the model and improving its rules, but other lexically dependent irregularities often
cannot be generalized.
Some irregularities are bound to particular lexemes or contexts, and cannot be accounted for by
general rules.
Ambiguity is present in all aspects of morphological processing and language processing at large.
Morphological parsing is not concerned with complete disambiguation of words in their context,
however; it can effectively restrict the set of valid interpretations of a given word form.
Morphological modeling also faces the problem of productivity and creativity in language, by which
unconventional but perfectly meaningful new words or new senses are coined. Usually, though,
words that are not licensed in some way by the lexicon of a morphological system will remain
completely unparsed. This unknown word problem is particularly severe in speech or writing that
gets out of the expected domain of the linguistic model, such as when special terms or foreign names
are involved in the discourse or when multiple languages or dialects are mixed together.
Morphological Models
There are many possible approaches to designing and implementing morphological models. Over
time, computational linguistics has witnessed the development of a number of formalisms and
frameworks, in particular grammars of different kinds and expressive power, with which to address
whole classes of problems in processing natural as well as formal languages. There are also many
approaches that do not resort to domain-specific programming. They, however, have to take care
of the runtime performance and efficiency of the computational model themselves. It is up to the
choice of the programming methods and the design style whether such models turn out to be pure,
intuitive, adequate, complete, reusable, elegant, or not. Let us now look at the most prominent types
of computational approaches to morphology.
Dictionary Lookup
Morphological parsing is a process by which word forms of a language are associated with
corresponding linguistic descriptions.
In this context, a dictionary is understood as a data structure that directly enables obtaining some
precomputed results, in our case word analyses.
The data structure can be optimized for efficient lookup, and the results can be shared.
Lookup operations are relatively simple and usually quick.
Dictionaries can be implemented, for instance, as lists, binary search trees, tries, hash tables, and
so on.
However, the issues that get encountered while working on this model are related to experiencing
unknown-Words and the size of the dictionary or corpus size
Because the set of associations between word forms and their desired descriptions is declared by
plain enumeration, the coverage of the model is finite and the generative potential of the language
is not exploited. Developing as well as verifying the association list is tedious, liable to errors, and
likely inefficient and inaccurate unless the data are retrieved automatically from large and reliable
linguistic resources. However, it deals easily with exceptions, and can implement even complex
morphology. For instance, dictionary-based approaches to Korean depend on a large dictionary of
all possible combinations of allomorphs and morphological alternations.
Finite-State Morphology
By finite-state morphological models, we mean those in which the specifications written by human
programmers are directly compiled into finite-state transducers. The two most popular tools
supporting this approach include XFST (Xerox Finite-State Tool) and LexTools.
Finite-state transducers are computational devices extending the power of finite-state automata.
They consist of a finite set of nodes connected by directed edges labeled with pairs of input and
output symbols. In such a network or graph, nodes are also called states, while edges are called
arcs. Traversing the network from the set of initial states to the set of final states along the arcs is
equivalent to reading the sequences of encountered input symbols and writing the sequences of
corresponding output symbols.
The set of possible sequences accepted by the transducer defines the input language; the set of
possible sequences emitted by the transducer defines the output language. For example, a finite-
state transducer could translate the infinite regular language consisting of the words vnuk, pravnuk,
prapravnuk, ... to the matching words in the infinite regular language defined by grandson, great-
grandson, great-great-grandson, ...
However, it needs Careful fine-tuning of lexicons, rules, and components, extending the code can
lead to unexpected interactions. It is error prone, complex and difficult
Unification-Based Morphology
Unification of feature structures can also fail, which means that the information in them is mutually
incompatible.
Functional Morphology
This group of morphological models includes not only the ones following the methodology of
functional morphology, but even those related to it, such as morphological resource grammars of
Grammatical Framework.
Functional morphology defines its models using principles of functional programming and type
theory. It treats morphological operations and processes as pure mathematical functions and
organizes the linguistic as well as abstract elements of a model into distinct types of values and type
classes.
Though functional morphology is not limited to modeling particular types of morphologies in human
languages, it is especially useful for fusional morphologies. Linguistic notions like paradigms, rules
and exceptions, grammatical categories and parameters, lexemes, morphemes, and morphs can
be represented intuitively and succinctly in this approach.
A functional morphology model can be compiled into finite-state transducers if needed, but can also
be used interactively in an interpreted mode, for instance. Computation within a model may exploit
lazy evaluation and employ alternative methods of efficient parsing, lookup, and so on.
Morphology Induction
We have not considered the problem of discovering and inducing word structure without the human
insight (i.e., in an unsupervised or semi-supervised manner).
The motivation for such approaches lies in the fact that for many languages, linguistic expertise
might be unavailable or limited, and implementations adequate to a purpose may not exist at all.
Automated acquisition of morphological and lexical information, even if not perfect, can be reused
for bootstrapping and improving the classical morphological models, too.
PART B
Finding the Structure of Documents
In human language, words and sentences do not appear randomly but usually have a structure. For
example, combinations of words form sentences—meaningful grammatical units, such as
statements, requests, and commands. Likewise, in written text, sentences form paragraphs—self-
contained units of discourse about a particular point or idea. Sentences may also be related to each
other by explicit discourse connectives such as therefore.
Given the ever-growing problem of written and spoken information overload, extracting the structure
of textual and audio documents is a meaningful and sometimes necessary first step in most speech
and language processing applications. Here, we discuss methods for finding the structure of
documents: Sentence boundary detection and Topic boundary detection.
These methods base their predictions on features of the input:
local characteristics that give evidence toward the presence or absence of a sentence or
topic boundary, such as:
a punctuation sign,
a pause in speech,
and a new word in a document.
Features are the core of classification approaches and require careful design and selection in order
to be successful and prevent overfitting and noise problems.
Sentence Boundary Detection
It is also known as sentence segmentation.
It is a task of deciding where sentences start and end.
It automatically segments a sequence of word tokens into sentence units.
In written text in English and some other languages, the beginning of a sentence is usually marked
In the first sentence, the abbreviation Dr. does not end a sentence, and in the second it does.
“This year has been difficult for both Hertz and Avis,” said Charles Finnie, carrental industry
analyst—yes, there is such a profession—at Alex. Brown & Sons.
An automatic method that outputs word boundaries as ending sentences according to the presence
of such punctuation marks would result in cutting some sentences incorrectly.
Challenges in Sentence boundary detection:
If text is Non-grammatical with missing punctuations then it can give rise to Ambiguous abbreviations
and capitalizations, errors in short message service (SMS) texts, and instant messaging (IM) texts.
In optical character recognition to translate images of handwritten, typewritten, or printed text or
spoken utterances into machine-editable text, the finding of sentence boundaries must deal with the
errors of those systems as well.
For conversational speech or text or multiparty meetings with ungrammatical sentences and
disfluencies, in most cases it is not clear where the boundaries are.
Two ways are there for sentence boundary detection as shorty explained below:
Code switching
use of words, phrases, or sentences from multiple languages by multilingual speakers
another problem that can affect the characteristics of sentences
e.g., Spanish uses the inverted question mark to precede questions, while English uses
direct question mark.
Spanish (¿ and ¡) – English (? and !) – Urdu, Arabic (؟and ! )
Conventional rule-based
It relies on patterns to identify potential ends of sentences and lists of abbreviations
for disambiguating them
For example if the word before the boundary is a known abbreviation, such as “Mr.” or
“Gov.,” the text is not segmented at that position even though some periods are
exceptions.
Although rules cover most of these cases, they do not address:
unknown abbreviations, abbreviations at the ends of sentences, or typos in the input
text.
Moreover, each language requires a specific set of rules.
Topic Boundary Detection
It is also known as discourse segmentation or text segmentation
It is the task of automatically dividing a stream of text or speech into topically homogeneous blocks.
The aim of topic segmentation is to find the boundaries where topics change.
If long documents can be segmented into shorter, topically coherent segments, then only the
segment that is about the user’s query could be retrieved.
During the late 1990s, the U.S. Defense Advanced Research Projects Agency (DARPA) initiated
the Topic Detection and Tracking (TDT) program to further the state of the art in finding and following
new topics in a stream of broadcast news stories (segmenting a news stream into individual stories).
The task of topic segmentation is inspired by discourse analysis. It is a problem without a very high
human agreement because of many natural-language-related issues and hence requires a good
definition of topic categories and their granularities.
For example, topics are not typically flat but occur in a semantic hierarchy. In text, topic boundaries
are usually marked with distinct segmentation cues, such as headlines and paragraph breaks.
These cues are absent in speech. However, speech provides other cues, such as pause duration
and speaker changes.
Methods
The methods that are used are: Generative is to learn each language and determine as to which
language the speech belongs to.
Discriminative is to determine the linguistic differences without learning any language – a much
easier task!
Generative:
The generative model is considered as a class of statistical models that can generate new data
instances.
This model is typically used to estimate probabilities, modeling data points and distinguishing
between classes based on these probabilities.
NLP
Traditional rule-based or Boolean logic systems (e.g. Dialog and Lexis-Nexis) are giving way
to statistical approaches (Markov models and stochastic context free grammars)\
Medical Diagnosis
QMR knowledge base, initially a heuristic expert systems for reasoning about diseases and
symptoms has been augmented with decision theoretic formulation
Genomics and Bioinformatics
Sequences represented as generative HMMs
Discriminative:
Discriminative model refers to a class of models used in statistical classification
It is especially used in supervised machine learning.
The model learns the boundary between classes or labels in a dataset.
The discriminative models have the advantage of being more robust to outliers.
Discriminative algorithms focus on modeling a direct solution.
No attempt to model underlying probability distributions.
Popular models
Logistic regression,
SVM
K-Nearest Neighbor
Decision Tree, Random Forest, etc.
Logistic regression
the logistic regression algorithm models a decision boundary. Then it decides on the outcome
of an observation based on where it stands relative to the decision boundary.
Complexity of the approaches
The approaches described here have advantages and disadvantages.
In a given context and under a set of observation features, one approach may be better than another.
These approaches can be rated in terms of complexity (time and memory) of their training and
prediction algorithms and in terms of their performance on real-world datasets.
Some may also require specific preprocessing, such as converting or normalizing continuous
features to discrete features.
DISCRIMINATIVE GENERATIVE
For sentence segmentation in speech, performance is usually evaluated using the error rate (ratio
of number of errors to the number of examples), F1-measure (the harmonic mean of recall and
precision, where recall is defined as the ratio of the number of correctly returned sentence
boundaries to the number of sentence boundaries in the reference annotations and precision is the
ratio of the number of correctly returned sentence boundaries to the number of all automatically
estimated sentence boundaries), and the National Institute of Standards and Technology (NIST)
error rate (number of candidates wrongly labeled divided by the number of actual boundaries).
𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑒𝑟𝑟𝑜𝑟𝑠
𝑒𝑟𝑟𝑜𝑟 𝑟𝑎𝑡𝑒 =
𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑒𝑥𝑎𝑚𝑝𝑙𝑒𝑠
((2∗𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛∗𝑅𝑒𝑐𝑎𝑙𝑙))
𝐹1 𝑆𝑐𝑜𝑟𝑒 =
((𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛+𝑅𝑒𝑐𝑎𝑙𝑙))