Text Analytics with Python
Natural Language Basics
Big data is data that is voluminous in size, has high velocity and veracity(many variety eg
text, tables). A billion records might be big but it is minute compared to the petabytes generated
by various sensors or even social media. A large volume of unstructured data is present across
various domains eg:
In social media – tweets, status updates, comments, articles, blogs, wikis etc.
Retails and e-commerce stores – new product information and metadata with
customer reviews and feedback.
Main challenges associated with textual data:
1. Text data is unstructured – It does not adhere to any predefined data schema or model
followed by relational databases. Organizations with enormous amounts of text data
resort to file-based databases like Hadoop Distributed File System (HDFS) databases like
MongoDB.
2. Analysing and trying to extracting useful patterns and insights from this data that would
beneficiary to the organization. Textual data does not follow regular syntax hence we
cannot use mathematical and statistical models. This resorts to us using specialised
techniques, algorithms and natural language processing to analyse text data.
Textual data is unstructured but it usually belongs to a specific language following a specific
syntax and semantics.
What is natural language?
It is one developed and evolved by humans through natural use and communication rather
than artificially, a computer language like Python. Example of natural languages include:
English, Swahili, Japanese.
Natural languages can be communicated in different forms eg: writing, speech, or even signs.
The Philosophy of Language
Main deals with the 4 questions and seeks to answer them:
The nature of meaning in language
The use of language
Language cognition
The relationship between language and reality
The nature of meaning in a language
It is concerned with the semantics of the language and the nature of meaning itself. Philosophers
seek to find out the meaning of any word or sentence and how it originated, came to be and how
different words in a language can be synonyms of each other and form relations.
It is also concerned with how syntax and structure pave way the way for semantics ie how words,
which have their own meanings, are structured to form meaningful sentences.
Linguistics is the study of language.
Can be expressed in linguistics between 2 humans as a sender and receiver. Sender tries to
express and communicate a message to the receiver. The receiver ends up understanding or
deducing from the context of the received message. From a non linguistic standpoint, things like
body language, prior experiences, and psychological effects are contributors to meaning of
language where each human perceives or infers meaning in their own way.
The Use of language
Concerned with how language is used as an entity in various scenarios and communications
between human beings. Includes: analysing speech including speaker’s , intent, tone, content and
actions involved in expressing a message. It also involves concepts such as the origin of
language of creation, and human cognitive activities such as language acquisition (responsible
for learning and usage of languages).
Language Cognition
Focuses on how the cognitive functions of the human brain are responsible for understanding and
interpreting language. Cognition tries to find out how the mind works in combining and relating
specific words into sentences and then into a meaningful message and what is the relation of
language to the thought process of the sender and receiver when they use the language to
communicate messages.
The relationship between language and reality
Explores the extent of truth of expressions originating from a language.
Philosophers of language try to measure how factual these expressions are and how they relate to
certain affairs of our world which are true.
1. Triangle of reference
used to explain how words convey meaning and ideas in the mind of the receiver and how that
meaning relates back to a real world fact or entity.
Also known as meaning of meaning model.
A symbol is denoted as a linguistic symbol, a word or an object that evokes thoughts in a person
mind. In this case: what is a couch, a piece of furniture that can be used for sitting on or lying
down and relaxing, something that gives us comfort .
Reference is the mental thought or idea that points to referent—the actual object in the real
world. In this example, the referent is the physical couch a person sees in front of them, while
the reference is their internal thought or concept of that couch. Essentially, reference connects
our internal ideas to external reality.
2. The Direction of fit
The word-to-world direction of fit talks about instances where the usage of language
can reflect reality. Indicates using words that match or relate to something that has or
happening in the real world. Eg: The Eiffel tower is really big, which accentuates a
fact in reality.
The world-to-word talks about the instances of usage of language can change reality.
Eg: I am going to take swim, where the person I is changing reality by going to take a
swim by representing the same in the sentence being communicated.
Summary:
A person perceives a referent (real-world object).
They create a symbol or word (mental representation) for it.
This symbol is communicated to another person.
The receiver forms their own mental representation of the world based on that symbol.
Thus, shared understanding is built through a cyclical exchange of symbols rooted in real-world
referents.
Language Acquisition and Usage
Language acquisition refers to the process by which humans use their cognitive abilities,
knowledge, and experience and to understand language based on hearing and perception and start
using it in terms of words, phrases, and sentences to communicate with other human beings.
Key Theories of Language Acquisition
Historical Views: Ancient theories often attributed language to divine origin or inheritance.
Philosophers like Plato suggested a "word-meaning mapping" as the foundation.
Behaviorist Theory (B.F. Skinner): A dominant modern theory posits that language is
learned through **operant conditioning**.
Process: Children produce sounds/words (behavior) and the consequences (rewards like
praise, or corrections) from adults strengthen or weaken those verbal behaviors.
Mechanism: Through repetition and reinforcement, correct word meanings and usages
become fixed in memory.
Core Idea: Language acquisition is primarily a behavioral process driven by imitation,
feedback, and environmental conditioning.
Chomsky's Critique of Behaviorism & The Innateness Hypothesis
1. Challenge to Skinner: Noam Chomsky argued that children cannot learn language solely
through imitation and reinforcement, as evidenced by their creation of incorrect but rule-
based forms like "goed" or "gived," which they never heard from adults.
2. Key Proposal: Children must unconsciously extract and apply grammatical rules and
patterns (syntax) from the language they hear, going beyond simple mimicry.
3. Language Acquisition Device (LAD): Chomsky proposed that humans possess an innate,
language-specific cognitive faculty—the LAD—that equips them with the inherent
ability to acquire language's structure.
4. Autonomy of Syntax: This is the idea that syntax (grammar) is a separate and
independent system from meaning (semantics). His famous nonsensical but grammatical
sentence, "Colorless green ideas sleep furiously," demonstrates that grammatical
correctness does not depend on logical meaning.
5. Conclusion: Chomsky's theory posits that language acquisition relies on innate, domain-
specific knowledge of language structure, not just general cognitive abilities or
behavioral conditioning. This debate remains central to linguistics.
Language Usage
There are many acts of language acts:
i. Locutionary acts are concerned with the delivery of the sentence when communicated
from one human to another human being when speaking it.
ii. Illocutionary acts focus on the actual semantics and significance of the sentence which
was communicated.
iii. Perlocutionary acts refers to the actual effects the communication had on its receiver,
which is more psychological or behavioral.
Eg: The phrase “Get me a book from the table” spoken by a father to his child.
This example illustrates speech act theory, dividing communication into three parts:
1. Locutionary Act: The literal act of uttering the words: "Get me the book from the table."
2. Illocutionary Act: The speaker's intention or force behind the words—in this case, a
directive meant to make the child perform the action.
3. Perlocutionary Act: The effect the utterance has on the listener, i.e., the child's action of
actually bringing the book.
In short: Saying words (Locution) → Intending a command (Illocution) → Causing an action
(Perlocution).
According to philosopher John Searle, there are 5 Illocutionary acts speech, as follows:
Assertiveness are speech acts that communicate how things are already existent in the
world, asserting a proposition that could be true or false.
o Example: Statements, claims, reports, and declarations of facts (Eg: “The world
revolves around the sun”)
o Direction of Fit: They follow the word-to-world direction, meaning the words
aim to match the state of reality.
Directives are speech acts that the sender communicates to the receiver asking or
directing them to do something. This represents a voluntary act which the receiver might
do in the future after receiving a directive from the sender.
o Involves: simple requests, orders, commands. Eg: Get me the book from the table
Commisives – commits them to the sender or speaker who utters them to some future
voluntary action or act.
o Includes: promises, oaths, pledges and vows.
o Eg: I promise to be there tomorrow for the ceremony .
Expressiveness reveal a speaker’s disposition or outlook toward a particular proposition
communicated through the message.
o Includes various forms of expression or emotion such as congratulatory, sarcastic,
etc.
o Example: Congratulations on graduating top of the class.
Declarations are powerful speech acts that have the capability to change the reality based
on the declared proposition in the message communicated by the sender/speaker.
o Direction of Fit: world-to-word and vice versa
o Example: I hereby declare him to be guilty of all charges.
Linguistics
It is defined as the scientific study of language, including form and syntax of language,
meaning, and semantics depicted by the usage of language and context of use.
Areas under linguistics:
1. Phonetics: Studies the physical properties of speech sounds (how they are produced,
transmitted, and perceived). A phenome/phone is smallest individual unit of human
speech in a specific language.
2. Phonology: Studies how sounds function and pattern within a specific language,
including phonemes (meaningful sound units), accents, and syllable structure.
3. Morphology: Studies the structure and meaning of morphemes—the smallest units of
language with meaning (e.g., words, prefixes, suffixes).
4. Syntax: Studies the rules and structure governing how words combine to form
grammatical phrases and sentences.
5. Semantics: Studies meaning in language.
Lexical Semantics: The meaning of individual words.
Compositional Semantics: How word meanings combine to form the meaning of
phrases and sentences.
6. Pragmatics: Studies how context (situation, speaker intent) affects the interpretation of
meaning, including indirect or hidden messages.
7. Discourse Analysis: Studies language beyond the sentence, analyzing structure and
meaning in conversations and extended texts (spoken, written, or signed).
8. Lexicon: Essentially the vocabulary of a language, studying the properties of words
(sound, meaning, part of speech, forms).
9. Stylistics: Studies linguistic style in writing and speech, including tone, voice, and
rhetorical choices.
10. Semiotics: Studies signs and symbols (linguistic and non-linguistic) and how they
communicate meaning, covering metaphor, analogy, and symbolism.
Language Syntax and Structure
Syntax and structure go hand in hand, where a set of specific rules, conventions, and principles
govern the way words are combined into phrases, phrases get combined in to clauses, and
clauses get combined into sentences.
Figure 1: A collection of words without any relation or structure
A random collection of words is difficult to interpret. Language requires syntax—the rules that
structure words into sentences. Syntax organizes words and uses their order and position to
convey clear, intended meaning.
Using the grammatical hierarchy (sentence > clause > phrase > word), the structure of a sentence
can be visually mapped as a tree diagram. This is done through shallow parsing, a technique
used to identify the main constituents or building blocks within a sentence.
Figure 2: Structured sentence following the hierarchical syntax
The tree diagram illustrates how a sentence is built from smaller units:
Leaf Nodes: Individual words.
Phrases: Combinations of words.
Clauses: Combinations of phrases.
Sentence: Formed by connecting clauses using conjunctions or other linking words.
This provides a structured view of syntax, preparing for a deeper analysis of syntactic
categories (the grammatical roles of each unit).
Words
A word is the smallest independent unit in a language that carries its own meaning.
A morpheme is the smallest meaningful unit, but it is not independent (e.g., prefixes,
suffixes). A single word can be made of multiple morphemes.
Purpose of POS Tagging:
To analyze syntax, words are classified into parts of speech (POS) tags (e.g., noun, verb). This
annotation reveals the major syntactic categories in a sentence.
Major categories of words:
i. Noun (N): Names an object, entity, or concept (e.g., fox, book).
ii. Verb (V): Describes an action, state, or occurrence (e.g., running, is).
iii. Adjective (ADJ): Modifies or describes a noun (e.g., brown, quick).
iv. Adverb (ADV): Modifies a verb, adjective, or other adverb, often indicating manner,
degree, or time (e.g., very, over).
Other Frequent Categories (Often Closed Classes): These have a finite set of words.
i. Determiner (DET): Articles and quantifiers (e.g., the, *a*).
ii. Pronoun (PRON): Replaces a noun (e.g., he, it).
iii. Conjunction (CONJ): Connects words, phrases, or clauses (e.g., and).
Key Concept:
i. Open Class: Vocabulary is infinite and accepts new words (N, V, ADJ, ADV).
ii. Closed Class: Vocabulary is fixed and rarely changes (e.g., pronouns, determiners).
The next level of grammatical analysis is the phrase, which builds from these tagged words.
Figure 3: Annotated words with their POS tags
Phrases
A phrase is a group of words that functions as a single grammatical unit within a clause or
sentence. Its central, most important word is called the head.
The Five Major Phrase Categories:
1. Noun Phrase (NP): Has a noun as its head. Acts as the subject or object of a verb
(e.g., the brown fox, dessert).
2. Verb Phrase (VP): Has a verb as its head. Can include the verb alone or the verb plus its
objects/complements (e.g., is quick, has started the engine). The exact structure
depends on the grammar used (constituency vs. dependency).
3. Adjective Phrase (ADJP): Has an adjective as its head. Modifies a noun or pronoun
(e.g., too quick).
4. Adverb Phrase (ADVP): Has an adverb as its head. Modifies a verb, adjective, or another
adverb (e.g., pretty soon).
5. Prepositional Phrase (PP): Has a preposition as its head and usually includes a noun
phrase. Acts as a modifier like an adjective or adverb (e.g., over the lazy dog).
Key Notes:
A phrase can sometimes be a single word if that word alone fulfills the phrasal role.
Shallow parsing is a technique used in NLP to automatically identify these phrases (and
POS tags) in a sentence.
The next level in the grammatical hierarchy is the clause, built from these phrases.
Clauses
A clause is a group of words containing a subject and a predicate (usually a verb phrase). It is
the next level above phrases in the grammatical hierarchy.
Two Main Types:
1. Independent (Main) Clause: Can stand alone as a complete sentence.
2. Dependent (Subordinate) Clause: Cannot stand alone; it depends on a main clause for
meaning and is often connected by a subordinating conjunction.
Syntactic Categories of Clauses (Based on Function):
Declarative: Makes a neutral statement (e.g., Grass is green).
Imperative: Gives a command, request, or instruction (e.g., Please do not talk).
Interrogative: Asks a question (e.g., Did you get my mail?).
Exclamative: Expresses strong emotion or surprise (e.g., What an amazing race!).
Relative: A type of dependent clause that modifies a noun (its antecedent) in another
part of the sentence (e.g., ...that he wanted a soda).
Applied to the Example:
The sentence "The brown fox is quick and he is jumping over the lazy dog" is divided by "and"
into two independent clauses. Both are declarative clauses, as each makes a neutral statement.
Grammar
Grammar is the set of rules governing the structure of language, including word order and
relationships in sentences. It evolves over time.
Two Main Classes:
1. Constituency Grammar (briefly mentioned for contrast): Focuses on grouping words into
nested phrases (NP, VP, etc.).
2. Dependency Grammar: The focus here—a word-based model that emphasizes
relationships between individual words rather than constituents.
Key Principles of Dependency Grammar:
The sentence has a root (usually the main verb), to which all other words connect.
Dependencies are labeled, directed, asymmetrical relationships between words
(e.g., subject, modifier).
It represents syntax as a Directed Acyclic Graph (DAG), showing relationships but not
necessarily word order.
Each word has exactly one incoming link (except the root).
Example from The brown fox is quick and he is jumping over the lazy dog:
det: Determiner relationship (e.g., fox → the).
amod: Adjectival modifier (e.g., fox → brown).
nsubj: Nominal subject (e.g., is → fox).
acomp: Adjective complement (e.g., is → quick).
prep/pobj: Prepositional modifier and its object (e.g., jumping → over → dog).
Dependency grammar provides a lean, relationship-focused representation of sentence
structure, often visualized as a graph. This contrasts with the phrase-based, hierarchical trees of
constituency grammar.
Figure 4: Dependency grammar based syntax tree with POS tags
Figure 5: Dependency grammar–based syntax tree annotated with dependency
Figure 6: Dependency grammar annotated graph for our sample sentence
Constituency Grammars
Core Idea: Constituency grammars (or phrase structure grammars) model sentences as
hierarchies of nested constituents (phrases and clauses).
Key Components:
Constituents: Meaningful units (words or groups) that act as independent or dependent
parts (e.g., NP, VP).
Phrase Structure Rules: Formal rules that define how constituents are built and ordered.
They are the engine of this grammar type.
o Format: Parent → Child1 Child2 ... (e.g., S → NP VP).
Fundamental Rules (from the text):
1. Sentence/Clause Rule:
S → NP VP
A sentence/clause divides into a Noun Phrase (subject) and a Verb Phrase (predicate).
2. Noun Phrase (NP) Rule:
NP → [DET] [ADJ] N [PP]
A noun phrase centers on a head noun (N). It can optionally include a determiner (DET),
adjectives (ADJ), and/or a prepositional phrase (PP).
These rules create a tree structure where larger units are recursively broken down into
smaller constituents.
Figure 7: Constituency syntax trees depicting structuring rules for noun phrases
Constituency Grammars (Verb Phrases & Recursion)
Verb Phrase (VP) Rule:
VP → V | MD [VP] [NP] [PP] [ADJP] [ADVP]
A verb phrase is headed by a verb (V) or a modal (MD).
It can be followed by optional constituents: another VP, NP, PP, Adjective Phrase (ADJP),
or Adverb Phrase (ADVP). This allows VPs to combine in various ways to express complex
predicates.
Key Property: Recursion
A constituent (like NP or VP) can appear on both sides of a production rule.
Example: An NP can contain a PP, which itself contains another NP (e.g., the book [on the
table]). This recursion enables the creation of infinitely long, nested phrases, a
fundamental feature of human language syntax.
Overall Structure:
Sentences are built via binary division (S → NP VP), where the NP and VP are then recursively
expanded using phrase structure rules, resulting in a hierarchical tree.
Figure 8: Constituency syntax trees depicting structuring rules for prepositional phrases
Recursion is a fundamental property of language syntax, where a constituent (e.g., NP,
PP) can be embedded inside another constituent of the same type.
In phrase structure rules, this is shown when a category appears on both sides of a rule
(e.g., NP → ... NP ...).
This property allows for the creation of infinitely long, nested phrases and complex
hierarchical trees, as seen in sentences with multiple embedded prepositional phrases
(e.g., The flying monkey [in the circus [on the trapeze [by the river]]]).
Constituency Grammar produces a hierarchical tree that:
1. Shows the order of words in the sentence.
2. Groups words into nested phrases and clauses (constituents).
3. Uses undirected edges to connect these constituents in a parent-child structure.
4. Clearly shows the binary division of a sentence into major constituents (e.g., two clauses
joined by a conjunction like "and").
Contrast with Dependency Grammar:
Dependency Grammar: Focuses on directed relationships between individual words,
forming a graph (DAG). It does not show word order or phrasal grouping.
Constituency Grammar: Focuses on hierarchical grouping of words into phrases,
showing both structure and word order.
Derived Frameworks: Constituency concepts form the basis for several formal grammars,
including Phrase Structure Grammar and Context-Free Grammar.
Word Order Topology
Typology is a field that specifically deals with trying to classify languages based on their syntax,
structure and functionality.
Word order topology is classifying languages based on their dominant words.
Primary word orders occur in clauses use the subject, verb, and an object.
There 6 major classes word orders languages like English follow the Subject-Verb-Object word
order class. An example would be the sentence He ate cake, where He is the subject, ate is the
verb, and cake is the object.