0% found this document useful (0 votes)
14 views35 pages

Understanding NLP and Text Data

Natural Language Processing (NLP) is a technology that enables machines to understand and interpret human language. It involves various components such as Natural Language Understanding (NLU) and Natural Language Generation (NGU), as well as processes like text pre-processing, model building, and sentiment analysis. Key techniques in NLP include tokenization, stop word removal, stemming, and lemmatization, which help in analyzing textual data effectively.

Uploaded by

GANESH DESHMUKH
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
14 views35 pages

Understanding NLP and Text Data

Natural Language Processing (NLP) is a technology that enables machines to understand and interpret human language. It involves various components such as Natural Language Understanding (NLU) and Natural Language Generation (NGU), as well as processes like text pre-processing, model building, and sentiment analysis. Key techniques in NLP include tokenization, stop word removal, stemming, and lemmatization, which help in analyzing textual data effectively.

Uploaded by

GANESH DESHMUKH
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Natural

Language
Processing
NATURAL
LANGUAGE
PROCESSING
❑NLP Is the Technology that is
used by machines to
understand, analyses,
manipulate and interpret
human’s language.
APPLICATIONS OF NLP

This Photo by Unknown Author is licensed under CC BY-SA-NC


COMPONNENTS OF NLP

NLU(NATURAL LANGUAGE UNDERSTANDING) NGU(NATURAL LANGUAGE GENERATION)


NLP PIPELINE
LIBRARIES FOR NLP
DATA COLLECTION

This Photo by Unknown Author is licensed under CC BY-SA This Photo by Unknown Author is licensed under CC BY-SA-NC
TEXT PRE-PROCESSISNG
❑ Text Cleaning : In-text Cleaning we do HTML tag removing, emoji handling,
Spelling Checker etc.
❑Basic Pre-processing : In basic preprocessing we do tokenization(word or
sent tokenization, stop word removal, removing digit, lower casing.
❑Advance Preprocessing : In this step we do POS tagging, Parsing, and
Conference resolution.
Featured Engineering
▪One Hot encoder

▪Bag of Word(BOW)

▪N-grams

▪Tf-ldf

▪Word2ved
MODEL BUILDING
Approaches of model building :
◦ Heuristic Approach
◦ Machine Learning Approach
◦ Deep learning Approach
◦ Cloud API.
MODEL EVALUATION

Intrinsic Evaluation Extrinsic Evaluation


Understanding
Textual
Concept
Textual data
❑Elements of text.
▪ Hierarchy of text
▪ Tokens
▪ Vocabulary
▪ Punctuation
▪ Part of speech
▪ Root of a word
▪ Base of a word
▪ Stop words
Hierarchy of Text
Natural language processing (NLP) is a subfield of computer science and
especially artificial intelligence. It is primarily concerned with providing computers with the
ability to process data encoded in natural language and is thus closely related to information
retrieval, knowledge representation and computational linguistics, a subfield of linguistics.
Tokenization
Sentence : - I am Nikhil Kumar. I am a good boy.

Sentence Tokenization : I am Nikhil Kumar. I am a good boy.

Word Tokenization : I am Nikhil Kumar I am good boy


STOP WORDS REMOVAL
oThe words which are generally filtered out before processing a natural language are called
stop words.

oThese are actually the most common words in any language (like articles, prepositions,
pronouns, conjunctions, etc) and does not add much information to the text. Examples of a
few stop words in English are “the”, “a”, "an", "so", "what".
N-grams
N-grams
➢N-grams are continues sequence of words or symbols or tokens in a document.

➢They can be defined as the neighboring sequences of items in a document.


N-gram Language Model
❑ After generalizing the above equation can be calculated as:
Stemming and Lemmatization.

Stemming is a technique used to extract the base form of the


words by removing affixes from them.
Lemmatization
Lemmatization technique is like stemming.

The output we will get after lemmatization is called ‘lemma.’

After lemmatization, we will be getting a valid word that means the same thing.
Mass is an intrinsic property of a body. It was traditionally
believed to be related to the quantity of matter in a body,

MASS

Involving a large number of people or


things
Mass is an intrinsic property of a body. It was traditionally
believed to be related to the quantity of matter in a body,

MASS

Involving a large number of people or


things
THIS IS A MOUSE.
Word Sense Disambiguation
❑Word sense disambiguation is an important is an important method of NLP by which the
meaning of a word is determined, which is used in a particular context.
COUNT
VECTORIZATION

Count vectorizer means


breaking down a sentence or
any text into words by
performing preprocessing
tasks like converting all words
to lowercase , Thus removing
Special characters.
SENTIMENT ANLAYSIS
What is sentiment analysis?
1 Standard Sentiment Analysis.
1. I love how Nikhil present the things --- Positive

2. I kind of like Nikhil’s sessions – Neutral

3. Nikhil’s Session are confusion for me ---- Negative


2 Fine Gradient Sentiment
Analysis
1. Nikhil’s Session are confusion for me ---- Negative

2. I kind of like Nikhil’s sessions – Neutral

3. Aww! I really hate Nikhil’s Session --- very negative


3 EMOTION DETECTION
1. I am blessed that I am having all of you --- Happiness

2. Your customer service is totally useless --- Angry

3. She left me for no reason -- sad


4 Aspect Based Sentiment
Analysis
1. Samsung S25 AI features are the best.

2. PS5 controller is the best.


5 Intent Detection
Very Frustrating LMS force closes when I try to login. Can anyone help ? --- Request for Assistance

You might also like