R-2023 PG Syllabus for 23CS6210T
VELAMMAL ENGINEERING COLLEGE, CHENNAI 66
(An Autonomous Institution, Affiliated to Anna University, Chennai)
Course code 23CS6210T Semester III
Category PROFESSIONAL ELECTIVE COURSE(PEC) L T P C
Course Title NATURAL LANGUAGE PROCESSING 2 0 2 3
COURSE OBJECTIVE
To understand basics of linguistics, probability and statistics
To study statistical approaches to NLP and understand sequence labeling
To outline different parsing techniques associated with NLP
To explore semantics of words and semantic role labeling of sentences
To understand discourse analysis, question answering and chatbots
COURSE OUTCOMES
CO. Blooms
No. Course Outcome level
At the end of the course students are able to
Understand basics of linguistics, probability and statistics associated with
CO 1 NLP C2
CO 2 Implement a Part-of-Speech Tagger C2
CO 3 Design and implement a sequence labeling problem for a given domain C3
Implement semantic processing tasks and simple document indexing and
CO 4 searching system using the concepts of NLP C4
CO 5 Implement a simple chatbot using dialogue system concepts C4
MAPPING WITH PROGRAM OUTCOMES
Program Outcome (POs)
PO1 PO2 PO3 PO4 PO5 PO6
1 - 2 3 1 1 -
2 2 2 2 3 - 3
3 3 - 3 3 - 3
4 1 - 2 3 - 3
5 1 - 2 3 - 3
Avg. 3 2 3 3 1 3
Note:- 1: Slight 2:Moderate 3: Substantial
Presented in 7th Board of Studies meeting held on 30-01-
2024(Approved) Passed in the 6th Academic Council meeting dated 10-
05-2024
Form No. CD 02 C [Link].00 Effective Date: 1/06/19
R-2023 PG Syllabus for 23CS6210T
VELAMMAL ENGINEERING COLLEGE, CHENNAI 66
(An Autonomous Institution, Affiliated to Anna University, Chennai)
SYLLABUS (Total contact hours = 60 Periods, No. of Credits = 3)
UNIT I INTRODUCTION 6
Natural Language Processing Components - Basics of Linguistics and Probability and Statistics
Words-Tokenization-Morphology-Finite State Automata
UNIT II STATISTICAL NLP AND SEQUENCE LABELING 6
N-grams and Language models Smoothing -Text classification- Naïve Bayes classifier
Evaluation - Vector Semantics TF-IDF - Word2Vec- Evaluating Vector Models -Sequence
Labeling Part of Speech Part of Speech Tagging -Named Entities Named Entity Tagging
UNIT III CONTEXTUAL EMBEDDING 6
Constituency Context Free Grammar Lexicalized Grammars- CKY Parsing Earley's algorithm-
Evaluating Parsers -Partial Parsing Dependency Relations- Dependency Parsing Transition Based
- Graph Based
UNIT IV COMPUTATIONAL SEMANTICS 6
Word Senses and WordNet Word Sense Disambiguation Semantic Role Labeling Proposition
Bank- FrameNet- Selectional Restrictions - Information Extraction - Template Filling
UNIT V DISCOURSE ANALYSIS AND SPEECH PROCESSING 6
Discourse Coherence Discourse Structure Parsing Centering and Entity Based Coherence
Question Answering Factoid Question Answering Classical QA Models Chatbots and
Dialogue systems Frame-based Dialogue Systems Dialogue State Architecture
TOTAL: 45 PERIODS
REFERENCES
1. Daniel Jurafsky and James [Link], and Language Processing: An Introduction to
Hall Series in Artificial Intelligence), 2020
2. Jacob Eisenstein. Language Processing MIT Press, 2019
3. Samuel Burns Language Processing: A Quick Introduction to NLP with Python and
NLTK, 2019
4. Christopher Manning, of Statistical Natural Language MIT Press,
2009.
Presented in 7th Board of Studies meeting held on 30-01-
2024(Approved) Passed in the 6th Academic Council meeting dated 10-
05-2024
Form No. CD 02 C [Link].00 Effective Date: 1/06/19
R-2023 PG Syllabus for 23CS6210T
VELAMMAL ENGINEERING COLLEGE, CHENNAI 66
(An Autonomous Institution, Affiliated to Anna University, Chennai)
TOTAL : 30
PERIODS SUGGESTED ACTIVITIES:
1. Probability and Statistics for NLP Problems
2. Carry out Morphological Tagging and Part-of-Speech Tagging for a sample text
3. Design a Finite State Automata for more Grammatical Categories
4. Problems associated with Vector Space Model
5. Hand Simulate the working of a HMM model
6. Examples for different types of work sense disambiguation
7. Give the design of a Chatbot
PRACTICAL EXERCISES: PERIODS : 30
1. Download nltk and packages. Use it to print the tokens in a document and the sentences
from it.
2. Include custom stop words and remove them and all stop words from a given document
using nltk or spaCY package
3. Implement a stemmer and a lemmatizer program.
4. Implement a simple Part-of-Speech Tagger
5. Write a program to calculate TFIDF of documents and find the cosine similarity
between any two documents.
6. Use nltk to implement a dependency parser.
7. Implement a semantic language processor that uses WordNet for semantic tagging.
8. Project - (in Pairs) Your project must use NLP concepts and apply them to some data.
a. Your project may be a comparison of several existing systems, or it may propose a new
system in which case you still must compare it to at least one other approach.
b. You are free to use any third-party ideas or code that you wish as long as it is publicly
available.
c. You must properly provide references to any work that is not your own in the writeup.
d. Project proposal You must turn in a brief project proposal. Your project proposal should
describe the idea behind your project. You should also briefly describe software you will
need to write, and papers (2-3) you plan to read.
List of Possible Projects
1. Sentiment Analysis of Product Reviews
2. Information extraction from News articles
Presented in 7th Board of Studies meeting held on 30-01-
2024(Approved) Passed in the 6th Academic Council meeting dated 10-
05-2024
Form No. CD 02 C [Link].00 Effective Date: 1/06/19
R-2023 PG Syllabus for 23CS6210T
VELAMMAL ENGINEERING COLLEGE, CHENNAI 66
(An Autonomous Institution, Affiliated to Anna University, Chennai)
3. Customer support bot
4. Language identifier
5. Media Monitor
6. Paraphrase Detector
7. Identification of Toxic Comment
8. Spam Mail Identification
Presented in 7th Board of Studies meeting held on 30-01-
2024(Approved) Passed in the 6th Academic Council meeting dated 10-
05-2024
Form No. CD 02 C [Link].00 Effective Date: 1/06/19