0% found this document useful (0 votes)
4 views87 pages

Machine Learning in NLP: Building Transformer-Based Natural Language Processing Applications (Part 1)

Uploaded by

Emnaa Hasnewi
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views87 pages

Machine Learning in NLP: Building Transformer-Based Natural Language Processing Applications (Part 1)

Uploaded by

Emnaa Hasnewi
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

MACHINE LEARNING IN NLP

Building Transformer-Based Natural Language Processing Applications


(Part 1)
FULL COURSE AGENDA
Part 1: Machine Learning in NLP
Lecture: NLP background and the role of DNNs leading to the
Transformer architecture
Lab: Tutorial-style exploration of a translation task using the
Transformer architecture

Part 2: Self-Supervision, BERT, and Beyond


Lecture: Discussion of how language models with self-
supervision have moved beyond the basic Transformer to BERT
and ever larger models
Lab: Practical hands-on guide to the NVIDIA NeMo API and
exercises to build a text classification task and a named
entity recognition task using BERT-based language models

Part 3: Production Deployment


Lecture: Discussion of production deployment considerations
and NVIDIA Triton Inference Server
Lab: Hands-on deployment of an example question answering
task to NVIDIA Triton

2
Part 1: Machine Learning in NLP
• Lecture
• What is NLP?
• Problem Formulation
• Text Representations
• Dimensionality Reduction
• Embeddings
• RNNs
• “Attention is All You Need”
• Lab
• Transformer Architecture
• BERT Model
• Pretraining BERT
3
FOUNDATION OF COUNTLESS
APPLICATIONS
4
NLP TASKS Automatic
Writing

General
Q&A

Intent Summary
Detection Generation

Common Natural
Word
Sense Language
Sense
Reasoning Processing
Automatic
Auto
Dialogue
Completion
Generation

Translation

Sentiment Code
Analysis Generation

And many more….


GLIMPSE OF WHAT IS POSSIBLE, TODAY…
6
Large NLP models powers:
o Multi-turn Information Retrieval for Q&A
Part 1: Machine Learning in NLP
• Lecture
• What is NLP?
• Problem Formulation
• Text Representations
• Dimensionality Reduction
• Embeddings
• RNNs
• “Attention is All You Need”
• Lab
• Transformer Architecture
• BERT Model
• Pretraining BERT
9
PROBLEM FORMULATION
10
MACHINE LEARNING
Discovering the discussed structures in text

Machine Learning
Text
Algorithm

11
MACHINE LEARNING
Discovering the discussed structures in text

Machine Learning
Text
Algorithm

12
MACHINE LEARNING
Design decisions

?
Problem formulation

? ? ? ? ? ?
Text Pre- Text Dimensionality Vector Machine Learning
Text Reweighting
processing Representation Reduction Comparison Algorithm

13
MACHINE LEARNING
All linear combinations feasible

?
Problem formulation

? ? ? ? ? ?
Text Pre- Text Dimensionality Vector Machine Learning
Text Reweighting
processing Representation Reduction Comparison Algorithm

GloVe Word2Vec

14
MACHINE LEARNING
In this class

Subset of Problem formulations

Text Pre- Text Dimensionality Vector Machine Learning


Text Reweighting
processing Representation Reduction Comparison Algorithm

Subset of
Subset of word representations approaches

15
Part 1: Machine Learning in NLP
• Lecture
• What is NLP?
• Problem Formulation
• Text Representations
• Dimensionality Reduction
• Embeddings
• RNNs
• “Attention is All You Need”
• Lab
• Transformer Architecture
• BERT Model
• Pretraining BERT
16
TEXT REPRESENTATIONS
The bag of words

• Bag of words/ngrams – feature per word/ngram

the cat sat on the mat

cat sat on the mat quic … |Vocabulary |

kly
• Various ways of choosing the values: Binary, Count, TF-IDF
1 1 1 2 1 0

17
THE BAG OF WORDS
Key challenges

Sparse Input (1-hot)


p >> n (overfitting!)

… … …

Word1
Word 1 Word
Word n n

No semantic generalization
lots of data required,
dog: 1 0 0 0 0 … 0 low accuracy

cat: 00100…0

18
DISTRIBUTED WORD
REPRESENTATIONS
19
DISTRIBUTIONAL HYPOTHESIS
The intuition

‘You can tell a word by the company it keeps’


Firth 1957

‘Distributional statements can cover all of the


material of a language without requiring support from
other types of information’
Harris 1954

‘The meaning of a word is its use in the language’


Wittgenstein 1953

‘The complete meaning of a word is always contextual,


and no study of meaning apart from context can be
taken seriously.’
Firth 1957

20
CO-OCCURRENCE PATTERNS
The latent information

21
CO-OCCURRENCE PATTERNS
Where to find them?

Possible relationships:

- Word to documents (very sparse and very wide)

- Word to word (very dense and compact)

- Word to user / person

- Word to user behaviour

- Word to product

- Word to custom feature (e.g. movie raking)

Not only metrices:

- Word to user to product


22
Part 1: Machine Learning in NLP
• Lecture
• What is NLP?
• Problem Formulation
• Text Representations
• Dimensionality Reduction
• Embeddings
• RNNs
• “Attention is All You Need”
• Lab
• Transformer Architecture
• BERT Model
• Pretraining BERT
23
DIMENSIONALITY REDUCTION
Rationale

The need for compact and computationally efficient representations

More robust notions of distance exposing the information captured by our distributional representation

24
LSA/LSI
25
LSA/LSI
Latent Semantic Analysis / Latent Semantic Indexing

26
LLSA/LSI
Truncated SVD

Terms x Documents

Dumais, Susan T., et al. "Using latent semantic analysis to improve access to textual information." Proceedings of the SIGCHI conference on 27
Human factors in computing systems. 1988.
LSA/LSI
Truncated SVD

Terms x Documents

K largest singular values

Dumais, Susan T., et al. "Using latent semantic analysis to improve access to textual information." Proceedings of the SIGCHI conference on 28
Human factors in computing systems. 1988.
LSA/LSI
Truncated SVD

Terms x Documents

K largest singular values

Latent Semantic Space

Dumais, Susan T., et al. "Using latent semantic analysis to improve access to textual information." Proceedings of the SIGCHI conference on 29
Human factors in computing systems. 1988.
LSA/LSI
Documents that are similar are closer

Landauer, Thomas K., Darrell Laham, and Marcia Derr. "From paragraph to graph: Latent semantic analysis for information visualization." Proceedings of the National Academy 30
of Sciences [Link] 1 (2004): 5214-5219.
LSA/LSI
Its so 1988

Dumais, Susan T., et al. "Using latent semantic analysis to improve access to textual information." Proceedings of the SIGCHI conference on
Human factors in computing systems. 1988.

31
DID WE MAKE FURTHER
PROGRESS?
32
STATUS AS OF 2010
Yes and No

Clustering Based
Representations
- Brown Clustering
- HMM-LDA
- CRF Chunker with
HMM
- …
It was not clear that you can combine unsupervised approaches
(i.e. embeddings) with supervised models

Distributed
Representations
Distributional - Collobert and
Weston embeddings
Representations
- LSA / LSI
-
-
HLBL embeddings

Text Unsupervised Machine Learning
-
-
pLSA
LDA
embedding Algorithm
- HAL
- ICA
- Random Indexing
- …

Turian, Joseph, Lev Ratinov, and Yoshua Bengio. "Word representations: a simple and general method for semi-supervised 33
learning." Proceedings of the 48th annual meeting of the association for computational linguistics. 2010.
Part 1: Machine Learning in NLP
• Lecture
• What is NLP?
• Problem Formulation
• Text Representations
• Dimensionality Reduction
• Embeddings
• RNNs
• “Attention is All You Need”
• Lab
• Transformer Architecture
• BERT Model
• Pretraining BERT
34
WHY NOT DO THE SAME
WITH NEURAL NETWORKS?
35
STATUS AS OF 2010
Not enough computational power

Turian, Joseph, Lev Ratinov, and Yoshua Bengio. "Word representations: a simple and general method for semi-supervised 36
learning." Proceedings of the 48th annual meeting of the association for computational linguistics. 2010.
WORD2VEC
37
WORD2VEC

Mikolov et al., 2013 (while at Google)

Linear model (trains quickly)

Two models for training embeddings in an unsupervised manner:

Continuous Bag-of-Words (CBOW) Skip-Gram

cat PAD the sat on


1-hot (|V|) 1-hot (|V|)

d-dimensional
Σ d-dimensional

d-dimensional E E E E E E E E d-dimensional

1-hot (|V|) 1-hot (|V|)


cat cat cat cat
PAD the sat on 38
GLOVE
39
GLOVE
The objective

To learn vectors for words such that their dot product is proportional to their probability of co-occurence

Pennington, J., Socher, R., & Manning, C. D. (2014, October). Glove: Global vectors for word representation. In Proceedings of the 2014 conference 40
on empirical methods in natural language processing (EMNLP) (pp. 1532-1543).
GLOVE
The objective

Pennington, J., Socher, R., & Manning, C. D. (2014, October). Glove: Global vectors for word representation. In Proceedings of the 2014 conference 41
on empirical methods in natural language processing (EMNLP) (pp. 1532-1543).
GLOVE
Properties

Comparative - Superlative Man - Woman

Pennington, J., Socher, R., & Manning, C. D. (2014, October). Glove: Global vectors for word representation. In Proceedings of the 2014 conference 42
on empirical methods in natural language processing (EMNLP) (pp. 1532-1543).
GLOVE
Not a distant past

Pennington, J., Socher, R., & Manning, C. D. (2014, October). Glove: Global vectors for word representation. In Proceedings of the 2014 conference
on empirical methods in natural language processing (EMNLP) (pp. 1532-1543).

43
USING THE EMBEDDINGS
44
THE APPROACH TO NLP
Unsupervised feature representation + Machine Learning models

Subset of Problem formulations

Text Pre- Text Dimensionality Vector Machine Learning


Text Reweighting
processing Representation Reduction Comparison Algorithm

Subset of
Subset of word representations approaches

45
THE APPROACH TO NLP
What ML model to choose

?
Subset of Problem formulations

Text Pre- Text Dimensionality Vector Machine Learning


Text Reweighting
processing Representation Reduction Comparison Algorithm

Subset of
Subset of word representations approaches

46
CLASSICAL APPROACHES
47
CLASSICAL APPROACHES
Very broad selection of tools

48
WHAT ABOUT FEATURE
ENGINEERING?
49
DEEP REPRESENTATION
LEARNING
50
DEEP REPRESENTATION LEARNING
Beyond distributional hypothesis

51
Part 1: Machine Learning in NLP
• Lecture
• What is NLP?
• Problem Formulation
• Text Representations
• Dimensionality Reduction
• Embeddings
• RNNs
• “Attention is All You Need”
• Lab
• Transformer Architecture
• BERT Model
• Pretraining BERT
52
RECURRENT NEURAL NETWORKS
Basic principles

y yt-1 yt yt+1 yt+2


st-1 st st+1 st+2
s

x xt-1 xt xt+1 xt+2

Unrolling in Time

53
LONG SHORT TERM (LSTM) CELL
Addressing problems of stability

54
CNNS
55
CONVOLUTIONAL NEURAL NETWORKS
Basic principles

Severyn, Aliaksei, and Alessandro Moschitti. "Unitn: Training deep convolutional neural network for twitter sentiment 56
classification." Proceedings of the 9th international workshop on semantic evaluation (SemEval 2015). 2015.
ATTENTION
57
WHAT ABOUT LONG SEQUENCES?
The challenge illustrated with SQuAD

58
The impact of attention mechanism on Question Answering performance
WHAT ABOUT LONG SEQUENCES?
The challenge

59
Bahdanau, D., Cho, K., & Bengio, Y. (2014). Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473.
ATTENTION
The mechanism

60
Bahdanau, D., Cho, K., & Bengio, Y. (2014). Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473.
ATTENTION
The mechanism

61
Bahdanau, D., Cho, K., & Bengio, Y. (2014). Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473.
ATTENTION
Examples

Wu, Y., Schuster, M., Chen, Z., Le, Q. V., Norouzi, M., Macherey, W., ... & Klingner, J. (2016). Google's neural machine translation system: Bridging the gap 62
between human and machine translation. arXiv preprint arXiv:1609.08144.
ATTENTION
Examples

Gehring, J., Auli, M., Grangier, D., Yarats, D., & Dauphin, Y. N. (2017, July). Convolutional sequence to sequence learning. In International 63
conference on machine learning (pp. 1243-1252). PMLR.
Part 1: Machine Learning in NLP
• Lecture
• What is NLP?
• Problem Formulation
• Text Representations
• Dimensionality Reduction
• Embeddings
• RNNs
• “Attention is All You Need”
• Lab
• Transformer Architecture
• BERT Model
• Pretraining BERT
64
ATTENTION IS ALL YOU NEED
Design

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. In Advances in neural 65
information processing systems (pp. 5998-6008).
ATTENTION IS ALL YOU NEED
Design

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. In Advances in neural 66
information processing systems (pp. 5998-6008).
WAS IT A BREAKTHROUGH
IN ITSELF?
67
ATTENTION IS ALL YOU NEED
Not a breakthrough in itself

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. In Advances in neural 68
information processing systems (pp. 5998-6008).
ATTENTION IS ALL YOU NEED
But …

“ … the Transformer can be trained significantly faster than


architectures based on recurrent or convolutional layers.”

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. In Advances in neural 69
information processing systems (pp. 5998-6008).
NEURAL EMBEDDINGS
70
FEATURE REUSE
The opportunity

71
IT WAS DIFFICULT TO
REUSE NLP EMBEDDINGS
72
SEMI-SUPERVISED SEQUENCE LEARNING
More complex representations

73
Dai, A. M., & Le, Q. V. (2015). Semi-supervised sequence learning. In Advances in neural information processing systems (pp. 3079-3087).
SEMI-SUPERVISED SEQUENCE LEARNING
More complex representations

74
Dai, A. M., & Le, Q. V. (2015). Semi-supervised sequence learning. In Advances in neural information processing systems (pp. 3079-3087).
SEMI-SUPERVISED SEQUENCE LEARNING
More complex representations

75
Dai, A. M., & Le, Q. V. (2015). Semi-supervised sequence learning. In Advances in neural information processing systems (pp. 3079-3087).
ELMO
Embeddings for Language Models

Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., & Zettlemoyer, L. (2018). Deep contextualized word representations. arXiv preprint 76
arXiv:1802.05365.
ELMO
Embeddings for Language Models

Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., & Zettlemoyer, L. (2018). Deep contextualized word representations. arXiv preprint 77
arXiv:1802.05365.
ULM-FIT
Universal Language Model Fine-Tuning for Text Classification

78
Howard, J., & Ruder, S. (2018). Universal language model fine-tuning for text classification. arXiv preprint arXiv:1801.06146.
TRANSFER LEARNING IN NLP
Not trivial to use and not universally applicable

Howard, J., & Ruder, S. (2018). Universal language model fine-tuning for text classification. arXiv preprint arXiv:1801.06146.

Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., & Zettlemoyer, L. (2018). Deep contextualized word representations. arXiv preprint
arXiv:1802.05365.

2013/
1957 1988 2010 2018
2014

Distributional The use of unsupervised First successes in


LSA / LSI Success of NN
Hypothesis embeddings transfer learning

79
THIS CREATED A
FOUNDATION FOR THE
NEW NLP MODELS
(DISCUSSED IN THE NEXT CLASS)
80
THE LAB
81
ATTENTION IS ALL YOU NEED
Deep dive into the transformer design

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. In Advances in neural 82
information processing systems (pp. 5998-6008).
83

BERT
How it relates to transformer and pretraining

83
IN THE NEXT CLASS…
84
SELF-SUPERVISION, BERT, AND BEYOND
Why did models start to work well? What does the future hold?

85
Part 1: Machine Learning in NLP
• Lecture
• What is NLP?
• Problem Formulation
• Text Representations
• Dimensionality Reduction
• Embeddings
• RNNs
• “Attention is All You Need”
• Lab
• Transformer Architecture
• BERT Model
• Pretraining BERT
86

You might also like