Machine Learning in NLP: Building Transformer-Based Natural Language Processing Applications (Part 1)
Machine Learning in NLP: Building Transformer-Based Natural Language Processing Applications (Part 1)
2
Part 1: Machine Learning in NLP
• Lecture
• What is NLP?
• Problem Formulation
• Text Representations
• Dimensionality Reduction
• Embeddings
• RNNs
• “Attention is All You Need”
• Lab
• Transformer Architecture
• BERT Model
• Pretraining BERT
3
FOUNDATION OF COUNTLESS
APPLICATIONS
4
NLP TASKS Automatic
Writing
General
Q&A
Intent Summary
Detection Generation
Common Natural
Word
Sense Language
Sense
Reasoning Processing
Automatic
Auto
Dialogue
Completion
Generation
Translation
Sentiment Code
Analysis Generation
Machine Learning
Text
Algorithm
11
MACHINE LEARNING
Discovering the discussed structures in text
Machine Learning
Text
Algorithm
12
MACHINE LEARNING
Design decisions
?
Problem formulation
? ? ? ? ? ?
Text Pre- Text Dimensionality Vector Machine Learning
Text Reweighting
processing Representation Reduction Comparison Algorithm
13
MACHINE LEARNING
All linear combinations feasible
?
Problem formulation
? ? ? ? ? ?
Text Pre- Text Dimensionality Vector Machine Learning
Text Reweighting
processing Representation Reduction Comparison Algorithm
GloVe Word2Vec
14
MACHINE LEARNING
In this class
Subset of
Subset of word representations approaches
15
Part 1: Machine Learning in NLP
• Lecture
• What is NLP?
• Problem Formulation
• Text Representations
• Dimensionality Reduction
• Embeddings
• RNNs
• “Attention is All You Need”
• Lab
• Transformer Architecture
• BERT Model
• Pretraining BERT
16
TEXT REPRESENTATIONS
The bag of words
kly
• Various ways of choosing the values: Binary, Count, TF-IDF
1 1 1 2 1 0
17
THE BAG OF WORDS
Key challenges
… … …
Word1
Word 1 Word
Word n n
No semantic generalization
lots of data required,
dog: 1 0 0 0 0 … 0 low accuracy
cat: 00100…0
18
DISTRIBUTED WORD
REPRESENTATIONS
19
DISTRIBUTIONAL HYPOTHESIS
The intuition
20
CO-OCCURRENCE PATTERNS
The latent information
21
CO-OCCURRENCE PATTERNS
Where to find them?
Possible relationships:
- Word to product
More robust notions of distance exposing the information captured by our distributional representation
24
LSA/LSI
25
LSA/LSI
Latent Semantic Analysis / Latent Semantic Indexing
26
LLSA/LSI
Truncated SVD
Terms x Documents
Dumais, Susan T., et al. "Using latent semantic analysis to improve access to textual information." Proceedings of the SIGCHI conference on 27
Human factors in computing systems. 1988.
LSA/LSI
Truncated SVD
Terms x Documents
Dumais, Susan T., et al. "Using latent semantic analysis to improve access to textual information." Proceedings of the SIGCHI conference on 28
Human factors in computing systems. 1988.
LSA/LSI
Truncated SVD
Terms x Documents
Dumais, Susan T., et al. "Using latent semantic analysis to improve access to textual information." Proceedings of the SIGCHI conference on 29
Human factors in computing systems. 1988.
LSA/LSI
Documents that are similar are closer
Landauer, Thomas K., Darrell Laham, and Marcia Derr. "From paragraph to graph: Latent semantic analysis for information visualization." Proceedings of the National Academy 30
of Sciences [Link] 1 (2004): 5214-5219.
LSA/LSI
Its so 1988
Dumais, Susan T., et al. "Using latent semantic analysis to improve access to textual information." Proceedings of the SIGCHI conference on
Human factors in computing systems. 1988.
31
DID WE MAKE FURTHER
PROGRESS?
32
STATUS AS OF 2010
Yes and No
Clustering Based
Representations
- Brown Clustering
- HMM-LDA
- CRF Chunker with
HMM
- …
It was not clear that you can combine unsupervised approaches
(i.e. embeddings) with supervised models
Distributed
Representations
Distributional - Collobert and
Weston embeddings
Representations
- LSA / LSI
-
-
HLBL embeddings
…
Text Unsupervised Machine Learning
-
-
pLSA
LDA
embedding Algorithm
- HAL
- ICA
- Random Indexing
- …
Turian, Joseph, Lev Ratinov, and Yoshua Bengio. "Word representations: a simple and general method for semi-supervised 33
learning." Proceedings of the 48th annual meeting of the association for computational linguistics. 2010.
Part 1: Machine Learning in NLP
• Lecture
• What is NLP?
• Problem Formulation
• Text Representations
• Dimensionality Reduction
• Embeddings
• RNNs
• “Attention is All You Need”
• Lab
• Transformer Architecture
• BERT Model
• Pretraining BERT
34
WHY NOT DO THE SAME
WITH NEURAL NETWORKS?
35
STATUS AS OF 2010
Not enough computational power
Turian, Joseph, Lev Ratinov, and Yoshua Bengio. "Word representations: a simple and general method for semi-supervised 36
learning." Proceedings of the 48th annual meeting of the association for computational linguistics. 2010.
WORD2VEC
37
WORD2VEC
d-dimensional
Σ d-dimensional
d-dimensional E E E E E E E E d-dimensional
To learn vectors for words such that their dot product is proportional to their probability of co-occurence
Pennington, J., Socher, R., & Manning, C. D. (2014, October). Glove: Global vectors for word representation. In Proceedings of the 2014 conference 40
on empirical methods in natural language processing (EMNLP) (pp. 1532-1543).
GLOVE
The objective
Pennington, J., Socher, R., & Manning, C. D. (2014, October). Glove: Global vectors for word representation. In Proceedings of the 2014 conference 41
on empirical methods in natural language processing (EMNLP) (pp. 1532-1543).
GLOVE
Properties
Pennington, J., Socher, R., & Manning, C. D. (2014, October). Glove: Global vectors for word representation. In Proceedings of the 2014 conference 42
on empirical methods in natural language processing (EMNLP) (pp. 1532-1543).
GLOVE
Not a distant past
Pennington, J., Socher, R., & Manning, C. D. (2014, October). Glove: Global vectors for word representation. In Proceedings of the 2014 conference
on empirical methods in natural language processing (EMNLP) (pp. 1532-1543).
43
USING THE EMBEDDINGS
44
THE APPROACH TO NLP
Unsupervised feature representation + Machine Learning models
Subset of
Subset of word representations approaches
45
THE APPROACH TO NLP
What ML model to choose
?
Subset of Problem formulations
Subset of
Subset of word representations approaches
46
CLASSICAL APPROACHES
47
CLASSICAL APPROACHES
Very broad selection of tools
48
WHAT ABOUT FEATURE
ENGINEERING?
49
DEEP REPRESENTATION
LEARNING
50
DEEP REPRESENTATION LEARNING
Beyond distributional hypothesis
51
Part 1: Machine Learning in NLP
• Lecture
• What is NLP?
• Problem Formulation
• Text Representations
• Dimensionality Reduction
• Embeddings
• RNNs
• “Attention is All You Need”
• Lab
• Transformer Architecture
• BERT Model
• Pretraining BERT
52
RECURRENT NEURAL NETWORKS
Basic principles
Unrolling in Time
53
LONG SHORT TERM (LSTM) CELL
Addressing problems of stability
54
CNNS
55
CONVOLUTIONAL NEURAL NETWORKS
Basic principles
Severyn, Aliaksei, and Alessandro Moschitti. "Unitn: Training deep convolutional neural network for twitter sentiment 56
classification." Proceedings of the 9th international workshop on semantic evaluation (SemEval 2015). 2015.
ATTENTION
57
WHAT ABOUT LONG SEQUENCES?
The challenge illustrated with SQuAD
58
The impact of attention mechanism on Question Answering performance
WHAT ABOUT LONG SEQUENCES?
The challenge
59
Bahdanau, D., Cho, K., & Bengio, Y. (2014). Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473.
ATTENTION
The mechanism
60
Bahdanau, D., Cho, K., & Bengio, Y. (2014). Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473.
ATTENTION
The mechanism
61
Bahdanau, D., Cho, K., & Bengio, Y. (2014). Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473.
ATTENTION
Examples
Wu, Y., Schuster, M., Chen, Z., Le, Q. V., Norouzi, M., Macherey, W., ... & Klingner, J. (2016). Google's neural machine translation system: Bridging the gap 62
between human and machine translation. arXiv preprint arXiv:1609.08144.
ATTENTION
Examples
Gehring, J., Auli, M., Grangier, D., Yarats, D., & Dauphin, Y. N. (2017, July). Convolutional sequence to sequence learning. In International 63
conference on machine learning (pp. 1243-1252). PMLR.
Part 1: Machine Learning in NLP
• Lecture
• What is NLP?
• Problem Formulation
• Text Representations
• Dimensionality Reduction
• Embeddings
• RNNs
• “Attention is All You Need”
• Lab
• Transformer Architecture
• BERT Model
• Pretraining BERT
64
ATTENTION IS ALL YOU NEED
Design
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. In Advances in neural 65
information processing systems (pp. 5998-6008).
ATTENTION IS ALL YOU NEED
Design
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. In Advances in neural 66
information processing systems (pp. 5998-6008).
WAS IT A BREAKTHROUGH
IN ITSELF?
67
ATTENTION IS ALL YOU NEED
Not a breakthrough in itself
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. In Advances in neural 68
information processing systems (pp. 5998-6008).
ATTENTION IS ALL YOU NEED
But …
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. In Advances in neural 69
information processing systems (pp. 5998-6008).
NEURAL EMBEDDINGS
70
FEATURE REUSE
The opportunity
71
IT WAS DIFFICULT TO
REUSE NLP EMBEDDINGS
72
SEMI-SUPERVISED SEQUENCE LEARNING
More complex representations
73
Dai, A. M., & Le, Q. V. (2015). Semi-supervised sequence learning. In Advances in neural information processing systems (pp. 3079-3087).
SEMI-SUPERVISED SEQUENCE LEARNING
More complex representations
74
Dai, A. M., & Le, Q. V. (2015). Semi-supervised sequence learning. In Advances in neural information processing systems (pp. 3079-3087).
SEMI-SUPERVISED SEQUENCE LEARNING
More complex representations
75
Dai, A. M., & Le, Q. V. (2015). Semi-supervised sequence learning. In Advances in neural information processing systems (pp. 3079-3087).
ELMO
Embeddings for Language Models
Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., & Zettlemoyer, L. (2018). Deep contextualized word representations. arXiv preprint 76
arXiv:1802.05365.
ELMO
Embeddings for Language Models
Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., & Zettlemoyer, L. (2018). Deep contextualized word representations. arXiv preprint 77
arXiv:1802.05365.
ULM-FIT
Universal Language Model Fine-Tuning for Text Classification
78
Howard, J., & Ruder, S. (2018). Universal language model fine-tuning for text classification. arXiv preprint arXiv:1801.06146.
TRANSFER LEARNING IN NLP
Not trivial to use and not universally applicable
Howard, J., & Ruder, S. (2018). Universal language model fine-tuning for text classification. arXiv preprint arXiv:1801.06146.
Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., & Zettlemoyer, L. (2018). Deep contextualized word representations. arXiv preprint
arXiv:1802.05365.
2013/
1957 1988 2010 2018
2014
79
THIS CREATED A
FOUNDATION FOR THE
NEW NLP MODELS
(DISCUSSED IN THE NEXT CLASS)
80
THE LAB
81
ATTENTION IS ALL YOU NEED
Deep dive into the transformer design
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. In Advances in neural 82
information processing systems (pp. 5998-6008).
83
BERT
How it relates to transformer and pretraining
83
IN THE NEXT CLASS…
84
SELF-SUPERVISION, BERT, AND BEYOND
Why did models start to work well? What does the future hold?
85
Part 1: Machine Learning in NLP
• Lecture
• What is NLP?
• Problem Formulation
• Text Representations
• Dimensionality Reduction
• Embeddings
• RNNs
• “Attention is All You Need”
• Lab
• Transformer Architecture
• BERT Model
• Pretraining BERT
86