0% found this document useful (0 votes)
13 views84 pages

Understanding Neural Networks and AI

machine learning

Uploaded by

Omis
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
13 views84 pages

Understanding Neural Networks and AI

machine learning

Uploaded by

Omis
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Version : 1.

14

SARI 6 Juin 2019


Powered by

Jean-Luc Parouty
Laboratoire SIMaP
[Link]
[ intelligence ]
« Ability to perceive or infer information,
and to retain it as knowledge to be appli
towards adaptive behaviors within an
environment or context »*

« Capacité de percevoir ou d'inférer l'information, et de la


conserver comme une connaissance à appliquer à des
comportements adaptatifs dans un environnement ou un
contexte donné »

4
*
Wikipedia 86
[ Méthode scientifique ]

5
86
[ Méthode scientifique ]

6
1
Jim Gray, 2007 [GRAY] 86
[ *-learning ]

Deep
Artificial Learning (DL)
Intelligence (AI)

Machine
Decision,
Learning (ML)
Prediction,
Classification,
etc. 7
86
8
86
9
86
1/ From the linear regression
to the first neuron

2/ Neural networks at the


heart of a controversy

3/ Neurons & data

4/ Conclusion
10
86
1/ From the linear regression
to the first neuron

...there's a little bit of math hidden behind the


neurons...

12
86
Linear regression

13
86
Linear regression

RMSE : Root Mean Square Error


Erreur quadratique moyenne

complexity in n3

14
Notebook [LAB1] 86
Gradient descent

δ loss
Iterations
Loss

δΘ

loss(Θ)

Best Θ Θ

15
86
Gradient descent

#i Loss Gradient Theta


0 +12.481 -6.777 -1.732 -3.388 +0.000
20 +4.653 -4.066 -1.039 -2.033 +0.346
40 +1.835 -2.440 -0.624 -1.220 +0.554
60 +0.821 -1.464 -0.374 -0.732 +0.679
80 +0.455 -0.878 -0.224 -0.439 +0.754
100 +0.324 -0.527 -0.135 -0.263 +0.799
120 +0.277 -0.316 -0.081 -0.158 +0.826
140 +0.260 -0.190 -0.048 -0.095 +0.842
160 +0.253 -0.114 -0.029 -0.057 +0.851
180 +0.251 -0.068 -0.017 -0.034 +0.857
200 +0.250 -0.041 -0.010 -0.020 +0.861

16
Notebook [LAB2] 86
Polynomial regression

17
Notebook [LAB3] 86
Logistic regression
A logistic regression is intended to provide a probability of belonging to a class.

Dataset : X characteristics
y probability of belonging

19
Notebook [LAB12] 86
Logistic regression
A logistic regression is intended to provide a probability of belonging to a class.

Dataset : X Observations Objective : Predict the class


y Classe X given, we want to predict y

20
Notebook [LAB12] 86
Logistic regression

80 % Determination of Θ
by a minimisation
Learning of the log loss J(Θ)
phase

Données
(X,y)

Evaluation
20 % phase

21
Notebook [LAB12] 86
Logistic regression

22
Notebook [LAB12] 86
Logistic regression
Determined by the minimisation
of a cost function J(Θ)

23
Notebook [LAB12] 86
Logistic regression

That’s an « artificial neuron » !


So, we have a neural network of… 1 neuron !
25
Notebook [LAB12] 86
Logistic regression

26
Notebook [LAB12] 86
Perceptron

Perceptron
Frank Rosenblatt
1958

Linear and binary classifier

28
F. Rosenblatt, 1958 [FROS] 86
Perceptron
Example
Iris plants dataset
Dataset from : Fisher, R.A. “The use of multiple measurements in
taxonomic problems” Annual Eugenics, 7, Part II, 179-188
(1936)
Length Width Iris Setosa (0/1)
x1 x2 y
1.4 1.4 1
1.6 1.6 1
1.4 1.4 1
1.5 1.5 1
1.4 1.4 1
4.7 4.7 0
4.5 4.5 0
4.9 4.9 0
4.0 4.0 0
4.6 4.6 0
(...)

29
Notebook [LAB13] 86
Perceptron

Linear classifier...

1969
Marvin Minsky, Seymour Papert
« Perceptrons : An Introduction to
Computational Geometry » 1

First AI winter…
(for neural networks)

30
1
Minsky, Marvin; Papert, Seymour, (1969) [MIPA] 86
2/ Neural networks at the
heart of a controversy

31
86
ro ig
rsy
nt e b
ve
Co Th

Modelling the brain : Making a mind :


« Penser s’apparente « Penser, c’est calculer des symboles qui
à un calcul massivement parallèle de ont à la fois une réalité matérielle et une
fonctions élémentaires. valeur sémantique de représentation »1
L’information est un signal avant L’information est une donnée
d’être un code »1 symbolique de haut niveau.

Connectionnism vs Symbolic
Modelling the brain Making a mind
Modéliser le cerveau Forger une opinion

Tout [homme] est [mortel]


[Socrate] est un [homme]
Donc [Socrate] est [mortel]

32
1
D Cardon, JP Cointet, A Mazieres, 2018 [LRDN] 86
ro ig
rsy
nt e b
ve
Co Th

Connectionnism vs Symbolic

Facts Rules and laws Expert

Model Rules and laws Special case

33
86
ro ig
rsy
nt e b
ve
Co Th

Evolution of the academic influence of connexionist and symbolic approaches 1


Ration of publications between connexionists and symbolists

65 522 publications
106 278 publications

34
1
D Cardon, JP Cointet, A Mazieres, 2018 [LRDN] 86
ro ig
rsy
nt e b
ve
Co Th

Evolution of the academic influence of connexionist and symbolic approaches 1

First concept of
artificial neural
network
McCulloch, Pitts Perceptron
1943 ONR
Rosenblatt
1957

Macy
conferences
1941-1960

65 522 publications
106 278 publications

35
1
D Cardon, JP Cointet, A Mazieres, 2018 [LRDN] 86
ro ig
rsy
nt e b
ve
Co Th

Evolution of the academic influence of connexionist and symbolic approaches 1

MIT (Minsky, Papert)


CM (Simon, Newel)
Stanford (McCarthy)
DARPA

Perceptrons
Minsky, Papert
1969

Artificial Intelligence
Darmouth workshop
McCarthy
1956
First AI Winter
Mansfield
amendment (1969)
65 522 publications
106 278 publications Lighthill
report (1973)
36
1
D Cardon, JP Cointet, A Mazieres, 2018 [LRDN] 86
ro ig
rsy
nt e b
ve
Co Th

Evolution of the academic influence of connexionist and symbolic approaches 1

Backprop. Convol. NN
Rumelhart LeCun
1986 1989

Extinction
First AI Winter of LISP
Mansfield Expert machines
65 522 publications
amendment (1969) systems
106 278 publications Lighthill
report (1973)
37
1
D Cardon, JP Cointet, A Mazieres, 2018 [LRDN] 86
Deep Neural Networks

39
86
Deep Neural Networks

Input layer Hidden layers Output layer


40
86
Deep Neural Networks

Optimisations :
Activation,
Gradient descent,
Back-propagation Dropout,
Regularization,
Learning process Etc. 41
86
Deep Neural Networks

1958

42
86
ro ig
rsy
nt e b
ve
Co Th

Evolution of the academic influence of connexionist and symbolic approaches 1

Backprop. Convol. NN
Rumelhart LeCun
1986 1989

Extinction
of LISP
Expert machines
65 522 publications
systems
106 278 publications

43
1
D Cardon, JP Cointet, A Mazieres, 2018 [LRDN] 86
ro ig
rsy
nt e b
ve
Co Th

Evolution of the academic influence of connexionist and symbolic approaches 1

Support Vector
Machine (SVM)
Vapnik
1995

2nd Winter
(For DL)

65 522 publications
106 278 publications

44
1
D Cardon, JP Cointet, A Mazieres, 2018 [LRDN] 86
rtoro boig
y?
vev f
resres
notn nde
CoC ETh

Performance Development1 Datasets for machine-learning2


100 000 000

ImageNet Reuters
10 000 000 SIFT10M
Open Images

COCO

1 000 000 Youtube comedy

PASCAL VOC
KDD
LabelMe
100 000
MNIST CIFAR-10
Caltech256
Letter Dataset
10 000 Caltech101

× 10 104
Mushroom
6 TIMIT

×
Flops datasets
1 000
1980 1985 1990 1995 2000 2005 2010 2015 2020
25 ans

Laboratoire Monde réel


Cas particulier
45
1
TOP500 List [TOP500] 2
Wikipedia [WKP1] 86
rors e
e rNo ve ig
evue ng
nys
thnt Ree b
oCfoTheTh

Images classification
Publications SVM vs DNN1 Top 5 error at ILSVRC3,4
18 000 18
AlexNet
SVM 16
16 000
DNN 14
AlexNet2 Clarifai
14 000 12
A. Krizhevsky, 10
I. Sutskever,
12 000 DN 8
G. Hinton N
GoogLeNet

6
10 000 2012 ILSVRC
Publications

ResNet
4 Trimps-Soushen
SE-ResNet
Top5 error : 26 %  15 % HUMAN
8 000 2
0
6 000 2011 2013 2015 2017
2nd Winter
4 000
(For DL)
2 000
Without mathematical
guarantee, DNN have proven to
0 be more effective in the face of
1990 1995 2000 2005 2010 2015 2020
the complexity of the real
world !
1
Web of Science [WOS1][WOS2]
2
AlexNet [ALEX] 4
Similar evolution in Natural language processing, translation, board games, etc. 46
3
ImageNet Large Scale Visual Recognition [ILSVRC] See : [Link], AlphaGo, AlphaZero, ... 86
3/ Neurons & data

47
86
Generative
Adversarial Basic
Network Classification
GAN DNN

Reinforcement 3/ Neurons & data Hight


Dimensionnal Data
learning (images, vidéos, …)
CNN

Sequences data Sparse data


(Time data, ...) (text, …)
RNN Embedding
48
86
Generative
Adversarial Basic
Network Classification
GAN DNN

Reinforcement 3/ Neurons & data Hight


Dimensionnal Data
learning (images, vidéos, …)
CNN

Sequences data Sparse data


(Time data, ...) (text, …)
RNN Embedding
49
86
Basic example / MNIST

Image Matrix Vector Input layer


28x28 pixels (28,28) (784) (784 neurons)
50
Notebook [LAB14.1] 86
Basic example / MNIST

Output layer Probability

51
Notebook [LAB14.1] 86
Basic example
Handwritten Digits
Recognition
MNIST dataset
Tensorflow, Jupyter lab


52
86
Generative
Adversarial Basic
Network Classification
GAN DNN

Reinforcement 3/ Neurons & data Hight


Dimensionnal Data
learning (images, vidéos, …)
CNN

Sequences data Sparse data


(Time data, ...) (text, …)
RNN Embedding
53
86
Convolutional Neural Networks (CNN)

24 M pixels 3 x 24 M neurons ?!
(r,v,b) 3x8 bits

10 000 70 M 100 Mds

1 000 000 700 M 250 Mds 54


86
Convolutional Neural Networks (CNN)

2D convolution
55
86
Convolutional Neural Networks (CNN)

3D convolution
56
86
Convolutional Neural Networks (CNN)

57
86
Image classification
with MobileNet v1
Trained model
TensorflowJS, Javascript


58
86
Generative
Adversarial Basic
Network Classification
GAN DNN

Reinforcement 3/ Neurons & data Hight


Dimensionnal Data
learning (images, vidéos, …)
CNN

Sequences data Sparse data


(Time data, ...) (text, …)
RNN Embedding
59
86
Word Embedding

« I've never seen a movie like this before. »


1 a 0 0 0 1 0 0 0 0
2 before 0 0 0 0 0 0 0 1
3 fantastic 0 0 0 0 0 0 0 0
4 i’ve 1 0 0 0 0 0 0 0
5 is 0 0 0 0 0 0 0 0
6 like 0 0 0 0 0 1 0 0
7 movie 0 0 0 0 1 0 0 0
8 never 0 1 0 0 0 0 0 0
9 seen 0 0 1 0 0 0 0 0
10 this 0 0 0 0 0 0 1 0
Sparse matrix

Dictionary = 80 000
Sentence = 300 Vectors = 24 M
60
86
Word Embedding

« movie» CBOW
1 a word2vec 1
0
2 before 0 SG
3 fantastic 0 « movie» GloVe2
4 i’ve 0
5 is 0 -2,03
6 like 0 15,12 ...
7 movie 1 -13,08
8 never 0
9 seen 0 Short ProtVec3
10 this 0 dense vector
based on
context

2
Jeffrey Pennington & all, (2014), [GLOVE]
1
Tomas Mikolov & all, (2013), [W3VEC] Training is performed on aggregated global word-word
CBOW : Continuous Bag of Words - Embedding based co-occurrence statistics.
on the prediction of the word according to its context. 3
Ehsaneddin Asgari, Mohammad R.K. Mofrad (2016), [PROTV] 61
SG : Skip-gram - Embedding based on context prediction from the word. Biological Sequences Representation 86
IMDB film review
classification
Word Embedding
Keras, jupyter lab

 87 % 62
86
Generative
Adversarial Basic
Network Classification
GAN DNN

Reinforcement 3/ Neurons & data Hight


Dimensionnal Data
learning (images, vidéos, …)
CNN

Sequences data Sparse data


(Time data, ...) (text, …)
RNN Embedding
63
86
Reccurent Neural Network (RNN)

...

X(t-4) X(t-3) X(t-2) X(t-1) X(t)

64
86
Reccurent Neural Network (RNN)

65
86
Reccurent Neural Network (RNN)

Unfold

66
86
Reccurent Neural Network (RNN)

Recurrent neuron Slow convergence,


Recurrent layer « Cell » Short memory,
Vanishing / exploding gradients 67
86
Reccurent Neural Network (RNN)

Long term

Short term

Long short-term memory (LSTM)1


Gated recurrent unit (GRU)2
68
1
Sepp Hochreiter, Jürgen Schmidhuber, (1997) [LSTM] 2
Kyunghyun Cho et al, (2014) [GRU] 86
Reccurent Neural Network (RNN)

Serie to serie Serie to vector


Example : Time serie prediction Example : Sentiment analysis

Vector to serie Encoder-decoder


Example : Image annotation Example : Language Translation

69
86
Reccurent Neural Network (RNN)

Time serie
prediction
RNN with LSTM cell
Tensorflow, jupyter lab


70
86
Generative
Adversarial Basic
Network Classification
GAN DNN

Reinforcement 3/ Neurons & data Hight


Dimensionnal Data
learning (images, vidéos, …)
CNN

Sequences data Sparse data


(Time data, ...) (text, …)
RNN Embedding
71
86
Reinforcement learning

72
86
Reinforcement learning

What actions can be taken to maximize rewards ?


73
86
Reinforcement learning

Inverted pendulum

Objective :
Keep the pendulum in balance,
in the centre of the stage

Impulse to
the left (-1)
Actions :
Impulse to
the right (+1)

74
86
Reinforcement learning

Inverted pendulum
Observations :
x Cart position
vx Cart velocity
Θ Pole angle
ωΘ Pole angular velocity

Rewards :
Based on keeping the bar in
balance for as long as possible,
while remaining in the centre of
the stage

75
86
Reinforcement learning

76
86
Reinforcement learning

77
86
Reinforcement learning

Reinforcement
learning
OpenAI/Gym Cartpole
with gradient policy


78
86
Generative
Adversarial Basic
Network Classification
GAN DNN

Reinforcement 3/ Neurons & data Hight


Dimensionnal Data
learning (images, vidéos, …)
CNN

Sequences data Sparse data


(Time data, ...) (text, …)
RNN Embedding
79
86
Generative Adversarial Network
GAN1 Use Cases :
● Photorealistic images generation
● Image to Image Translation
● Increasing Image Resolution
● Text to Image Generation
● Video / Frame prediction
● Etc.

Counterfeiter Expert
(Generator) (Discriminator)

80
1
Ian J. Goodfellow & all, (2014), « Generative Adversarial Networks » [GAN] 86
Generative Adversarial Network

81
86
Generative Adversarial Network

Generative
Adversarial
Network
Photorealistic generation


82
86
4/ Conclusion

83
86
Conclusion

Complex but
Great accessible tools Very significant
opportunities and techniques and rapid progress
Science
it works !
Open Data
Source

AlphaFold won 13th Critical Assessment of Structure Prediction (CASP) 84


Prediction of the 3D structure of proteins from a given amino acid sequence as input. 86
Conclusion

« (…) Due to our concerns about


malicious applications of the
technology, we are not releasing the
trained model.(...) »
[Link]

Algorithmes, la bombe à retardement


Editions Les Arènes
Cathy O'Neil

Major societal « San Francisco Bans Facial


impacts Recognition Technology »
New York Times
May 14, 2019

COMMENT PERMETTRE À L’HOMME


DE GARDER LA MAIN1 ?
Les enjeux éthiques des algorithmes et de
l’intelligence artificielle
SYNTHÈSE DU DÉBAT PUBLIC ANIMÉ PAR LA CNIL DANS LE CADRE DE LA MISSION
DE RÉFLEXION ÉTHIQUE CONFIÉE PAR LA LOI POUR UNE RÉPUBLIQUE NUMÉRIQUE
85
1
Report available on the CNIL website 86
Références
[JGRAY] Gray, J. (2001), from « The Fourth Paradigm: Data-Intensive [AMAZ] Antoine Mazieres (2016) Thèse : « Cartographie de l’apprentissage
Scientific Discovery » Tony Hey, Stewart Tansley, Kristin Tolle artificiel et de ses algorithmes » Université Paris 7 Denis Diderot,
(2009). Published by Microsoft Research. <hal-01771655>
ISBN: 978-0-9825442-0-4 [TOP500] Statistics on top 500 high-performance computers. (2018)
[MCPIT] McCulloch, Warren; Walter Pitts (1943). "A Logical Calculus of « Exponential growth of supercomputing power as recorded by the
Ideas Immanent in Nervous Activity". Bulletin of Mathematical TOP500 list ». [Link]
Biophysics. 5 (4): 115–133. doi:10.1007/BF02478259 [WKP1] Wikipedia/en. (2018) « List of datasets for machine-learning
[DHEBB] Hebb, D. O. (1949). « The Organization of Behavior: A research ». [Link]
Neuropsychological Theory. » New York: Wiley and Sons. [WOS1] Core database : TS=("support vector machine*" OR ("SVM" AND
ISBN 9780471367277. "classification") OR ("SVM" AND "regression") OR ("SVM" AND
[FROS] Rosenblatt, Frank. (1958). « The perceptron: A probabilistic "classifier") OR "support vector network*" OR ("SVM" AND "kernel
model for information storage and organization in the brain. » trick*"))
Psychological Review, 65(6), 386-408. [WOS2] Core database : TS=("deep learning" OR "deep neural network*" OR
[MIPA] Minsky, Marvin; Papert, Seymour. (1969). « Perceptrons : An ("DNN" AND "neural network*") OR "convolutional neural network*"
Introduction to Computational Geometry », MIT Press OR ("CNN" AND "neural network*") OR "recurrent neural network*"
OR ("LSTM" AND "neural network*") OR ("RNN*" AND "neural
[DRUM] Rumelhart, David E.; Hinton, Geoffrey E.; Williams, Ronald J. network*") )
(1986). « Learning representations by back-propagating
errors ». Nature. 323 (6088): 533–536. doi:10.1038/323533a0. [ALEX] A. Krizhevsky, I. Sutskever, G. Hinton. (2012). « ImageNet
Classification with Deep Convolutional Neural Networks »
[YLEC1] Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. doi: 10.1145/3065386
Hubbard, L. D. Jackel, « Backpropagation Applied to
Handwritten Zip Code Recognition », AT&T Bell Laboratories [ILSVRC] ImageNet Large Scale Visual Recognition Challenges
[Link]
[LRDN] Dominique Cardon, Jean-Philippe Cointet, Antoine Mazieres. [Link]
(2018). « La revanche des neurones », Réseaux, La Découverte, 5
(211), <10.3917/res.211.0173>. <hal-01925644> [MOBIN] Howard, Andrew G. et al. (2017) “MobileNets: Efficient
Convolutional Neural Networks for Mobile Vision Applications.”
[Link] 86
86
Références
[W2VEC] Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, Jeffrey [GAN] Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu,
Dean (2013), « Distributed Representations of Words and David Warde-Farley, Sherjil Ozair, Aaron Courville, Yoshua Bengio,
Phrases and their Compositionality », (2014), « Generative Adversarial Networks »
[Link] [Link]
[GLOVE] Jeffrey Pennington, Richard Socher, Christopher D. Manning [CNIL] Comment permettre à l’homme de garder la main ?
(2014) « GloVe: Global Vectors for Word Representation », Synthèse du débat public animé par la cnil dans le cadre de la
[Link] mission de réflexion éthique confiée par la loi pour une république
[P2VEC] Ehsaneddin Asgari, Mohammad R.K. Mofrad, (2016), « ProtVec: numérique.
A Continuous Distributed Representation of Biological [Link]
ur-les-enjeux-ethiques-des-algorithmes-et-de
Sequences »,
[Link]
[LSTM] Sepp Hochreiter, Jürgen Schmidhuber, (1997), « Long Short-
Term Memory,
[Link]
[GRU] Cho, Kyunghyun; van Merrienboer, Bart; Gulcehre, Caglar;
Bahdanau, Dzmitry; Bougares, Fethi; Schwenk, Holger; Bengio,
Yoshua (2014), « Learning Phrase Representations using RNN
Encoder-Decoder for Statistical Machine Translation ».
[Link]
[CARTP] AG Barto, RS Sutton and CW Anderson, (1983), « Neuronlike
Adaptive Elements That Can Solve Difficult Learning Control
Problem », IEEE Transactions on Systems, Man, and Cybernetics,
1983

87
86
Notebooks Notebooks
[LAB1] 01 Regression Liné[Link] [LAB22.2] Word Embedding – Basic
[LAB2] 02 Descente de [Link] TripAdvisor CBOW Embedding with Gensim
[LAB12] 12 Regression [Link] [LAB22.3] Word Embedding – IMDB*
[LAB1] Regression linéaire IMDB film review classification with Keras
Exemple de régression linéaire avec résolution directe [LAB21.3] Time series prediction with RNN*
[LAB2] Gradient descent Prediction of a time serie with LSTM RNN using
Simple gradient descent example Tensofflow
[LAB12] Logistic Regression [LAB19.5] CartPole with Policy gradients*
Logistic Regression with Gradient Descent using CartPole game (from Gym) with gradient policy using
TensorFlow Tensorflow
[LAB12.1] Activation functions
Example of activation functions
[LAB13] Simple Perceptron
IRIS classification with a simple perceptron, using

Illustrations
sklearn
[LAB14.1] Deep Neural Network*
MNIST Example with Tensor Flow
[WEB1] Image classification with MobileNet v1* Illustrations from Wikimedia Commons, the free media repository.
Image classification with MobileNet using tensorflow js
"Morondava - 28" by Olivier Lejade is licensed under CC BY-SA 2.0
[WEB2] Object detection with coco-ssd*
Object detection with coco-ssd/mobilenet using "straight ahead" by HarisDrako is licensed under CC BY-NC-ND 3.0
tensorflow js

88
86
[Link]

binder
[Link]

Attribution-NonCommercial-NoDerivatives 4.0 International (CC BY-NC-ND 4.0)


[Link] 89
86
Thanks !

90
86
91
86

You might also like