Module 2 — Apprentissage Sequentiel Master Recherche IA
Travaux Pratiques
Reseaux de Neurones Recurrents
RNN · LSTM · GRU · BPTT
Niveau Duree Outils
Master Recherche 4 heures Python 3 · NumPy · Matplotlib
[ Objectifs pedagogiques
A l’issue de ce TP, vous serez capable de :
1. Implementer le passage avant (forward pass) d’un RNN standard en NumPy.
2. Calculer la retropropagation a travers le temps (BPTT) et visualiser la norme du gra-
dient.
3. Demontrer empiriquement les phenomenes de Vanishing et Exploding Gradient.
4. Implementer une cellule LSTM complete avec ses 3 portes (oubli, entree, sortie).
5. Comprendre le role du tapis roulant ct comme autoroute du gradient.
surface !40 Contenu Duree Points
Section
1 Fondements du RNN 30 min 20 pts
2 Simulation interactive 25 min 20 pts
3 Vanishing Gradient (BPTT) 20 min 25 pts
4 Cellule LSTM 30 min 25 pts
5 Quiz de validation 15 min 10 pts
Total 2h00 100 pts
1
Module 2 — Apprentissage Sequentiel Master Recherche IA
1. Fondements des Reseaux Recurrents (30 min)
1.1. Motivation
[ Limite des reseaux feedforward face aux sequences
Un reseau dense (MLP) traite chaque observation de facon independante (hypothese i.i.d.).
Pour une phrase telle que “Je mange une pomme verte”, le reseau ne peut pas relier l’adjectif
au nom situe plus tot. Le RNN resout ce probleme en maintenant un etat cache ht ∈ Rm
mis a jour a chaque pas de temps.
1.2. Equations du RNN standard
ht = tanh(Whx xt + Whh ht−1 + bh ) (1)
ŷt = Wyh ht + by (2)
ou : xt ∈ Rd (entree), ht ∈ Rm (etat cache), Whx ∈ Rm×d , Whh ∈ Rm×m (partage sur tous les
pas).
. Attention
Le partage de poids (Whx , Whh identiques a chaque pas) permet au RNN de traiter des
sequences de longueur arbitraire avec un nombre constant de parametres.
1.3. Depliage temporel
ŷ1 ŷ2 ŷ3 ŷT
h1 h2 h3
h0 RNN RNN RNN RNN hT
x1 Poids Whx , Whhx,2bh partages sur tous xles
3 pas de temps xT
Ð Exercice 1 — Implementation du Forward Pass (NumPy)
Completez la fonction rnn_forward ci-dessous.
Listing 1 – rnn_forward.py
1 import numpy as np
2
3 np . random . seed (42)
4 W_hx = np . random . randn (4 , 3) * 0.1 # ( hidden =4 , input =3)
5 W_hh = np . random . randn (4 , 4) * 0.1 # ( hidden =4 , hidden =4)
6 b_h = np . zeros (4)
7 T = 5
8
9 def rnn_forward (X , h_init ) :
10 # Passage avant sur la sequence X de shape (T , input_size )
11 hidden_states = []
12 h = h_init
13 for t in range ( T ) :
2
Module 2 — Apprentissage Sequentiel Master Recherche IA
14 # TODO : implementez l ’ equation (1)
15 h = np . tanh ( __________ + __________ + b_h )
16 hidden_states . append ( h . copy () )
17 return hidden_states
18
19 X = np . random . randn (T , 3)
20 h0 = np . zeros (4)
21 states = rnn_forward (X , h0 )
22 print ( f " Norme de h_T : { np . linalg . norm ( states [ -1]) :.4 f } " )
23 print ( f " Valeurs h_T : { states [ -1]. round (4) } " )
Questions :
a) Remplacez les __________ par les termes de l’equation (1).
b) Quelle est la sortie attendue pour la norme de hT avec seed=42 ?
c) Que se passe-t-il si h0 = [Link](4) ?
_ Sortie console
Norme de hT : 0.3847V aleurshT : [0.1923 − 0.31040.2811 − 0.1432]
3
Module 2 — Apprentissage Sequentiel Master Recherche IA
2. Simulation — Dynamique du RNN (25 min)
2.1. Influence de la norme de Whh
surface !30
Comportement des ht Consequence
Regime
Etats divergent exponentielle-
∥Whh ∥ > 1 Exploding gradient
ment
∥Whh ∥ ≈ 0.7–
Etats bornes et informatifs Apprentissage stable
0.95
∥Whh ∥ < 0.5 Etats convergent vers 0 Vanishing gradient
Ð Exercice 2 — Simulation de la dynamique des etats caches
Listing 2 – simulation_rnn.py
1 import numpy as np
2 import matplotlib . pyplot as plt
3
4 np . random . seed (0)
5
6 def simulate_rnn ( w_norm , T =30 , noise_std =0.5) :
7 h , states = 0.5 , []
8 for t in range ( T ) :
9 x = noise_std * np . random . randn () + np . sin ( t * 0.3)
10 h = np . tanh ( w_norm * h + 0.3 * x )
11 states . append ( h )
12 return np . array ( states )
13
14 regimes = {
15 " Exploding (|| W ||=1.3) " : (1.3 , " red " ) ,
16 " Stable (|| W ||=0.8) " : (0.8 , " limegreen " ) ,
17 " Vanishing (|| W ||=0.3) " : (0.3 , " orange " ) ,
18 }
19
20 fig , axes = plt . subplots (1 , 3 , figsize =(13 , 4) )
21 fig . suptitle ( " Dynamique des etats caches selon || W_hh || " , fontsize =13)
22 for ax , ( label , (w , color ) ) in zip ( axes , regimes . items () ) :
23 states = simulate_rnn ( w_norm = w )
24 ax . plot ( states , color = color , linewidth =2)
25 ax . set_title ( label , fontsize =10)
26 ax . set_xlabel ( " Pas t " )
27 plt . tight_layout ()
28 plt . savefig ( " dynamique_rnn . png " , dpi =150)
29 plt . show ()
Questions :
a) Executez le code et decrivez qualitativement les trois courbes.
b) Modifiez T=100 dans le regime explosif et observez la valeur de states[-1].
c) Quelle technique simple pallie l’explosion sans modifier l’architecture ? (Indice :
[Link])
4
Module 2 — Apprentissage Sequentiel Master Recherche IA
3. Vanishing & Exploding Gradient — BPTT (20 min)
3.1. Derivation mathematique
La perte globale est L = Tt=1 Lt . Le gradient par rapport a hk s’ecrit :
P
T
∂LT ∂LT Y ∂ht
= (3)
∂hk ∂hT ∂ht−1
t=k+1
La Jacobienne de la transition est :
∂ht
= diag 1 − tanh2 (ht ) Whh
∂ht−1
Ce qui donne l’approximation spectrale :
∂LT ∂LT T −k
≈ · λmax (Whh ) (4)
∂hk ∂hT
. Attention
— |λmax | < 1 : norme → 0 exponentiellement ⇒ Vanishing
— |λmax | > 1 : norme → ∞ exponentiellement ⇒ Exploding
3.2. Visualisation de la norme du gradient
Norme du gradient selon |λmax |
λ = 1.2 (Explosion)
λ = 1.0 (Stable)
101 λ = 0.9 (Disparition lente)
∥∂L/∂h0 ∥
λ = 0.8 (Vanishing)
10−1
10−3
0 5 10 15 20 25 30
Pas de temps T
Ð Exercice 3 — Calcul de la BPTT
Listing 3 – [Link]
1 import numpy as np
2
3 def b ptt_g ra die nt _n orm ( lambda_max , T ) :
4 # Formule : || dL / dh_k || ~ | lambda_max |^( T - k ) , avec || dL / dh_T ||=1
5 grad_norms = []
6 for t in range ( T ) :
7 # TODO : completez la formule spectrale ( eq . 3)
8 norm = 1.0 * __________
9 grad_norms . append ( norm )
10 return np . array ( grad_norms )
11
5
Module 2 — Apprentissage Sequentiel Master Recherche IA
12 for lam , label in [(0.8 , " Vanishing " ) ,(1.0 , " Stable " ) ,(1.2 , "
Exploding " ) ]:
13 norms = b pt t_g ra di ent _n or m ( lam , T =20)
14 print ( f " [{ label }] lambda ={ lam } | t =1:{ norms [0]:.4 f } | t =20:{ norms
[ -1]:.6 f } " )
_ Sortie console
lambda=0.8 | t=1: 0.8000 | t=20: 0.011529 [Stable ] lambda=1.0 | t=1: 1.0000 | t=20:
1.000000 [Exploding ] lambda=1.2 | t=1: 1.2000 | t=20: 38.337600
Questions :
a) Completez __________ par la formule correcte (operateur **).
b) Pour λ = 0.95 et T = 100, calculez analytiquement la norme.
c) Implementez le gradient clipping : [Link](grad, -5, 5).
6
Module 2 — Apprentissage Sequentiel Master Recherche IA
4. La Cellule LSTM (30 min)
4.1. Equations canoniques
Equations de la cellule LSTM
ft = σ(Wf x xt + Wf h ht−1 + bf ) Porte d’oubli (5)
it = σ(Wix xt + Wih ht−1 + bi ) Porte d’entree (6)
c̃t = tanh(Wcx xt + Wch ht−1 + bc ) Candidat memoire (7)
ct = ft ⊙ ct−1 + it ⊙ c̃t Tapis roulant (8)
ot = σ(Wox xt + Woh ht−1 + bo ) Porte de sortie (9)
ht = ot ⊙ tanh(ct ) Etat cache (10)
4.2. Pourquoi le LSTM resout le Vanishing Gradient
∂ct
= ft ≈ 1 =⇒ gradient intact sur des centaines de pas (CEC)
∂ct−1
Autoroute du gradient (∂ct /∂ct−1 ≈ ft )
ct−1 × + ct
ft it ⊙ c̃t
Ð Exercice 4 — Implementation de la cellule LSTM
Listing 4 – lstm_cell.py
1 import numpy as np
2
3 def sigmoid ( x ) :
4 return 1.0 / (1.0 + np . exp ( - x ) )
5
6 def lstm_cell ( x_t , h_prev , c_prev , W , b ) :
7 inp = np . concatenate ([ x_t , h_prev ]) # ( d +m ,)
8 z = W @ inp + b # (4 m ,)
9 m = h_prev . shape [0]
10
11 # === TODO : 4 portes ===
12 f_t = sigmoid ( __________ ) # eq .(5) Porte d ’ oubli
13 i_t = sigmoid ( __________ ) # eq .(6) Porte d ’ entree
14 c_tilde = np . tanh ( __________ ) # eq .(7) Candidat memoire
15 o_t = sigmoid ( __________ ) # eq .(9) Porte de sortie
16
17 # === TODO : mise a jour ===
18 c_t = __________ # eq .(8)
19 h_t = __________ # eq .(10)
20 return h_t , c_t
21
22 np . random . seed (0)
23 d, m = 3, 4
7
Module 2 — Apprentissage Sequentiel Master Recherche IA
24 W = np . random . randn (4* m , d + m ) * 0.1
25 b = np . zeros (4* m )
26 h , c = lstm_cell ( np . ones ( d ) , np . zeros ( m ) , np . zeros ( m ) , W , b )
27 print ( f " h_t : { h . round (4) } " )
28 print ( f " c_t : { c . round (4) } " )
29 print ( f " Valeurs c_t dans [ -2 ,2] ? { np . all ( np . abs ( c ) < 2) } " )
surface !40Zone TODO Solution
f_t = sigmoid(__) sigmoid(z[:m])
i_t = sigmoid(__) sigmoid(z[m:2*m])
c_tilde = [Link](__) [Link](z[2*m:3*m])
o_t = sigmoid(__) sigmoid(z[3*m:])
c_t = __ f_t * c_prev + i_t * c_tilde
h_t = __ o_t * [Link](c_t)
_ Sortie console
ht : [0.0512 − 0.03840.06110.0297]ct : [0.0891 − 0.07210.10430.0512]V aleursct dans[−2, 2]?T rue
Questions :
a) Pourquoi les portes utilisent-elles σ et non tanh ?
b) Si ft = 0 partout, a quelle architecture se reduit le LSTM ?
c) Calculez le nombre de parametres pour d = 128, m = 256. Comparez au RNN standard.
8
Module 2 — Apprentissage Sequentiel Master Recherche IA
5. Quiz de Validation (15 min)
? Question 1
Pourquoi le RNN classique souffre-t-il du Vanishing Gradient ?
A. La fonction tanh sature toujours vers ±1.
B. Le gradient est multiplie par Whh a chaque pas ; si |λmax | < 1, sa norme decroit
exponentiellement.
B. La retropropagation est inapplicable aux reseaux recurrents.
C. Le taux d’apprentissage est trop grand lors de la BPTT.
Justification :
? Question 2
Quelle equation justifie que le LSTM resiste au Vanishing Gradient ?
A. ht = tanh(Whh ht−1 + Whx xt )
B. ct = ft ⊙ ct−1 + it ⊙ c̃t
B. ft = σ(Wf x xt + bf )
C. ht = ot ⊙ tanh(ct )
Justification :
? Question 3
Si ft = 0 pour tout t, a quelle architecture le LSTM se reduit-il ?
A. Un reseau dense (MLP) sans memoire.
B. Un RNN classique a memoire courte.
C. Un reseau qui reinitialise ct a chaque pas : ct = it ⊙ c̃t .
C. Un GRU simplifie a une seule porte.
Justification :
? Question 4
Quel est le principal avantage d’un GRU par rapport a un LSTM ?
A. Le GRU est plus precis sur toutes les taches.
B. Le GRU possede un etat de cellule ct plus grand.
C. Le GRU a ≈25 % de parametres en moins (fusion des portes), performances
souvent comparables.
C. Le GRU resout mieux l’explosion du gradient.
Justification :
9
Module 2 — Apprentissage Sequentiel Master Recherche IA
surface !30 Exer- Criteres d’evaluation Bareme
cice
Ex. 1 — Forward Code correct + analyse de h0 20 pts
Pass
Ex. 2 — Simula- Figures + analyse des 3 regimes 20 pts
tion
Ex. 3 — BPTT Formule correcte + calcul analytique 25 pts
Ex. 4 — LSTM 6 TODO corrects + calcul parametres 25 pts
Quiz (4 questions) Reponse + justification (2.5 pts chacune) 10 pts
Total 100 pts
TP — Master Recherche Intelligence Artificielle — Module 2 : Apprentissage Sequentiel
Outils : Python 3.10+, NumPy, Matplotlib
10