🧠 Optimiseur Adam : Lissage + Accélération
*Adam (short for Adaptive Moment Estimation)** optimizer combines the
strengths of two other well-known techniques—
**Momentum** and **RMSprop**—to deliver a powerful method for adjusting
the learning rates of parameters during training.
Adam is highly effective, especially when working with large datasets and
complex models, because it is memory-efficient and adapts the learning rate
dynamically for each parameter.
🧩 Contexte
Dans l’apprentissage profond, l’optimisation est essentielle pour ajuster les poids
du modèle. Les techniques de lissage (smoothing) et accélération (acceleration)
visent à rendre cette descente plus stable et rapide.
🎯 Lissage (Smoothing)
But : Réduire les variations brusques pour stabiliser l’apprentissage.
🧠 Moyenne exponentielle des gradients (ex: Momentum).
🧪 Réduction du bruit dans les données (ex: label smoothing).
⚡ Accélération (Acceleration)
But : Aller plus vite vers le minimum sans se perdre.
🏃♂️ Momentum : ajoute une inertie à la descente.
🔍 Nesterov : anticipe la prochaine position.
🤖 Méthodes adaptatives : ajustent dynamiquement le pas selon la direction.
🔄 Complémentarité
Ces techniques sont souvent combinées (ex : Adam = smoothing + acceleration)
pour un apprentissage rapide et stable.
Adam (Adaptive Moment Estimation) est un optimiseur qui combine les deux
approches :
📉 Lissage avec un suivi de la moyenne des gradients.
🚀 Accélération grâce à l’adaptation du pas d’apprentissage.
🔬 Comment fonctionne Adam ?
Adam s'appuie sur deux estimations :
1. Moyenne des gradients (momentum → 1ʳᵉ estimation / moment) → lissage.
2. Moyenne des gradients au carré (RMSprop → 2ᵉ moment) → adaptation.
⚙️ Formules de base
🔁 1. Moyenne exponentielle des gradients (Momentum) :
m t = β 1 ⋅ m t−1 + (1 − β 1 ) ⋅ ∇L(w t )
🧮 2. Moyenne exponentielle des gradients au carré (RMSprop)
:
2
v t = β 2 ⋅ v t−1 + (1 − β 2 ) ⋅ (∇L(w t ))
✅ 3. Correction du biais (important au début) :
Bias correction: Since both and are initialized at zero, they tend to be
mt vt
biased toward zero, especially during the initial steps. To correct this bias, Adam
computes the bias-corrected estimates:
mt vt
m
^ t = , v
^t =
t t
1 − β 1 − β
1 2
Comme β
t
quand
→ 0 , la correction du biais disparaît progressivement,
t → ∞
Donc la correction devient négligeable
Elle est utile seulement au début pour compenser l’effet de l’initialisation à zéro
🔄 4. Mise à jour des poids :
^ t
m
w t+1 = w t − α ⋅
√v
^ + ϵ
t
📌 Paramètres par défaut
α (alpha) = 0.001 → taux d’apprentissage initial.
β₁ = 0.9 → pour la moyenne des gradients.
β₂ = 0.999 → pour la moyenne des carrés.
ε = 10⁻⁸ → évite la division par zéro.
⚖️ Pourquoi Adam est puissant ?
Adam addresses several challenges of gradient descent optimization:
Dynamic learning rates: Each parameter has its own adaptive learning rate
based on past gradients and their magnitudes. This helps the optimizer avoid
oscillations and get past local minima more effectively.
Bias correction: By adjusting for the initial bias when the first and second
moment estimates are close to zero, Adam helps prevent early-stage
instability.
Efficient performance: Adam typically requires fewer hyperparameter tuning
adjustments compared to other optimization algorithms like SGD, making it a
more convenient choice for most problems.
Avantage Explication
✅ Adaptatif Chaque poids a son propre taux d’apprentissage.
✅ Stable Lissage des gradients évite les oscillations.
✅ Peu de Fonctionne bien sans tuning excessif.
réglages
Avantage Explication
✅ Rapide Converge souvent plus vite que SGD classique.
Performance of Adam
In comparison to other optimizers like SGD (Stochastic Gradient Descent) and
momentum-based SGD, Adam outperforms them significantly in terms of both
training time and convergence accuracy. Its ability to adjust the learning rate per
parameter, combined with the bias-correction mechanism, leads to faster
convergence and more stable optimization. This makes Adam especially useful in
complex models with large datasets, as it avoids slow convergence and
instability while reaching the global minimum.
In practice, Adam often achieves superior results with minimal tuning, making it a
go-to optimizer for deep learning tasks.
🎯 Visualisation mentale
Imagine un randonneur (le modèle) :
SGD : descend avec des pas constants, parfois instables.
Momentum : il court dans la descente, porté par l’élan.
RMSprop : il adapte la taille de ses pas à chaque pente.
Adam : il court, mais ajuste son élan et ses pas en temps réel.
🧪 Quand utiliser Adam ?
Adam est particulièrement utile :
Dans des réseaux profonds (CNN, RNN, Transformers)
Quand les gradients sont bruités ou instables
Si l’on souhaite une optimisation rapide sans tuning excessif
Il est donc le choix par défaut dans beaucoup de bibliothèques (TensorFlow,
PyTorch).
🔗 Références utiles
DigitalOcean - Adam Tutorial
GeeksforGeeks - RMSprop & AdaGrad