RMSprop and Adadelta
RMSprop is an optimizer that divides the learning rate by the square root of an exponentially weighted average of the squared gradients, so the learning rate adapts like Adagrad's without shrinking towards zero.
Last updated: 05 Oct, 2026 · TensorFlow 2 / Keras
Adagrad adds every squared gradient into αt. In a deep network the sum becomes huge, η′ becomes negligible and wt ≈ wt−1. RMSprop and Adadelta keep the adaptive learning rate but stop the denominator from growing without limit.
Replacing the sum with an exponentially weighted average
The video treats Adadelta and RMSprop together, since they work in a similar way. η′ keeps its shape, with a new term Sdw where αt was:
This is the exponentially weighted average from the momentum lesson, applied to the squared gradient. With β = 0.95 the board writes Sdwt = 0.95·Sdwt−1 + 0.05·(∂L/∂wt−1)². The new squared gradient only gets 5 percent of the say, so Sdw never bumps up by a very large value. It follows the recent size of the gradients instead of their total, and η′ decreases slowly and stays under control. (The video says "h dw" for this term; the board writes Sdw.)

Dropping the learning rate with Adadelta
Adadelta goes one step further, as the notes in the materials describe. It keeps the same EWA of g², and it also keeps an EWA of the squared updates, Dw. The ratio of the two replaces η, so the original Adadelta has no learning rate to tune:
So the two are similar but not the same: RMSprop still needs η, Adadelta does not.
Running RMSprop and Adadelta in NumPy
The RMSprop loop
One line differs from the Adagrad loop: += becomes the EWA.
s_dw = 0.0 # S_dw starts at 0
for t in range(1, steps + 1):
g = gradient(w)
s_dw = beta * s_dw + (1 - beta) * g ** 2 # EWA of the squared gradient
lr_t = lr / np.sqrt(s_dw + eps) # η′ = η / √(S_dw + ε)
w = w - lr_t * gThe Adadelta loop
s_dw, d_w = 0.0, 0.0 # EWA of g² and EWA of the squared updates
for t in range(1, steps + 1):
g = gradient(w)
s_dw = rho * s_dw + (1 - rho) * g ** 2
step = np.sqrt(d_w + eps) / np.sqrt(s_dw + eps) * g # no learning rate anywhere
d_w = rho * d_w + (1 - rho) * step ** 2
w = w - stepRacing to the far target
The Adagrad lesson's failure case again: |w − 50| from w = 0, gradient −1 until the target. Adagrad and RMSprop both use η = 0.1 and RMSprop uses the board's β = 0.95; Adadelta uses ρ = 0.95 and ε = 1e-6, and no η.
import numpy as np
import matplotlib.pyplot as plt
def gradient(w):
return np.sign(w - 50.0) # the slope of |w − 50|
def run(kind, steps=10_000, lr=0.1, beta=0.95):
w, acc, d_w, reached, rates = 0.0, 0.0, 0.0, None, []
for t in range(1, steps + 1):
g = gradient(w)
if kind == "Adadelta":
acc = beta * acc + (1 - beta) * g ** 2
step = np.sqrt(d_w + 1e-6) / np.sqrt(acc + 1e-6) * g
d_w = beta * d_w + (1 - beta) * step ** 2
else:
acc = acc + g ** 2 if kind == "Adagrad" else beta * acc + (1 - beta) * g ** 2
step = lr / np.sqrt(acc + 1e-8) * g
w -= step
rates.append(abs(step / g) if g else np.nan) # the learning rate in force
if reached is None and w >= 49.5:
reached = t
return w, reached, rates
for kind in ("Adagrad", "RMSprop", "Adadelta"):
w, reached, rates = run(kind)
print(f"{kind:8}: η′ at steps 1 / 100 / 10,000 = {rates[0]:.4f} / {rates[99]:.4f} / {rates[-1]:.4f}"
f" w = {w:.2f} reached 50 at step {reached}")
plt.plot(rates, label=kind)
plt.yscale("log"); plt.xscale("log"); plt.xlabel("step"); plt.ylabel("learning rate in force")
plt.title("Adagrad keeps shrinking, RMSprop levels off, Adadelta grows"); plt.legend(); plt.show()Adagrad : η′ at steps 1 / 100 / 10,000 = 0.1000 / 0.0100 / 0.0010 w = 19.85 reached 50 at step None RMSprop : η′ at steps 1 / 100 / 10,000 = 0.4472 / 0.1003 / 0.1000 w = 49.91 reached 50 at step 474 Adadelta: η′ at steps 1 / 100 / 10,000 = 0.0045 / 0.0052 / 0.0229 w = 49.98 reached 50 at step 4438

What the three runs show
- Adagrad stalls near w = 20 with η′ falling to 0.001, as in the previous lesson.
- RMSprop reaches 50 in under 500 steps. Sdw starts at 0, so its first η′ is large (0.1/√0.05 ≈ 0.45); then Sdw settles at 1, the size of the recent squared gradients, and η′ settles at 0.1.
- Adadelta gets there with no learning rate. Its steps start tiny, set by √ε, and grow as its record of past updates grows, so it is slower here, but it arrives.
- Near 50 the gradient flips sign every time w crosses the target, so RMSprop keeps stepping across it by about 0.1: the gradient itself is still the raw, unsmoothed one.
That last point is what the video says comes next: RMSprop has the adaptive learning rate but has lost the smoothing of SGD with momentum. Combining both is the Adam optimizer.
RMSprop and Adadelta in Keras
Keras calls the board's β rho:
import keras
opt = keras.optimizers.RMSprop(learning_rate=0.001, rho=0.9)
opt = keras.optimizers.Adadelta(learning_rate=1.0, rho=0.95) # 1.0 matches the original paperKeras still gives Adadelta a learning_rate that multiplies the step; with 1.0 the update is the paper's, which has no learning rate. Its default is 0.001, which makes every step a thousand times smaller than the paper's, so set it yourself. RMSprop's defaults are the ones in the code: 0.001 and rho=0.9, with ε inside the square root as on the board.
Adagrad vs RMSprop vs Adadelta
| Adagrad | RMSprop | Adadelta | |
|---|---|---|---|
| Denominator | αt: sum of all g² | Sdw: EWA of g² | Sdw: EWA of g² |
| Grows without limit | yes | no | no |
| Learning rate η | needed | needed | not needed (EWA of past updates instead) |
| Long training | stalls | keeps learning | keeps learning |
| Smoothing of the gradient | no | no | no |
Where you use RMSprop
- Recurrent networks, where RMSprop was a common choice before Adam.
- Reinforcement learning, where the gradients change size as the agent improves; DeepMind's DQN used RMSprop.
- Any model where Adagrad stalls: the change is one line.
Related
- Previous: Adagrad
- Next: Adam optimizer
- Change
beta=0.95to 0.5 inrun: RMSprop's first η′ drops to about 0.14 and it still reaches 50, at step 495. - Change it to 0.999: Sdw builds up so slowly that RMSprop's η′ starts at 3.16 and is still 0.32 at step 100, so it reaches 50 by step 72.
- Change Adadelta's
1e-6to1e-4in both places: its first steps are ten times bigger and it reaches 50 much sooner.
This is what real progress feels like.