Deep LearningTensorFlow 2.21 / Keras 3 · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
41 small wins to finish your pathNext lesson →

RMSprop and Adadelta

RMSprop is an optimizer that divides the learning rate by the square root of an exponentially weighted average of the squared gradients, so the learning rate adapts like Adagrad's without shrinking towards zero.

Last updated: 05 Oct, 2026 · TensorFlow 2 / Keras

Adagrad adds every squared gradient into αt. In a deep network the sum becomes huge, η′ becomes negligible and wt ≈ wt−1. RMSprop and Adadelta keep the adaptive learning rate but stop the denominator from growing without limit.

From Adagrad's problem to RMSprop · from the Deep Learning In-depth Tutorials in 5 Hours video · 195:38 to 199:47

Replacing the sum with an exponentially weighted average

The video treats Adadelta and RMSprop together, since they work in a similar way. η′ keeps its shape, with a new term Sdw where αt was:

This is the exponentially weighted average from the momentum lesson, applied to the squared gradient. With β = 0.95 the board writes Sdwt = 0.95·Sdwt−1 + 0.05·(∂L/∂wt−1)². The new squared gradient only gets 5 percent of the say, so Sdw never bumps up by a very large value. It follows the recent size of the gradients instead of their total, and η′ decreases slowly and stays under control. (The video says "h dw" for this term; the board writes Sdw.)

Over the training steps with a gradient of 1 at every step: Adagrad's alpha t is a straight line that keeps rising, while RMSprop's S dw rises quickly and levels off at 1; below, Adagrad's eta prime keeps falling while RMSprop's eta prime levels off.

Dropping the learning rate with Adadelta

Adadelta goes one step further, as the notes in the materials describe. It keeps the same EWA of g², and it also keeps an EWA of the squared updates, Dw. The ratio of the two replaces η, so the original Adadelta has no learning rate to tune:

So the two are similar but not the same: RMSprop still needs η, Adadelta does not.

Running RMSprop and Adadelta in NumPy

The RMSprop loop

One line differs from the Adagrad loop: += becomes the EWA.

python
s_dw = 0.0                                    # S_dw starts at 0
for t in range(1, steps + 1):
    g = gradient(w)
    s_dw = beta * s_dw + (1 - beta) * g ** 2  # EWA of the squared gradient
    lr_t = lr / np.sqrt(s_dw + eps)           # η′ = η / √(S_dw + ε)
    w = w - lr_t * g

The Adadelta loop

python
s_dw, d_w = 0.0, 0.0                          # EWA of g² and EWA of the squared updates
for t in range(1, steps + 1):
    g = gradient(w)
    s_dw = rho * s_dw + (1 - rho) * g ** 2
    step = np.sqrt(d_w + eps) / np.sqrt(s_dw + eps) * g   # no learning rate anywhere
    d_w = rho * d_w + (1 - rho) * step ** 2
    w = w - step

Racing to the far target

The Adagrad lesson's failure case again: |w − 50| from w = 0, gradient −1 until the target. Adagrad and RMSprop both use η = 0.1 and RMSprop uses the board's β = 0.95; Adadelta uses ρ = 0.95 and ε = 1e-6, and no η.

ExampleAdagrad, RMSprop and Adadelta on the far target, run with NumPy
import numpy as np
import matplotlib.pyplot as plt

def gradient(w):
    return np.sign(w - 50.0)                  # the slope of |w − 50|

def run(kind, steps=10_000, lr=0.1, beta=0.95):
    w, acc, d_w, reached, rates = 0.0, 0.0, 0.0, None, []
    for t in range(1, steps + 1):
        g = gradient(w)
        if kind == "Adadelta":
            acc = beta * acc + (1 - beta) * g ** 2
            step = np.sqrt(d_w + 1e-6) / np.sqrt(acc + 1e-6) * g
            d_w = beta * d_w + (1 - beta) * step ** 2
        else:
            acc = acc + g ** 2 if kind == "Adagrad" else beta * acc + (1 - beta) * g ** 2
            step = lr / np.sqrt(acc + 1e-8) * g
        w -= step
        rates.append(abs(step / g) if g else np.nan)      # the learning rate in force
        if reached is None and w >= 49.5:
            reached = t
    return w, reached, rates

for kind in ("Adagrad", "RMSprop", "Adadelta"):
    w, reached, rates = run(kind)
    print(f"{kind:8}: η′ at steps 1 / 100 / 10,000 = {rates[0]:.4f} / {rates[99]:.4f} / {rates[-1]:.4f}"
          f"   w = {w:.2f}   reached 50 at step {reached}")
    plt.plot(rates, label=kind)
plt.yscale("log"); plt.xscale("log"); plt.xlabel("step"); plt.ylabel("learning rate in force")
plt.title("Adagrad keeps shrinking, RMSprop levels off, Adadelta grows"); plt.legend(); plt.show()
The learning rate in force over 10,000 steps on log axes: Adagrad's falls in a straight line the whole time, RMSprop's starts high and levels off near 0.1, and Adadelta's starts tiny and keeps growing.

What the three runs show

  • Adagrad stalls near w = 20 with η′ falling to 0.001, as in the previous lesson.
  • RMSprop reaches 50 in under 500 steps. Sdw starts at 0, so its first η′ is large (0.1/√0.05 ≈ 0.45); then Sdw settles at 1, the size of the recent squared gradients, and η′ settles at 0.1.
  • Adadelta gets there with no learning rate. Its steps start tiny, set by √ε, and grow as its record of past updates grows, so it is slower here, but it arrives.
  • Near 50 the gradient flips sign every time w crosses the target, so RMSprop keeps stepping across it by about 0.1: the gradient itself is still the raw, unsmoothed one.

That last point is what the video says comes next: RMSprop has the adaptive learning rate but has lost the smoothing of SGD with momentum. Combining both is the Adam optimizer.

RMSprop and Adadelta in Keras

Keras calls the board's β rho:

python
import keras

opt = keras.optimizers.RMSprop(learning_rate=0.001, rho=0.9)
opt = keras.optimizers.Adadelta(learning_rate=1.0, rho=0.95)   # 1.0 matches the original paper

Keras still gives Adadelta a learning_rate that multiplies the step; with 1.0 the update is the paper's, which has no learning rate. Its default is 0.001, which makes every step a thousand times smaller than the paper's, so set it yourself. RMSprop's defaults are the ones in the code: 0.001 and rho=0.9, with ε inside the square root as on the board.

Adagrad vs RMSprop vs Adadelta

AdagradRMSpropAdadelta
Denominatorαt: sum of all g²Sdw: EWA of g²Sdw: EWA of g²
Grows without limityesnono
Learning rate ηneededneedednot needed (EWA of past updates instead)
Long trainingstallskeeps learningkeeps learning
Smoothing of the gradientnonono

Where you use RMSprop

  • Recurrent networks, where RMSprop was a common choice before Adam.
  • Reinforcement learning, where the gradients change size as the agent improves; DeepMind's DQN used RMSprop.
  • Any model where Adagrad stalls: the change is one line.
Watch out. Sdw starts at 0, so the first few η′ values are large: here the first step was 0.45 instead of 0.1. On a real network that can push the weights far in the first iterations. Adam fixes this with a bias correction.
Try it yourself
  • Change beta=0.95 to 0.5 in run: RMSprop's first η′ drops to about 0.14 and it still reaches 50, at step 495.
  • Change it to 0.999: Sdw builds up so slowly that RMSprop's η′ starts at 3.16 and is still 0.32 at step 100, so it reaches 50 by step 72.
  • Change Adadelta's 1e-6 to 1e-4 in both places: its first steps are ten times bigger and it reaches 50 much sooner.
PreviousAdagrad

This is what real progress feels like.