Deep LearningTensorFlow 2.21 / Keras 3 · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
41 small wins to finish your pathNext lesson →

Transfer learning with VGG16

Transfer learning is a technique that reuses a network already trained on a large dataset, here VGG16 trained on ImageNet, as a fixed feature extractor and trains only a small new output layer for your own classes.

Last updated: 05 Oct, 2026 · TensorFlow 2 / Keras

The CIFAR-10 model in CNN in Keras learned all its filters from 50,000 images. With a few thousand images that is not enough, but the filters a big network learned on ImageNet's 1.2 million photos (edges, textures, parts) suit most photos. The video says that the number of layers is learned from transfer learning and the ImageNet competition, and promises transfer learning for a later session. The Transfer Learning VGG16 notebook in the Advanced-CNN-Architectures materials repo does it on cotton plant disease photos with 4 classes, run on TensorFlow 2.2.0 in September 2020. Its cells and saved outputs follow.

Loading VGG16 without its top

VGG16 is a CNN of 13 convolution layers, all 3×3 with same padding, in five blocks separated by 2×2 max pooling, followed by three dense layers. include_top=False drops the three dense layers, which were built for ImageNet's 1,000 classes, and keeps the convolutional base. weights='imagenet' downloads the trained filters (58,889,256 bytes). Every image is resized to 224×224, and IMAGE_SIZE + [3] adds the 3 RGB channels.

The notebook first prints its TensorFlow version and then imports the layers, the Model class, VGG16 and glob. Some imports, such as VGG19, Lambda and Sequential, are never used.

ExampleOutput from the Transfer Learning VGG16 notebook (TensorFlow 2.2)
import tensorflow as tf
print(tf.__version__)
python
# import the libraries as shown below

from tensorflow.keras.layers import Input, Lambda, Dense, Flatten
from tensorflow.keras.models import Model
from tensorflow.keras.applications.vgg16 import VGG16
from tensorflow.keras.applications.vgg19 import VGG19
from tensorflow.keras.preprocessing import image
from tensorflow.keras.preprocessing.image import ImageDataGenerator,load_img
from tensorflow.keras.models import Sequential
import numpy as np
from glob import glob
#import matplotlib.pyplot as plt
python
# re-size all the images to this
IMAGE_SIZE = [224, 224]

train_path = 'Datasets/train'
valid_path = 'Datasets/test'
ExampleOutput from the Transfer Learning VGG16 notebook (TensorFlow 2.2)
# Import the VGG16 library as shown below and add preprocessing layer to the front of VGG
# Here we will be using imagenet weights

vgg16 = VGG16(input_shape=IMAGE_SIZE + [3], weights='imagenet', include_top=False)

Freezing the convolutional base

Setting trainable = False on every layer of the base means backpropagation leaves the ImageNet filters as they are; only the layers added next will learn.

python
# don't train existing weights
for layer in vgg16.layers:
    layer.trainable = False

Adding a new classification head

The number of classes is read from the training folders: one folder per class.

python
  # useful for getting number of output classes
folders = glob('Datasets/train/*')
ExampleOutput from the Transfer Learning VGG16 notebook (TensorFlow 2.2)
folders

The head is a Flatten on the base's output and a Dense layer with one softmax unit per folder. Keras' functional API joins them: Model(inputs=vgg16.input, outputs=prediction).

python
# our layers - you can add more if you want
x = Flatten()(vgg16.output)

prediction = Dense(len(folders), activation='softmax')(x)

# create a model object
model = Model(inputs=vgg16.input, outputs=prediction)
ExampleOutput from the Transfer Learning VGG16 notebook (TensorFlow 2.2)
# view the structure of the model
model.summary()
A frozen VGG16 base with ImageNet weights, five convolution blocks that shrink the image from 224 by 224 to 7 by 7 with 512 channels and 14,714,688 parameters, feeds a new trainable head: a flatten to 25,088 values and a 4-class softmax Dense layer with 100,356 parameters.

The summary shows the ideas of this part at full scale. Each 3×3 convolution with same padding keeps the size (224 + 2 − 3 + 1 = 224), and each 2×2 pool halves it: 224 → 112 → 56 → 28 → 14 → 7. block1_conv1 has 3·3·3·64 + 64 = 1,792 parameters and block1_conv2 3·3·64·64 + 64 = 36,928. The flattened 7×7×512 = 25,088 values feed Dense(4): 25,088·4 + 4 = 100,356 trainable parameters, against 14,714,688 frozen ones.

Augmenting the training images

ImageDataGenerator rescales the pixels to 0-1 and, for the training images only, applies random changes each time an image is used: a shear of up to 0.2 degrees (shear_range is an angle), a zoom of up to 20% and a horizontal flip. The network sees a slightly different leaf every epoch, which works like more data and reduces overfitting. The test images are only rescaled.

python
# Use the Image Data Generator to import the images from the dataset
from tensorflow.keras.preprocessing.image import ImageDataGenerator

train_datagen = ImageDataGenerator(rescale = 1./255,
                                   shear_range = 0.2,
                                   zoom_range = 0.2,
                                   horizontal_flip = True)

test_datagen = ImageDataGenerator(rescale = 1./255)
ExampleOutput from the Transfer Learning VGG16 notebook (TensorFlow 2.2)
# Make sure you provide the same target size as initialied for the image size
training_set = train_datagen.flow_from_directory('Datasets/train',
                                                 target_size = (224, 224),
                                                 batch_size = 32,
                                                 class_mode = 'categorical')
ExampleOutput from the Transfer Learning VGG16 notebook (TensorFlow 2.2)
test_set = test_datagen.flow_from_directory('Datasets/test',
                                            target_size = (224, 224),
                                            batch_size = 32,
                                            class_mode = 'categorical')

Training the new head

Compiling uses categorical cross-entropy, because class_mode='categorical' gives one-hot labels (Categorical cross-entropy), and Adam. fit_generator trains for 20 epochs of 61 steps: 1,951 images in batches of 32.

python
# tell the model what cost and optimization method to use
model.compile(
  loss='categorical_crossentropy',
  optimizer='adam',
  metrics=['accuracy']
)
ExampleOutput from the Transfer Learning VGG16 notebook (TensorFlow 2.2), epochs 4 to 14 cut
# fit the model
# Run the cell. It will take some time to execute
r = model.fit_generator(
  training_set,
  validation_data=test_set,
  epochs=20,
  steps_per_epoch=len(training_set),
  validation_steps=len(test_set)
)

Training accuracy climbs from 0.38 to about 0.72 to 0.77. Validation accuracy jumps between 0.50 and 0.83 from one epoch to the next because the test folder holds only 18 images: one image more or less is 5.6 percentage points. The last epoch ends at 0.6914 training and 0.7778 validation accuracy. Its first line is TensorFlow 2.2's warning that fit_generator is deprecated in favour of fit.

Reading the notebook's stale cells

The notebook's later cells save the model and predict the 18 test images (the argmax gives one class index per image, such as 3, 2, 0). After that, several cells were copied from another project and do not belong to this model:

  • Loading the wrong file: a cell loads model_resnet50.h5, while the notebook saved model_vgg16.h5.
  • A variable shown before it exists: img_data is printed before the cell that creates it.
  • Another dataset's path: the test image is read from Datasets/Test/Coffee/, not from the cotton folders.
  • Preprocessing twice: preprocess_input is never imported and is applied after x/255; VGG16's own preprocessing expects pixels from 0 to 255.
  • Two outputs from a 4-class model: the prediction prints [[0.9745471, 0.0254529]], which only a 2-class model gives.
  • Blank plot files: plt.savefig runs after plt.show(), which has already cleared the figure.

Updating the notebook for Keras 3

TensorFlow 2.16 and later ship Keras 3, where several of the notebook's calls are gone or deprecated: fit_generator is removed (use fit), ImageDataGenerator is replaced by image_dataset_from_directory plus augmentation layers, the tf.compat.v1 session for GPU memory is not needed (tf.config.experimental.set_memory_growth does that), and models save as .keras rather than .h5.

Data and the frozen base

python
import keras
from keras import layers
from keras.applications import VGG16
from keras.applications.vgg16 import preprocess_input

train = keras.utils.image_dataset_from_directory("Datasets/train", image_size=(224, 224),
                                                 batch_size=32, label_mode="categorical")
test = keras.utils.image_dataset_from_directory("Datasets/test", image_size=(224, 224),
                                                batch_size=32, label_mode="categorical")
base = VGG16(weights="imagenet", include_top=False, input_shape=(224, 224, 3))
base.trainable = False                                   # freeze the whole base

Augmentation layers and the head

python
inputs = keras.Input(shape=(224, 224, 3))
x = layers.RandomFlip("horizontal")(inputs)             # horizontal_flip=True
x = layers.RandomZoom(0.2)(x)                             # zoom_range=0.2
x = layers.RandomShear(x_factor=0.2)(x)                   # shears up to 20% of the width
x = base(preprocess_input(x), training=False)            # VGG16's own scaling of 0-255 pixels
outputs = layers.Dense(4, activation="softmax")(layers.Flatten()(x))
model = keras.Model(inputs, outputs)
model.compile(optimizer="adam", loss="categorical_crossentropy", metrics=["accuracy"])
model.fit(train, validation_data=test, epochs=20)
model.save("model_vgg16.keras")

The augmentation layers are active only during training, and preprocess_input replaces the 1/255 rescale with the scaling VGG16 was trained on. RandomShear measures the shear as a fraction of the image width, not as an angle, so 0.2 here shears far more than the notebook's 0.2 degrees; lower it for a gentler shear. This version is shown, not run: it needs TensorFlow and the cotton dataset.

Counting VGG16's parameters

The summary's numbers come from the formulas of this part. The program also counts the ImageNet top that include_top=False dropped, and shows the horizontal flip on a tiny array.

ExampleRun with NumPy
import numpy as np

blocks = [(64, 2), (128, 2), (256, 3), (512, 3), (512, 3)]    # (filters, conv layers) per VGG16 block
size, c_in, base = 224, 3, 0
for b, (k, convs) in enumerate(blocks, 1):
    for _ in range(convs):
        base += 3 * 3 * c_in * k + k                            # f*f*C*k + k; 'same' padding keeps the size
        c_in = k
    size //= 2                                                  # 2x2 max pool, stride 2
    print(f"block {b}: {size}x{size}x{k}")
flat = size * size * c_in
head = flat * 4 + 4                                             # Dense(4) on the flattened output
print("frozen base:", base, " flatten:", flat, " new head:", head, " total:", base + head)
top = (flat * 4096 + 4096) + (4096 * 4096 + 4096) + (4096 * 1000 + 1000)
print("the dropped ImageNet top:", top, " full VGG16:", base + top)

leaf = np.arange(9).reshape(3, 3)                               # a tiny 3x3 'image'
print("horizontal flip:")
print(np.fliplr(leaf))

What the counts show

  • The five blocks end at 112, 56, 28, 14 and 7 pixels, with 64 up to 512 channels, as in the saved summary.
  • The frozen base has 14,714,688 parameters and the new head 100,356; the total, 14,815,044, matches the summary's Total params.
  • The dropped ImageNet top has 123,642,856 parameters, so the full VGG16 has 138,357,544: almost 90% of it sits in the three dense layers.
  • A horizontal flip reverses each row, 0 1 2 becoming 2 1 0, which is what horizontal_flip=True does to a photo.

Transfer learning vs training from scratch

From scratch (CIFAR-10 CNN)Transfer learning (VGG16)
Filterslearned from random valuestaken from ImageNet, frozen
Trainable parameters122,570, all of them100,356 of 14,815,044
Training images50,0001,951
Input size32×32×3224×224×3
Training costevery layer, every steponly the head; the base runs forward only

Where you use transfer learning

  • Small image datasets such as plant disease, medical scans or factory defects, where a few thousand labelled photos are all there is.
  • Fast baselines: a frozen base and a new head train in minutes and often beat a model trained from scratch.
  • Fine-tuning: after the head has trained, unfreeze the last block with a small learning rate to adapt the deepest filters to the new images.
Watch out. Preprocess every image the same way in training and in prediction. The notebook trains on pixels / 255 and later predicts on pixels / 255 passed through preprocess_input as well; the model then sees inputs unlike anything it trained on, and its predictions are wrong with no error message.
Try it yourself
  • Change the 4 in head = flat * 4 + 4 to the number of classes of your own dataset and read the new head size.
  • Replace the last pool's 7×7 flatten with global average pooling (512 values) and recompute the head: 512·4 + 4.
  • Flip leaf vertically with np.flipud and think about whether an upside-down cotton leaf is still a fair training image.
PreviousCNN in Keras

This is what real progress feels like.