Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your path

Fine-tuning transformers with Hugging Face

Fine-tuning a transformer is continuing the training of a pretrained model on a smaller labelled dataset for one task, so the model keeps what it learned in pretraining and only has to learn the task on top.

Last updated: 07 Oct, 2026 · Transformers 5.19 · Datasets 5.1 · NumPy

Sentiment analysis of Kindle reviews built a classifier from TF-IDF and Word2Vec features. A pretrained encoder from BERT, GPT and T5 already knows English; fine-tuning teaches it the labels. The example fine-tunes DistilBERT, a 6-layer distilled BERT (hidden size 768, about 66 million parameters), on the IMDb movie reviews, following the text-classification guide in the Transformers documentation.

The code below is shown, not run here: it downloads the IMDb dataset and the DistilBERT checkpoint, a few hundred megabytes, and training needs PyTorch and ideally a GPU (a free Colab GPU works). What each step returns is described in words under it; the numbers you will get depend on your run.

Installing transformers and datasets

bash
pip install transformers datasets accelerate torch

transformers holds the models, tokenizers and the Trainer; datasets loads and maps datasets; accelerate is what the Trainer uses to place work on the GPU; torch is the deep-learning library underneath.

Loading the IMDb reviews

python
from datasets import load_dataset

imdb = load_dataset("stanfordnlp/imdb")    # DatasetDict: train, test, unsupervised
imdb["train"][0]                           # {"text": "...", "label": 0}   0 = negative, 1 = positive

load_dataset returns a DatasetDict with a train and a test split of 25,000 labelled reviews each, plus an unlabelled unsupervised split. Each row has two fields, text and label (0 for negative, 1 for positive).

Tokenizing with the checkpoint's own tokenizer

python
from transformers import AutoTokenizer

checkpoint = "distilbert/distilbert-base-uncased"
tokenizer = AutoTokenizer.from_pretrained(checkpoint)

def preprocess(batch):
    return tokenizer(batch["text"], truncation=True)   # cut reviews at the model's 512 tokens

tokenized = imdb.map(preprocess, batched=True)        # adds input_ids and attention_mask columns

AutoTokenizer loads the WordPiece tokenizer DistilBERT was trained with (see Subword tokenization (BPE and WordPiece)). map(..., batched=True) runs preprocess on batches of rows and returns a new DatasetDict with two extra columns per review: input_ids, the token indices with [CLS] and [SEP] added, and attention_mask, 1 for every real token. truncation=True cuts reviews longer than the model's 512 positions.

Padding each batch

python
from transformers import DataCollatorWithPadding

data_collator = DataCollatorWithPadding(tokenizer=tokenizer)   # pads each batch to its longest review

Reviews have different lengths. The collator pads each batch only to the longest review in that batch and sets the attention mask to 0 on the padding: the padding mask from masked self-attention, built for you.

Loading the model with a new classification head

python
from transformers import AutoModelForSequenceClassification

id2label = {0: "NEGATIVE", 1: "POSITIVE"}
label2id = {"NEGATIVE": 0, "POSITIVE": 1}
model = AutoModelForSequenceClassification.from_pretrained(
    checkpoint, num_labels=2, id2label=id2label, label2id=label2id)

This loads DistilBERT's pretrained encoder and puts a new, randomly initialised classification head on top with two outputs. The library prints a warning that some weights (the head's pre_classifier and classifier layers) were newly initialised and that the model should be trained: that is expected, and it is what fine-tuning fixes. Calling the model returns an output object whose logits has one row per review and two columns.

Measuring accuracy

python
import numpy as np

def compute_metrics(eval_pred):
    logits, labels = eval_pred
    predictions = np.argmax(logits, axis=1)
    return {"accuracy": round(float((predictions == labels).mean()), 4)}

The Trainer calls this after each evaluation with the logits and the true labels of the test split. The documentation's version uses the evaluate library's accuracy metric; the NumPy line computes the same number.

Setting up TrainingArguments and the Trainer

python
from transformers import TrainingArguments, Trainer

args = TrainingArguments(
    output_dir="imdb-distilbert", learning_rate=2e-5,
    per_device_train_batch_size=16, per_device_eval_batch_size=16,
    num_train_epochs=2, weight_decay=0.01,
    eval_strategy="epoch", save_strategy="epoch", load_best_model_at_end=True)

trainer = Trainer(model=model, args=args,
                  train_dataset=tokenized["train"], eval_dataset=tokenized["test"],
                  processing_class=tokenizer, data_collator=data_collator,
                  compute_metrics=compute_metrics)

TrainingArguments holds the run's settings: a small learning rate (2e-5, so the pretrained weights move gently), batch size 16, two epochs, and evaluation and saving at the end of each epoch. The tokenizer is passed as processing_class, so it is saved with the model.

Training, evaluating and saving

python
result = trainer.train()        # TrainOutput: global_step, training_loss, metrics
scores = trainer.evaluate()     # dict: eval_loss, eval_accuracy, eval_runtime, ...
trainer.save_model("imdb-distilbert")

train() runs the epochs with the AdamW optimizer, logs the training loss as it goes, and returns a TrainOutput holding the number of steps, the average training loss and timing metrics. evaluate() returns a dictionary with eval_loss, eval_accuracy (from compute_metrics, with the eval_ prefix added) and runtime figures. save_model writes the weights, the configuration and the tokenizer to the folder.

Predicting with a pipeline

python
from transformers import pipeline

classifier = pipeline("text-classification", model="imdb-distilbert")
classifier("This was a masterpiece from beginning to end.")   # [{"label": ..., "score": ...}]

The pipeline loads the saved folder and returns a list with one dictionary per input, holding the predicted label (“POSITIVE” or “NEGATIVE”, from id2label) and its softmax score.

A labelled dataset goes through the checkpoint's tokenizer to input ids and an attention mask, a pretrained body gets a new head, the Trainer runs the loss and AdamW for a few epochs, and the fine-tuned model labels new text through a text-classification pipeline.

Checking the metric function and the training budget

Two parts of the setup can be checked without a download: the metric function, on a stand-in batch of six reviews' logits, and the arithmetic of the run.

ExampleRun with NumPy (stand-in logits)
import math
import numpy as np

# a stand-in evaluation batch: 6 reviews, 2 logits each (NEGATIVE, POSITIVE)
logits = np.array([[2.1, -1.3], [-0.4, 1.8], [0.3, 0.1], [-2.2, 2.5], [1.1, 0.9], [-0.7, -0.2]])
labels = np.array([0, 1, 1, 1, 0, 0])
print(compute_metrics((logits, labels)))          # the function the Trainer will call

probs = np.exp(logits) / np.exp(logits).sum(axis=1, keepdims=True)
print("P(POSITIVE) per review:", np.round(probs[:, 1], 3))

# what the Trainer will do on IMDb: 25,000 training reviews, batch 16, 2 epochs, one device
steps_per_epoch = math.ceil(25_000 / 16)
print("steps per epoch:", steps_per_epoch, " total steps:", 2 * steps_per_epoch)

# new weights the classification head adds on top of DistilBERT (hidden size 768, 2 labels)
head = (768 * 768 + 768) + (768 * 2 + 2)          # pre_classifier + classifier
print("new head parameters:", f"{head:,}")

What the check shows

  • Accuracy 0.6667: argmax of the six logit rows gives [0, 1, 0, 1, 0, 1] against the labels [0, 1, 1, 1, 0, 0], so four of six are right.
  • P(POSITIVE) is the softmax of each row: 0.9 for [−0.4, 1.8], 0.45 for [0.3, 0.1]; the pipeline's score is the probability of the winning label, so [0.3, 0.1] would be reported as NEGATIVE with 0.55.
  • 1,563 steps per epoch, 3,126 in total: 25,000 reviews in batches of 16 (the last batch is partly full), for two epochs on one device.
  • The new head has 592,130 weights, against about 66 million pretrained ones: fine-tuning updates all of them, but only the head starts from random values.

Fine-tuning vs feature extraction vs training from scratch

Fine-tuningFeature extractionTraining from scratchTF-IDF + logistic regression
Starts frompretrained weightspretrained weights, frozenrandom weightsword counts
Trainsall layers + new headonly the headeverythingone linear model
Data neededthousands of labelshundreds to thousandsmillions of examplesthousands
Hardwarea GPU helpsa CPU can do itmany GPUsa CPU
Knows word order and contextyesyesonly what the data teachesno

Where you use fine-tuning

  • Text classification: sentiment, spam, topic or intent labels on your own data.
  • Token classification: named entity recognition with AutoModelForTokenClassification.
  • Domain adaptation: legal, medical or support text, where a general model misses the vocabulary.
Watch out. Older tutorials pass evaluation_strategy= to TrainingArguments and tokenizer= to the Trainer. Current releases use eval_strategy= and processing_class=; the old names fail. And always load the tokenizer from the same checkpoint as the model.
Try it yourself
  • Change the labels in the stand-in batch so that all six predictions are right and check the accuracy.
  • Recompute the steps for per_device_train_batch_size=32 and three epochs.
  • Change num_labels to 3 and recompute the head's weight count (the classifier becomes 768 × 3 + 3).
PreviousBERT, GPT and T5

Little by little, you're building something great.