Fine-tuning transformers with Hugging Face
Fine-tuning a transformer is continuing the training of a pretrained model on a smaller labelled dataset for one task, so the model keeps what it learned in pretraining and only has to learn the task on top.
Last updated: 07 Oct, 2026 · Transformers 5.19 · Datasets 5.1 · NumPy
Sentiment analysis of Kindle reviews built a classifier from TF-IDF and Word2Vec features. A pretrained encoder from BERT, GPT and T5 already knows English; fine-tuning teaches it the labels. The example fine-tunes DistilBERT, a 6-layer distilled BERT (hidden size 768, about 66 million parameters), on the IMDb movie reviews, following the text-classification guide in the Transformers documentation.
The code below is shown, not run here: it downloads the IMDb dataset and the DistilBERT checkpoint, a few hundred megabytes, and training needs PyTorch and ideally a GPU (a free Colab GPU works). What each step returns is described in words under it; the numbers you will get depend on your run.
Installing transformers and datasets
pip install transformers datasets accelerate torchtransformers holds the models, tokenizers and the Trainer; datasets loads and maps datasets; accelerate is what the Trainer uses to place work on the GPU; torch is the deep-learning library underneath.
Loading the IMDb reviews
from datasets import load_dataset
imdb = load_dataset("stanfordnlp/imdb") # DatasetDict: train, test, unsupervised
imdb["train"][0] # {"text": "...", "label": 0} 0 = negative, 1 = positiveload_dataset returns a DatasetDict with a train and a test split of 25,000 labelled reviews each, plus an unlabelled unsupervised split. Each row has two fields, text and label (0 for negative, 1 for positive).
Tokenizing with the checkpoint's own tokenizer
from transformers import AutoTokenizer
checkpoint = "distilbert/distilbert-base-uncased"
tokenizer = AutoTokenizer.from_pretrained(checkpoint)
def preprocess(batch):
return tokenizer(batch["text"], truncation=True) # cut reviews at the model's 512 tokens
tokenized = imdb.map(preprocess, batched=True) # adds input_ids and attention_mask columnsAutoTokenizer loads the WordPiece tokenizer DistilBERT was trained with (see Subword tokenization (BPE and WordPiece)). map(..., batched=True) runs preprocess on batches of rows and returns a new DatasetDict with two extra columns per review: input_ids, the token indices with [CLS] and [SEP] added, and attention_mask, 1 for every real token. truncation=True cuts reviews longer than the model's 512 positions.
Padding each batch
from transformers import DataCollatorWithPadding
data_collator = DataCollatorWithPadding(tokenizer=tokenizer) # pads each batch to its longest reviewReviews have different lengths. The collator pads each batch only to the longest review in that batch and sets the attention mask to 0 on the padding: the padding mask from masked self-attention, built for you.
Loading the model with a new classification head
from transformers import AutoModelForSequenceClassification
id2label = {0: "NEGATIVE", 1: "POSITIVE"}
label2id = {"NEGATIVE": 0, "POSITIVE": 1}
model = AutoModelForSequenceClassification.from_pretrained(
checkpoint, num_labels=2, id2label=id2label, label2id=label2id)This loads DistilBERT's pretrained encoder and puts a new, randomly initialised classification head on top with two outputs. The library prints a warning that some weights (the head's pre_classifier and classifier layers) were newly initialised and that the model should be trained: that is expected, and it is what fine-tuning fixes. Calling the model returns an output object whose logits has one row per review and two columns.
Measuring accuracy
import numpy as np
def compute_metrics(eval_pred):
logits, labels = eval_pred
predictions = np.argmax(logits, axis=1)
return {"accuracy": round(float((predictions == labels).mean()), 4)}The Trainer calls this after each evaluation with the logits and the true labels of the test split. The documentation's version uses the evaluate library's accuracy metric; the NumPy line computes the same number.
Setting up TrainingArguments and the Trainer
from transformers import TrainingArguments, Trainer
args = TrainingArguments(
output_dir="imdb-distilbert", learning_rate=2e-5,
per_device_train_batch_size=16, per_device_eval_batch_size=16,
num_train_epochs=2, weight_decay=0.01,
eval_strategy="epoch", save_strategy="epoch", load_best_model_at_end=True)
trainer = Trainer(model=model, args=args,
train_dataset=tokenized["train"], eval_dataset=tokenized["test"],
processing_class=tokenizer, data_collator=data_collator,
compute_metrics=compute_metrics)TrainingArguments holds the run's settings: a small learning rate (2e-5, so the pretrained weights move gently), batch size 16, two epochs, and evaluation and saving at the end of each epoch. The tokenizer is passed as processing_class, so it is saved with the model.
Training, evaluating and saving
result = trainer.train() # TrainOutput: global_step, training_loss, metrics
scores = trainer.evaluate() # dict: eval_loss, eval_accuracy, eval_runtime, ...
trainer.save_model("imdb-distilbert")train() runs the epochs with the AdamW optimizer, logs the training loss as it goes, and returns a TrainOutput holding the number of steps, the average training loss and timing metrics. evaluate() returns a dictionary with eval_loss, eval_accuracy (from compute_metrics, with the eval_ prefix added) and runtime figures. save_model writes the weights, the configuration and the tokenizer to the folder.
Predicting with a pipeline
from transformers import pipeline
classifier = pipeline("text-classification", model="imdb-distilbert")
classifier("This was a masterpiece from beginning to end.") # [{"label": ..., "score": ...}]The pipeline loads the saved folder and returns a list with one dictionary per input, holding the predicted label (“POSITIVE” or “NEGATIVE”, from id2label) and its softmax score.
Checking the metric function and the training budget
Two parts of the setup can be checked without a download: the metric function, on a stand-in batch of six reviews' logits, and the arithmetic of the run.
import math
import numpy as np
# a stand-in evaluation batch: 6 reviews, 2 logits each (NEGATIVE, POSITIVE)
logits = np.array([[2.1, -1.3], [-0.4, 1.8], [0.3, 0.1], [-2.2, 2.5], [1.1, 0.9], [-0.7, -0.2]])
labels = np.array([0, 1, 1, 1, 0, 0])
print(compute_metrics((logits, labels))) # the function the Trainer will call
probs = np.exp(logits) / np.exp(logits).sum(axis=1, keepdims=True)
print("P(POSITIVE) per review:", np.round(probs[:, 1], 3))
# what the Trainer will do on IMDb: 25,000 training reviews, batch 16, 2 epochs, one device
steps_per_epoch = math.ceil(25_000 / 16)
print("steps per epoch:", steps_per_epoch, " total steps:", 2 * steps_per_epoch)
# new weights the classification head adds on top of DistilBERT (hidden size 768, 2 labels)
head = (768 * 768 + 768) + (768 * 2 + 2) # pre_classifier + classifier
print("new head parameters:", f"{head:,}"){'accuracy': 0.6667}
P(POSITIVE) per review: [0.032 0.9 0.45 0.991 0.45 0.622]
steps per epoch: 1563 total steps: 3126
new head parameters: 592,130What the check shows
- Accuracy 0.6667: argmax of the six logit rows gives [0, 1, 0, 1, 0, 1] against the labels [0, 1, 1, 1, 0, 0], so four of six are right.
- P(POSITIVE) is the softmax of each row: 0.9 for [−0.4, 1.8], 0.45 for [0.3, 0.1]; the pipeline's
scoreis the probability of the winning label, so [0.3, 0.1] would be reported as NEGATIVE with 0.55. - 1,563 steps per epoch, 3,126 in total: 25,000 reviews in batches of 16 (the last batch is partly full), for two epochs on one device.
- The new head has 592,130 weights, against about 66 million pretrained ones: fine-tuning updates all of them, but only the head starts from random values.
Fine-tuning vs feature extraction vs training from scratch
| Fine-tuning | Feature extraction | Training from scratch | TF-IDF + logistic regression | |
|---|---|---|---|---|
| Starts from | pretrained weights | pretrained weights, frozen | random weights | word counts |
| Trains | all layers + new head | only the head | everything | one linear model |
| Data needed | thousands of labels | hundreds to thousands | millions of examples | thousands |
| Hardware | a GPU helps | a CPU can do it | many GPUs | a CPU |
| Knows word order and context | yes | yes | only what the data teaches | no |
Where you use fine-tuning
- Text classification: sentiment, spam, topic or intent labels on your own data.
- Token classification: named entity recognition with
AutoModelForTokenClassification. - Domain adaptation: legal, medical or support text, where a general model misses the vocabulary.
evaluation_strategy= to TrainingArguments and tokenizer= to the Trainer. Current releases use eval_strategy= and processing_class=; the old names fail. And always load the tokenizer from the same checkpoint as the model.Related
- Previous: BERT, GPT and T5
- See also: Sentiment analysis of Kindle reviews
- Reference: Text classification in the Transformers documentation
- Reference: Fine-tuning with the Trainer
- Change the labels in the stand-in batch so that all six predictions are right and check the accuracy.
- Recompute the steps for
per_device_train_batch_size=32and three epochs. - Change
num_labelsto 3 and recompute the head's weight count (the classifier becomes 768 × 3 + 3).
Little by little, you're building something great.