Precision, recall and F-beta
Precision and recall are classification metrics read from the confusion matrix: precision is the share of predicted positives that are right, recall is the share of actual positives that were found, and the F-beta score combines the two.
Last updated: 05 Oct, 2026 · scikit-learn 1.9.1
Accuracy, from Confusion matrix, counts all right answers together. On imbalanced data that can hide a model which never finds the rare class at all.
Seeing accuracy fail on imbalanced data
Say the output has 900 zeros and 100 ones. That is an imbalanced, or biased, data set. 600 zeros and 400 ones counts as balanced: a gap that small may not affect many algorithms. Now build a model that outputs 0 for every input. Its accuracy is 900 / 1000 = 90%.
90% looks like a good accuracy, yet the model only ever says 0. "You should only not be dependent on accuracy": the video turns to precision, recall and the F-score.

Defining precision and recall
- Recall: out of all the actual positives, how many were predicted correctly. It is also called the true positive rate (TPR) or sensitivity. Recall gives priority to the false negatives: to raise it, reduce FN.
- Precision: out of all the predicted positives, how many are truly positive. Precision gives priority to the false positives: to raise it, reduce FP.

Choosing between precision and recall
- Spam classification uses precision. A false positive marks a real mail as spam and hides it from you, so false positives are the costly mistake.
- Cancer detection uses recall. If a person has cancer and the model predicts not, a false negative, "that scenario is very dangerous". If the model wrongly says cancer, the person goes for further tests and finds out. So false negatives get the priority. Predicting whether a person has diabetes works the same way: a diabetic person told "no diabetes" is the blunder, while a person wrongly told "diabetes" gets a second opinion and a check.
- "Tomorrow the stock market is going to crash" depends on who uses it. For people, missing a real crash (FN) is the costly mistake; for companies, a false alarm (FP) is very bad. When both kinds of mistake matter, use the F-score.
Combining precision and recall with F-beta
- β = 1, the F1 score. When false positives and false negatives are equally important. The formula becomes 2 · P · R / (P + R), the harmonic mean, the same shape as 2xy / (x + y).
- β = 0.5, the F0.5 score. When a false positive is more important than a false negative: decreasing β gives more weight to precision.
- β = 2, the F2 score. When a false negative is more important than a false positive: increasing β gives more weight to recall.

β is the deciding parameter that picks between the F1, F2 and F0.5 scores. Note where β² sits: it multiplies precision in the denominator too. Written without it, as (1 + β²) · P · R / (P + R), the formula is no longer an F-score: for β = 0.5 and the numbers below it gives 1.25 × 0.6 × 0.75 / 1.35 = 0.417 instead of 0.625, below both precision and recall.
Computing the metrics in code
precision_score, recall_score and fbeta_score
Each takes the actual and the predicted labels, in that order, and treats class 1 as the positive class.
from sklearn.metrics import precision_score, recall_score, f1_score, fbeta_score
precision_score(y, y_hat) # TP / (TP + FP)
recall_score(y, y_hat) # TP / (TP + FN)
f1_score(y, y_hat) # beta = 1
fbeta_score(y, y_hat, beta=2) # beta > 1 leans on recallThe video's seven predictions
import numpy as np
from sklearn.metrics import precision_score, recall_score, f1_score, fbeta_score
y = np.array([0, 1, 0, 1, 1, 0, 1]) # actual
y_hat = np.array([1, 1, 0, 1, 1, 1, 0]) # predicted
p = precision_score(y, y_hat)
r = recall_score(y, y_hat)
print("precision:", round(p, 3), " recall:", round(r, 3))
print("F1 by hand:", round(2 * p * r / (p + r), 3), " f1_score:", round(f1_score(y, y_hat), 3))
for beta in [0.5, 1, 2]:
by_hand = (1 + beta ** 2) * p * r / (beta ** 2 * p + r)
print(f"F{beta}: by hand {by_hand:.3f}, fbeta_score {fbeta_score(y, y_hat, beta=beta):.3f}")precision: 0.6 recall: 0.75 F1 by hand: 0.667 f1_score: 0.667 F0.5: by hand 0.625, fbeta_score 0.625 F1: by hand 0.667, fbeta_score 0.667 F2: by hand 0.714, fbeta_score 0.714
The model that always predicts 0
import numpy as np
from sklearn.metrics import accuracy_score, precision_score, recall_score
y = np.array([0] * 900 + [1] * 100) # 900 zeros, 100 ones
always_zero = np.zeros(1000, dtype=int) # a model that always predicts 0
print("accuracy: ", accuracy_score(y, always_zero))
print("recall: ", recall_score(y, always_zero))
print("precision:", precision_score(y, always_zero, zero_division=0))accuracy: 0.9 recall: 0.0 precision: 0.0
What the metrics show
- Precision 0.6 and recall 0.75. Of the 5 students predicted 1, 3 are right (3/5); of the 4 actual 1s, 3 were found (3/4).
- F1 is 0.667. The harmonic mean sits between 0.6 and 0.75, closer to the lower one.
- F0.5 is 0.625 and F2 is 0.714. F0.5 moves toward precision (0.6), F2 toward recall (0.75): β decides which mistake weighs more.
- The always-0 model scores 0.9 accuracy and 0.0 recall. It finds none of the 100 ones. Precision is 0 / 0 here, which scikit-learn reports as 0 when
zero_division=0is set.
Precision vs recall
| Precision | Recall | |
|---|---|---|
| Formula | TP / (TP + FP) | TP / (TP + FN) |
| Question | of the predicted 1s, how many are right? | of the actual 1s, how many were found? |
| Reduces | false positives | false negatives |
| Other names | true positive rate, sensitivity | |
| Example from the video | spam classification | has cancer or not |
| F-beta that favours it | F0.5 | F2 |
Where you use precision, recall and F-beta
- Imbalanced classes. Fraud, rare diseases, defects: accuracy is high for a useless model, recall is not.
- When one mistake costs more. Pick precision, recall, F0.5 or F2 by which mistake is worse, as with spam and cancer.
- As a tuning score.
scoring="f1"in GridSearchCV, used in Logistic regression in scikit-learn.
zero_division. Use pos_label=0 to score class 0 instead.Related
- Previous: Confusion matrix
- Next: Logistic regression in scikit-learn
- Reference: scikit-learn: Precision, recall and F-measures
- Change the last predicted value from 0 to 1. Recall becomes 4/4 = 1.0 and precision 4/6 = 0.667.
- Add
pos_label=0to precision_score and recall_score to score class 0 instead of class 1. - Run
fbeta_score(y, y_hat, beta=3)and check that it is closer to the recall than F2 is.
Every expert started right here.