Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

TF-IDF

TF-IDF (term frequency, inverse document frequency) is a text representation that weights each word in a sentence by how often it appears there and how rare it is across all the sentences, so words that appear everywhere count less.

Last updated: 07 Oct, 2026 · NLTK 3.10 · scikit-learn 1.9

Bag of words (BoW) gives every present word the same 1. TF-IDF keeps the same fixed-size vector but changes the values: a word found in every sentence does little to tell the sentences apart, so it gets a smaller weight.

The TF-IDF formulas and the term frequency table · from the Complete NLP Machine Learning in One Shot video · 2:31:18 to 2:34:48

Calculating the term frequency

The video reuses the bag of words sentences after lower-casing and stopword removal: S1 “good boy”, S2 “good girl” and S3 “boy girl good”. TF-IDF has two parts. The first is the term frequency:

Term frequency, as on the board

For S1, good appears once among two words, so TF = 1/2; boy is also 1/2; girl does not appear, so TF = 0/2 = 0. S2 gives good 1/2, boy 0, girl 1/2. S3 has three words, so each of good, boy and girl gets 1/3.

Calculating the inverse document frequency

The second part looks at the whole corpus. N is the number of sentences, and df(w) is the number of sentences that contain the word:

Inverse document frequency, natural log (log base e)

With N = 3: good is in all three sentences, so IDF = ln(3/3) = ln 1 = 0. boy is in S1 and S3, so IDF = ln(3/2) = 0.405. girl is in S2 and S3, also ln(3/2) = 0.405. The rarer the word, the larger its IDF.

Multiplying TF by IDF

Each sentence's TF times the word's IDF

Multiplying each sentence's column of TF values by the IDF column gives the final vectors:

goodboygirl
Sentence 11/2 × 0 = 01/2 × ln(3/2) = 0.20270
Sentence 21/2 × 0 = 001/2 × ln(3/2) = 0.2027
Sentence 31/3 × 0 = 01/3 × ln(3/2) = 0.13521/3 × ln(3/2) = 0.1352

So “good boy” becomes the vector [0, 0.2027, 0] over good, boy and girl. Every sentence becomes a vector as long as the vocabulary, like bag of words, but with weights in place of 1s.

Computing the textbook TF-IDF with NumPy

The code follows the board's two formulas: TF as count divided by sentence length, IDF as the natural log of N over df, and no other step.

ExampleFrom the video, run with NumPy
import numpy as np

sentences = ["good boy", "good girl", "boy girl good"]
words = ["good", "boy", "girl"]
tokens = [s.split() for s in sentences]
np.set_printoptions(precision=4, suppress=True)

tf = np.array([[t.count(w) / len(t) for w in words] for t in tokens])   # rows S1-S3, columns good, boy, girl
df = np.array([sum(w in t for t in tokens) for w in words])            # sentences that contain each word
idf = np.log(len(sentences) / df)                                       # natural log, as on the board

print("TF:\n", tf)
print("df:", df, " idf:", idf)
print("TF-IDF:\n", tf * idf)
  • TF matches the board's table: 0.5 for the words of the two-word sentences, 0.3333 in S3.
  • df = [3 2 2], so the idf of good is 0 and the idf of boy and girl is 0.4055.
  • TF-IDF is 0.2027 for boy in S1 and girl in S2, 0.1352 for boy and girl in S3, and 0 for good everywhere.
Word importance in TF-IDF · from the Complete NLP Machine Learning in One Shot video · 2:41:11 to 2:44:15

Capturing word importance

This is the advantage of TF-IDF over bag of words. In bag of words, good and boy both get 1 in “good boy”: equal importance. TF-IDF looks at the whole paragraph. A word present in every sentence should get less importance, because it does not single out any sentence.

good is in all three sentences, so its TF-IDF is 0 everywhere and the model ignores it. boy carries S1, girl carries S2, and both boy and girl carry S3. Each sentence's vector is driven by the words that make it different from the others.

Computing TF-IDF with scikit-learn's TfidfVectorizer

TfidfVectorizer uses a different formula from the board, so its numbers differ. By default it makes three changes:

  • TF is the raw count, not the count divided by the sentence length.
  • The idf is smoothed and has 1 added (smooth_idf=True): as if one extra document contained every word, then +1 so that no word gets an idf of 0.
  • Each row is scaled to length 1 (norm="l2"), which cancels the effect of sentence length.
scikit-learn's default: smooth idf with +1, then L2 normalisation of each row

The vectorizer's settings

python
TfidfVectorizer()                    # smooth_idf=True, norm="l2", sublinear_tf=False
TfidfVectorizer(smooth_idf=False)    # idf = ln(N / df) + 1
TfidfVectorizer(norm=None)           # keep count x idf without scaling the rows
tfidf.idf_                           # one idf value per column
ExampleFrom the video, run on scikit-learn 1.9.1
import numpy as np
from sklearn.feature_extraction.text import CountVectorizer, TfidfVectorizer

sentences = ["good boy", "good girl", "boy girl good"]
np.set_printoptions(precision=4, suppress=True)

tfidf = TfidfVectorizer()
X = tfidf.fit_transform(sentences)
print(tfidf.get_feature_names_out())
print("idf_:", tfidf.idf_)
print(X.toarray())

# the same numbers by hand: counts x smooth idf, then each row divided by its length
counts = CountVectorizer().fit_transform(sentences).toarray()
df = (counts > 0).sum(axis=0)
idf = np.log((1 + 3) / (1 + df)) + 1
raw = counts * idf
print(raw / np.linalg.norm(raw, axis=1, keepdims=True))

print("smooth_idf=False:", TfidfVectorizer(smooth_idf=False).fit(sentences).idf_)

What scikit-learn's numbers show

  • The columns are alphabetical: boy, girl, good.
  • idf_ is [1.2877 1.2877 1.]: ln(4/3) + 1 for boy and girl, ln(4/4) + 1 = 1 for good. good keeps an idf of 1 instead of 0.
  • S1 “good boy” is [0.7898, 0, 0.6134]: boy 0.7898 and good 0.6134. good still ranks below boy, as on the board, but it is no longer 0.
  • The hand calculation (counts × smooth idf, each row divided by its length) gives exactly the same matrix.
  • With smooth_idf=False the idf is ln(N/df) + 1: 1.4055 for boy and girl and 1.0 for good. The +1 stays in every setting.

Textbook TF-IDF vs scikit-learn TF-IDF

The board's TF-IDF next to scikit-learn's: TF times ln of N over df gives 0, 0.2027, 0 for sentence 1 and 0, 0.1352, 0.1352 for sentence 3 with good at 0, while scikit-learn's smooth idf plus 1 with L2 normalisation gives good 0.6134 and boy 0.7898 for sentence 1.
Textbook (the board)scikit-learn default
TFcount / words in the sentenceraw count
IDFln(N / df)ln((1 + N) / (1 + df)) + 1
Row scalingnoneL2: each row has length 1
idf of good (df = 3, N = 3)01.0
S1 “good boy” over good, boy, girl[0, 0.2027, 0][0.6134, 0.7898, 0]
S3 “boy girl good”[0, 0.1352, 0.1352][0.4813, 0.6198, 0.6198]

Both give boy and girl more weight than good wherever they appear. When a hand calculation and a library disagree, name the formula each one uses before deciding that one is wrong.

TF-IDF on the SMS spam messages

The video's TF-IDF practical cleans the 5,572 SMS messages like the bag of words one, but reduces words with the WordNet lemmatizer instead of the Porter stemmer, then keeps the 100 most frequent words in TfidfVectorizer(max_features=100).

ExampleFrom the video, run on scikit-learn 1.9.1
import re
import nltk
import pandas as pd
from nltk.corpus import stopwords
from nltk.stem import WordNetLemmatizer
from sklearn.feature_extraction.text import TfidfVectorizer

nltk.download("stopwords", quiet=True)
nltk.download("wordnet", quiet=True)
url = "https://raw.githubusercontent.com/krishnaik06/Avgword2vec-Implementation/main/smsspamcollection/SMSSpamCollection"
messages = pd.read_csv(url, sep="\t", names=["label", "message"])

lemmatizer = WordNetLemmatizer()
stop = set(stopwords.words("english"))
corpus = []
for message in messages["message"]:
    review = re.sub("[^a-zA-Z]", " ", message).lower().split()
    corpus.append(" ".join(lemmatizer.lemmatize(word) for word in review if word not in stop))
print(corpus[0])

tfidf = TfidfVectorizer(max_features=100)
X = tfidf.fit_transform(corpus)
names = tfidf.get_feature_names_out()
for j in sorted(X[0].nonzero()[1]):
    print(f"{names[j]:6} column {j:2}  tf-idf {X[0, j]:.3f}  idf {tfidf.idf_[j]:.4f}")
  • The lemmatized first message keeps real words: crazy, available, bugis, amore.
  • Message 0 has four of the 100 features: go 0.434, got 0.461, great 0.544 and wat 0.550. Each value is the count times the smooth idf, with the row scaled to length 1.
  • wat has the largest idf of the four (4.9910), so it is the rarest across the messages and gets the largest weight.

Listing the advantages and disadvantages of TF-IDF

  • Intuitive: frequent in this sentence and rare elsewhere means important.
  • Fixed-size input, the size of the vocabulary, as with bag of words.
  • Word importance is captured, which bag of words cannot do.
  • Sparsity still exists: most values are 0.
  • Out of vocabulary: a test word missing from the training vocabulary is ignored.
  • Still no word order or meaning: without n-grams, “not good” is two separate columns, and good and great remain unrelated.

Where you use TF-IDF

  • Search and ranking: a query word that is rare in the collection counts more when scoring documents.
  • Text classification with Logistic regression or naive Bayes on spam, reviews or tickets.
  • Keyword extraction: the highest TF-IDF words of a document summarise what is special about it.
Watch out. No setting of TfidfVectorizer reproduces the board's pure ln(N/df): its idf always adds 1, so a word in every document keeps a weight above 0. To check a hand calculation, use TfidfVectorizer(norm=None, smooth_idf=False): each value is then the raw count × (ln(N/df) + 1).
Try it yourself
  • Add a fourth sentence, “good school”, to the NumPy example. Does the idf of good stay 0?
  • In the scikit-learn example, set TfidfVectorizer(norm=None) and compare S1 with counts × idf_.
  • Use np.log10 instead of np.log in the NumPy example. Which values change, and does the order of the words change?
PreviousN-grams

This is what real progress feels like.