TF-IDF
TF-IDF (term frequency, inverse document frequency) is a text representation that weights each word in a sentence by how often it appears there and how rare it is across all the sentences, so words that appear everywhere count less.
Last updated: 07 Oct, 2026 · NLTK 3.10 · scikit-learn 1.9
Bag of words (BoW) gives every present word the same 1. TF-IDF keeps the same fixed-size vector but changes the values: a word found in every sentence does little to tell the sentences apart, so it gets a smaller weight.
Calculating the term frequency
The video reuses the bag of words sentences after lower-casing and stopword removal: S1 “good boy”, S2 “good girl” and S3 “boy girl good”. TF-IDF has two parts. The first is the term frequency:
For S1, good appears once among two words, so TF = 1/2; boy is also 1/2; girl does not appear, so TF = 0/2 = 0. S2 gives good 1/2, boy 0, girl 1/2. S3 has three words, so each of good, boy and girl gets 1/3.
Calculating the inverse document frequency
The second part looks at the whole corpus. N is the number of sentences, and df(w) is the number of sentences that contain the word:
With N = 3: good is in all three sentences, so IDF = ln(3/3) = ln 1 = 0. boy is in S1 and S3, so IDF = ln(3/2) = 0.405. girl is in S2 and S3, also ln(3/2) = 0.405. The rarer the word, the larger its IDF.
Multiplying TF by IDF
Multiplying each sentence's column of TF values by the IDF column gives the final vectors:
| good | boy | girl | |
|---|---|---|---|
| Sentence 1 | 1/2 × 0 = 0 | 1/2 × ln(3/2) = 0.2027 | 0 |
| Sentence 2 | 1/2 × 0 = 0 | 0 | 1/2 × ln(3/2) = 0.2027 |
| Sentence 3 | 1/3 × 0 = 0 | 1/3 × ln(3/2) = 0.1352 | 1/3 × ln(3/2) = 0.1352 |
So “good boy” becomes the vector [0, 0.2027, 0] over good, boy and girl. Every sentence becomes a vector as long as the vocabulary, like bag of words, but with weights in place of 1s.
Computing the textbook TF-IDF with NumPy
The code follows the board's two formulas: TF as count divided by sentence length, IDF as the natural log of N over df, and no other step.
import numpy as np
sentences = ["good boy", "good girl", "boy girl good"]
words = ["good", "boy", "girl"]
tokens = [s.split() for s in sentences]
np.set_printoptions(precision=4, suppress=True)
tf = np.array([[t.count(w) / len(t) for w in words] for t in tokens]) # rows S1-S3, columns good, boy, girl
df = np.array([sum(w in t for t in tokens) for w in words]) # sentences that contain each word
idf = np.log(len(sentences) / df) # natural log, as on the board
print("TF:\n", tf)
print("df:", df, " idf:", idf)
print("TF-IDF:\n", tf * idf)TF: [[0.5 0.5 0. ] [0.5 0. 0.5 ] [0.3333 0.3333 0.3333]] df: [3 2 2] idf: [0. 0.4055 0.4055] TF-IDF: [[0. 0.2027 0. ] [0. 0. 0.2027] [0. 0.1352 0.1352]]
- TF matches the board's table: 0.5 for the words of the two-word sentences, 0.3333 in S3.
- df = [3 2 2], so the idf of good is 0 and the idf of boy and girl is 0.4055.
- TF-IDF is 0.2027 for boy in S1 and girl in S2, 0.1352 for boy and girl in S3, and 0 for good everywhere.
Capturing word importance
This is the advantage of TF-IDF over bag of words. In bag of words, good and boy both get 1 in “good boy”: equal importance. TF-IDF looks at the whole paragraph. A word present in every sentence should get less importance, because it does not single out any sentence.
good is in all three sentences, so its TF-IDF is 0 everywhere and the model ignores it. boy carries S1, girl carries S2, and both boy and girl carry S3. Each sentence's vector is driven by the words that make it different from the others.
Computing TF-IDF with scikit-learn's TfidfVectorizer
TfidfVectorizer uses a different formula from the board, so its numbers differ. By default it makes three changes:
- TF is the raw count, not the count divided by the sentence length.
- The idf is smoothed and has 1 added (
smooth_idf=True): as if one extra document contained every word, then +1 so that no word gets an idf of 0. - Each row is scaled to length 1 (
norm="l2"), which cancels the effect of sentence length.
The vectorizer's settings
TfidfVectorizer() # smooth_idf=True, norm="l2", sublinear_tf=False
TfidfVectorizer(smooth_idf=False) # idf = ln(N / df) + 1
TfidfVectorizer(norm=None) # keep count x idf without scaling the rows
tfidf.idf_ # one idf value per columnimport numpy as np
from sklearn.feature_extraction.text import CountVectorizer, TfidfVectorizer
sentences = ["good boy", "good girl", "boy girl good"]
np.set_printoptions(precision=4, suppress=True)
tfidf = TfidfVectorizer()
X = tfidf.fit_transform(sentences)
print(tfidf.get_feature_names_out())
print("idf_:", tfidf.idf_)
print(X.toarray())
# the same numbers by hand: counts x smooth idf, then each row divided by its length
counts = CountVectorizer().fit_transform(sentences).toarray()
df = (counts > 0).sum(axis=0)
idf = np.log((1 + 3) / (1 + df)) + 1
raw = counts * idf
print(raw / np.linalg.norm(raw, axis=1, keepdims=True))
print("smooth_idf=False:", TfidfVectorizer(smooth_idf=False).fit(sentences).idf_)['boy' 'girl' 'good'] idf_: [1.2877 1.2877 1. ] [[0.7898 0. 0.6134] [0. 0.7898 0.6134] [0.6198 0.6198 0.4813]] [[0.7898 0. 0.6134] [0. 0.7898 0.6134] [0.6198 0.6198 0.4813]] smooth_idf=False: [1.4055 1.4055 1. ]
What scikit-learn's numbers show
- The columns are alphabetical: boy, girl, good.
- idf_ is [1.2877 1.2877 1.]: ln(4/3) + 1 for boy and girl, ln(4/4) + 1 = 1 for good. good keeps an idf of 1 instead of 0.
- S1 “good boy” is [0.7898, 0, 0.6134]: boy 0.7898 and good 0.6134. good still ranks below boy, as on the board, but it is no longer 0.
- The hand calculation (counts × smooth idf, each row divided by its length) gives exactly the same matrix.
- With
smooth_idf=Falsethe idf is ln(N/df) + 1: 1.4055 for boy and girl and 1.0 for good. The +1 stays in every setting.
Textbook TF-IDF vs scikit-learn TF-IDF
| Textbook (the board) | scikit-learn default | |
|---|---|---|
| TF | count / words in the sentence | raw count |
| IDF | ln(N / df) | ln((1 + N) / (1 + df)) + 1 |
| Row scaling | none | L2: each row has length 1 |
| idf of good (df = 3, N = 3) | 0 | 1.0 |
| S1 “good boy” over good, boy, girl | [0, 0.2027, 0] | [0.6134, 0.7898, 0] |
| S3 “boy girl good” | [0, 0.1352, 0.1352] | [0.4813, 0.6198, 0.6198] |
Both give boy and girl more weight than good wherever they appear. When a hand calculation and a library disagree, name the formula each one uses before deciding that one is wrong.
TF-IDF on the SMS spam messages
The video's TF-IDF practical cleans the 5,572 SMS messages like the bag of words one, but reduces words with the WordNet lemmatizer instead of the Porter stemmer, then keeps the 100 most frequent words in TfidfVectorizer(max_features=100).
import re
import nltk
import pandas as pd
from nltk.corpus import stopwords
from nltk.stem import WordNetLemmatizer
from sklearn.feature_extraction.text import TfidfVectorizer
nltk.download("stopwords", quiet=True)
nltk.download("wordnet", quiet=True)
url = "https://raw.githubusercontent.com/krishnaik06/Avgword2vec-Implementation/main/smsspamcollection/SMSSpamCollection"
messages = pd.read_csv(url, sep="\t", names=["label", "message"])
lemmatizer = WordNetLemmatizer()
stop = set(stopwords.words("english"))
corpus = []
for message in messages["message"]:
review = re.sub("[^a-zA-Z]", " ", message).lower().split()
corpus.append(" ".join(lemmatizer.lemmatize(word) for word in review if word not in stop))
print(corpus[0])
tfidf = TfidfVectorizer(max_features=100)
X = tfidf.fit_transform(corpus)
names = tfidf.get_feature_names_out()
for j in sorted(X[0].nonzero()[1]):
print(f"{names[j]:6} column {j:2} tf-idf {X[0, j]:.3f} idf {tfidf.idf_[j]:.4f}")go jurong point crazy available bugis n great world la e buffet cine got amore wat go column 22 tf-idf 0.434 idf 3.9387 got column 25 tf-idf 0.461 idf 4.1833 great column 26 tf-idf 0.544 idf 4.9343 wat column 89 tf-idf 0.550 idf 4.9910
- The lemmatized first message keeps real words: crazy, available, bugis, amore.
- Message 0 has four of the 100 features: go 0.434, got 0.461, great 0.544 and wat 0.550. Each value is the count times the smooth idf, with the row scaled to length 1.
- wat has the largest idf of the four (4.9910), so it is the rarest across the messages and gets the largest weight.
Listing the advantages and disadvantages of TF-IDF
- Intuitive: frequent in this sentence and rare elsewhere means important.
- Fixed-size input, the size of the vocabulary, as with bag of words.
- Word importance is captured, which bag of words cannot do.
- Sparsity still exists: most values are 0.
- Out of vocabulary: a test word missing from the training vocabulary is ignored.
- Still no word order or meaning: without n-grams, “not good” is two separate columns, and good and great remain unrelated.
Where you use TF-IDF
- Search and ranking: a query word that is rare in the collection counts more when scoring documents.
- Text classification with Logistic regression or naive Bayes on spam, reviews or tickets.
- Keyword extraction: the highest TF-IDF words of a document summarise what is special about it.
TfidfVectorizer reproduces the board's pure ln(N/df): its idf always adds 1, so a word in every document keeps a weight above 0. To check a hand calculation, use TfidfVectorizer(norm=None, smooth_idf=False): each value is then the raw count × (ln(N/df) + 1).Related
- Previous: N-grams
- Next: Cosine similarity for documents
- Reference: Tf-idf term weighting in the scikit-learn user guide
- Add a fourth sentence, “good school”, to the NumPy example. Does the idf of good stay 0?
- In the scikit-learn example, set
TfidfVectorizer(norm=None)and compare S1 with counts ×idf_. - Use
np.log10instead ofnp.login the NumPy example. Which values change, and does the order of the words change?
This is what real progress feels like.