Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Cosine similarity for documents

Cosine similarity is a measure of how alike two vectors are, computed as the cosine of the angle between them: 1 when they point the same way, 0 when they are at right angles.

Last updated: 07 Oct, 2026 · NLTK 3.10 · scikit-learn 1.9

Bag of words and TF-IDF turn each document into a vector. Cosine similarity compares two such vectors, which is how a search engine ranks documents and how a system finds messages that say the same thing.

Defining cosine similarity

Cosine similarity: the dot product divided by the product of the lengths

The value runs from −1 to 1. Count and TF-IDF vectors have no negative values, so for documents it runs from 0 (no shared words) to 1 (the same direction). Cosine distance is 1 − cosine similarity.

Cosine distance near 0 and movie recommendations · from the Complete NLP Machine Learning in One Shot video · 3:07:32 to 3:08:59

Reading the angle between two vectors

The video turns similarity into a distance: distance = 1 − cosine similarity. Two vectors 45° apart have a cosine of cos 45° = 1/√2 = 0.7071, so their distance is 1 − 0.7071 = 0.29: fairly similar.

At 90° the cosine is 0 and the distance is 1 − 0 = 1: the two vectors share no direction, which for count vectors means no word in common. At 0° the cosine is 1 and the distance is 1 − 1 = 0: the two vectors point the same way.

Recommendations work the same way. If movies are vectors over features such as action, comic and comedy, Avengers and Iron Man point almost the same way, so someone who watched Avengers is shown Iron Man.

Two vectors at 0, 45, 90 and 180 degrees have a cosine similarity of 1, 0.7071, 0 and minus 1, so the cosine distance is 0, 0.29, 1 and 2.

For vectors that can hold negative values, such as word embeddings, two vectors can also point in opposite directions: 180°, a cosine of −1 and a distance of 2.

Computing cosine similarity with NumPy

The formula is one line of NumPy. The example checks the 45° and 90° cases, the bag of words vectors of “The food is good” and “The food is not good”, and a sentence compared with itself repeated twice.

ExampleFrom the video, run with NumPy
import numpy as np

def cosine(a, b):
    return a @ b / (np.linalg.norm(a) * np.linalg.norm(b))

a, b = np.array([1, 0]), np.array([1, 1])          # 45 degrees apart
print(f"45 degrees: similarity {cosine(a, b):.4f}, distance {1 - cosine(a, b):.4f}")
print("90 degrees:", cosine(np.array([1, 0]), np.array([0, 1])))

v1 = np.array([1, 1, 1, 0, 1])                    # the food is good      (the, food, is, not, good)
v2 = np.array([1, 1, 1, 1, 1])                    # the food is not good
sim = cosine(v1, v2)
print(f"food sentences: similarity {sim:.4f}, angle {np.degrees(np.arccos(sim)):.2f} degrees")

short, long = np.array([1, 1, 0]), np.array([2, 2, 0])   # "good boy" and "good boy good boy"
print("same words twice: cosine", round(cosine(short, long), 4), " euclidean", round(np.linalg.norm(short - long), 4))
  • 45 degrees gives 0.7071 and a distance of 0.2929, the 0.29 of the board.
  • 90 degrees gives 0.0.
  • The two food sentences score 0.8944, an angle of 26.57°: similar by this measure, although one says the opposite of the other.
  • Repeating a sentence doubles its counts but keeps the direction, so the cosine is 1.0 while the Euclidean distance is 1.4142. Cosine ignores document length.

Comparing documents with scikit-learn's cosine_similarity

cosine_similarity(X) returns the similarity of every pair of rows at once. The example compares the three sentences of the TF-IDF lesson as bag of words, as scikit-learn's TF-IDF and as the board's textbook TF-IDF.

ExampleFrom the video, run on scikit-learn 1.9.1
import numpy as np
from sklearn.feature_extraction.text import CountVectorizer, TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity, linear_kernel

sentences = ["good boy", "good girl", "boy girl good"]
bow = CountVectorizer().fit_transform(sentences)
tfidf = TfidfVectorizer().fit_transform(sentences)
textbook = np.array([[0, 0.5, 0], [0, 0, 0.5], [0, 1 / 3, 1 / 3]]) * np.log(3 / np.array([3, 2, 2]))

print("bag of words:\n", cosine_similarity(bow).round(4))
print("scikit-learn tf-idf:\n", cosine_similarity(tfidf).round(4))
print("textbook tf-idf:\n", cosine_similarity(textbook).round(4))
print("tf-idf dot product equals its cosine:", np.allclose(linear_kernel(tfidf), cosine_similarity(tfidf)))

What the similarity matrices show

  • The diagonal is 1: each sentence compared with itself.
  • Bag of words: “good boy” and “good girl” share good, so they score 0.5; each two-word sentence and “boy girl good” score 0.8165.
  • scikit-learn TF-IDF lowers the pair that shares only good to 0.3762, because good has the smallest weight.
  • Textbook TF-IDF gives good a weight of 0, so “good boy” and “good girl” have nothing in common: 0.0. Each of them and “boy girl good” score 0.7071, against 0.7848 with scikit-learn's TF-IDF.
  • TfidfVectorizer's rows already have length 1, so a plain dot product (linear_kernel) gives the same matrix as cosine_similarity.

Ranking SMS messages for a search query

Search is cosine similarity between a query vector and every document vector. The example fits TF-IDF on the 5,572 SMS messages with scikit-learn's English stopwords, turns a query into a vector with the same vocabulary, and prints the three closest messages.

ExampleFrom the video, run on scikit-learn 1.9.1
import pandas as pd
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity

url = "https://raw.githubusercontent.com/krishnaik06/Avgword2vec-Implementation/main/smsspamcollection/SMSSpamCollection"
messages = pd.read_csv(url, sep="\t", names=["label", "message"])

tfidf = TfidfVectorizer(stop_words="english")
X = tfidf.fit_transform(messages["message"])
query = tfidf.transform(["win a cash prize"])
scores = cosine_similarity(query, X).ravel()

for i in scores.argsort()[::-1][:3]:
    print(f"{scores[i]:.3f}  {messages['label'][i]:4}  {messages['message'][i][:55]}")

All three top messages are spam about a cash prize. The search used no labels: the closest vectors are the messages that share the query's rarer words, cash and prize.

Cosine similarity vs Euclidean distance

Cosine similarityEuclidean distance
Measuresthe angle between the vectorsthe straight-line gap between their tips
“good boy” vs the same text twice1.0 (same direction)1.4142 (different lengths)
Range for count or TF-IDF vectors0 to 1, larger is more similar0 upwards, smaller is more similar
Common usecomparing documents and word vectorspoints in a low-dimensional space

Where you use cosine similarity

  • Search: ranking documents by their similarity to a query.
  • Duplicate detection: flagging support tickets or questions that ask the same thing.
  • Word vectors: finding the nearest words to a word, which is what Word2Vec does with most_similar.
Watch out. A document whose words are all out of vocabulary becomes an all-zero vector, and the cosine formula divides by its length of 0. NumPy returns nan with a warning; scikit-learn's cosine_similarity returns 0. Check for empty vectors before ranking.
Try it yourself
  • Change b to np.array([0, 1]) and then np.array([-1, 0]) in the NumPy example. What similarity and distance do you get?
  • Replace the query with "are you coming home tonight" in the search example. Are the top messages ham or spam?
  • Try the query "claim your free prize now". Why do three very short ham messages containing “free” come first?
PreviousTF-IDF

Every expert started right here.