Cosine similarity for documents
Cosine similarity is a measure of how alike two vectors are, computed as the cosine of the angle between them: 1 when they point the same way, 0 when they are at right angles.
Last updated: 07 Oct, 2026 · NLTK 3.10 · scikit-learn 1.9
Bag of words and TF-IDF turn each document into a vector. Cosine similarity compares two such vectors, which is how a search engine ranks documents and how a system finds messages that say the same thing.
Defining cosine similarity
The value runs from −1 to 1. Count and TF-IDF vectors have no negative values, so for documents it runs from 0 (no shared words) to 1 (the same direction). Cosine distance is 1 − cosine similarity.
Reading the angle between two vectors
The video turns similarity into a distance: distance = 1 − cosine similarity. Two vectors 45° apart have a cosine of cos 45° = 1/√2 = 0.7071, so their distance is 1 − 0.7071 = 0.29: fairly similar.
At 90° the cosine is 0 and the distance is 1 − 0 = 1: the two vectors share no direction, which for count vectors means no word in common. At 0° the cosine is 1 and the distance is 1 − 1 = 0: the two vectors point the same way.
Recommendations work the same way. If movies are vectors over features such as action, comic and comedy, Avengers and Iron Man point almost the same way, so someone who watched Avengers is shown Iron Man.
For vectors that can hold negative values, such as word embeddings, two vectors can also point in opposite directions: 180°, a cosine of −1 and a distance of 2.
Computing cosine similarity with NumPy
The formula is one line of NumPy. The example checks the 45° and 90° cases, the bag of words vectors of “The food is good” and “The food is not good”, and a sentence compared with itself repeated twice.
import numpy as np
def cosine(a, b):
return a @ b / (np.linalg.norm(a) * np.linalg.norm(b))
a, b = np.array([1, 0]), np.array([1, 1]) # 45 degrees apart
print(f"45 degrees: similarity {cosine(a, b):.4f}, distance {1 - cosine(a, b):.4f}")
print("90 degrees:", cosine(np.array([1, 0]), np.array([0, 1])))
v1 = np.array([1, 1, 1, 0, 1]) # the food is good (the, food, is, not, good)
v2 = np.array([1, 1, 1, 1, 1]) # the food is not good
sim = cosine(v1, v2)
print(f"food sentences: similarity {sim:.4f}, angle {np.degrees(np.arccos(sim)):.2f} degrees")
short, long = np.array([1, 1, 0]), np.array([2, 2, 0]) # "good boy" and "good boy good boy"
print("same words twice: cosine", round(cosine(short, long), 4), " euclidean", round(np.linalg.norm(short - long), 4))45 degrees: similarity 0.7071, distance 0.2929 90 degrees: 0.0 food sentences: similarity 0.8944, angle 26.57 degrees same words twice: cosine 1.0 euclidean 1.4142
- 45 degrees gives 0.7071 and a distance of 0.2929, the 0.29 of the board.
- 90 degrees gives 0.0.
- The two food sentences score 0.8944, an angle of 26.57°: similar by this measure, although one says the opposite of the other.
- Repeating a sentence doubles its counts but keeps the direction, so the cosine is 1.0 while the Euclidean distance is 1.4142. Cosine ignores document length.
Comparing documents with scikit-learn's cosine_similarity
cosine_similarity(X) returns the similarity of every pair of rows at once. The example compares the three sentences of the TF-IDF lesson as bag of words, as scikit-learn's TF-IDF and as the board's textbook TF-IDF.
import numpy as np
from sklearn.feature_extraction.text import CountVectorizer, TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity, linear_kernel
sentences = ["good boy", "good girl", "boy girl good"]
bow = CountVectorizer().fit_transform(sentences)
tfidf = TfidfVectorizer().fit_transform(sentences)
textbook = np.array([[0, 0.5, 0], [0, 0, 0.5], [0, 1 / 3, 1 / 3]]) * np.log(3 / np.array([3, 2, 2]))
print("bag of words:\n", cosine_similarity(bow).round(4))
print("scikit-learn tf-idf:\n", cosine_similarity(tfidf).round(4))
print("textbook tf-idf:\n", cosine_similarity(textbook).round(4))
print("tf-idf dot product equals its cosine:", np.allclose(linear_kernel(tfidf), cosine_similarity(tfidf)))bag of words: [[1. 0.5 0.8165] [0.5 1. 0.8165] [0.8165 0.8165 1. ]] scikit-learn tf-idf: [[1. 0.3762 0.7848] [0.3762 1. 0.7848] [0.7848 0.7848 1. ]] textbook tf-idf: [[1. 0. 0.7071] [0. 1. 0.7071] [0.7071 0.7071 1. ]] tf-idf dot product equals its cosine: True
What the similarity matrices show
- The diagonal is 1: each sentence compared with itself.
- Bag of words: “good boy” and “good girl” share good, so they score 0.5; each two-word sentence and “boy girl good” score 0.8165.
- scikit-learn TF-IDF lowers the pair that shares only good to 0.3762, because good has the smallest weight.
- Textbook TF-IDF gives good a weight of 0, so “good boy” and “good girl” have nothing in common: 0.0. Each of them and “boy girl good” score 0.7071, against 0.7848 with scikit-learn's TF-IDF.
- TfidfVectorizer's rows already have length 1, so a plain dot product (
linear_kernel) gives the same matrix ascosine_similarity.
Ranking SMS messages for a search query
Search is cosine similarity between a query vector and every document vector. The example fits TF-IDF on the 5,572 SMS messages with scikit-learn's English stopwords, turns a query into a vector with the same vocabulary, and prints the three closest messages.
import pandas as pd
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity
url = "https://raw.githubusercontent.com/krishnaik06/Avgword2vec-Implementation/main/smsspamcollection/SMSSpamCollection"
messages = pd.read_csv(url, sep="\t", names=["label", "message"])
tfidf = TfidfVectorizer(stop_words="english")
X = tfidf.fit_transform(messages["message"])
query = tfidf.transform(["win a cash prize"])
scores = cosine_similarity(query, X).ravel()
for i in scores.argsort()[::-1][:3]:
print(f"{scores[i]:.3f} {messages['label'][i]:4} {messages['message'][i][:55]}")0.716 spam Win a £1000 cash prize or a prize worth £5000 0.418 spam You have WON a guaranteed £1000 cash or a £2000 prize.T 0.389 spam This is the 2nd attempt to contract U, you have won thi
All three top messages are spam about a cash prize. The search used no labels: the closest vectors are the messages that share the query's rarer words, cash and prize.
Cosine similarity vs Euclidean distance
| Cosine similarity | Euclidean distance | |
|---|---|---|
| Measures | the angle between the vectors | the straight-line gap between their tips |
| “good boy” vs the same text twice | 1.0 (same direction) | 1.4142 (different lengths) |
| Range for count or TF-IDF vectors | 0 to 1, larger is more similar | 0 upwards, smaller is more similar |
| Common use | comparing documents and word vectors | points in a low-dimensional space |
Where you use cosine similarity
- Search: ranking documents by their similarity to a query.
- Duplicate detection: flagging support tickets or questions that ask the same thing.
- Word vectors: finding the nearest words to a word, which is what Word2Vec does with
most_similar.
nan with a warning; scikit-learn's cosine_similarity returns 0. Check for empty vectors before ranking.Related
- Previous: TF-IDF
- Next: Word embeddings
- See also: K nearest neighbours (KNN)
- Reference: cosine_similarity in the scikit-learn API
- Change
btonp.array([0, 1])and thennp.array([-1, 0])in the NumPy example. What similarity and distance do you get? - Replace the query with
"are you coming home tonight"in the search example. Are the top messages ham or spam? - Try the query
"claim your free prize now". Why do three very short ham messages containing “free” come first?
Every expert started right here.