Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

CBOW (continuous bag of words)

CBOW (continuous bag of words) is the Word2Vec architecture that predicts a centre word from the words around it, learning one vector per word as it trains.

Last updated: 07 Oct, 2026 · gensim 4.4 · NumPy

Word2Vec learns from windows of text without labels. CBOW is the first of its two ways to turn a window into a training example: the surrounding words go in, and the word in the middle is what the network must predict.

CBOW windows on the iNeuron sentence · from the Complete NLP Machine Learning in One Shot video · 3:13:21 to 3:16:19

Sliding a window over the corpus

The video's corpus is one sentence: “iNeuron company is related to data science”. A real corpus has millions of words, but one line is enough to see how the data is made.

A window size of 5 takes five words at a time. The centre word is the output, and the two words before it and the two after it are the input, so the model learns both the backward and the forward context. Moving the window one step at a time gives three training rows:

Input (context words)Output (centre word)
iNeuron, company, related, tois
company, is, to, datarelated
is, related, data, scienceto

Any size can be used; the video picks an odd one so that the centre word has the same number of words on each side. In gensim, the window parameter counts the words on one side, so the video's window of 5 (two each side) is window=2 in gensim.

Encoding the seven words as one-hot vectors

The text cannot go into the network directly, so each word becomes a one-hot vector (One-hot encoding for text) over the vocabulary of 7 words: iNeuron is [1 0 0 0 0 0 0], company [0 1 0 0 0 0 0], is [0 0 1 0 0 0 0], related [0 0 0 1 0 0 0] and to [0 0 0 0 1 0 0].

Passing the context through the CBOW network

CBOW is a small fully connected neural network. Its parts, as Word2Vec defines them:

  • Input: the one-hot vectors of the four context words, 7 numbers each.
  • Input matrix W (7 × N): one shared matrix for every context word. A one-hot vector times W picks out one row, that word's vector.
  • Hidden layer h: the average of the four picked rows, N numbers. There is no bias and no activation function.
  • Output matrix W′ (N × 7) turns h into one score per vocabulary word, and a softmax turns the scores into probabilities that sum to 1.
  • Loss: the cross-entropy −log p(centre word). Backpropagation updates W and W′ until the loss is small (Softmax and Backpropagation and weight update cover both steps).
CBOW with C context words: average the context rows of W, score with W′, softmax over the vocabulary
CBOW on the video's corpus: the windows of five words give context iNeuron, company, related, to for the centre is, then company, is, to, data for related and is, related, data, science for to; each context word's one-hot vector picks a row of one shared 7 by N matrix W, the rows are averaged into h, and h times W prime and a softmax over 7 words should give is a high probability.

After training, the vector of a word is its row of W. N is the vector size, the number of values each word gets, and it is chosen on its own: Google's model uses 300. The window decides which words count as context; the vector size decides how many numbers each word gets. They are separate settings.

Training CBOW on the three windows with NumPy

The forward pass

python
h = W[ctx].mean(axis=0)          # average the context words' rows of W
u = h @ W_out                    # a score for each of the 7 words
p = np.exp(u - u.max())
p = p / p.sum()                  # softmax

One gradient step

python
e = p.copy()
e[c] -= 1                        # p minus the one-hot vector of the true centre word
grad_h = W_out @ e
W_out -= lr * np.outer(h, e)
W[ctx] -= lr * grad_h / len(ctx) # each context row gets its share of the gradient
ExampleFrom the video, run with NumPy
import numpy as np

words = ["ineuron", "company", "is", "related", "to", "data", "science"]
V, N, half = len(words), 5, 2                # 7 words, vector size 5, 2 words each side

windows = []                                 # (context indices, centre index)
for c in range(half, V - half):
    windows.append(([j for j in range(c - half, c + half + 1) if j != c], c))
for ctx, c in windows:
    print([words[j] for j in ctx], "->", words[c])

rng = np.random.default_rng(42)
W = rng.normal(0, 0.1, (V, N))               # input matrix: one row per word
W_out = rng.normal(0, 0.1, (N, V))           # output matrix

def forward(ctx):
    h = W[ctx].mean(axis=0)                  # average of the context rows: no bias, no activation
    u = h @ W_out                            # one score per vocabulary word
    p = np.exp(u - u.max())
    return h, p / p.sum()                    # softmax: probabilities that sum to 1

h, p = forward(windows[0][0])
print("before training: p(is) =", round(p[2], 4), " sum of p =", round(p.sum(), 4))

lr = 0.5
for epoch in range(300):
    loss = 0.0
    for ctx, c in windows:
        h, p = forward(ctx)
        loss -= np.log(p[c])                 # cross-entropy for the true centre word
        e = p.copy()
        e[c] -= 1                            # gradient of the loss with respect to the scores
        grad_h = W_out @ e
        W_out -= lr * np.outer(h, e)
        W[ctx] -= lr * grad_h / len(ctx)
    if epoch in (0, 299):
        print(f"epoch {epoch + 1}: loss {loss:.4f}")

for ctx, c in windows:
    h, p = forward(ctx)
    print(f"{words[c]:8} predicted {words[p.argmax()]:8} p = {p.max():.4f}")
print("vector of ineuron:", W[0].round(3), " W shape:", W.shape, " weights:", W.size + W_out.size)

What the training run shows

  • The three windows are the board's rows: four context words in, the centre word out.
  • Before training, p(is) is close to 1/7 = 0.1429, a near-uniform guess, and the softmax output sums to 1.
  • The loss falls from 5.8419 to 0.0038 over 300 passes, and every window then predicts its own centre word with a probability above 0.998.
  • The vector of ineuron is 5 numbers, its row of W (7 × 5). The two matrices hold 70 weights in all.

Training CBOW with gensim

gensim's Word2Vec with sg=0 is CBOW. The loop trains the same sentence three times, changing the window and the vector size, and prints the shapes.

ExampleRun on gensim 4.4.0
from gensim.models import Word2Vec

sentence = [["ineuron", "company", "is", "related", "to", "data", "science"]]

for window, size in [(2, 5), (2, 10), (3, 5)]:
    model = Word2Vec(sentence, sg=0, window=window, vector_size=size, min_count=1,
                     seed=42, workers=1, epochs=50)
    print(f"window={window} vector_size={size}: wv['ineuron'] has shape {model.wv['ineuron'].shape}, "
          f"input matrix {model.wv.vectors.shape}, output matrix {model.syn1neg.shape}")
print("negative sampling:", model.negative, " hierarchical softmax:", model.hs, " cbow_mean:", model.cbow_mean)
  • vector_size changes the shape: 5 gives vectors of 5 numbers, 10 gives 10.
  • window changes nothing in the shape: window=3 still gives 5 numbers per word.
  • The input matrix is (7, 5), one row per vocabulary word, and the output matrix has the same shape.
  • gensim trains with negative sampling (negative=5, hs=0) and averages the context vectors (cbow_mean=1).

Window size vs vector size

windowvector_size
Decideswhich neighbours count as contexthow many numbers each word gets
The video's example5 words in all, 2 on each side5
In gensimwindow=2 (one side)vector_size=5
Larger valuemore topical, related-word similaritymore capacity, needs more data
Google News modela separate setting300

Where you use CBOW

  • Large corpora: CBOW makes one training example per window, so it trains faster than skip-gram.
  • Frequent words: it gives slightly better vectors for common words.
  • gensim's default: Word2Vec(sentences) trains CBOW unless sg=1 is passed.
Watch out. A bigger window does not make the vectors longer. To get 300-number vectors, set vector_size=300; window only changes which words are treated as neighbours.
Try it yourself
  • Set N = 2 in the NumPy example. Does training still reach a low loss?
  • Change half = 2 to half = 1. How many windows are there now, and what are their contexts?
  • In the gensim example, add (5, 20) to the list and check the shapes.
PreviousWord2Vec

Slow is fine. Stopping is the only problem.