CBOW (continuous bag of words)
CBOW (continuous bag of words) is the Word2Vec architecture that predicts a centre word from the words around it, learning one vector per word as it trains.
Last updated: 07 Oct, 2026 · gensim 4.4 · NumPy
Word2Vec learns from windows of text without labels. CBOW is the first of its two ways to turn a window into a training example: the surrounding words go in, and the word in the middle is what the network must predict.
Sliding a window over the corpus
The video's corpus is one sentence: “iNeuron company is related to data science”. A real corpus has millions of words, but one line is enough to see how the data is made.
A window size of 5 takes five words at a time. The centre word is the output, and the two words before it and the two after it are the input, so the model learns both the backward and the forward context. Moving the window one step at a time gives three training rows:
| Input (context words) | Output (centre word) |
|---|---|
| iNeuron, company, related, to | is |
| company, is, to, data | related |
| is, related, data, science | to |
Any size can be used; the video picks an odd one so that the centre word has the same number of words on each side. In gensim, the window parameter counts the words on one side, so the video's window of 5 (two each side) is window=2 in gensim.
Encoding the seven words as one-hot vectors
The text cannot go into the network directly, so each word becomes a one-hot vector (One-hot encoding for text) over the vocabulary of 7 words: iNeuron is [1 0 0 0 0 0 0], company [0 1 0 0 0 0 0], is [0 0 1 0 0 0 0], related [0 0 0 1 0 0 0] and to [0 0 0 0 1 0 0].
Passing the context through the CBOW network
CBOW is a small fully connected neural network. Its parts, as Word2Vec defines them:
- Input: the one-hot vectors of the four context words, 7 numbers each.
- Input matrix W (7 × N): one shared matrix for every context word. A one-hot vector times W picks out one row, that word's vector.
- Hidden layer h: the average of the four picked rows, N numbers. There is no bias and no activation function.
- Output matrix W′ (N × 7) turns h into one score per vocabulary word, and a softmax turns the scores into probabilities that sum to 1.
- Loss: the cross-entropy −log p(centre word). Backpropagation updates W and W′ until the loss is small (Softmax and Backpropagation and weight update cover both steps).
After training, the vector of a word is its row of W. N is the vector size, the number of values each word gets, and it is chosen on its own: Google's model uses 300. The window decides which words count as context; the vector size decides how many numbers each word gets. They are separate settings.
Training CBOW on the three windows with NumPy
The forward pass
h = W[ctx].mean(axis=0) # average the context words' rows of W
u = h @ W_out # a score for each of the 7 words
p = np.exp(u - u.max())
p = p / p.sum() # softmaxOne gradient step
e = p.copy()
e[c] -= 1 # p minus the one-hot vector of the true centre word
grad_h = W_out @ e
W_out -= lr * np.outer(h, e)
W[ctx] -= lr * grad_h / len(ctx) # each context row gets its share of the gradientimport numpy as np
words = ["ineuron", "company", "is", "related", "to", "data", "science"]
V, N, half = len(words), 5, 2 # 7 words, vector size 5, 2 words each side
windows = [] # (context indices, centre index)
for c in range(half, V - half):
windows.append(([j for j in range(c - half, c + half + 1) if j != c], c))
for ctx, c in windows:
print([words[j] for j in ctx], "->", words[c])
rng = np.random.default_rng(42)
W = rng.normal(0, 0.1, (V, N)) # input matrix: one row per word
W_out = rng.normal(0, 0.1, (N, V)) # output matrix
def forward(ctx):
h = W[ctx].mean(axis=0) # average of the context rows: no bias, no activation
u = h @ W_out # one score per vocabulary word
p = np.exp(u - u.max())
return h, p / p.sum() # softmax: probabilities that sum to 1
h, p = forward(windows[0][0])
print("before training: p(is) =", round(p[2], 4), " sum of p =", round(p.sum(), 4))
lr = 0.5
for epoch in range(300):
loss = 0.0
for ctx, c in windows:
h, p = forward(ctx)
loss -= np.log(p[c]) # cross-entropy for the true centre word
e = p.copy()
e[c] -= 1 # gradient of the loss with respect to the scores
grad_h = W_out @ e
W_out -= lr * np.outer(h, e)
W[ctx] -= lr * grad_h / len(ctx)
if epoch in (0, 299):
print(f"epoch {epoch + 1}: loss {loss:.4f}")
for ctx, c in windows:
h, p = forward(ctx)
print(f"{words[c]:8} predicted {words[p.argmax()]:8} p = {p.max():.4f}")
print("vector of ineuron:", W[0].round(3), " W shape:", W.shape, " weights:", W.size + W_out.size)['ineuron', 'company', 'related', 'to'] -> is ['company', 'is', 'to', 'data'] -> related ['is', 'related', 'data', 'science'] -> to before training: p(is) = 0.1428 sum of p = 1.0 epoch 1: loss 5.8419 epoch 300: loss 0.0038 is predicted is p = 0.9989 related predicted related p = 0.9984 to predicted to p = 0.9989 vector of ineuron: [-1.159 -1.59 1.514 -0.573 -0.925] W shape: (7, 5) weights: 70
What the training run shows
- The three windows are the board's rows: four context words in, the centre word out.
- Before training, p(is) is close to 1/7 = 0.1429, a near-uniform guess, and the softmax output sums to 1.
- The loss falls from 5.8419 to 0.0038 over 300 passes, and every window then predicts its own centre word with a probability above 0.998.
- The vector of ineuron is 5 numbers, its row of W (7 × 5). The two matrices hold 70 weights in all.
Training CBOW with gensim
gensim's Word2Vec with sg=0 is CBOW. The loop trains the same sentence three times, changing the window and the vector size, and prints the shapes.
from gensim.models import Word2Vec
sentence = [["ineuron", "company", "is", "related", "to", "data", "science"]]
for window, size in [(2, 5), (2, 10), (3, 5)]:
model = Word2Vec(sentence, sg=0, window=window, vector_size=size, min_count=1,
seed=42, workers=1, epochs=50)
print(f"window={window} vector_size={size}: wv['ineuron'] has shape {model.wv['ineuron'].shape}, "
f"input matrix {model.wv.vectors.shape}, output matrix {model.syn1neg.shape}")
print("negative sampling:", model.negative, " hierarchical softmax:", model.hs, " cbow_mean:", model.cbow_mean)window=2 vector_size=5: wv['ineuron'] has shape (5,), input matrix (7, 5), output matrix (7, 5) window=2 vector_size=10: wv['ineuron'] has shape (10,), input matrix (7, 10), output matrix (7, 10) window=3 vector_size=5: wv['ineuron'] has shape (5,), input matrix (7, 5), output matrix (7, 5) negative sampling: 5 hierarchical softmax: 0 cbow_mean: 1
- vector_size changes the shape: 5 gives vectors of 5 numbers, 10 gives 10.
- window changes nothing in the shape: window=3 still gives 5 numbers per word.
- The input matrix is (7, 5), one row per vocabulary word, and the output matrix has the same shape.
- gensim trains with negative sampling (
negative=5,hs=0) and averages the context vectors (cbow_mean=1).
Window size vs vector size
| window | vector_size | |
|---|---|---|
| Decides | which neighbours count as context | how many numbers each word gets |
| The video's example | 5 words in all, 2 on each side | 5 |
| In gensim | window=2 (one side) | vector_size=5 |
| Larger value | more topical, related-word similarity | more capacity, needs more data |
| Google News model | a separate setting | 300 |
Where you use CBOW
- Large corpora: CBOW makes one training example per window, so it trains faster than skip-gram.
- Frequent words: it gives slightly better vectors for common words.
- gensim's default:
Word2Vec(sentences)trains CBOW unlesssg=1is passed.
vector_size=300; window only changes which words are treated as neighbours.Related
- Previous: Word2Vec
- Next: Skip-gram
- Reference: models.word2vec in the gensim documentation
- Set
N = 2in the NumPy example. Does training still reach a low loss? - Change
half = 2tohalf = 1. How many windows are there now, and what are their contexts? - In the gensim example, add
(5, 20)to the list and check the shapes.
Slow is fine. Stopping is the only problem.