Skip-gram
Skip-gram is the Word2Vec architecture that predicts the words around a centre word from the centre word, the reverse of CBOW.
Last updated: 07 Oct, 2026 · gensim 4.4 · NumPy
CBOW (continuous bag of words) took the context words in and predicted the centre word. Skip-gram keeps the same windows and swaps the two sides.
Reversing CBOW: the centre word in, the context words out
The video takes the same corpus, “iNeuron company is related to data science”, and the same window size of 5. Everything stays the same except that the input and the output change places: the centre word is the input, and the four words around it are the output.
| Input (centre word) | Output (context words) |
|---|---|
| is | iNeuron, company, related, to |
| related | company, is, to, data |
| to | is, related, data, science |
Turning each window into training pairs
Skip-gram trains on (centre, context) pairs: from the first window, (is, iNeuron), (is, company), (is, related) and (is, to). The three windows of the board give 12 pairs. Taking every word of the sentence as a centre, with up to two neighbours on each side, gives more, because the words at the ends have fewer neighbours.
Passing the centre word through the skip-gram network
- Input: the one-hot vector of the centre word, for is [0 0 1 0 0 0 0].
- Input matrix W (7 × N): the one-hot vector picks out the row of is, and that row is the hidden layer h. No bias, no activation.
- Output matrix W′ (N × 7) and a softmax give a probability for every vocabulary word.
- Loss: the same probabilities are scored against each context word, so the loss is the sum of −log p(context) over the four neighbours. For the target iNeuron, y is [1 0 0 0 0 0 0].
The board draws one output box per context word. They all use the same W′; there is one input matrix and one output matrix for the whole model, as in CBOW.
Training skip-gram on the sentence with NumPy
The example builds the pairs and trains the two matrices with a full softmax. Each step takes one centre word with all its context words at once, as in the board's picture of one input and four outputs. Then it asks which words the model expects around is.
import numpy as np
words = ["ineuron", "company", "is", "related", "to", "data", "science"]
V, N, half = len(words), 5, 2
board = [] # the pairs from the board's three windows
for c in range(half, V - half):
board += [(c, j) for j in range(c - half, c + half + 1) if j != c]
contexts = {c: [j for j in range(max(0, c - half), min(V, c + half + 1)) if j != c] for c in range(V)}
print("pairs from the board's windows:", len(board), " pairs from every position:", sum(map(len, contexts.values())))
print([(words[c], words[j]) for c, j in board[:4]])
rng = np.random.default_rng(42)
W = rng.normal(0, 0.1, (V, N))
W_out = rng.normal(0, 0.1, (N, V))
def probs(c):
u = W[c] @ W_out # h is the centre word's row of W
p = np.exp(u - u.max())
return p / p.sum()
lr = 0.2
for epoch in range(300):
loss = 0.0
for c, ctx in contexts.items():
h, p = W[c].copy(), probs(c)
loss -= np.log(p[ctx]).sum() # one term per context word
e = len(ctx) * p
e[ctx] -= 1 # gradient of the summed loss with respect to the scores
W[c] -= lr * (W_out @ e)
W_out -= lr * np.outer(h, e)
if epoch in (0, 299):
print(f"epoch {epoch + 1}: loss {loss:.4f}")
print("lowest possible loss:", round(sum(len(x) * np.log(len(x)) for x in contexts.values()), 4))
p = probs(words.index("is"))
top = np.argsort(p)[::-1][:4]
print("most likely neighbours of 'is':", [(words[k], round(float(p[k]), 3)) for k in top], " total:", round(float(p[top].sum()), 3))pairs from the board's windows: 12 pairs from every position: 22
[('is', 'ineuron'), ('is', 'company'), ('is', 'related'), ('is', 'to')]
epoch 1: loss 42.7980
epoch 300: loss 27.3981
lowest possible loss: 25.9998
most likely neighbours of 'is': [('to', 0.376), ('related', 0.306), ('company', 0.175), ('ineuron', 0.134)] total: 0.991What the skip-gram run shows
- 12 pairs come from the board's three windows and 22 from every position of the sentence.
- The first four pairs all have is as the input: one training example per context word.
- The loss falls from 42.80 to 27.40, close to the lowest possible 26.00 but never 0: the same input, is, has four different correct answers, so the best the softmax can do is share its probability between them.
- The four most likely neighbours of is are its context words to, related, company and ineuron, holding 0.991 of the probability between them. After 300 passes the shares are still unequal; the lowest loss would give each of them 0.25.
Choosing skip-gram in gensim
In gensim the only change from CBOW is sg=1; the vectors have the same shape.
from gensim.models import Word2Vec
sentence = [["ineuron", "company", "is", "related", "to", "data", "science"]]
cbow = Word2Vec(sentence, sg=0, window=2, vector_size=5, min_count=1, seed=42, workers=1)
skip = Word2Vec(sentence, sg=1, window=2, vector_size=5, min_count=1, seed=42, workers=1)
print("CBOW sg =", cbow.sg, " vector:", cbow.wv["is"].shape)
print("skip-gram sg =", skip.sg, " vector:", skip.wv["is"].shape)CBOW sg = 0 vector: (5,) skip-gram sg = 1 vector: (5,)
CBOW vs skip-gram
The comparison below is the one given on the original word2vec project page by its authors.
| CBOW | Skip-gram | |
|---|---|---|
| Predicts | the centre word from its context | each context word from the centre word |
| Training examples per window | 1 | one per context word |
| Speed | several times faster to train | slower |
| Strength | slightly better accuracy for frequent words | works well with a small amount of training data; represents rare words and phrases well |
| gensim | sg=0 (the default) | sg=1 |
Where you use skip-gram
- Small corpora, such as a few thousand domain documents.
- Rare words that matter, such as drug names in medical notes or product codes in tickets.
- Phrases: the 2013 paper that introduced phrase vectors and negative sampling, “Distributed Representations of Words and Phrases and their Compositionality”, builds on skip-gram.
window=5 each word can produce up to 10 pairs. On the same text it trains several times slower than CBOW, so budget the time on a large corpus.Related
- Previous: CBOW (continuous bag of words)
- Next: Training Word2Vec with gensim
- Reference: The word2vec project page
- Set
half = 1in the NumPy example. How many pairs are there now, and which neighbours does is get? - Print
probs(words.index("ineuron"))after training. Why are only two words likely? - Train the gensim model with
sg=1, window=1and compare it withwindow=2: does the vector shape change?
This is what real progress feels like.