Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Skip-gram

Skip-gram is the Word2Vec architecture that predicts the words around a centre word from the centre word, the reverse of CBOW.

Last updated: 07 Oct, 2026 · gensim 4.4 · NumPy

CBOW (continuous bag of words) took the context words in and predicted the centre word. Skip-gram keeps the same windows and swaps the two sides.

Skip-gram is CBOW with input and output swapped · from the Complete NLP Machine Learning in One Shot video · 3:29:55 to 3:31:00

Reversing CBOW: the centre word in, the context words out

The video takes the same corpus, “iNeuron company is related to data science”, and the same window size of 5. Everything stays the same except that the input and the output change places: the centre word is the input, and the four words around it are the output.

Input (centre word)Output (context words)
isiNeuron, company, related, to
relatedcompany, is, to, data
tois, related, data, science

Turning each window into training pairs

Skip-gram trains on (centre, context) pairs: from the first window, (is, iNeuron), (is, company), (is, related) and (is, to). The three windows of the board give 12 pairs. Taking every word of the sentence as a centre, with up to two neighbours on each side, gives more, because the words at the ends have fewer neighbours.

Passing the centre word through the skip-gram network

  • Input: the one-hot vector of the centre word, for is [0 0 1 0 0 0 0].
  • Input matrix W (7 × N): the one-hot vector picks out the row of is, and that row is the hidden layer h. No bias, no activation.
  • Output matrix W′ (N × 7) and a softmax give a probability for every vocabulary word.
  • Loss: the same probabilities are scored against each context word, so the loss is the sum of −log p(context) over the four neighbours. For the target iNeuron, y is [1 0 0 0 0 0 0].
The skip-gram loss for one centre word: one term per context word

The board draws one output box per context word. They all use the same W′; there is one input matrix and one output matrix for the whole model, as in CBOW.

Skip-gram on the same corpus: the one-hot vector of is picks its row of the 7 by N matrix W, that row times W prime and a softmax over 7 words should give high probability to the context words iNeuron, company, related and to; the table shows is, related and to as inputs with their four context words as outputs.

Training skip-gram on the sentence with NumPy

The example builds the pairs and trains the two matrices with a full softmax. Each step takes one centre word with all its context words at once, as in the board's picture of one input and four outputs. Then it asks which words the model expects around is.

ExampleFrom the video, run with NumPy
import numpy as np

words = ["ineuron", "company", "is", "related", "to", "data", "science"]
V, N, half = len(words), 5, 2

board = []                                   # the pairs from the board's three windows
for c in range(half, V - half):
    board += [(c, j) for j in range(c - half, c + half + 1) if j != c]
contexts = {c: [j for j in range(max(0, c - half), min(V, c + half + 1)) if j != c] for c in range(V)}
print("pairs from the board's windows:", len(board), " pairs from every position:", sum(map(len, contexts.values())))
print([(words[c], words[j]) for c, j in board[:4]])

rng = np.random.default_rng(42)
W = rng.normal(0, 0.1, (V, N))
W_out = rng.normal(0, 0.1, (N, V))

def probs(c):
    u = W[c] @ W_out                         # h is the centre word's row of W
    p = np.exp(u - u.max())
    return p / p.sum()

lr = 0.2
for epoch in range(300):
    loss = 0.0
    for c, ctx in contexts.items():
        h, p = W[c].copy(), probs(c)
        loss -= np.log(p[ctx]).sum()          # one term per context word
        e = len(ctx) * p
        e[ctx] -= 1                           # gradient of the summed loss with respect to the scores
        W[c] -= lr * (W_out @ e)
        W_out -= lr * np.outer(h, e)
    if epoch in (0, 299):
        print(f"epoch {epoch + 1}: loss {loss:.4f}")
print("lowest possible loss:", round(sum(len(x) * np.log(len(x)) for x in contexts.values()), 4))

p = probs(words.index("is"))
top = np.argsort(p)[::-1][:4]
print("most likely neighbours of 'is':", [(words[k], round(float(p[k]), 3)) for k in top], " total:", round(float(p[top].sum()), 3))

What the skip-gram run shows

  • 12 pairs come from the board's three windows and 22 from every position of the sentence.
  • The first four pairs all have is as the input: one training example per context word.
  • The loss falls from 42.80 to 27.40, close to the lowest possible 26.00 but never 0: the same input, is, has four different correct answers, so the best the softmax can do is share its probability between them.
  • The four most likely neighbours of is are its context words to, related, company and ineuron, holding 0.991 of the probability between them. After 300 passes the shares are still unequal; the lowest loss would give each of them 0.25.

Choosing skip-gram in gensim

In gensim the only change from CBOW is sg=1; the vectors have the same shape.

ExampleRun on gensim 4.4.0
from gensim.models import Word2Vec

sentence = [["ineuron", "company", "is", "related", "to", "data", "science"]]
cbow = Word2Vec(sentence, sg=0, window=2, vector_size=5, min_count=1, seed=42, workers=1)
skip = Word2Vec(sentence, sg=1, window=2, vector_size=5, min_count=1, seed=42, workers=1)
print("CBOW sg =", cbow.sg, " vector:", cbow.wv["is"].shape)
print("skip-gram sg =", skip.sg, " vector:", skip.wv["is"].shape)

CBOW vs skip-gram

The comparison below is the one given on the original word2vec project page by its authors.

CBOWSkip-gram
Predictsthe centre word from its contexteach context word from the centre word
Training examples per window1one per context word
Speedseveral times faster to trainslower
Strengthslightly better accuracy for frequent wordsworks well with a small amount of training data; represents rare words and phrases well
gensimsg=0 (the default)sg=1

Where you use skip-gram

  • Small corpora, such as a few thousand domain documents.
  • Rare words that matter, such as drug names in medical notes or product codes in tickets.
  • Phrases: the 2013 paper that introduced phrase vectors and negative sampling, “Distributed Representations of Words and Phrases and their Compositionality”, builds on skip-gram.
Watch out. Skip-gram makes a training pair for every context word, so with gensim's default window=5 each word can produce up to 10 pairs. On the same text it trains several times slower than CBOW, so budget the time on a large corpus.
Try it yourself
  • Set half = 1 in the NumPy example. How many pairs are there now, and which neighbours does is get?
  • Print probs(words.index("ineuron")) after training. Why are only two words likely?
  • Train the gensim model with sg=1, window=1 and compare it with window=2: does the vector shape change?

This is what real progress feels like.