Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Parts of speech (POS) tagging

Part-of-speech (POS) tagging is the task of labelling each word in a sentence with its grammatical class, such as noun, verb or adjective, using a tag set like the Penn Treebank's NN, VB and JJ.

Last updated: 07 Oct, 2026 · NLTK 3.10 · scikit-learn 1.9

Lemmatization needs to know whether a word is a verb or a noun to find its lemma. POS tagging supplies that, and it is useful on its own wherever the grammar of a sentence matters.

Downloading the tagger

python
import nltk
nltk.download("punkt_tab")
nltk.download("averaged_perceptron_tagger_eng")    # the English tagger

Reading the Penn Treebank tags

The POS tag list · from the Complete NLP Machine Learning in One Shot video · 1:34:40 to 1:35:49

NLTK categorises each word of a sentence into one of these parts of speech automatically: CC coordinating conjunction, CD cardinal digit, DT determiner, EX existential there, FW foreign word, IN preposition, JJ adjective, JJR comparative adjective, and so on. Personal pronouns such as I, he and she get the tag PRP; adverbs such as very and silently RB, a comparative adverb such as better RBR, a superlative RBS. The exercise: tag the sentence "Taj Mahal is a beautiful Monument".

The full list, from the video's notebook:

TagMeaningExample
CCcoordinating conjunctionand, but
CDcardinal digitthree, 1857
DTdeterminera, the
EXexistential therethere is
FWforeign word
INpreposition or subordinating conjunctionin, because
JJadjectivebig
JJRadjective, comparativebigger
JJSadjective, superlativebiggest
LSlist marker1)
MDmodalcould, will
NNnoun, singulardesk
NNSnoun, pluraldesks
NNPproper noun, singularHarrison
NNPSproper noun, pluralAmericans
PDTpredeterminerall the kids
POSpossessive endingparent's
PRPpersonal pronounI, he, she
PRP$possessive pronounmy, his, hers
RBadverbvery, silently
RBRadverb, comparativebetter
RBSadverb, superlativebest
RPparticlegive up
TOtoto go to the store
UHinterjectionerrrrrrrrm
VBverb, base formtake
VBDverb, past tensetook
VBGverb, gerund or present participletaking
VBNverb, past participletaken
VBPverb, present, not 3rd person singulartake
VBZverb, 3rd person singular presenttakes
WDTwh-determinerwhich
WPwh-pronounwho, what
WP$possessive wh-pronounwhose
WRBwh-adverbwhere, when

Tagging the Taj Mahal sentence

nltk.pos_tag takes a list of words and returns a list of (word, tag) pairs. Split the sentence first, with .split() or word_tokenize:

python
nltk.pos_tag("Taj Mahal is a beautiful Monument".split())    # a list of (word, tag) pairs
ExampleFrom the video, run on NLTK 3.10.3
import nltk

print(nltk.pos_tag("Taj Mahal is a beautiful Monument".split()))
print(nltk.pos_tag(nltk.word_tokenize("Taj Mahal is a beautiful Monument")))
Taj Mahal is a beautiful Monument tagged: Taj and Mahal are NNP, proper nouns, is is VBZ, a is DT, beautiful is JJ and Monument is NN, a singular noun.

Taj and Mahal are NNP, proper nouns, because they name one place; Monument is NN, a common noun; beautiful is JJ, an adjective; is is VBZ, a verb in the third person singular present; a is DT, a determiner. Both ways of splitting give the same tags.

Passing a string instead of a list

Give pos_tag the sentence itself and today's NLTK stops with an error:

ExampleRun on NLTK 3.10.3
import nltk

nltk.pos_tag("Taj Mahal is a beautiful Monument")

A string is a sequence of characters, so tagging it would label T, a and j one by one. NLTK checks the input and asks for a list of strings: pass .split() or word_tokenize output.

Tagging the Kalam speech

The video then tags the speech from Stopwords sentence by sentence. The tagger reads each word's neighbours to choose a tag, so here every sentence keeps its stopwords: pos_tag_sents tags a list of tokenized sentences in one call. The second half of the example removes the stopwords first, for comparison.

python
sentences = nltk.sent_tokenize(paragraph)
tagged = nltk.pos_tag_sents([nltk.word_tokenize(s) for s in sentences])    # one list per sentence
ExampleFrom the video, run on NLTK 3.10.3
import nltk
from nltk.corpus import stopwords

sentences = nltk.sent_tokenize(paragraph)
tagged = nltk.pos_tag_sents([nltk.word_tokenize(s) for s in sentences])
print(tagged[0])
print(tagged[2])

stop = set(stopwords.words("english"))
kept = [w for w in nltk.word_tokenize(sentences[2]) if w.lower() not in stop]
full = dict(tagged[2])
print([(w, full[w], t) for w, t in nltk.pos_tag(kept) if full[w] != t])

What the speech's tags show

  • "I have three visions for India.": I is PRP, a personal pronoun; three CD, a cardinal number; visions NNS, a plural noun; India NNP, a singular proper noun; have VBP; for IN.
  • Proper nouns: Alexander, Greeks, Turks and Moguls are NNP. NNPS is the plural form, as in Americans or Indians.
  • Tags are a model's guesses: British and French get JJ, the adjective tag, although "the British" here names a people. The tagger picks the most likely tag from the word and its neighbours, and it is not always right.
  • Without the stopwords, three tags change: British and French become NNP, and looted, a past-tense verb (VBD) in the full sentence, becomes JJ, an adjective, which is wrong. Removing the small words changes the context the tagger reads, so tag first and filter afterwards.

Penn Treebank tags vs WordNet POS

Penn Treebank tagsMeaningWordNet POS for lemmatize
NN, NNS, NNP, NNPSnounsn
VB, VBD, VBG, VBN, VBP, VBZverbsv
JJ, JJR, JJSadjectivesa
RB, RBR, RBSadverbsr
DT, IN, CC, PRP, CD and the restfunction words, numbersno WordNet entry; n is the safe default

Where you use POS tagging

  • Lemmatization: the tag picks the right lemma, as in Lemmatization.
  • Named entity recognition: the NER chunker reads the tags, as the next lesson shows.
  • Extracting phrases: adjective-noun pairs (JJ NN) such as "beautiful Monument" give the aspects reviews talk about.
Watch out. The tagger is a statistical model, an averaged perceptron trained on newspaper text. It makes mistakes, and more of them on short fragments, lower-cased text or text with the stopwords removed. Tag full, cleanly cased sentences.
Try it yourself
  • Tag "I will book a flight" and "I read a book" and compare the tags of book.
  • Lower-case the Taj Mahal sentence before tagging and see which tags change.
  • Run nltk.help.upenn_tagset("VB.*") after nltk.download("tagsets_json") to read NLTK's own description of the verb tags.
PreviousStopwords

This is what real progress feels like.