Parts of speech (POS) tagging
Part-of-speech (POS) tagging is the task of labelling each word in a sentence with its grammatical class, such as noun, verb or adjective, using a tag set like the Penn Treebank's NN, VB and JJ.
Last updated: 07 Oct, 2026 · NLTK 3.10 · scikit-learn 1.9
Lemmatization needs to know whether a word is a verb or a noun to find its lemma. POS tagging supplies that, and it is useful on its own wherever the grammar of a sentence matters.
Downloading the tagger
import nltk
nltk.download("punkt_tab")
nltk.download("averaged_perceptron_tagger_eng") # the English taggerReading the Penn Treebank tags
NLTK categorises each word of a sentence into one of these parts of speech automatically: CC coordinating conjunction, CD cardinal digit, DT determiner, EX existential there, FW foreign word, IN preposition, JJ adjective, JJR comparative adjective, and so on. Personal pronouns such as I, he and she get the tag PRP; adverbs such as very and silently RB, a comparative adverb such as better RBR, a superlative RBS. The exercise: tag the sentence "Taj Mahal is a beautiful Monument".
The full list, from the video's notebook:
| Tag | Meaning | Example |
|---|---|---|
CC | coordinating conjunction | and, but |
CD | cardinal digit | three, 1857 |
DT | determiner | a, the |
EX | existential there | there is |
FW | foreign word | |
IN | preposition or subordinating conjunction | in, because |
JJ | adjective | big |
JJR | adjective, comparative | bigger |
JJS | adjective, superlative | biggest |
LS | list marker | 1) |
MD | modal | could, will |
NN | noun, singular | desk |
NNS | noun, plural | desks |
NNP | proper noun, singular | Harrison |
NNPS | proper noun, plural | Americans |
PDT | predeterminer | all the kids |
POS | possessive ending | parent's |
PRP | personal pronoun | I, he, she |
PRP$ | possessive pronoun | my, his, hers |
RB | adverb | very, silently |
RBR | adverb, comparative | better |
RBS | adverb, superlative | best |
RP | particle | give up |
TO | to | to go to the store |
UH | interjection | errrrrrrrm |
VB | verb, base form | take |
VBD | verb, past tense | took |
VBG | verb, gerund or present participle | taking |
VBN | verb, past participle | taken |
VBP | verb, present, not 3rd person singular | take |
VBZ | verb, 3rd person singular present | takes |
WDT | wh-determiner | which |
WP | wh-pronoun | who, what |
WP$ | possessive wh-pronoun | whose |
WRB | wh-adverb | where, when |
Tagging the Taj Mahal sentence
nltk.pos_tag takes a list of words and returns a list of (word, tag) pairs. Split the sentence first, with .split() or word_tokenize:
nltk.pos_tag("Taj Mahal is a beautiful Monument".split()) # a list of (word, tag) pairsimport nltk
print(nltk.pos_tag("Taj Mahal is a beautiful Monument".split()))
print(nltk.pos_tag(nltk.word_tokenize("Taj Mahal is a beautiful Monument")))[('Taj', 'NNP'), ('Mahal', 'NNP'), ('is', 'VBZ'), ('a', 'DT'), ('beautiful', 'JJ'), ('Monument', 'NN')]
[('Taj', 'NNP'), ('Mahal', 'NNP'), ('is', 'VBZ'), ('a', 'DT'), ('beautiful', 'JJ'), ('Monument', 'NN')]Taj and Mahal are NNP, proper nouns, because they name one place; Monument is NN, a common noun; beautiful is JJ, an adjective; is is VBZ, a verb in the third person singular present; a is DT, a determiner. Both ways of splitting give the same tags.
Passing a string instead of a list
Give pos_tag the sentence itself and today's NLTK stops with an error:
import nltk
nltk.pos_tag("Taj Mahal is a beautiful Monument")Traceback (most recent call last):
File "main.py", line 3, in <module>
nltk.pos_tag("Taj Mahal is a beautiful Monument")
TypeError: tokens: expected a list of strings, got a stringA string is a sequence of characters, so tagging it would label T, a and j one by one. NLTK checks the input and asks for a list of strings: pass .split() or word_tokenize output.
Tagging the Kalam speech
The video then tags the speech from Stopwords sentence by sentence. The tagger reads each word's neighbours to choose a tag, so here every sentence keeps its stopwords: pos_tag_sents tags a list of tokenized sentences in one call. The second half of the example removes the stopwords first, for comparison.
sentences = nltk.sent_tokenize(paragraph)
tagged = nltk.pos_tag_sents([nltk.word_tokenize(s) for s in sentences]) # one list per sentenceimport nltk
from nltk.corpus import stopwords
sentences = nltk.sent_tokenize(paragraph)
tagged = nltk.pos_tag_sents([nltk.word_tokenize(s) for s in sentences])
print(tagged[0])
print(tagged[2])
stop = set(stopwords.words("english"))
kept = [w for w in nltk.word_tokenize(sentences[2]) if w.lower() not in stop]
full = dict(tagged[2])
print([(w, full[w], t) for w, t in nltk.pos_tag(kept) if full[w] != t])[('I', 'PRP'), ('have', 'VBP'), ('three', 'CD'), ('visions', 'NNS'), ('for', 'IN'), ('India', 'NNP'), ('.', '.')]
[('From', 'IN'), ('Alexander', 'NNP'), ('onwards', 'NNS'), (',', ','), ('the', 'DT'), ('Greeks', 'NNP'), (',', ','), ('the', 'DT'), ('Turks', 'NNP'), (',', ','), ('the', 'DT'), ('Moguls', 'NNP'), (',', ','), ('the', 'DT'), ('Portuguese', 'NNP'), (',', ','), ('the', 'DT'), ('British', 'JJ'), (',', ','), ('the', 'DT'), ('French', 'JJ'), (',', ','), ('the', 'DT'), ('Dutch', 'NNP'), (',', ','), ('all', 'DT'), ('of', 'IN'), ('them', 'PRP'), ('came', 'VBD'), ('and', 'CC'), ('looted', 'VBD'), ('us', 'PRP'), (',', ','), ('took', 'VBD'), ('over', 'RP'), ('what', 'WP'), ('was', 'VBD'), ('ours', 'PRP'), ('.', '.')]
[('British', 'JJ', 'NNP'), ('French', 'JJ', 'NNP'), ('looted', 'VBD', 'JJ')]What the speech's tags show
- "I have three visions for India.": I is PRP, a personal pronoun; three CD, a cardinal number; visions NNS, a plural noun; India NNP, a singular proper noun; have VBP; for IN.
- Proper nouns: Alexander, Greeks, Turks and Moguls are NNP. NNPS is the plural form, as in Americans or Indians.
- Tags are a model's guesses: British and French get JJ, the adjective tag, although "the British" here names a people. The tagger picks the most likely tag from the word and its neighbours, and it is not always right.
- Without the stopwords, three tags change: British and French become NNP, and looted, a past-tense verb (VBD) in the full sentence, becomes JJ, an adjective, which is wrong. Removing the small words changes the context the tagger reads, so tag first and filter afterwards.
Penn Treebank tags vs WordNet POS
| Penn Treebank tags | Meaning | WordNet POS for lemmatize |
|---|---|---|
| NN, NNS, NNP, NNPS | nouns | n |
| VB, VBD, VBG, VBN, VBP, VBZ | verbs | v |
| JJ, JJR, JJS | adjectives | a |
| RB, RBR, RBS | adverbs | r |
| DT, IN, CC, PRP, CD and the rest | function words, numbers | no WordNet entry; n is the safe default |
Where you use POS tagging
- Lemmatization: the tag picks the right lemma, as in Lemmatization.
- Named entity recognition: the NER chunker reads the tags, as the next lesson shows.
- Extracting phrases: adjective-noun pairs (JJ NN) such as "beautiful Monument" give the aspects reviews talk about.
Related
- Previous: Stopwords
- Next: Named entity recognition (NER)
- Notebook: 8-Parts Of Speech Tagging
- Reference: nltk.tag and the Penn Treebank tag set
- Tag "I will book a flight" and "I read a book" and compare the tags of book.
- Lower-case the Taj Mahal sentence before tagging and see which tags change.
- Run
nltk.help.upenn_tagset("VB.*")afternltk.download("tagsets_json")to read NLTK's own description of the verb tags.
This is what real progress feels like.