Named entity recognition (NER)
Named entity recognition (NER) is the task of finding the names of real-world things in text, such as people, organisations and places, and labelling each one with its entity type.
Last updated: 07 Oct, 2026 · NLTK 3.10 · scikit-learn 1.9
Parts of speech (POS) tagging labels every word with its grammar. NER builds on those tags to find which words, alone or in groups, name something.
Downloading the chunker
import nltk
for resource in ["punkt_tab", "averaged_perceptron_tagger_eng", "maxent_ne_chunker_tab", "words"]:
nltk.download(resource)Tagging named entities with ne_chunk
The video's sentence: "The Eiffel Tower was built from 1887 to 1889 by French engineer Gustave Eiffel, whose company specialized in building metal frameworks and structures." POS tagging says which words are nouns; named entity tags say which are names, and of what kind.
The entity types the video lists, with the notebook's examples: person, place or location (India), date (September, 24-09-1989), time (4:30pm), money (1 million dollar), organization (iNeuron Private Limited) and percent (20%, twenty percent). So the Eiffel Tower should come out as a place, Gustave Eiffel as a person, and an amount such as $1 million as money.
In NLTK the steps are the ones from the earlier lessons: word_tokenize the sentence, pos_tag the words into tag_elements, then pass them to nltk.ne_chunk, NLTK's recommended named entity chunker. It needs its model downloaded, because it chunks the tagged words with a trained model.
words = nltk.word_tokenize(sentence)
tag_elements = nltk.pos_tag(words)
tree = nltk.ne_chunk(tag_elements) # a Tree: entities are subtreesThe clip downloads maxent_ne_chunker; NLTK 3.10 loads maxent_ne_chunker_tab and the words list, as in the download cell above. The video shows the tree with .draw(), which opens a desktop window; print(tree) shows the same tree as text and also works in Colab.
Chunking the Eiffel Tower sentence
sentence = "The Eiffel Tower was built from 1887 to 1889 by French engineer Gustave Eiffel, whose company specialized in building metal frameworks and structures."
import nltk
words = nltk.word_tokenize(sentence)
tag_elements = nltk.pos_tag(words)
tree = nltk.ne_chunk(tag_elements)
print(tree)
for subtree in tree.subtrees(lambda t: t.label() != "S"):
print(subtree.label(), "->", " ".join(w for w, t in subtree.leaves()))(S The/DT (ORGANIZATION Eiffel/NNP Tower/NNP) was/VBD built/VBN from/IN 1887/CD to/TO 1889/CD by/IN (GPE French/JJ) engineer/NN (PERSON Gustave/NNP Eiffel/NNP) ,/, whose/WP$ company/NN specialized/VBD in/IN building/NN metal/NN frameworks/NNS and/CC structures/NNS ./.) ORGANIZATION -> Eiffel Tower GPE -> French PERSON -> Gustave Eiffel
Reading the entity tree
- S is the whole sentence. Plain words hang under it with their POS tags (was/VBD, 1887/CD); each entity is a subtree with a label.
- PERSON: Gustave Eiffel, the engineer, correct.
- GPE: French. GPE means geo-political entity, a country, city or state; the ACE scheme the chunker was trained on labels a nationality word such as French as GPE.
- ORGANIZATION: Eiffel Tower is a mistake. The tower is a structure; the chunker's types include FACILITY for buildings and monuments, but its model chose ORGANIZATION.
- 1887 and 1889 stay plain CD numbers. NLTK's chunker has six entity types, PERSON, ORGANIZATION, GPE, LOCATION, FACILITY and GSP (geo-social-political group), and none of them is a date.
Seeing where the chunker breaks
Two more runs: a second sentence with a date and a company, and the Eiffel sentence in lower case.
sentence = "The Eiffel Tower was built from 1887 to 1889 by French engineer Gustave Eiffel, whose company specialized in building metal frameworks and structures."
import nltk
s2 = "Sundar Pichai became the CEO of Google in August 2015, and the company is based in Mountain View, California."
for text in [s2, sentence.lower()]:
tree = nltk.ne_chunk(nltk.pos_tag(nltk.word_tokenize(text)))
print([(t.label(), " ".join(w for w, _ in t.leaves())) for t in tree.subtrees(lambda t: t.label() != "S")])
binary = nltk.ne_chunk(nltk.pos_tag(nltk.word_tokenize(sentence)), binary=True)
print([" ".join(w for w, _ in t.leaves()) for t in binary.subtrees(lambda t: t.label() == "NE")])[('PERSON', 'Sundar'), ('PERSON', 'Pichai'), ('ORGANIZATION', 'CEO'), ('GPE', 'Google'), ('GPE', 'Mountain View'), ('GPE', 'California')]
[]
['Eiffel Tower', 'French', 'Gustave Eiffel']- The second sentence shows several errors at once: Sundar and Pichai come out as two separate PERSON entities, CEO as an ORGANIZATION, and Google as a GPE instead of an organisation. Mountain View and California are correct GPEs, and August 2015 is not tagged.
- The lower-cased sentence gives an empty list: the chunker leans on capital letters, so lower-casing before NER removes every entity.
binary=Truedrops the types and marks each entity as NE: Eiffel Tower, French and Gustave Eiffel.
Finding entities with spaCy
The video's list includes dates, times, money and percentages, which NLTK's chunker cannot label. Trained pipelines such as spaCy's English models use the OntoNotes scheme of 18 types, which has DATE, TIME, MONEY, PERCENT and NORP (nationalities and religious or political groups) as well as PERSON, ORG, GPE and FAC for buildings. The model is a separate download:
import spacy # pip install spacy
nlp = spacy.load("en_core_web_sm") # python -m spacy download en_core_web_sm
doc = nlp(sentence)
print([(ent.text, ent.label_) for ent in doc.ents]) # (text, label) for each entityIt prints a list of (text, label) pairs, one per entity spaCy finds, with labels from those 18 types.
NLTK ne_chunk vs spaCy NER
| NLTK ne_chunk | spaCy (en_core_web_sm) | |
|---|---|---|
| Entity types | 6 (ACE): PERSON, ORGANIZATION, GPE, LOCATION, FACILITY, GSP | 18 (OntoNotes), including DATE, TIME, MONEY, PERCENT, NORP |
| Input | POS-tagged tokens | Raw text: nlp(text) |
| Model | A maximum entropy classifier | A neural network pipeline |
| Dates and money | Not labelled | Labelled |
| Output | A Tree with entity subtrees | doc.ents, spans with .label_ |
Where you use NER
- Reading documents: names, companies, places and amounts from invoices, contracts and CVs.
- Search and news: tag articles with the people and organisations they mention.
- Chatbots: pick the city and the date out of "book a flight to Paris on Friday".
Related
- Previous: Parts of speech (POS) tagging
- Next: One-hot encoding for text
- Notebook: 9-Named Entity Recognition
- Reference: NLTK book, chapter 7: extracting information from text and spaCy named entities
- Remove "French engineer" from the sentence, as the notebook's version does, and see which entity disappears.
- Chunk "Mumbai is the capital of Maharashtra and Tata Steel is based there." and check each label.
- Count the entities of each type in the Kalam speech with
collections.Counter.
Every expert started right here.