Text cleaning and normalisation
Text cleaning is the preprocessing step that removes noise from raw text, such as HTML, links and stray symbols, and normalisation rewrites what remains into one consistent form, such as lower case with contractions expanded, so the same word always becomes the same token.
Last updated: 07 Oct, 2026 · NLTK 3.10 · scikit-learn 1.9
Real text is messy. The board notes list the first preprocessing stage as tokenization, lower-casing the words and regular expressions, before stemming, lemmatization and stopwords. Tokenization with NLTK covered the first item; this lesson covers the other two.
The examples are real messages from the SMS spam dataset in the course materials, the one the spam classifiers in part 5 learn from, and a sentence from the Kalam speech used in the stopwords lesson.
Lower-casing the words
To a computer, "Free", "FREE" and "free" are three different strings, so they would become three vocabulary entries with their counts split. str.lower() merges them. It also merges words you may want apart, such as the name "Apple" and the fruit "apple"; for most classification tasks the merge helps, and for named entities, which rely on capital letters, it hurts.
"WINNER!! As a valued network customer".lower() # 'winner!! as a valued network customer'Removing HTML, links and symbols with regular expressions
A regular expression is a pattern for text, and re.sub(pattern, replacement, text) replaces every match. The patterns used here:
| Pattern | Matches | Replaced with |
|---|---|---|
<[^>]+> | an HTML tag such as <p> | a space |
https?://\S+|www\.\S+ | a link up to the next space | the token url |
\d[\d,]* | a number such as 100,000 or 87121 | the token num |
[^a-z\s] | anything that is not a lower-case letter or a space | a space |
Two more fixes come from the standard library. html.unescape turns entities such as & and < back into & and <; the SMS dataset is full of them. Curly apostrophes (’) become straight ones ('), because the tokenizer treats them differently.
Straightening curly apostrophes
The Kalam speech is typed with curly apostrophes, and word_tokenize splits "Isn’t" into Isn, ’ and t. With a straight apostrophe it gives Is and n't, the Penn Treebank split that keeps the negation as one token:
from nltk.tokenize import word_tokenize
print(word_tokenize("Isn’t this incorrect? My second vision for India’s development."))
print(word_tokenize("Isn't this incorrect? My second vision for India's development."))['Isn', '’', 't', 'this', 'incorrect', '?', 'My', 'second', 'vision', 'for', 'India', '’', 's', 'development', '.'] ['Is', "n't", 'this', 'incorrect', '?', 'My', 'second', 'vision', 'for', 'India', "'s", 'development', '.']
Expanding contractions
"don't", "I'm" and "we're" each hide two words, and after punctuation is stripped they turn into fragments: don t, i m, we re. Expanding them first with a small dictionary keeps whole words, including the "not" that decides whether a review is positive or negative. \b in the pattern is a word boundary, so "i'm" is replaced only as a whole word.
CONTRACTIONS = {"don't": "do not", "i'm": "i am", "we're": "we are"}
for short, full in CONTRACTIONS.items():
text = re.sub(r"\b" + re.escape(short) + r"\b", full, text)Putting the steps in order
The order matters: unescape before removing symbols, lower-case before expanding contractions (the dictionary is lower case), and expand contractions before stripping punctuation. The steps on one example each:
Cleaning real SMS messages
The whole function, run on three messages from the dataset and the Kalam sentence:
import re
import html
CONTRACTIONS = {"don't": "do not", "isn't": "is not", "i'm": "i am", "we're": "we are",
"i'll": "i will", "there's": "there is", "won't": "will not", "can't": "cannot"}
def clean(text):
text = html.unescape(text) # & -> &
text = re.sub(r"<[^>]+>", " ", text) # HTML tags
text = re.sub(r"https?://\S+|www\.\S+", " url ", text) # links
text = text.replace("\u2019", "'").lower() # straight quotes, lower case
for short, full in CONTRACTIONS.items():
text = re.sub(r"\b" + re.escape(short) + r"\b", full, text)
text = re.sub(r"\d[\d,]*", " num ", text) # every number -> num
text = re.sub(r"[^a-z\s]", " ", text) # letters only
return " ".join(text.split()) # one space between words
messages = [
"URGENT! You have won a 1 week FREE membership in our £100,000 Prize Jackpot! Txt the word: CLAIM to No: 81010 T&C www.dbuk.net LCCLTD POBOX 4403LDNW1A7RW18",
"I'm back & we're packing the car now, I'll let you know if there's room",
"Nah I don't think he goes to usf, he lives around here though",
"<p>Isn’t this incorrect?</p>",
]
for m in messages:
print(clean(m))urgent you have won a num week free membership in our num prize jackpot txt the word claim to no num t c url lccltd pobox num ldnw num a num rw num i am back we are packing the car now i will let you know if there is room nah i do not think he goes to usf he lives around here though is not this incorrect
What the cleaning changed
- The spam message keeps its signal words, urgent, won, free, prize, claim, and the link becomes url. Every number, including the prize and the short code, becomes num, so "won num" looks the same in every prize message. The postcode-like code at the end breaks into num ldnw num a num rw num: noise that a stricter rule could drop.
- & became & and was then removed with the other symbols.
- I'm, we're, I'll and there's became i am, we are, i will and there is.
- don't became do not, so the negation survives as the word not.
- The Kalam sentence lost its tags and gained a straight apostrophe, and Isn't became is not.
Seeing the wrong order fail
Strip punctuation before unescaping and expanding, and the entity and the contractions turn into junk tokens:
import re
text = "I'm back & we're packing"
print(re.sub(r"[^a-z\s]", " ", text.lower()).split())['i', 'm', 'back', 'amp', 'we', 're', 'packing']
The output has amp, a fragment of the entity, and m and re, fragments of the contractions. All three would become vocabulary entries of their own.
Cleaning vs normalisation
| Cleaning | Normalisation | |
|---|---|---|
| Goal | Remove what is not language | Write the same word one way |
| Examples | HTML tags, links, stray symbols, extra spaces | Lower case, straight quotes, expanded contractions, num for numbers |
| Effect on the vocabulary | Drops junk tokens | Merges variants of one word |
| Later steps of the same kind | Stopword removal | Stemming and lemmatization |
Where you use text cleaning
- Spam and sentiment classifiers with bag of words or TF-IDF, where every distinct string is a column.
- Scraped web text, which arrives with tags, entities and links.
- Search indexes, so that a query in any case finds the same documents.
A-z is not the same as A-Za-z. The range from A to z also covers [ \ ] ^ _ and the backtick, so re.sub('[^a-zA-z]', ' ', 'update_now') keeps the underscore. Write [^a-zA-Z].Related
- Previous: Tokenization with NLTK
- Next: Stemming
- Reference: re and html in the Python documentation
- Add
"can't"to a message and check thatcleangives cannot. - Change the number rule to replace numbers with nothing instead of num and compare the spam message.
- Remove the
.lower()call and runcleanagain: which words disappear, and why?
Every expert started right here.