Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Text cleaning and normalisation

Text cleaning is the preprocessing step that removes noise from raw text, such as HTML, links and stray symbols, and normalisation rewrites what remains into one consistent form, such as lower case with contractions expanded, so the same word always becomes the same token.

Last updated: 07 Oct, 2026 · NLTK 3.10 · scikit-learn 1.9

Real text is messy. The board notes list the first preprocessing stage as tokenization, lower-casing the words and regular expressions, before stemming, lemmatization and stopwords. Tokenization with NLTK covered the first item; this lesson covers the other two.

The examples are real messages from the SMS spam dataset in the course materials, the one the spam classifiers in part 5 learn from, and a sentence from the Kalam speech used in the stopwords lesson.

Lower-casing the words

To a computer, "Free", "FREE" and "free" are three different strings, so they would become three vocabulary entries with their counts split. str.lower() merges them. It also merges words you may want apart, such as the name "Apple" and the fruit "apple"; for most classification tasks the merge helps, and for named entities, which rely on capital letters, it hurts.

python
"WINNER!! As a valued network customer".lower()    # 'winner!! as a valued network customer'

A regular expression is a pattern for text, and re.sub(pattern, replacement, text) replaces every match. The patterns used here:

PatternMatchesReplaced with
<[^>]+>an HTML tag such as <p>a space
https?://\S+|www\.\S+a link up to the next spacethe token url
\d[\d,]*a number such as 100,000 or 87121the token num
[^a-z\s]anything that is not a lower-case letter or a spacea space

Two more fixes come from the standard library. html.unescape turns entities such as &amp; and &lt; back into & and <; the SMS dataset is full of them. Curly apostrophes (’) become straight ones ('), because the tokenizer treats them differently.

Straightening curly apostrophes

The Kalam speech is typed with curly apostrophes, and word_tokenize splits "Isn’t" into Isn, ’ and t. With a straight apostrophe it gives Is and n't, the Penn Treebank split that keeps the negation as one token:

ExampleRun on NLTK 3.10.3
from nltk.tokenize import word_tokenize

print(word_tokenize("Isn’t this incorrect? My second vision for India’s development."))
print(word_tokenize("Isn't this incorrect? My second vision for India's development."))

Expanding contractions

"don't", "I'm" and "we're" each hide two words, and after punctuation is stripped they turn into fragments: don t, i m, we re. Expanding them first with a small dictionary keeps whole words, including the "not" that decides whether a review is positive or negative. \b in the pattern is a word boundary, so "i'm" is replaced only as a whole word.

python
CONTRACTIONS = {"don't": "do not", "i'm": "i am", "we're": "we are"}
for short, full in CONTRACTIONS.items():
    text = re.sub(r"\b" + re.escape(short) + r"\b", full, text)

Putting the steps in order

The order matters: unescape before removing symbols, lower-case before expanding contractions (the dictionary is lower case), and expand contractions before stripping punctuation. The steps on one example each:

Six cleaning steps with a before and after for each: HTML entities unescaped, tags and links removed, quotes straightened and lower-cased, contractions expanded, numbers replaced by num, and only letters kept with single spaces.

Cleaning real SMS messages

The whole function, run on three messages from the dataset and the Kalam sentence:

ExampleRun on NLTK 3.10.3
import re
import html

CONTRACTIONS = {"don't": "do not", "isn't": "is not", "i'm": "i am", "we're": "we are",
                "i'll": "i will", "there's": "there is", "won't": "will not", "can't": "cannot"}

def clean(text):
    text = html.unescape(text)                              # &amp; -> &
    text = re.sub(r"<[^>]+>", " ", text)                    # HTML tags
    text = re.sub(r"https?://\S+|www\.\S+", " url ", text)  # links
    text = text.replace("\u2019", "'").lower()              # straight quotes, lower case
    for short, full in CONTRACTIONS.items():
        text = re.sub(r"\b" + re.escape(short) + r"\b", full, text)
    text = re.sub(r"\d[\d,]*", " num ", text)                # every number -> num
    text = re.sub(r"[^a-z\s]", " ", text)                    # letters only
    return " ".join(text.split())                           # one space between words

messages = [
    "URGENT! You have won a 1 week FREE membership in our £100,000 Prize Jackpot! Txt the word: CLAIM to No: 81010 T&C www.dbuk.net LCCLTD POBOX 4403LDNW1A7RW18",
    "I'm back &amp; we're packing the car now, I'll let you know if there's room",
    "Nah I don't think he goes to usf, he lives around here though",
    "<p>Isn’t this incorrect?</p>",
]
for m in messages:
    print(clean(m))

What the cleaning changed

  • The spam message keeps its signal words, urgent, won, free, prize, claim, and the link becomes url. Every number, including the prize and the short code, becomes num, so "won num" looks the same in every prize message. The postcode-like code at the end breaks into num ldnw num a num rw num: noise that a stricter rule could drop.
  • &amp; became & and was then removed with the other symbols.
  • I'm, we're, I'll and there's became i am, we are, i will and there is.
  • don't became do not, so the negation survives as the word not.
  • The Kalam sentence lost its tags and gained a straight apostrophe, and Isn't became is not.

Seeing the wrong order fail

Strip punctuation before unescaping and expanding, and the entity and the contractions turn into junk tokens:

ExampleRun on NLTK 3.10.3
import re

text = "I'm back &amp; we're packing"
print(re.sub(r"[^a-z\s]", " ", text.lower()).split())

The output has amp, a fragment of the entity, and m and re, fragments of the contractions. All three would become vocabulary entries of their own.

Cleaning vs normalisation

CleaningNormalisation
GoalRemove what is not languageWrite the same word one way
ExamplesHTML tags, links, stray symbols, extra spacesLower case, straight quotes, expanded contractions, num for numbers
Effect on the vocabularyDrops junk tokensMerges variants of one word
Later steps of the same kindStopword removalStemming and lemmatization

Where you use text cleaning

  • Spam and sentiment classifiers with bag of words or TF-IDF, where every distinct string is a column.
  • Scraped web text, which arrives with tags, entities and links.
  • Search indexes, so that a query in any case finds the same documents.
Watch out. In a character class, A-z is not the same as A-Za-z. The range from A to z also covers [ \ ] ^ _ and the backtick, so re.sub('[^a-zA-z]', ' ', 'update_now') keeps the underscore. Write [^a-zA-Z].
Try it yourself
  • Add "can't" to a message and check that clean gives cannot.
  • Change the number rule to replace numbers with nothing instead of num and compare the spam message.
  • Remove the .lower() call and run clean again: which words disappear, and why?

Every expert started right here.