StatisticsSciPy 1.18 · pandas 3.0 · statsmodels 0.15 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
57 small wins to finish your pathNext lesson →

Types of variables

A variable is a characteristic recorded for each member of a dataset whose value can change from one member to the next; it is quantitative when its values are numbers and qualitative when they are categories.

Last updated: 07 Oct, 2026 · SciPy 1.18

A sample, chosen with one of the Sampling techniques, gives a table: one row per member, one column per variable. The type of each column decides which summaries and charts make sense, so it is the first thing to check in any dataset.

Defining a variable

The video's examples are height and weight. Height takes values such as 182, 178, 168, 150, 160 and 170 cm; weight takes 78, 99, 100, 60 and 50 kg. A variable does not take any value at all: a blood group is always one of eight, and a height is never negative. It takes one value from its set for each member.

Quantitative and qualitative, discrete and continuous variables · from the Complete Statistics for Data Science in 6 Hours video · 25:51 to 29:33

Measuring quantitative variables

A quantitative variable is measured as a number: age, weight, height. The numbers mean amounts, so arithmetic on them is meaningful: you can add weights, subtract heights and average ages. How much arithmetic is allowed depends on the variable's scale, as Measurement scales shows (a temperature of 20 °C is not twice as warm as 10 °C).

Grouping qualitative variables

A qualitative or categorical variable puts each member into a category: gender (male, female), blood group (A positive, A negative, ...), T-shirt size (small, medium, large, XL). Adding or averaging the categories means nothing; you count how many members fall in each.

A categorical variable can also be made from a number, by cutting it into bands. The video bins IQ scores into "less", "medium" and "good". Banding keeps the order of the numbers, so the bands are ordered categories, like the T-shirt sizes: small comes before medium before large. Categories with an order are ordinal; categories without one, like blood groups, are nominal.

Telling discrete from continuous variables

Quantitative variables split in two. A discrete variable takes separate, countable values with gaps between them. The video's examples are counts: the number of bank accounts a person has (2, 3, 4, 5, 6, 7, never 2.5) and the number of children in a family.

A continuous variable can take any value in a range, limited only by how precisely it is measured: a height of 172.5 cm, 162 cm or 163.5 cm; a weight of 100 kg, 99.5 kg or 99.75 kg; rainfall of 1.1, 1.25 or 1.35 inches.

Counts are the usual discrete variables, and a discrete scale can also have steps of a half: shoe sizes 7, 7.5 and 8 are discrete, because nothing lies between the sizes on offer. And a continuous variable is often recorded in rounded steps: age is continuous, but a dataset usually stores it in whole years.

A variable is either quantitative, a number, or qualitative, a category; quantitative variables are discrete, such as the number of bank accounts, or continuous, such as height 172.5 cm; qualitative variables are nominal, such as blood group or pin code, or ordinal, such as T-shirt size S, M, L, XL.

The notes end with four to classify. Marital status is categorical. The length of the Nile and the duration of a movie are continuous quantitative. IQ is quantitative, reported as whole numbers on a continuous scale.

Reading variable types in pandas

pandas stores each column with a dtype that hints at its type: float64 for decimals, int64 for whole numbers, category or str for labels. The tips dataset has all three kinds:

ExampleRun on seaborn 0.13.2
import seaborn as sns

tips = sns.load_dataset("tips")
print(tips.head(3).to_string())
print()
kinds = tips.dtypes.astype(str).to_frame("dtype")
kinds["distinct values"] = tips.nunique()
print(kinds.to_string())

Reading the tips columns

  • total_bill and tip are float64 with many distinct values: amounts of money, continuous quantitative variables (recorded to the cent).
  • size is int64 with 6 distinct values: the number of people at the table, a discrete count.
  • sex, smoker, day and time are category with 2 to 4 values: qualitative variables. Days of the week have an order; smoker or not does not.

Spotting a number that is a category

A dtype is a hint, not the answer. A pin code is stored as a whole number, yet it is a label: 560034 is not "more" than 110001 in any useful sense. The video leaves pin code as a question; it is a categorical variable. Treated as a number, it gives an average with no meaning:

ExampleRun on pandas 3.0.6
import pandas as pd

shops = pd.DataFrame({"pin_code": [560001, 560034, 110001, 400050, 560001]})
print("mean pin code:", shops["pin_code"].mean())           # a number with no meaning

shops["pin_code"] = shops["pin_code"].astype("category")     # store it as a label
print(shops["pin_code"].value_counts().to_string())

As a category, the right summary is a count per code: 560001 appears twice. Phone numbers, roll numbers and customer ids are the same kind of variable.

Discrete vs continuous variables

DiscreteContinuous
ValuesSeparate, countable, with gapsAny value in a range
Typical sourceCountingMeasuring
The video's examplesBank accounts, children in a familyHeight, weight, rainfall
Other examplesNumber of employees, shoe size, people at a tableSong length, river length, blood pressure
Usual chartBar chart when there are few values, histogram when manyHistogram

Where you use variable types

  • Choosing a chart: a bar chart for categories and a histogram for numbers, as Bar charts and pie charts and Histograms show.
  • Choosing a summary: an average for a quantitative variable, counts and the most common category for a qualitative one.
  • Preparing data for machine learning: numbers are scaled, categories are encoded as columns of 0s and 1s, which Measurement scales shows.
Watch out. Numbers that are labels, such as pin codes, roll numbers and phone numbers, are categorical. pandas reads them as int64, and describe() will report their mean and standard deviation without complaint. Convert them with astype("category") before summarizing.
Try it yourself
  • Print tips.select_dtypes("number").columns and tips.select_dtypes("category").columns to split the columns by kind.
  • Print tips["size"].value_counts().sort_index(). Why can there be no table of 2.5 people?
  • Run shops.describe() before and after the astype("category") line and compare what pandas reports.

Every expert started right here.