StatisticsSciPy 1.18 · pandas 3.0 · statsmodels 0.15 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
57 small wins to finish your pathNext lesson →

Frequency distribution

A frequency distribution is a table that lists each value or class of a variable together with its frequency, the number of times it occurs in the data.

Last updated: 07 Oct, 2026 · SciPy 1.18

Once the scale of a variable is known, the first descriptive step is to count. A frequency table turns a long list of values into a few rows, and every chart in this part is drawn from one.

Frequency distribution and cumulative frequency · from the Complete Statistics for Data Science in 6 Hours video · 34:47 to 36:51

Counting the flowers

The video's sample dataset is a list of nine flowers of three types: Rose, Lily, Sunflower, Rose, Lily, Sunflower, Rose, Lily, Lily. To show it at a glance, count each type. Rose appears 3 times, Lily 4 times and Sunflower 2 times. The three counts are the frequencies, and a table of them is a frequency distribution. Bar charts and pie charts are drawn from this table, as Bar charts and pie charts shows.

A frequency table of nine flowers: Rose 3, Lily 4 and Sunflower 2, with relative frequencies 0.333, 0.444 and 0.222, and a cumulative column 3, 7 and 9 built by adding 4 and then 2, so the last running total is the sample size 9.

Adding relative frequency

A relative frequency divides each frequency by the number of values n, so the column adds up to 1 (or 100%). It makes tables of different sizes comparable: 4 lilies out of 9 and 40 out of 90 are the same 0.444.

Relative frequency of the flowers, n = 9

Adding cumulative frequency

A cumulative frequency is a running total down the table. The video starts with the 3 roses, adds the 4 lilies to get 7, adds the 2 sunflowers to get 9, and the last value is the total number of flowers.

The flowers are nominal, so their rows could be listed in any order: Lily, Rose, Sunflower would give 4, 7, 9 instead. Only the last total means something on its own. Cumulative frequency becomes useful on ordered values, where a running total answers "how many are at most this?"

Grouping numbers into classes

A numeric variable with many different values is counted in classes, ranges of equal width. The video's histogram data, 16 ages, falls into five classes of width 10. Each class is written [10, 20), which includes 10 and excludes 20, so a value on a boundary is counted once; Histograms covers that choice. On these ordered classes the cumulative column is meaningful: 10 of the 16 ages are below 40.

Building frequency tables in pandas

Counting the flowers with value_counts

python
import pandas as pd

flowers = pd.Series(["Rose", "Lily", "Sunflower", "Rose", "Lily",
                     "Sunflower", "Rose", "Lily", "Lily"])
counts = flowers.value_counts(sort=False)    # sort=False keeps the order of the data

Grouping the ages into classes with pd.cut

python
ages = pd.Series([10, 12, 14, 18, 24, 26, 30, 35, 36, 37, 40, 41, 42, 43, 50, 51])
classes = pd.cut(ages, bins=range(10, 70, 10), right=False)   # [10, 20), [20, 30), ...
age_counts = classes.value_counts(sort=False)

Printing the flower and age tables

ExampleFrom the video, run on pandas 3.0.6
table = pd.DataFrame({"frequency": counts})
table["relative"] = (counts / counts.sum()).round(3)
table["cumulative"] = counts.cumsum()
print(table.to_string())
print()
age_table = pd.DataFrame({"frequency": age_counts})
age_table["cumulative"] = age_counts.cumsum()
age_table["cumulative relative"] = (age_counts.cumsum() / len(ages)).round(3)
print(age_table.to_string())

Reading the two tables

  • Rose 3, Lily 4, Sunflower 2: the board's frequency table, with cumulative 3, 7 and 9.
  • Relative 0.333, 0.444, 0.222: the shares of the nine flowers; they add up to 1 (0.999 after rounding).
  • Ages 4, 2, 4, 4, 2: the counts in [10, 20) up to [50, 60), which add up to 16.
  • Cumulative relative 0.625 at [30, 40): 62.5% of the ages are below 40, the kind of statement a percentile makes, as Percentiles and percentile rank shows.

Drawing the cumulative frequency curve

Plotting the cumulative share against each class's upper boundary gives an ogive, which reads off how many values lie below any age:

ExampleRun on matplotlib 3.11.2
import matplotlib.pyplot as plt

edges = [10, 20, 30, 40, 50, 60]
cum_share = [0] + (age_counts.cumsum() / len(ages)).tolist()
plt.figure(figsize=(7, 4))
plt.plot(edges, cum_share, marker="o")
plt.title("Cumulative relative frequency of the ages (ogive)")
plt.xlabel("Age")
plt.ylabel("Share of ages below")
plt.grid(alpha=0.3)
plt.show()
print([round(v, 3) for v in cum_share])
The ogive of the 16 ages rises from 0 at age 10 through 0.25 at 20, 0.375 at 30, 0.625 at 40 and 0.875 at 50 to 1 at 60.

Frequency vs relative frequency vs cumulative frequency

FrequencyRelative frequencyCumulative frequency
What it isCount of each value or classCount ÷ nRunning total of the counts
Adds up ton1Its last row equals n
Flowers3, 4, 20.333, 0.444, 0.2223, 7, 9
Needs ordered rowsNoNoYes, to mean anything
In pandasvalue_counts()value_counts(normalize=True)value_counts(sort=False).cumsum()

Where you use frequency distributions

  • Exploring a new dataset: value_counts() on every categorical column shows its categories, typos and rare values.
  • Checking class balance in machine learning: the relative frequency of the labels shows whether one class is rare, which changes how a model is trained and scored.
  • Grouped tables for reports: customers by age band or orders by price band, the counts a histogram draws.
Watch out. value_counts() leaves out missing values by default (dropna=True), so the frequencies can add up to less than the number of rows. Pass dropna=False to count the missing values as their own row.
Try it yourself
  • The notes use colours instead: ["Green", "Red", "Yellow", "Green", "Red", "Yellow", "Green", "Red"]. Is the table Green 3, Red 3, Yellow 2?
  • Print flowers.value_counts(normalize=True) to get the relative frequencies in one call.
  • Change right=False to right=True, include_lowest=True in pd.cut. Which classes change, and which ages moved?

You understood something today that you didn't yesterday.