StatisticsSciPy 1.18 · pandas 3.0 · statsmodels 0.15 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
57 small wins to finish your pathNext lesson →

Statistics logoStatistics

Statistics is the science of collecting, organizing, analyzing and interpreting data, so that a decision can rest on evidence instead of a guess.

Last updated: 07 Oct, 2026

The opening example of the video is a classroom of maths students and their first-semester marks: 84, 86, 78, 72, 75, 65, 80, 81, 92, 95, 96 and 97. Two kinds of question can be asked about it. "What is the average mark in this class?" is answered from the 12 marks alone: their mean is 83.42. "Are the marks in this class similar to the marks of all the maths classrooms in the college?" is about students whose marks were never collected, so the answer has to be inferred from one class, with some uncertainty. The first question belongs to descriptive statistics and the second to inferential statistics, and the lessons teach both.

Defining statistics and data

What statistics is · from the Complete Statistics for Data Science in 6 Hours video · 4:03 to 7:39

The video starts from a short definition: statistics is the science of collecting, organizing and analyzing data, and the reason to do it is better decision making. A product team, a hospital or a government has more data than anyone can read, and statistics turns it into an answer it can act on. Fuller definitions add two more steps, interpreting the results and presenting them, and much of statistics is about saying how sure an answer is.

Data are facts or pieces of information recorded about something. The video's examples are the IQ scores of a class and the ages of the students in a class, such as 30, 25, 24, 23, 27 and 28. Data can be numbers, like these ages, or categories, like a blood group or a city, and the difference decides which methods apply, as Types of variables shows.

  • Descriptive statistics organizes and summarizes the data you have: tables, charts and numbers such as the mean.
  • Inferential statistics uses a sample to draw conclusions about the larger population it came from, and measures how uncertain those conclusions are. Descriptive and inferential statistics compares the two.

Following the learning path

The course as ten parts in reading order with every lesson listed: what statistics is, tables and charts, centre and spread, percentiles and box plots, the normal distribution, probability, probability distributions, estimation and testing, statistical tests, and relationships; parts one to four describe data, five to seven are probability models, and eight to ten draw conclusions from samples.

The parts follow the video's order with one change: the probability distributions come before estimation, because confidence intervals and tests are built on the normal distribution, the t distribution and the central limit theorem. Parts 1 to 4 describe data. Parts 5 to 7 are the probability models that describe how data varies by chance. Parts 8 to 10 use those models to reach conclusions from samples. The table lists every lesson, part by part.

PartLessonsStarts with
1. What statistics isInstalling Python for statistics · Descriptive and inferential statistics · Population and sample · Sampling techniques · Types of variables · Measurement scalesInstalling Python for statistics
2. Describing data with tables and chartsFrequency distribution · Bar charts and pie charts · Histograms · Probability density function (PDF) and KDEFrequency distribution
3. Centre and spreadMean, median and mode · Variance and standard deviation · Sample variance and why n − 1 · Range, MAD and coefficient of variation · Skewness and kurtosisMean, median and mode
4. Percentiles, quartiles and box plotsPercentiles and percentile rank · Quartiles and the interquartile range · Five-number summary and box plot · Outlier detection with IQR and z-scorePercentiles and percentile rank
5. The normal distributionNormal (Gaussian) distribution · Empirical rule (68-95-99.7) · Z-score and the standard normal distribution · Standardization and normalization · Z-table and normal probabilities · Normality tests (Q-Q plot and Shapiro-Wilk)Normal (Gaussian) distribution
6. ProbabilityProbability basics · Addition rule of probability · Multiplication rule and independent events · Conditional probability and Bayes' theorem · Permutations and combinationsProbability basics
7. Probability distributionsRandom variables and probability distributions · Bernoulli distribution · Binomial distribution · Poisson distribution · Uniform distribution · Log-normal distribution · Power law and Pareto distribution · Central limit theoremRandom variables and probability distributions
8. Estimation and hypothesis testingPoint estimates and standard error · Confidence intervals · Hypothesis testing · P-value · Significance level, one-tailed and two-tailed tests · Type I and Type II errorsPoint estimates and standard error
9. Statistical testsOne-sample z-test · One-sample t-test and the t distribution · Two-sample and paired t-tests · Z-test for a proportion · Chi-square goodness-of-fit test · Chi-square test of independence · One-way ANOVA (F-test) · Choosing a statistical testOne-sample z-test
10. Relationships between variablesCovariance · Pearson correlation coefficient · Spearman rank correlation · Statistics interview questionsCovariance

Who this course is for

  • Beginners in data science, data analysis and business intelligence. The video is aimed at these roles, and every method is shown on small numbers you can check by hand before Python does it.
  • Interview preparation. The video poses the questions interviewers ask, such as "what is the difference between ordinal and nominal data?" and "what is the difference between a bar chart and a histogram?", and the lessons answer them.
  • Machine learning learners. Scaling, train and test splits, model metrics and A/B tests all rest on these ideas; the Machine Learning course builds on them.

Preparing what you need

  • Python basics: variables, lists, loops, functions and importing a library.
  • School maths: averages, fractions, percentages, squares and square roots. Every formula is worked out on the video's numbers before it is used.
  • Python 3.12 or newer on Windows, macOS or Linux, or Google Colab in a browser. Installing Python for statistics sets up both.
  • No API keys and no GPU. The data is typed into the code or loads from seaborn over the internet.

Reading a lesson

Each lesson opens with a one-sentence definition and a clip from the video where the video teaches the idea. The text under a clip follows it, with the video's example and numbers, then turns the idea into a formula and works it out. Lessons with something to compute end in Python code whose printed output comes from a real run, and every lesson closes with a few changes to try.

Following the video and its notes

The lessons follow Complete Statistics For Data Science In 6 hours (2022, 5 hours 28 minutes), edited from a seven-day live series held in January 2022. The theory is written on a digital board, and the practical parts run in Jupyter notebooks.

The video's description links its materials, which serve as the notes for this course: the Complete Statistics folder of The Grand Complete Data Science Materials on GitHub. It holds five PDFs of handwritten board notes, a PDF of statistics interview questions with answers, the Stats_Practicals notebook (finding outliers with z-scores and the IQR) and a hypothesis-testing notebook (chi-square tests, t-tests and correlation) with its saved outputs. The notes and standard statistics add topics the video does not teach, such as Sample variance and why n − 1, Conditional probability and Bayes' theorem, Two-sample and paired t-tests, One-way ANOVA (F-test) and Choosing a statistical test.

The Python libraries have changed since 2022. These are the behaviours the lessons point out where they matter:

The video (2022)Today (NumPy 2.5, pandas 3.0, SciPy 1.18, seaborn 0.13)
sns.boxplot(data) with the data passed by positionPass it by name, sns.boxplot(x=data), to get a horizontal box of one variable
df.corr() on a table with text columnsRaises an error unless you pass numeric_only=True
chi2_contingency(table) on a 2 × 2 tableApplies Yates' continuity correction by default (correction=True)
np.var and np.std next to pandasNumPy divides by n (ddof=0); pandas .var() and .std() divide by n − 1 (ddof=1)
np.random.choice(data, n) for a random sampleSamples with replacement unless you pass replace=False
Percentiles by handnp.percentile uses linear interpolation by default, which can differ from a hand method on small data
Back toAll courses

Every expert started right here.