Foundations of Data Science slides 📂 Introduction · 2 of 10 39 min read

What Is Descriptive Statistics? Centre, Spread & Shape

Before you predict anything, you have to describe what you have. Descriptive statistics compresses thousands of rows into a few numbers and charts across three pillars — centre (mean, median, mode), spread (range, variance, std, IQR), and shape (skewness, kurtosis). This tutorial also covers the four data types, population vs sample, the five-number summary and box plot — with animated diagrams.

📋

What Is Descriptive Statistics?

Before you predict anything, you have to describe what you have. Descriptive statistics compresses thousands of rows into a handful of numbers and charts that tell you where the data sits, how it spreads, and what shape it takes.
Centre Spread Shape Data Types

Press Next → or use ← → arrow keys

Section 01

The Intuition — The Board Meeting

50,000 patient records in three numbers
A hospital manager walks into a board meeting with a spreadsheet of 50,000 patient records — names, ages, diagnoses, costs, outcomes. Nobody wants to scroll 50,000 rows. So she distils it: average age 47, average stay 3.2 days, success rate 94.3%. Three numbers, and the whole room understands the year at a glance.

That compression is descriptive statistics — the branch of statistics that organises, summarises, and describes a dataset with a few well-chosen numbers and charts.
💡
Describe First, Then Decide

Every analysis — every machine-learning model — starts here. You can't clean, model, or draw conclusions from data you haven't first described. Descriptive statistics is the mandatory first look at any dataset.

Section 01 · Two Branches

Descriptive vs Inferential

📋
Descriptive
describes what you have
Summarises the data in hand — averages, spreads, charts. No guessing beyond the dataset; every statement is a fact about this data.
🔮
Inferential
predicts the unseen
Uses a sample to draw conclusions about a larger population you didn't measure — with a stated margin of uncertainty.
🪜
The Order
describe → infer
Description always comes first. You must understand the data you have before you can responsibly generalise beyond it.
🗳️
One Example, Both Branches

Survey 1,000 voters: describing them ("62% of those polled prefer Party A") is a fact about the sample. Inferring ("Party A leads nationally, ±3%") is a claim about 46 million voters you never asked. Same data — different jobs, different confidence.

Section 02 · The Map

The Three Pillars

Descriptive Statistics CENTRE "what's typical?" Mean Median Mode the balance / middle / peak SPREAD "how consistent?" Range · IQR Variance Std Deviation how far values scatter SHAPE "what does it look like?" Skewness Kurtosis Modality symmetry & tails
🏛️
Three Questions Cover A Dataset

Centre tells you where the data sits. Spread tells you how tightly it clusters. Shape tells you whether it's symmetric, skewed, or riddled with outliers. Answer all three and you've genuinely described the data.

Section 03 · Pillar 1

Centre — Where Does The Data Sit?

Mean
Σx / n
The balance point. Uses every value — but a single outlier can drag it away from typical.
↔️
Median
middle value
Sort and take the centre. Ignores how extreme the extremes are — robust to outliers and skew.
👑
Mode
most frequent
The value that appears most. The only centre that works on categories like blood type or colour.
🩺
Eight Waiting Times

For clinic waits of 12, 18, 25, 8, 34, 21, 15, 19 minutes, the mean is 19 minutes. A gap between mean and median would flag skew — but here they're close, so the average is a fair summary of a typical wait.

Section 04 · Pillar 2

Spread — How Consistent Is It?

Range
max − min
Simplest spread. Fast, but driven entirely by the two most extreme values.
Variance (sample)
s² = Σ(xᵢ − x̄)² / (n − 1)
Average squared distance from the mean. In squared units, so hard to read directly.
Standard deviation
s = √s²
Variance back in the original units — the everyday measure of spread.
Interquartile range
IQR = Q3 − Q1
The spread of the middle 50% — outlier-resistant, and the backbone of the box plot.
📏
Centre Alone Lies

Two teams can average the same score while one is wildly inconsistent and the other rock-steady. The mean can't tell them apart — spread can. Always report a measure of spread alongside the centre.

Section 04 · Diagram

Same Average, Very Different Data

same mean Low spread · consistent small σ High spread · inconsistent large σ
🎯
Spread Is The Second Half Of The Story

Both rows share the same mean, marked by the blue line — yet the top is tightly clustered and the bottom is all over the place. Standard deviation is what separates "reliably around 19 minutes" from "anywhere from 2 to 40." Centre + spread together, always.

Section 05 · Pillar 3

Shape — Symmetry, Tails & Outliers

📐
Skewness
asymmetry
Which way the tail leans. Positive = long right tail (incomes); negative = long left tail; zero = symmetric.
⛰️
Kurtosis
tailedness
How heavy the tails are. High kurtosis means more extreme outliers than a normal bell would predict.
🔔
Modality
number of peaks
One peak (unimodal), two (bimodal), or more — often a hint that two groups are mixed together.
🔍
Shape Warns You Before You Model

Skewness tells you whether to trust the mean; kurtosis warns of outlier-heavy tails; a second peak hints that two populations are hiding in one column. These are exactly the red flags that decide how you'll clean and model the data.

Section 06 · Diagram

The Five-Number Summary & Box Plot

increasing value → outlier min Q1 median (Q2) Q3 max IQR = Q3 − Q1 (middle 50%)
📦
Five Numbers Tell The Whole Story

Minimum, Q1, median, Q3, maximum — the five-number summary — pack centre, spread, and skew into one picture. The box is the middle 50% (the IQR), the line inside is the median, the whiskers reach the typical range, and any point beyond them is flagged as an outlier. It's the fastest way to compare groups side by side.

Section 07 · Data Types

The Four Data Types — A Ladder

more information → Nominal unordered categories · blood type, colour, gender valid: mode, frequency, % Ordinal ordered, unequal gaps · ratings, pain 1–10 valid: + median, percentiles Interval equal gaps, no true zero · °C, IQ, year valid: + mean, std dev Ratio equal gaps + true zero · height, weight, age, ₹ valid: ALL statistics
🌡️
The Thermometer Test

To tell interval from ratio, ask: does zero mean "none of it exists"? 0 kg = no weight → ratio. 0 °C is just a cold day, not "no temperature" → interval. Each rung up the ladder unlocks more valid statistics — but never borrow a statistic from a rung above your data.

Section 07 · Reference

Matching Statistics To Data Types

TypePropertiesExamplesValid statistics
NominalCategories, no orderBlood type, colour, genderMode, frequency, %
OrdinalOrdered, unequal gapsRatings, pain scale, grade+ Median, percentiles
IntervalEqual gaps, no true zero°C, IQ, calendar year+ Mean, std dev
RatioEqual gaps, true zeroHeight, weight, age, salaryAll statistics
⚠️
The Classic Mistake: Averaging Ordinal Data

Reporting a "mean satisfaction of 7.3" on a 1–10 scale looks precise but is misleading — the gap between 3 and 4 isn't guaranteed to equal the gap between 8 and 9. For ordinal data, report the median and the mode, not the mean.

Section 08 · Population vs Sample

Who Are You Actually Describing?

Population (everyone)
N · μ · σ
Every member of the target group. Variance divides by N.
Sample (a subset)
n · x̄ · s
A selection used to represent the whole. Variance divides by (n − 1).
Why (n − 1)? Bessel's Correction

A sample's spread systematically under-estimates the population's. Dividing by n−1 instead of n nudges the estimate up to compensate. On the eight waiting times, the population variance is 444/8 = 55.5, but the sample variance is 444/7 = 63.4.

🗄️
When In Doubt, Treat It As A Sample

Even a 10-million-row database is usually a sample of all the data you could have collected. Unless you've truly captured every possible case, use the sample formulas — divide by (n−1).

Section 09 · Worked Example

Describing Eight Waiting Times

Clinic waits (minutes): 12, 18, 25, 8, 34, 21, 15, 19. Run the full descriptive pass.

19Mean (min)
18.5Median
26Range (34−8)
63.4Sample variance
≈7.96Sample std dev
+Mild right skew
🧮
One Dataset, A Complete Picture

Centre: a typical wait is ~19 minutes. Spread: a standard deviation of ~8 means most waits fall roughly 11–27 minutes. Shape: the mean sitting just above the median hints at a mild right skew from the odd long wait. Three pillars, one honest summary.

Section 10 · Code

One Line To Describe Everything

import pandas as pd

df = pd.DataFrame({
    'wait_min': [12, 18, 25, 8, 34, 21, 15, 19],
    'blood_type': ['A', 'O', 'B', 'AB', 'O', 'A', 'O', 'B']  # nominal
})

df['wait_min'].describe()      # count, mean, std, min, 25%, 50%, 75%, max
df['wait_min'].std()           # sample std (ddof=1 by default)
df['blood_type'].value_counts()  # frequencies for the nominal column
df['blood_type'].mode()[0]      # 'O' — the only valid "centre" here
🐼
describe() Is Your First Look

df.describe() returns count, mean, standard deviation, min, the three quartiles, and max in one call — the whole five-number summary plus centre and spread. For categorical columns reach for value_counts() and mode() instead; describe()'s numeric stats don't apply.

Section 11 · Applications

Descriptive Statistics Everywhere

🏥
Healthcare
Patient demographics, average stay, and success rates summarise hospital performance.
🗳️
Polling
A sample of 1,000 voters describes preferences before any national inference is drawn.
🏢
Business
Employee satisfaction and quality checks summarised per shift to catch problems early.
🎓
Education
Class mean, median, and spread of exam scores reveal how a cohort really performed.
📺
Streaming
Monthly view counts across the entire logged population — a true population, not a sample.
🤖
Machine Learning
EDA — describing every feature's centre, spread, and shape — is step one of any ML pipeline.
Section 12 · Golden Rules

Six Rules For Honest Description

🏅 Descriptive Statistics, Distilled
1Identify the data type first. Nominal, ordinal, interval, or ratio — it dictates which statistics are legal.
2Match the statistic to the type. Never take the mean of nominal data, or average an ordinal scale.
3Report centre AND spread — and glance at shape. One number never describes a dataset.
4Know population vs sample and use the right divisor — (n−1) for samples.
5Describe before you infer. Understand the data in hand before generalising beyond it.
6Visualize. A histogram or box plot reveals skew and outliers that summary numbers hide.
Wrap-Up

You Can Now Describe Any Dataset

centreMean · median · mode
spreadRange · var · σ · IQR
shapeSkew · kurtosis
4 typesNominal→ratio
n−1Sample divisor
5-numBox plot summary
🎯
The Through-Line

Descriptive statistics compresses raw data into centre, spread, and shape — but only after you've identified the data type and decided whether you hold a population or a sample. Nail the description, and every model and inference that follows rests on solid ground.

📚
Where To Go Next

Drill into each pillar: mean, median & mode for centre; variance and standard deviation for spread; skewness and kurtosis for shape. Then step up to inferential statistics — confidence intervals and hypothesis tests.

📋 End of tutorial · Press to review, or click Restart