Mean or median: which should you report for skewed data?

Mean or median: which should you report for skewed data?

When the mean and median disagree, which one tells the truth? Three quick NumPy experiments: skewed salaries, one extreme value, and two groups where the mean describes almost nobody. Tested with NumPy 2.5 and SciPy 1.18.

TL;DR — For skewed data, report the median when you want the “typical” value, and the mean when you need totals (budgets, payroll, total sales). In our test, 63% of people earned less than the mean salary, and one extreme value nearly doubled the mean while moving the median by 29. If the data has two groups, neither number is typical: look at the histogram.

Tested on 2026-10-06 with NumPy 2.5.3, SciPy 1.18.1, Matplotlib 3.11.2, Python 3.12.

Key points

  • The mean is pulled by the tail. In right-skewed data (incomes, prices, response times) the mean sits above the median.
  • The median ignores extreme values. One huge value can move the mean a lot and the median almost not at all.
  • The mean is the right choice for totals. Total = mean × count. The median times the count gives the wrong total.
  • A trimmed mean is a middle ground. It drops a fixed share of the smallest and largest values before averaging.
  • With two groups, both can describe nobody. Only 2.3% of our two-group data was near the mean.

How far apart are the mean and median in skewed data?

Far enough to change the story. I generated 1,000 salaries from a lognormal distribution, a common shape for incomes, and compared the two numbers. I also added one CEO, checked a 10% trimmed mean, made a two-group dataset, and computed a payroll total.

import numpy as np
from scipy import stats

rng = np.random.default_rng(7)

# 1. Right-skewed data: 1,000 salaries
salaries = rng.lognormal(mean=np.log(50_000), sigma=0.6, size=1_000)
print(f"salaries     mean {salaries.mean():>9,.0f}  median {np.median(salaries):>9,.0f}")
print(f"share of people earning less than the mean: {(salaries < salaries.mean()).mean():.0%}")

# 2. One extreme value
with_ceo = np.append(salaries, 50_000_000)
print(f"+ one CEO    mean {with_ceo.mean():>9,.0f}  median {np.median(with_ceo):>9,.0f}")
print(f"+ one CEO    10% trimmed mean {stats.trim_mean(with_ceo, 0.1):,.0f}")

# 3. Two groups: who is "typical"?
two_groups = np.concatenate([rng.normal(35, 6, 500), rng.normal(70, 6, 500)])
m, med = two_groups.mean(), np.median(two_groups)
print(f"two groups   mean {m:.1f}  median {med:.1f}")
print(f"share within 5 of the mean: {(abs(two_groups - m) < 5).mean():.1%}")
print(f"share within 5 of 35:       {(abs(two_groups - 35) < 5).mean():.1%}")

# 4. When the mean is the right answer: totals
print(f"payroll = mean x n: {salaries.mean() * len(salaries):,.0f} vs sum {salaries.sum():,.0f}")
print(f"payroll from the median x n: {np.median(salaries) * len(salaries):,.0f}")

Output:

salaries     mean    56,165  median    48,005
share of people earning less than the mean: 63%
+ one CEO    mean   106,059  median    48,034
+ one CEO    10% trimmed mean 51,244
two groups   mean 52.5  median 53.0
share within 5 of the mean: 2.3%
share within 5 of 35:       28.1%
payroll = mean x n: 56,164,625 vs sum 56,164,625
payroll from the median x n: 48,004,851

The mean salary is 56,165 and the median is 48,005. Saying “the average person earns 56k” is misleading here: 63% of people earn less than that.

What does one extreme value do?

It nearly doubles the mean. Adding a single 50-million salary to 1,000 normal ones moved the mean from 56,165 to 106,059. The median moved from 48,005 to 48,034.

A 10% trimmed mean drops the lowest 10% and highest 10% before averaging. With the CEO included it gave 51,244. That is close to the median but still uses most of the data, which is why it is popular for noisy measurements.

What if the data has two groups?

Then neither number is “typical”. The two-group data has peaks near 35 and 70. Its mean (52.5) and median (53.0) land in the gap between them:

WindowShare of the data
Within 5 of the mean (52.5)2.3%
Within 5 of the first peak (35)28.1%

The chart shows both datasets with the mean and median marked:

import numpy as np
import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt

# same data as above (same seed, same order of draws)
rng = np.random.default_rng(7)
salaries = rng.lognormal(mean=np.log(50_000), sigma=0.6, size=1_000)
two_groups = np.concatenate([rng.normal(35, 6, 500), rng.normal(70, 6, 500)])

plt.style.use("dark_background")
fig, axes = plt.subplots(1, 2, figsize=(12, 4))
panels = [(salaries / 1000, "Salaries (thousands)"), (two_groups, "Two groups")]
for ax, (x, title) in zip(axes, panels):
    ax.hist(x, bins=40, color="#5fa8d3", edgecolor="#1b1b1b")
    ax.axvline(x.mean(), color="#f4a261", lw=2.5, label=f"mean {x.mean():.1f}")
    ax.axvline(np.median(x), color="#e9edc9", lw=2.5, ls="--", label=f"median {np.median(x):.1f}")
    ax.set_title(title)
    ax.legend()
fig.tight_layout()
fig.savefig("mean_vs_median.png", dpi=100)

Two histograms. Left: right-skewed salaries in thousands with the median at 48.0 and the mean at 56.2, to the right of the peak. Right: two separate groups peaking near 35 and 70, with the mean at 52.5 and the median at 53.0 both falling in the nearly empty gap between them

For data like this, report the groups separately, or show the histogram. A single summary number hides the most important fact: there are two kinds of values.

When is the mean the right answer?

When you need a total. Total payroll is the mean salary times the number of people: 56,164,625, exactly the sum. Using the median gives 48,004,851, which underestimates the budget by about 8.2 million.

QuestionReport
What does a typical person earn / pay / wait?Median
How much will this cost in total?Mean (or just the sum)
Noisy measurements with a few bad readingsTrimmed mean or median
Data with two or more groupsSplit the groups, or show the histogram
Roughly symmetric data without outliersEither. They are almost equal

Common mistakes

  • Reporting only the mean for incomes, prices or latencies. Report the median, and the mean if totals matter.
  • Treating “mean > median” as proof of skew. It is a hint. Look at a histogram to see the actual shape.
  • Dropping outliers to make the mean look right. If an extreme value is real, the median handles it without deleting data.
  • Averaging across groups that should be separate. Two groups give a mean that describes neither.

Sources

Related: What is a histogram? · How many bins should a histogram have?