VizSoup

Summary statistics calculator

Paste a list of numbers, or a table and pick a column. Every statistic updates as you type, and where textbooks and software disagree on a definition, each version is shown.

Sample: 40 invented delivery times in minutes.

Count (n)

40

Mean

35.425

Median

33

Mode

33

Std. dev. (sample)

10.879

Std. dev. (population)

10.7422

Variance (sample)

118.3532

Std. error of mean

1.7201

Minimum

26

Maximum

95

Range

69

Sum

1,417

Coeff. of variation

30.7%

Skewness

4.405

Quartiles, three ways

MethodQ1Q3IQR
Linear, inclusive 30 37.25 7.25
Linear, exclusive 30 37.75 7.75
Tukey's hinges 30 37.5 7.5

Linear, inclusive: Excel QUARTILE.INC, Google Sheets QUARTILE, R type 7, NumPy default.Linear, exclusive: Excel QUARTILE.EXC, Minitab, SPSS, R type 6.Tukey's hinges: medians of the lower and upper halves, the middle value included in both when n is odd.

Outliers (1.5 × IQR fences)

Lower fence 19.125, upper fence 48.125

Outside the fences: 95.

How to use it

Paste numbers one per line, separated by commas or tabs, or paste a whole table from a spreadsheet and choose the column. Blank cells and text are skipped rather than counted as zero, and values such as 1,250, $40 or (12) are read as numbers. The count at the top left tells you how many values were actually used; check it against what you expected.

The formulas

Why quartiles have three answers

A quartile falls between two data points more often than on one, and there are several defensible ways to decide where. Hyndman and Fan's 1996 paper catalogued nine definitions in use in statistical software. Three cover almost everything you will meet:

Worked example

The sample is 40 invented delivery times. They sum to 1,417 minutes, so the mean is 35.43. The median is 33, and 33 is also the mode, appearing four times. The mean is higher than the median because of a single 95-minute delivery, and the skewness of 4.40 says the same thing more formally.

The first quartile is 30 under all three methods. The third quartile is 37.25 (inclusive), 37.5 (Tukey) or 37.75 (exclusive), giving an IQR between 7.25 and 7.75. Using the inclusive quartiles, the fences are 30 − 1.5 × 7.25 = 19.125 and 37.25 + 1.5 × 7.25 = 48.125, and only the 95 lies outside.

Delete the 95 and the mean drops to 33.90 while the median stays at 33. The sample standard deviation halves, from 10.88 to 5.07. The fences also tighten, to 20.25 and 46.25, and now the 47-minute order is flagged. That is a known property of the rule: removing one outlier and re-running can always reveal another. Decide what to exclude once, for a stated reason, rather than repeating the test until nothing is left.

Limits

These statistics describe the numbers you paste; they do not test hypotheses or give confidence intervals beyond the standard error. With fewer than about ten values, the standard deviation, skewness and quartiles are all unstable, and small changes to one value move them a lot. To see the distribution's shape rather than summarise it, use the histogram maker. For statistics by group (the mean per region, the median per product), use the pivot table.

Common questions

Should I use the sample or the population standard deviation?

Use the sample version (dividing by n − 1) when your numbers are a sample from something larger and you want to describe that larger thing, which is nearly always. Use the population version (dividing by n) only when the list is the complete set you care about, such as every employee in a small company. Excel calls them STDEV.S and STDEV.P.

Why do Excel and my calculator give different quartiles?

Because there is no single agreed definition. Excel's QUARTILE.INC, Google Sheets and R's default interpolate one way, QUARTILE.EXC and Minitab another, and many textbooks and calculators use the medians of the two halves. On large datasets they converge; on small ones they can differ noticeably. This page shows all three so you can match whichever your course or software uses.

When should I report the median instead of the mean?

When the data is skewed or has outliers, as incomes, house prices, waiting times and response times usually do. The mean is pulled towards the long tail; the median is the middle value and barely moves. If the two differ by a lot, report both and say why.

Does an IQR outlier mean the value is a mistake?

No. The 1.5 × IQR rule, from John Tukey, flags values worth looking at, not values to delete. In a normal distribution about 0.7% of values fall outside the fences by chance alone. Check whether an outlier is an error, such as a typo or a wrong unit, before removing it, and say so when you do.

What does the skewness number mean?

It measures lopsidedness. Zero means symmetric, positive means a longer tail to the right (a few large values), negative a longer tail to the left. As a rough guide, values beyond about ±1 indicate strong skew. The figure here is the adjusted Fisher–Pearson coefficient, which matches Excel's SKEW function.

Other tools