VizSoup

Histogram maker

Paste a column of numbers, or a whole table and pick the column. The number of bins is a choice rather than a fact, so five standard rules are computed side by side and you choose.

Sample: 40 invented delivery times in minutes, including one very late order.

Histogram of MinutesHistogram with 14 bins.05101520Count25 to 30: 830 to 35: 1635 to 40: 940 to 45: 545 to 50: 150 to 55: 055 to 60: 060 to 65: 065 to 70: 070 to 75: 075 to 80: 080 to 85: 085 to 90: 090 to 95: 12535455565758595Minutes
40 values drawn in 14 bins of width 5, from 25 to 95.

What each rule suggests

RuleFormulaBins
Sturgesk = ⌈log₂ n⌉ + 17
Freedman–Diaconish = 2 × IQR / ∛n17
Scotth = 3.49 × s / ∛n7
Square rootk = ⌈√n⌉7
Ricek = ⌈2 × ∛n⌉7

Counts per bin

From (incl.)ToCount
25308
303516
35409
40455
45501
50550
55600
60650
65700
70750
75800
80850
85900
90951

How to use it

Paste one number per line, or paste a table and choose the column. Blank cells and text are skipped, and numbers written with thousands separators or currency signs are read correctly. The bin-rule menu shows how many bins each rule recommends for your data. Pick one, or set your own count, and the chart and the table of counts update together. Hover over a bar to see its range and count.

The five bin rules

Here n is the number of values, s the sample standard deviation and IQR the interquartile range (third quartile minus first). Rules that give a width are turned into a count by dividing the data range by that width and rounding up.

Worked example

The sample has 40 delivery times with a median of 33 minutes and one order at 95 minutes. For Sturges, log₂ 40 is 5.32, rounded up to 6, plus one gives 7 bins. For Freedman–Diaconis, the IQR is 37.25 − 30 = 7.25, and ∛40 is 3.42, so the width is 2 × 7.25 / 3.42 = 4.24 minutes. The range is 95 − 26 = 69, so 69 / 4.24 gives 17 bins. Scott uses the standard deviation of 10.88 instead: 3.49 × 10.88 / 3.42 = 11.1 minutes, which gives 7 bins.

The difference matters. At 7 bins, rounded to a width of 10, nearly everything lands in two bars (8 orders from 20 to 30 and 25 from 30 to 40). At the Freedman–Diaconis width, rounded to 5, the peak from 30 to 35 and the tail stretching to 50 both become visible. Scott's rule gives few bins here precisely because the one late order inflates the standard deviation. The IQR does not notice it.

Readable edges

By default the bin width is rounded to 1, 2, 2.5 or 5 times a power of ten, and the first edge is placed on a multiple of that width. Edges such as 25, 30, 35 are easier to read and to quote than 25.76, 30.00, 34.24. The cost is that the number of bins may differ by one or two from the rule's suggestion. Tick "Exact bin count" to split the range from the minimum to the maximum into exactly the chosen number of equal bins instead.

Limits

A histogram's shape depends on the bin width and on where the first edge falls. With fewer than about 20 values, almost any shape can appear by chance; a plain list or a dot plot is more honest at that size. Dates and times of day are not supported as numbers. Convert them to minutes or days first. If you need the numbers behind the shape, the summary statistics calculator gives the mean, median, quartiles and outlier fences for the same column.

Common questions

How many bins should a histogram have?

There is no single correct number, which is why this page shows five rules. Sturges suits small, roughly bell-shaped data. Freedman–Diaconis is the safest general choice because it uses the interquartile range and is not thrown off by outliers. Try two or three settings: a real feature of the data survives a change of bin width, an artefact does not.

What is the difference between a histogram and a bar chart?

A bar chart compares separate categories, and the gaps between bars say the categories are distinct. A histogram shows how one continuous measurement is distributed, so its bars touch: each covers a range of values, and the ranges meet end to end. The bar heights here are counts of values in each range.

Which bin does a value on the boundary go into?

The upper one. Every bin includes its lower edge and excludes its upper edge, so with bins of width 5 a value of 30 is counted in 30–35, not 25–30. The only exception is the last bin, which also includes its upper edge so the maximum is never lost.

Why does the chart have a different number of bins from the rule?

Because the edges are rounded to readable numbers. A rule might ask for a width of 4.24, which would give edges like 25.76 and 30.00. The tool rounds the width to 1, 2, 2.5 or 5 times a power of ten and starts on a multiple of it, which can add or remove a bin. Tick "Exact bin count" to use the rule's number exactly.

My data has one huge outlier and everything is squashed into one bar. What do I do?

That is what a histogram should show: the outlier really is far from everything else. To see the shape of the bulk, remove the outlier from the pasted data and redraw, and say that you did. The summary statistics tool flags values outside the 1.5 × IQR fences, which is a reasonable place to start deciding.

Other tools