Distributions and heavy tails

How to choose a histogram bin width

Choose a histogram bin width with a documented rule, usually Freedman–Diaconis for skewed or outlier-prone data and Scott's rule for data close to normal, then check that your conclusion survives nearby bin counts on the same range. No rule finds a true bin width. Each one proposes a starting point from the sample size and spread, and the bars still depend on where the edges fall.

Updated · 5 min read

On this page
  1. Freedman–Diaconis, Scott's rule and Sturges compared
  2. How many bins should a histogram have
  3. What a bin-width sensitivity check should look for
  4. Bin width changes the fit as well as the picture
  5. What to report
  6. Choosing bins in Autoplot
  7. Frequently asked questions

Freedman–Diaconis, Scott's rule and Sturges compared

Most rules give a bin width h; the number of bins follows as k = ⌈(max − min) / h⌉ over the range you fix. Here n is the number of included observations, IQR the interquartile range and σ the sample standard deviation.

RuleWidth or countGood forCaveat
Freedman–Diaconish = 2 · IQR · n^(−1/3)Skewed data, outliersMany bins when a long tail stretches the range; fails when ties make the IQR zero
Scott's ruleh = 3.49 · σ · n^(−1/3)Roughly normal dataOutliers inflate σ and widen every bin
Sturges' rulek = ⌈log2(n)⌉ + 1Small, near-normal samplesToo few bins for large or non-normal samples
Square rootk = ⌈√n⌉A quick first lookIgnores spread and shape

NumPy's bins="auto" takes the smaller width of Sturges and Freedman–Diaconis, which in practice means Sturges for small samples and Freedman–Diaconis for large ones. It is a sensible default, and it still knows nothing about the scientific meaning of a feature.

How many bins should a histogram have

As many as the data can support around the structure you care about, and no more. A useful procedure:

  1. Decide which observations belong in the histogram. Remove non-finite values, record scientific exclusions and say whether out-of-range values were excluded or merely hidden.
  2. Fix the range. Changing it changes the represented population and the width of every equal-width bin.
  3. Compute a starting width with one of the rules above and convert it to a bin count for that range.
  4. Render at least one coarser and one finer version, for example half and double the count, on the same range.
  5. Shift the edges by a fraction of a bin. A narrow cluster can be split by an edge in one view and sit inside one bar in the next.
  6. Keep the conclusions that survive steps 4 and 5, and freeze the final edges before fitting anything to the bars.

The measurement's resolution is a floor. Bins narrower than the instrument can resolve, or than a rounding step in the data, produce combs and empty intervals that describe the recording, not the variable.

Autoplot Histogram controls showing the cell diameter variable, 72 bins, bar display, automatic xmin and xmax, and a Gaussian fit
The bin count and natural-value range sit on the card next to the fit settings. They define the bars before any styling, so they belong in the figure's methods note.

What a bin-width sensitivity check should look for

Compare the views for the features you intend to report, not for overall smoothness.

  • Modes: a second peak that appears at only one bin count, or vanishes with a small edge shift, is not yet a result.
  • Gaps and shoulders: a gap that persists across counts may be real; one that appears only with narrow bins is often sampling noise.
  • Skew and tails: wide bins compress a tail into one or two bars. If the tail is the question, read log binning or use a cumulative view from PDF vs CDF vs CCDF, which needs no bins at all.

Very wide bins hide variation; very narrow bins amplify noise. When a feature survives only in a narrow window of choices, inspect the raw observations before describing it.

Comparing groups

Compute one range and one edge array for all groups. Two histograms with separate automatic ranges can have the same bin count and still cover different intervals, so equal-looking bars are not comparable. If the groups differ in size, compare proportions or densities rather than counts; the histogram normalization guide sets out which.

Autoplot density histogram of cell diameter with narrow blue bars, a PDF vertical axis, and separate Gaussian and mixture fit curves
At 72 bins the two modes are clear and a single Gaussian visibly misses them. Whether the dip between the modes survives 36 and 144 bins on the same range decides whether it belongs in the paper.

Bin width changes the fit as well as the picture

A curve fitted to a histogram is fitted to the bin centres and heights, not to the observations. Changing the bin count or the range changes those points, so it can change the fitted parameters and their errors.

Refit at the neighbouring bin counts from your sensitivity check. If a parameter moves by more than its reported error, the binning is part of the estimate and should be reported as such. For the model-checking side, see how to fit a curve to data and how to interpret residual plots.

What to report

  • Source variable, units and the number of included observations.
  • Exclusions: non-finite values, scientific filters and the range limits.
  • The rule used as a starting point, the final bin count and the width, or the full edge array if bins are unequal.
  • The normalisation of the vertical axis.
  • For a fitted histogram: the fit settings, and the number of non-empty bins that entered the fit.
In Autoplot

Choosing bins in Autoplot

The Histogram card (Run Analysis, Histogram Fit, in an X&Y figure) is in the free plan. It takes a Bins value, clamped to 1–1000, and natural-value xmin and xmax fields; left empty, they use the data minimum and maximum. Edges are linearly spaced across that range, heights are always density-normalised, and centres are arithmetic midpoints. Negative values and zero are allowed.

Empty bins are drawn but do not enter the Gaussian-mixture or custom fit, which needs at least three non-empty bins per Gaussian component. Because bins and range stay on the card, a sensitivity check is a matter of changing one field and refitting.

Autoplot does not compute Freedman–Diaconis, Scott or Sturges for you. Compute the width with NumPy's histogram_bin_edges, or ask the assistant, which writes and runs the Python on your Mac, then type the resulting count into Bins with the range fixed. For bins of unequal width, positive data across decades use the log-binned PDF card instead; see the features page.

Frequently asked questions

Frequently asked questions
How many bins should a histogram have?There is no single answer. Start from Freedman–Diaconis or Scott's rule for your range, then check that the features you report survive half and double that count.
Is Freedman–Diaconis better than Sturges' rule?For large or skewed samples, usually yes. Sturges depends only on sample size and gives too few bins as n grows; Freedman–Diaconis scales with the interquartile range and resists outliers.
Should two histograms I compare use the same bins?Yes. Use one range and one set of edges for both, otherwise the bars cover different intervals and cannot be compared directly.

Sources

Autoplot statements are based on the app's documentation for the Histogram card and its binning, checked on 9 October 2026, and the features page. NumPy's histogram_bin_edges reference gives the estimator formulas, the auto rule and their limits; its histogram reference defines edges and density. The NIST/SEMATECH handbook's histogram page describes the histogram as a summary of a univariate distribution whose classes are set by the user or by a rule.

Try it on your own data.

Autoplot's Histogram card is in the free plan. It keeps the bin count and range on the card, so a sensitivity check is a matter of changing one number and comparing.

Decide what the bar heights mean →
Download free for Mac Mac App Store

Apple silicon · macOS 15 or later · Free plan · Plus free for 30 days

Send to my Mac

Autoplot runs on a Mac. Send the link to yourself, or open it in the Mac App Store.