Distributions and heavy tails

Histogram normalization: counts, proportions and density

Histogram normalization decides what the height of a bar means. Counts give observations per bin, proportions give each bin's share of the sample, and probability density gives that share per unit of x, so that bar areas, not heights, sum to one. A bar needs a denominator: choose it from the question, and label the axis with the quantity and its units.

Updated · 4 min read

On this page
  1. The four histogram normalizations compared
  2. Density vs frequency histogram
  3. How to normalize a histogram
  4. Comparing groups and overlaying models
  5. Normalisation in Autoplot
  6. Frequently asked questions

The four histogram normalizations compared

For bin i with count nᵢ and width wᵢ, out of n included observations:

QuantityBar heightSums toAnswers
Count (frequency)nᵢn, over the heightsHow many observations fell in this interval
Proportion (relative frequency)nᵢ / n1, over the heightsWhat share of the sample fell here
Percentage100 · nᵢ / n100%, over the heightsThe same share, in percent
Probability densitynᵢ / (n · wᵢ)1, over height × widthHow concentrated the sample is per unit of x

Counts, proportions and percentages differ only by a constant, so with equal-width bins they draw the same shape on a different axis. Density divides by width as well, which changes the shape as soon as bins differ in width.

Density vs frequency histogram

A frequency histogram answers "how many" or "what share". A probability density histogram answers "how much per unit of x", and it is the only one of the four that estimates a probability density function.

Three consequences follow from the division by width:

  • Density has units. If x is a diameter in micrometres, density is in µm⁻¹. Label the axis that way.
  • Heights can exceed 1. Narrow bins on a variable with a small range produce large densities. Only the area is bounded.
  • Heights change with bin width. Halving the width redistributes the same probability over narrower intervals, so individual bars can rise even though nothing about the sample changed.

With unequal widths, density is the only honest choice. A count or proportion axis makes a wide bin look important partly because it covers more of x; dividing by width removes that effect. This is why log binning always reports density.

Autoplot cell-diameter histogram with density-normalised blue bars, a PDF vertical axis, and Gaussian and mixture curves
The PDF axis label marks a density view: probability is the area of each bar. With diameter in micrometres, the axis is in inverse micrometres, and the fitted curves are on the same scale.

How to normalize a histogram

  1. Fix the edges and the included sample. Record the total n after removing non-finite values, range limits and scientific exclusions.
  2. Count the observations in each bin.
  3. For proportions, divide every count by n. For percentages, multiply the proportions by 100.
  4. For density, divide each count by n and by that bin's own width.
  5. Check the total: counts sum to n, proportions to 1, percentages to 100, and Σ density × width to 1, apart from rounding.
  6. Label the axis with the quantity and, for density, its units.

In NumPy and Matplotlib

np.histogram(x, bins=edges) returns counts; adding density=True returns density, dividing by each bin's width even when edges are unequal. For proportions, pass weights=np.ones(len(x)) / len(x) and leave density off. Matplotlib's hist takes the same arguments. NumPy normalises over the observations inside the edges, so values outside the range are silently left out of n.

Comparing groups and overlaying models

Use one range and one edge array for every group; the histogram bin width guide covers how to choose them. Then pick the normalisation from the comparison:

  • Counts when absolute numbers matter and the groups had comparable observation opportunities.
  • Proportions or percentages when the groups differ in size and the question is each interval's share.
  • Density when you compare shapes across different bin widths, or overlay a fitted density.

No normalisation repairs incompatible sampling. A group recorded for twice as long, filtered by another threshold or truncated to a different range answers a different question, whatever the axis says.

A probability density model integrates to one, so it does not sit on a count axis. Either plot the histogram as density, or scale the model by n · w for equal-width bins and say so. Weighted histograms need the same care: state what the weights mean and which total they are normalised by.

In Autoplot

Normalisation in Autoplot

Autoplot's Histogram card, in the free plan, is always density-normalised. There is no count, proportion or percentage toggle: the bars and the values passed to the Gaussian-mixture or custom fit are densities from NumPy's density=True, plotted against arithmetic bin centres. The log-binned PDF card is density-normalised in the same way.

With Pre-binned data, the card accepts centres and densities you supply. It cannot check that they share edges, units or normalisation, so verify that before fitting.

If a figure needs counts or percentages, compute them in your own environment, or ask the assistant to write the Python, which runs on your Mac and stays with the project. Do not relabel the density axis. For two-dimensional data, the Joint Density Map in a heat map figure (Plus) does offer Density, Probability or Count per cell; see how to make a scientific heat map and the features page.

Autoplot Histogram controls showing a source variable, 72 bins, automatic range, bars, and no separate normalisation selector
The card controls the bins, range and display. Normalisation is fixed to density, so a different axis label would not turn the bars into counts.

Frequently asked questions

Frequently asked questions
Should a density histogram sum to one?The areas should, not the heights. Multiply each density by its bin width and add them up; the result is one, apart from rounding and any observations outside the edges.
Why are my density values greater than 1?Because density is probability per unit of x. With narrow bins or a variable measured on a small scale, heights above 1 are normal; the total area is still one.
When should I use a frequency histogram instead of density?When readers need absolute numbers, the bins have equal width and you are not overlaying a probability density. Use proportions instead when the compared groups differ in size.

Sources

Autoplot statements are based on the app's documentation for the Histogram card, binning and density, and the Joint Density Map, checked on 9 October 2026, and the features page. NumPy's histogram reference defines counts, weights and density normalisation. Matplotlib's histogram bins, density and weight example shows counts, 1/N weights and density side by side, and why only density keeps its shape when bin widths change. The NIST/SEMATECH handbook's histogram page describes the relative histogram, normalised either by n or by n times the class width.

Try it on your own data.

Autoplot's free Histogram card draws a density-normalised histogram with the bin count and range on the card, ready for a Gaussian-mixture or custom fit.

Choose the bin width →
Download free for Mac Mac App Store

Apple silicon · macOS 15 or later · Free plan · Plus free for 30 days

Send to my Mac

Autoplot runs on a Mac. Send the link to yourself, or open it in the Mac App Store.