Distributions and heavy tails

How to make a CCDF plot

To make a CCDF plot, sort the positive observations, give each one the fraction of the sample at or above it, P(X ≥ x), and plot those points on log–log axes. No bins are involved, so every observation keeps its place. Build and read the empirical curve first; a fitted tail model is a separate layer with its own range.

Updated · 5 min read

On this page
  1. How to plot a CCDF
  2. Ties and the P(X ≥ x) convention
  3. How to read a CCDF on log–log axes
  4. Setting the fit bounds
  5. What to report with a CCDF plot
  6. CCDF plots in Autoplot
  7. Frequently asked questions

How to plot a CCDF

  1. Define the population. Keep finite, strictly positive observations, and record how many you removed and why; zero and negative values have no place on a log axis.
  2. Sort the n included values in ascending order, x₍₁₎ ≤ x₍₂₎ ≤ … ≤ x₍ₙ₎.
  3. Assign the i-th sorted value the share of observations at or above it: P(X ≥ x₍ᵢ₎) = (n − i + 1) / n. The smallest value gets 1, the largest 1/n.
  4. Plot P against x with both axes logarithmic, as points rather than a joined line.
  5. Note n on the figure or in the caption. The lowest point is always 1/n, so the depth of the tail is set by the sample size.

In Python

After x = np.sort(x[x > 0]), the empirical CCDF is p = 1 - np.arange(len(x)) / len(x), which runs from 1 to 1/n; plot it with ax.loglog(x, p, "."). SciPy's ecdf(x).sf returns the survival function P(X > x) instead, which reaches zero at the largest value and drops that point from a log plot.

Ties and the P(X ≥ x) convention

With continuous data the convention barely matters. With repeated values it does: P(X ≥ x) and P(X > x) differ at every tie, and a ranking that gives each tied value its own step draws a short vertical run where a strict survival function would draw one point.

Rounded or integer-valued data make this visible as staircases at the low end. Decide whether to keep one point per observation or one per distinct value, and say which.

Autoplot Observed CCDF card showing the waiting_time_s variable and Power Law fit type
The card takes one positive variable and a fit family. The empirical curve does not depend on the family; only the overlaid model does.

How to read a CCDF on log–log axes

Equal distances on the axes are equal ratios. The upper left of the plot holds the whole sample at P = 1; each decade down the vertical axis holds ten times fewer observations. The last decade above 1/n rests on fewer than ten values.

Neighbouring points share almost the entire sample, so the curve is smoother than the information it carries. Do not read wiggles in the lowest decade as structure, and do not judge the fit there by eye.

What you seeWhat it can meanWhat to check
A straight segment over several decadesConsistent with a power-law tail over that rangeSlope is −(α−1); test against lognormal and truncated alternatives
A downward bend at large xCutoff, finite system size, or a truncated distributionWhether the bend survives a larger sample or a different xmax
Gentle curvature throughoutLognormal-like or stretched-exponential behaviourA power-law fit to a curving CCDF depends heavily on xmin
Steps or plateausTies, rounding or discrete valuesThe measurement resolution

To decide between those readings, read how to fit a power law to data, which covers the tail fit, the choice of xmin and the evidence a claim needs. To decide whether a CCDF is the right view at all, see PDF vs CDF vs CCDF; to see the same tail as a density, build a log-binned histogram.

Autoplot observed CCDF shown as blue points on log10 axes with a red power-law tail fit
An empirical CCDF in blue with a power-law tail model in red. In the lowest half-decade the points drift above the line; that region rests on the fewest observations, so it is where a deviation is both most visible and least certain.

Setting the fit bounds

A tail model is fitted from a lower cutoff, xmin, upward, optionally below an upper cutoff, xmax. Both are analysis choices, not display settings.

  • xmin sets where the tail begins. Leaving it at the smallest value fits the whole sample, which is rarely what a tail claim means. A Kolmogorov–Smirnov scan over candidate cutoffs gives a defensible starting point; the power-law fitting guide describes it.
  • xmax removes observations above it. Use it for a known instrument limit, not to trim an inconvenient end of the tail.
  • The tail count, the number of observations at or above xmin, decides how much the fit can say. A stable-looking line through a few dozen points is not a stable estimate.

Many tools, Autoplot among them, take the bounds as log10 values. Typing 3 means 1,000; report the natural value.

What to report with a CCDF plot

  • Source, total count and included positive count, with the reason for every exclusion.
  • The convention, P(X ≥ x) or P(X > x), and how ties were handled.
  • Any xmax applied before the curve was built.
  • For a fitted tail: the family, xmin and xmax in natural units, the tail count, the fitted parameters and which exponent convention they follow.
  • Anything display-only, such as an offset or an extrapolated fit line.
In Autoplot

CCDF plots in Autoplot

The Log-space CCDF card (Run Analysis, Log-Space CCDF Fit, in an X&Y figure) is part of Plus. It keeps values greater than zero, needs at least ten of them, applies an optional xmax first, and builds the empirical P(X ≥ x) from the full filtered sample with ranks from n/n down to 1/n. Tied values each keep their own point. Above 10,000 points, the stored and plotted curve is deterministically downsampled to 10,000.

The fit family is Power Law, Truncated PL, Log-Normal or Stretched Exponential, fitted with the powerlaw library on a subsample of at most 1,000 values. log₁₀ xmin and log₁₀ xmax take log10 values. A blank xmin means the smallest positive value, not an estimated cutoff, and xmin bounds the fit and the tail count without cropping the empirical points. The model line is scaled by the tail fraction so it sits on the full-sample curve; the Results block reports the exponent, the KS distance, the range and that normalisation.

The card does not export its points, and it does not run a goodness-of-fit test or a lognormal comparison. For a cutoff, run the Xmin diagnostic card; for the tests, use your own Python environment or ask the assistant to write the analysis and run it on your Mac. See the features page and pricing.

Frequently asked questions

Frequently asked questions
Should a CCDF be plotted on log–log axes?For heavy-tailed data, yes: log–log axes spread the tail over decades and turn a power law into a straight line. For a light-tailed variable, a log vertical axis alone, or linear axes, may be clearer.
Why does my CCDF look smoother than my histogram?Because each point accumulates all larger observations, so neighbouring points share most of their data. The smoothness is not extra information, and the lowest points remain noisy.

Sources

Autoplot statements are based on the app's documentation for the Log-space CCDF fit and its CCDF model algorithm, checked on 9 October 2026, and the features page. SciPy's ecdf reference describes empirical CDF and survival-function estimates. Newman, Power laws, Pareto distributions and Zipf's law, explains why cumulative plots are preferred for heavy tails. Clauset, Shalizi and Newman, Power-law distributions in empirical data, sets out the fitted-range, goodness-of-fit and model-comparison requirements. The powerlaw package is the library behind the tail fit.

Try it on your own data.

Autoplot's Log-space CCDF card builds the empirical P(X ≥ x) from your raw values and overlays a tail model. It is part of Plus, which is free for the first 30 days.

Fit a power law to the tail →
Download free for Mac Mac App Store

Apple silicon · macOS 15 or later · Free plan · Plus free for 30 days

Send to my Mac

Autoplot runs on a Mac. Send the link to yourself, or open it in the Mac App Store.