Scale-Free Distributions: Stability, Power Laws, and Extreme Events
Complex Systems · 2/8 · Series index · Notation
The finite-variance baseline and MGF method are developed in Central limit theorem.
Scale-Free Distributions: Stability, Power Laws, and Extreme Events
Stable distributions describe how a distribution can retain its shape under addition and rescaling. Power laws describe tails without a characteristic scale. Together, these ideas explain why a normal approximation may fail and how to calculate probabilities and moments for heavy-tailed variables.
Stable Distributions and Domains of Attraction: What Survives Addition?
Let be independent and identically distributed, and define
A distribution is stable if addition preserves its shape up to a shift and a change of scale. More precisely, for every , there are constants such that
The symbol means “equal in distribution”; it does not mean that the two sides take identical values in each realization. If one can choose , the distribution is strictly stable. Stability is an exact property for every finite ; attraction is a limiting property as . For example, the uniform distribution is not stable, but its standardized sums approach a normal distribution, so it belongs to the normal distribution's domain of attraction.
Stability of the Normal Distribution
If , then
Independence allows the expectation of a product to factor into the product of expectations. Therefore,
For a standard normal distribution, , the identity holds. It shows that the width of the sum increases by a factor of . The width of the sample mean decreases by a factor of .
Self-Similarity Must Include Density Normalization
If , then
The factor comes from the change of variable and ensures that the density still integrates to . For standard normal summands,
Common mistake: the expression illustrates rescaling the argument of a function. It cannot be used directly as the density of , because that would omit the Jacobian. For nonnormal , the unstandardized sum also generally does not converge to a fixed normal distribution: it must first be centered and rescaled.
Why Do Domains of Attraction Matter?
Macroscopic observations are often sums of many small contributions. If different microscopic distributions approach the same limit under the same rescaling, predicting the macroscopic shape does not require knowing every microscopic detail. This is the probabilistic starting point for universality. Independence, dependence, and tail conditions must still be checked.
How Can the Assumptions of the Central Limit Theorem Be Relaxed?
The ordinary i.i.d. version assumes and , and concludes that
The following three changes require separate consideration. “There are many terms” alone does not justify applying the CLT.
Nonidentical Distributions: The Lyapunov CLT
Suppose the are independent but may have different means and variances . Define
If there exists such that the Lyapunov condition holds,
then
The condition says that the aggregate contribution of large deviations, measured by a higher moment, becomes small relative to the overall fluctuation scale. It prevents a few terms from completely dominating the sum. Here is the standard deviation of the sum, not the sum of the individual standard deviations. The Lyapunov condition is sufficient, but it is not necessary for every possible CLT.
Example: Independent Bernoulli Variables with Different Parameters
Let the be independent, and suppose for some . Choose . Since ,
The standardized sum therefore approaches a normal distribution even when the differ. To answer such a question, write down the means, variances, and ; choose ; bound the ratio; and state the limiting distribution.
Weak Dependence: The Coarse-Graining Intuition
If dependence persists only over short distances, grouping many observations into large blocks can make sufficiently separated blocks nearly independent, allowing a normal limit to survive. Condition: “small correlation” or “correlation tending to zero” alone is not sufficient to guarantee a CLT. Appropriate mixing or other dependence conditions are still needed. Being uncorrelated is also not the same as being independent.
For a stationary sequence, let . Then the exact variance is
When the covariances are absolutely summable and suitable CLT conditions hold, one typically obtains
The normalization must use the long-run variance . Long-range dependence can change the scaling and may even change the limiting distribution.
Infinite Mean or Variance
The assumptions of the ordinary finite-variance CLT fail, so a normal approximation cannot be used automatically. The Cauchy distribution in the next section provides a concrete counterexample. More generally, normalized sums may approach a non-Gaussian stable distribution. However, “the assumptions of this theorem fail” does not mean that “every possible form of a normal limit is impossible.”
Cauchy / Lorentz: Averaging More Does Not Always Improve Accuracy
The Cauchy, or Lorentz, distribution centered at with scale has density
Its center is its median and mode, not its expectation. Normalization can be checked using
.
Why Does the Cauchy Mean Not Exist?
The positive and negative halves of the integral must be checked separately:
The negative half-line contributes , so the ordinary expectation is the undefined expression . The symmetric principal value
does not establish the existence of the expectation. Also, , so . The logarithmic antiderivative has coefficient .
Why Switch to the Characteristic Function?
For this distribution, the moment-generating function diverges at every nonzero real . The characteristic function always exists, because :
The inversion formula requires appropriate integrability conditions. Those conditions hold here because the Cauchy characteristic function is integrable. The forward and inverse Fourier transforms use opposite signs, with a factor in the inverse transform.
For the Cauchy distribution, . Verification: the inverse transform gives an elementary check without requiring contour integration:
The Key Calculation for Independent Sums
Thus , whereas
The width of the sample-mean distribution does not shrink with sample size. For example, whatever the value of , . This does not contradict the usual law of large numbers, because the Cauchy distribution does not satisfy the sufficient condition of a finite first absolute moment.
Lévy Stable Distributions, the Generalized CLT, and Scaling Exponents
A symmetric stable family centered at zero has characteristic function
Here is the stability index, also associated with the tail exponent, while controls scale and has the same physical dimensions as . In this stable-family formula, is a scale parameter, distinct from its use for a density-tail exponent. A location parameter of zero does not imply that the mean exists: a symmetric stable distribution has no finite mean when . General stable distributions may also have location and skewness parameters.
Independence and uniqueness of characteristic functions give
The density and sample mean therefore scale as
This derivation explains three important cases:
- : the normal distribution. Comparing with the normal characteristic function gives . The width of the sample mean decreases as .
- : the Cauchy distribution. The sum scales by , while the scale of the sample mean remains unchanged.
- : the mean exists, but the fluctuation scale of the sample mean decreases as , more slowly than in the normal case. For , this scale instead increases.
Here “width” may be measured by a quantile-based scale such as the interquartile range. It cannot mean a standard deviation that does not exist. For example, with , the scale of is and that of is .
Power-Law Tails and Moments
Symmetric non-Gaussian stable distributions satisfy
and, for , if and only if . Thus every non-Gaussian stable distribution has a divergent second moment. In the symmetric case, a finite mean exists only when . Important endpoint: when , the tail is Gaussian. Substituting into the non-Gaussian power-law expression to obtain is incorrect. The normal distribution has finite absolute moments of every order.
The Generalized Central Limit Theorem
If the tails of i.i.d. random variables place them in the domain of attraction of a stable distribution, constants can be chosen so that
For regular power-law tails with index , the typical growth of is , possibly accompanied by a slowly varying factor in more general cases. When , the usual centering is . The earlier equality for every finite applies to exactly symmetric stable distributions.
Stable Distributions and Other Distribution Families
“Stable distributions are either normal or have power-law tails” describes the stable family. It cannot be extended to “all observations of complex systems must be normal or power-law distributed.” Empirical distributions may also have exponential tails, lognormal tails, mixture shapes, or finite cutoffs. The reverse claim, “every power-law distribution is strictly stable,” is also false: the Pareto distribution is generally not stable, although for suitable tail indices it can belong to a stable domain of attraction.
Calculating with Power Laws: Exponents, Tail Probabilities, and Moments
To avoid an off-by-one error in the exponent, write
Here is the density exponent, while is the cumulative-tail exponent. For , normalization gives
A pure power law cannot be integrable at both ends of the entire interval . A lower cutoff or another modification at small scales is therefore essential.
What Does “Scale-Free” Mean?
Within the range where the power law holds,
Multiplying the threshold by a fixed factor produces the same relative reduction in probability, regardless of the starting value . The tail has no single exponential decay scale of the kind found in . This does not mean that the distribution has no scale parameters at all, or that its mean must fail to exist. The Pareto lower cutoff , a stable distribution's scale parameter, and quantile-based widths can all exist.
From the Density to the Complementary CDF
Consequently, on logarithmic axes,
The PDF slope is ; the CCDF slope is ; the CDF itself is generally not a straight line. Identify the vertical-axis quantity before reading off an exponent. With logarithmic binning, raw bin counts also include the bin width, which increases with ; divide by bin width to estimate a density. A straight line is a diagnostic clue for a power law, but approximate linearity over a finite range is not a rigorous proof.
One Integral Gives the Moment Condition
For ,
At , the integral is of the form , which diverges logarithmically; equality does not count as finite. A finite mean therefore requires , and a finite second moment and variance require . For a positive random variable with , the expectation is . For a symmetric distribution with two heavy tails, the mean may instead be undefined. These are different statements.
Finite-System Cutoffs
If , all positive-order moments are finite, even though a broad intermediate range may still follow a power law. For ,
Thus a cutoff removes mathematically infinite moments, but their estimated values can remain highly sensitive to the upper limit and to extreme observations.
Pareto Wealth Distributions and Microscopic Exchange Models
The Pareto model has parameters :
For , the CDF follows by integrating , and the tail probability is . For any , the moment is
The mean, median, and variance are
The final expression follows by simplifying . Do not substitute into this formula and interpret a negative result as a “variance.” For , the variance is infinite. For , a finite mean already fails to exist.
The Gini Coefficient and Lorenz Curve
Assume . Let be the cumulative population share counted from the poorest individuals upward. The Lorenz curve gives the share of wealth held by those individuals:
Under perfect equality, . The Gini coefficient is twice the area between the equality line and the Lorenz curve:
A smaller means a heavier tail and greater concentration of wealth. For an ideal Pareto distribution with , the population Gini coefficient cannot be calculated directly from this finite-mean definition.
Worked Example: A Pareto Distribution with
Take . Then
The second moment and variance diverge, while the mean remains finite. The density has log-log slope , and the CCDF has slope . Interpretation: is one possible model of a heavy wealth tail. It does not imply that all countries, periods, or wealth ranges follow exactly this distribution. A finite sample also cannot prove that the true population variance of wealth is infinite.
Why Consider an Agent-Based Model?
Wealth-exchange models describe how trading and saving rules generate a distribution of wealth across individuals. When two individuals are selected,
This closed model conserves total wealth, much as colliding molecules exchange energy. A simulation must also constrain its trading rules so that it does not inadvertently produce wealth values that the model forbids, such as negative wealth. Conservation alone does not imply a power law: the exchange and saving rules determine the eventual distribution.
Kinetic theory and wealth exchange have the following correspondence:
| Kinetic theory of gases | Wealth-exchange model |
|---|---|
| Particles | Individuals or agents |
| Kinetic energy | Wealth |
| Collisions | Trades |
| Geometric dimension | Effective parameter |
The corresponding temperature scales are
The dimensionless variables are and , respectively. A Gamma-distribution shape function may be written ; it is distinct from the stable-distribution scale parameter . Simple exchange rules can generate population-level statistics with an approximately Gibbs-like body and a Pareto-like upper tail.
Earthquakes and Aftershocks: Identify the Variable Before Calling It a Power Law
Earthquake magnitude, amplitude, cumulative event counts, and probability densities are different quantities. Distinguishing them is essential because a change of variable changes the tail exponent and can change whether the mean is finite.
The Gutenberg–Richter Law
For a fixed region, observation period, and completeness threshold , the cumulative number of earthquakes satisfies
Normalizing within the set of events with gives
The magnitude density and mean are therefore
The tail of magnitude is exponential, not a power law. For and a finite lower threshold, this magnitude distribution has a finite mean.
Using the simplified relation , the amplitude instead satisfies
The same result follows from , where . Omitting this Jacobian changes the power-law exponent by one. The amplitude-tail index is . In the ideal untruncated case with , the amplitude mean diverges, while the magnitude mean remains finite. Actual magnitude definitions also include instrumental, distance, and other corrections, so is not a complete measurement formula.
The meaning of the cumulative count can be checked against the USGS description of the Gutenberg–Richter law. The probability transformation and mean above then follow directly from the equations.
Omori's Law for Aftershocks
An aftershock decay law , with , describes the event rate per unit time. A commonly used regularized form is
where removes the singularity at . When , the expected event count over is
A rate is not a normalized probability density. The relation alone does not establish that “the mean aftershock time is infinite, so aftershocks never stop.” If a separate probability model is defined by for , normalization requires , and a finite mean requires . With , a model over an infinite time interval cannot even be normalized. This is another reason to handle exponent boundaries carefully.
Consequences of Having No Characteristic Scale
The St Petersburg paradox. The first head of a fair coin appears on toss , where , and the payout is . Then
Nevertheless, : a finite outcome almost surely is entirely consistent with an infinite expectation. For example, . Pricing only by risk-neutral expected break-even value gives no finite fair entry fee. Wealth limits, payout caps, risk preferences, or utility functions change the pricing problem. If is the toss on which the first head occurs, the number of preceding tails is ; consistent indexing is essential.
Extreme events are not automatically outliers to discard. A heavy-tailed distribution assigns much greater probabilities to large events than a Gaussian model does. If very large and small events arise from the same mechanism, removing the largest observations simply because of their size systematically underestimates tail risk. Biological extinctions and large market crashes illustrate extreme events. In financial models, option prices can be sensitive to distribution tails and volatility. The methodological lesson is: check data quality and the generating mechanism before deciding that a point is anomalous; “rare” is not equivalent to “bad data.” A single histogram or market crash does not rigorously establish an exact power law for the entire tail.
Econophysics: returns, memory, and volatility. The efficient-market hypothesis, Brownian motion, normality of returns, existence of variance, and short or long memory are distinct questions. The efficient-market hypothesis does not itself imply normally distributed returns. GARCH is a family of models for a conditional variance that changes over time, allowing volatility clustering: large fluctuations tend to be followed by further large fluctuations. A volatility smile describes implied volatility, inferred through the same underlying pricing model, varying with strike price.
Structures across scales and the motivation for modeling. Cosmic superstructures raise questions about patterns across scales. Avalanche distributions in systems of coupled pendulums show how local interactions can generate a population-level power law. Simple mechanisms can therefore explain collective phenomena and statistical patterns without predicting every individual catastrophe. A power law alone does not uniquely identify the microscopic mechanism.
Self-Test
Check whether you can solve the following questions independently.
-
Let for . Find and the CCDF, and determine whether the mean and variance are finite.
Answer: , , and . The variance diverges.
-
An empirical CCDF has log-log slope . What is the density-tail exponent? Does the mean of the positive random variable exist?
Answer: and . Under the ideal infinite-tail model, the mean is .
-
Let be independent standard Cauchy variables, with . What are the scales of and ?
Answer: and , respectively. The formula does not apply.
-
A symmetric stable distribution has index . If the sample size increases from to , how does the typical width of the sample mean change?
Answer: it is multiplied by . There is no finite standard deviation to rescale in this way.
-
Does mean that the magnitude density is proportional to ?
Answer: no. is a shifted exponential variable. If , the amplitude CCDF is proportional to and its density to .
Problem-solving sequence: define the random variable and its support; identify whether the given quantity is a PDF, CDF, CCDF, or event rate; put the exponents into a consistent notation; check moments and independence; then decide whether a CLT applies and how to normalize. These five steps expose the most common errors: an exponent off by one, a nonexistent moment, or a missing scale factor.
← Central limit theorem · Series index · Scale-free networks →

