What is Central Limit Theorem: Why Sums of Random Variables Become Normal

Why Tutoring - California Graduate Tutor
Central Limit Theorem: Why Sums Become Normal
Submit Homework

The central limit theorem (CLT) is one of the most important results in statistics tutoring. It explains why averages and sums of random variables behave normally, even when the original data are not normal. This page breaks down the intuition, conditions, and exam‑relevant implications of the CLT.

The central limit theorem states that the standardized sum of independent, identically distributed random variables converges in distribution to a normal distribution.

\[ \frac{\sum_{i=1}^n X_i – n\mu}{\sigma\sqrt{n}} \xrightarrow{d} N(0,1) \]

Why does the CLT work? Because when many small, independent contributions add together, their combined effect smooths out irregularities in the individual distributions. The normal distribution emerges as the universal limit due to how moment generating functions and characteristic functions behave under convolution. This is why the normal distribution appears everywhere in nature, finance, and data science.

  1. Start with i.i.d. random variables. Each has mean \(\mu\) and variance \(\sigma^2\).
  2. Form the sum or sample mean. \(S_n = \sum X_i\), \(\bar{X}_n = S_n / n\).
  3. Standardize the sum. Subtract the mean and divide by the standard deviation.
  4. Analyze the characteristic function. The characteristic function of the standardized sum converges to that of a standard normal.
  5. Apply Lévy’s continuity theorem. Convergence of characteristic functions implies convergence in distribution.
  6. Interpret the result. Large samples behave normally regardless of the original distribution (with mild conditions).

Suppose \(X_i\) are Bernoulli(0.3). Mean = 0.3, variance = 0.21.

Let \(n = 100\). Then the standardized sum is:

\[ Z = \frac{S_{100} – 30}{\sqrt{21}} \]

Even though Bernoulli variables are discrete and skewed, \(Z\) is approximately standard normal by the CLT.

  • Thinking the CLT requires normal data (it does not).
  • Applying the CLT to small samples without checking skewness.
  • Forgetting that independence is required.
  • Confusing convergence in distribution with convergence in probability.

The CLT underlies confidence intervals, hypothesis tests, regression theory, and nearly all large‑sample approximations. It appears in every graduate statistics exam and is foundational for modern statistical inference.

This idea connects directly to:

Speak Directly to a Tutor — Send Your Message Below

No call centers. No delays. Your message goes straight to the tutor.

Get help with probability theory, distributions, asymptotics, hypothesis testing, and mathematical statistics.