R Bloggers iconR BloggersSep 15, 2026 ~7 min source read

How large a sample for the CLT to give you approximately normal sample means?

Simon Ellis runs simulations showing that for skewed populations the sample mean can remain far from normal even for sample sizes much larger than the textbook rule-of-thumb, and that bootstrap intervals can improve coverage when asymptotic normality fails.

Sample size needed for central limit theorem to kick in by @ellis2013nz

Share this story

Send the public story page.

Useful takeaways from this story.

For symmetric (normal) populations the sample mean looks normal at very small n, but for skewed populations (e.g., log-normal) sample means can remain non-normal even at n = 100.

A later comparison shows bias-corrected and accelerated (BCa) bootstrap confidence intervals give better coverage than naive asymptotic intervals, particularly at small sample sizes.

# Short answer The central limit theorem (CLT) guarantees that sample means converge to a normal distribution as sample size increases, but "how large" depends on the population. For symmetric, light-tailed populations the mean behaves normally for very small n. For heavily skewed populations, simulations here show that even n = 100 can leave the distribution of sample means visibly non-normal and that standard 95% CIs based on asymptotic normality can miss the true mean much more than 5% of the time.

# What the author did

  • calculates the empirical coverage of 95% confidence intervals built by assuming the sample mean is normally distributed (using the sample SD to estimate the standard error).

The simulations use a large population (N = 1e6) and many repetitions (default 10,000) to estimate the distribution of the sample mean precisely.

  • From a standard normal population with n = 5, the AD statistic was 0.45 and the 95% CI built under asymptotic normality covered the true mean about 87% of the time. Ellis notes this is because he deliberately avoids using the t-distribution (which would be more appropriate in that specific case) to test reliance on asymptotic normality alone.
  • From a heavily skewed log-normal population (generated as exp(N(0,1))), a sample size of n = 5 shows clear non-normality of the sample mean. Increasing sample size to n = 100 improves things but still does not produce a "satisfactory" normal distribution of the sample means according to the author's graphical checks.

# Wider sweep of simulations

# Follow-up: bootstrap vs asymptotic CIs A short sequel tests whether bootstrap confidence intervals help. The follow-up compares the traditional asymptotic 95% CI to a bias-corrected and accelerated (BCa) bootstrap interval across the same skewed distributions and sample sizes. Results reported in that sequel: the BCa bootstrap gives substantially better coverage than naive CLT-based intervals, especially at smaller n, though it is not claimed as a universal fix for every scenario.

# Practical takeaway Don't rely on a fixed rule-of-thumb like n = 30 for the CLT to make the sample mean approximately normal in every situation. For skewed or heavy-tailed populations you may need much larger n for asymptotic normality to be a good approximation. When sample size is limited, use diagnostic plots or tests on the sampling distribution (or simulation if you can generate a plausible population), and consider bootstrap CIs (BCa) as a practical alternative to naive normal-based intervals.

# Code and reproducibility

More context around this story.

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app