Learn & Understand

Statistical Power: The Fuller Picture Behind Sample Size

In a hurry? Skip straight to the numbers.

Open the Data Sample Size Calculator →

The companion calculator uses a standard formula to turn a target confidence level and margin of error into a required sample size for a survey. That is one important use of sample-size planning, but it addresses only estimation precision. For experiments and A/B tests, whose goal is to detect an effect rather than estimate a proportion, a richer framework governs how big the sample must be, built around statistical power. Understanding power and the levers that determine it explains why so many studies fail to find real effects, and why sample size cannot be chosen in isolation.

Estimation vs Detection

There are two distinct sample-size questions, and conflating them causes trouble. One is estimation: how large a sample do I need to estimate a proportion within a certain margin of error, the question the survey formula answers. The other is detection: how large a sample do I need to reliably detect a real effect if one exists, the question that matters for experiments and A/B tests. The second is governed by statistical power, and it involves considerations the margin-of-error formula does not capture.

The Interlocking Levers

Statistical power, the probability of detecting a real effect, is bound up with several quantities that trade off against each other.

The interlocking levers of experiment design
LeverEffect on required sample size
Larger effect to detectSmaller sample needed
Higher power (fewer missed effects)Larger sample needed
Stricter significance (fewer false positives)Larger sample needed
Noisier dataLarger sample needed

These are linked: fix any three and the fourth is determined. This is why sample size cannot be chosen alone, it depends on how small an effect you want to catch, how confident you want to be of catching it, how strict you are about false alarms, and how noisy the data is. The survey formula fixes the analogues of some of these implicitly; power analysis makes them all explicit.

Effect Size Is the Crux

The single most important, and most often neglected, input is the effect size, how large a difference you are trying to detect. Small effects require dramatically larger samples to detect reliably than large ones, and this relationship is steep. A study designed to catch a big, obvious effect can be small; one hoping to detect a subtle effect needs to be large, often far larger than intuition suggests. Failing to think about effect size in advance is why so many experiments are the wrong size, either wastefully large or hopelessly small for the effect they seek.

The Cost of Underpowered Studies

An underpowered study, one too small to reliably detect the effect it seeks, is worse than it appears. It frequently fails to find a real effect, wasting the entire effort, and, more insidiously, when such a study does report a significant result, that result is more likely to be a fluke or an exaggeration, because only unusually large sample estimates cross the significance line in a small study. Chronic underpowering is a major contributor to unreliable, non-reproducible findings. Adequate power is not a nicety; it is what makes a study's conclusions trustworthy, in both directions.

The Danger of Peeking

A final trap specific to experiments: the sample size should be decided in advance and adhered to. Repeatedly checking results as data accumulates and stopping as soon as significance appears, peeking, dramatically inflates the false-positive rate, because random fluctuations will eventually cross the line by chance. This is one form of the broader problem of p-hacking. Fixing the sample size beforehand through a power analysis, and analyzing once, protects the validity of the result. Proper sample-size planning is thus not just about having enough data, but about committing to a design before the data can tempt you.

Planning Sample Size Fully

Use the calculator for the estimation question of survey precision, and for experiments extend to the power framework: decide the effect size worth detecting, the power and significance you require, and account for noise, recognizing these levers determine the sample together. Avoid underpowered studies and pre-commit to the sample size to resist peeking. The calculation sizes an estimate; understanding statistical power is what properly sizes an experiment.

Ready to Put This Into Practice?

Now that you understand how it works, plug in your own numbers and get an instant, accurate result.

Use the Data Sample Size Calculator Now →