Two-sample t test power calculation
n = 25
delta = 9.26214
sd = 8.9
sig.level = 0.05
power = 0.95
alternative = two.sided
NOTE: n is number in *each* group
Chapter 19: Power Tests
Shih Chien University
2026-08-03
Before collecting data, two questions decide an experiment’s fate:
R’s power.* family answers both. The vocabulary: significance level = Type I error probability (false positive); power = 1 − Type II error probability (1 − false negative). Each function computes whichever parameter you leave as NULL.
Testing an antidepressant, with depression measured on the HAMD scale (0–48). Assume sd = 8.9 (from FDA materials), significance 0.05, desired power 0.95. With 25 subjects per arm, what group difference would be detectable?
Two-sample t test power calculation
n = 25
delta = 9.26214
sd = 8.9
sig.level = 0.05
power = 0.95
alternative = two.sided
NOTE: n is number in *each* group
A difference of at least 9.26 HAMD points. Double the subjects:
Two-sample t test power calculation
n = 50
delta = 6.480487
sd = 8.9
sig.level = 0.05
power = 0.95
alternative = two.sided
NOTE: n is number in *each* group
Now 6.48 suffices. This is experiment design in action: trade sample size against detectable effect before spending a dollar on data collection.
| Argument | Description |
|---|---|
n |
observations per group |
delta |
true difference in means |
sd |
true standard deviation |
sig.level |
significance level (Type I error probability) |
power |
1 − Type II error probability |
type |
two-sample / one-sample / paired |
alternative |
two-sided or one-sided |
strict |
strict interpretation in the two-sided case |
Supply four of n, delta, sd, sig.level, power; leave the fifth NULL — that’s what gets computed.
For experiments measured with prop.test:
p1/p2 are the success probabilities in the two groups; the rest as before. Again: four of n, p1, p2, sig.level, power given, the NULL one computed.
ESPN flashes stats like “3-for-10 with two on and two out.” Suppose a batter hits .300 in some situation but .260 otherwise — how many situational at-bats make .300 significant (α = 0.05, power 0.95, one-sided)?
Two-sample comparison of proportions power calculation
n = 2724.482
p1 = 0.26
p2 = 0.3
sig.level = 0.05
power = 0.95
alternative = one.sided
NOTE: n is number in *each* group
Over 2,724 at-bats — per group. Flip the question: with n = 10, what’s achievable?
[1] 0.9256439
[1] 0.07393654
Significance level 0.93, or power 0.07. The book’s verdict: most situational statistics are nonsense — now you can prove it.
For multi-group designs analyzed with ANOVA:
| Argument | Description |
|---|---|
groups |
number of groups |
n |
observations per group |
between.var / within.var |
variance between / within groups |
sig.level / power |
as before |
Here five of the six parameters are specified; the one left NULL is calculated.
Balanced one-way analysis of variance power calculation
groups = 4
n = 11.92613
between.var = 1
within.var = 3
sig.level = 0.05
power = 0.8
NOTE: n is number in each group
Tip
Habit for every empirical project: run the power calculation before the experiment. It converts wishful thinking (“we’ll see if it’s significant”) into design (“we need n per group to detect the effect we care about”).
Copyright. These slides are adapted from R in a Nutshell: A Desktop Quick Reference (2nd ed.) by Joseph Adler, O’Reilly Media. All rights reserved by the original author and publisher.
Non-commercial use only. These materials are strictly for educational purposes and may not be used for commercial gain.
Attribution. Any reproduction, distribution, or use of these materials must properly credit the original source.
R in a Nutshell: A Desktop Quick Reference