Volume 8 · eight days

No analysis can rescue a study that was too small.

An underpowered study does not give you a cautious answer; it gives you an unstable one, with inflated effect sizes when it happens to reach significance and no information when it does not. This volume covers analytic power where formulas exist, simulation-based power where they do not — including mixed models and SEM — sensitivity analysis after the fact, sequential designs, and multiverse analysis as a robustness result rather than an apology.

8
days, before you collect anything
80
per cent power is a floor, not a target
1
simulation recipe covers every design
Day 77 45 minutes · the argument, before the arithmetic

Why an underpowered study cannot be saved by analysis

Power is the probability of detecting an effect that is really there. Low power has three consequences, and only the first is widely known. You miss real effects. But also: the significant findings you do get are systematically overestimated, because only the luckiest samples clear the threshold. And a null result from a small study is uninformative — it cannot distinguish "no effect" from "not enough data".

# Simulate a real effect of d = 0.35 and see what "significant" studies report. set.seed(2026) one <- function(n, d = 0.35) { x <- rnorm(n, 0); y <- rnorm(n, d) tt <- t.test(y, x) c(p = tt$p.value, d_obs = (mean(y) - mean(x)) / sd(c(x, y))) } small <- t(replicate(5000, one(20))) large <- t(replicate(5000, one(200))) rbind( "n = 20 per group" = c(power = mean(small[, "p"] < .05), mean_d_when_sig = mean(small[small[, "p"] < .05, "d_obs"])), "n = 200 per group" = c(power = mean(large[, "p"] < .05), mean_d_when_sig = mean(large[large[, "p"] < .05, "d_obs"])) ) |> round(3)#> power mean_d_when_sig #> n = 20 per group 0.223 0.744 #> n = 200 per group 0.884 0.381 #> ← true effect is 0.35. The small study, WHEN significant, reports 0.74 — #> more than double. This is why literatures shrink on replication.
ConsequenceWhat it does to your PhD
Missed real effectsThree years of work reported as "no significant difference", with no way to argue for the null.
Inflated published estimatesYour effect size feeds someone's power analysis, and their study fails.
Uninformative nullsCannot claim absence without equivalence testing or a Bayes factor (volume 7 day 72).
Unstable conclusionsEstimates swing wildly between subsamples; robustness checks contradict each other.
Pressure toward flexibilityUnderpowered studies create the incentive to keep analysing until something appears.
The thing to internalise
Power is not a statistical formality you satisfy for an ethics committee. It is the design question: what is the smallest effect this study could reliably detect, and is that effect scientifically interesting? If the answer is "only effects larger than anything ever reported in this field", the study as designed cannot answer its question, and no amount of sophisticated modelling afterwards will change that. Better to know in week one than in year three.
Do this now · 15 minutes
Run the simulation above with the effect size you actually expect and the n you can realistically recruit. Write the two numbers down: your power, and the average effect you would report if you got lucky. Take both to your next supervision meeting.
Day 78 50 minutes · when a formula exists

pwr and WebPower: the closed-form cases

For simple designs the answer is a formula, and R gives it in one line. The rule that makes power analysis honest: specify the smallest effect size of interest, not the effect you hope for and certainly not the effect a previous small study reported. Pilot-study effect sizes are among the least reliable numbers in science — they are exactly the inflated estimates day 77 demonstrated.

library(pwr) # n for a two-group comparison pwr.t.test(d = 0.4, power = 0.80, sig.level = 0.05, type = "two.sample") # ...and for a within-subject design (note how much smaller) pwr.t.test(d = 0.4, power = 0.80, type = "paired") pwr.anova.test(k = 3, f = 0.25, power = 0.80) # one-way ANOVA pwr.r.test(r = 0.25, power = 0.80) # correlation pwr.f2.test(u = 3, f2 = 0.09, power = 0.80) # regression, 3 predictors pwr.2p.test(h = ES.h(0.30, 0.45), power = 0.80) # two proportions # The whole curve is more useful than a single number plot(pwr.t.test(d = 0.4, power = 0.80, type = "two.sample"))#> Two-sample t test power calculation #> n = 99.08 #> d = 0.4 #> sig.level = 0.05 #> power = 0.8 #> alternative = two.sided #> NOTE: n is number in *each* group ← so 198 participants total #> #> Paired t test power calculation #> n = 51.01 ← same effect, half the people library(WebPower) wp.rmanova(n = NULL, ng = 2, nm = 3, f = 0.25, nscor = 0.8, power = 0.80, type = 2) # within-between interaction wp.mediation(n = NULL, power = 0.80, a = 0.35, b = 0.35, cp = 0.10) wp.logistic(n = NULL, p0 = 0.30, p1 = 0.45, power = 0.80, family = "normal") # Choosing the smallest effect of interest, three defensible routes: # 1. A meta-analytic estimate, bias-adjusted (volume 7 day 76) # 2. A clinically or practically meaningful change on your scale # 3. What the design can afford, reported as a sensitivity analysis (day 81)#> Repeated-measures ANOVA analysis #> n f ng nm nscor alpha power #> 43.81204 0.25 2 3 0.8 0.05 0.8 #> NOTE: n is the total sample size (nm = number of measurements)
Three ways power analyses go wrong
Using the pilot's effect size. Small-sample estimates are inflated and unstable; a pilot tells you about feasibility and procedures, not effect magnitude. Powering for the effect you want. Power for the smallest effect that would still matter, then your study is informative either way. Ignoring your real design. Powering a t-test when you will run a 2×3 mixed ANOVA with three covariates gives a number with no relationship to your actual analysis — which is what day 79 fixes.
Do this now · 20 minutes
Compute the required n for your primary analysis at three effect sizes: the meta-analytic estimate, half of it, and the smallest effect you would call meaningful. Plot the power curve. That plot goes into your preregistration.
Day 79 60 minutes · the recipe for every design

Simulation: generate, analyse, repeat, count

There is no formula for your 2×3 mixed design with two covariates, a moderated mediation and 15% attrition — and there does not need to be. Simulation-based power is four steps and covers every design you will ever run: write code that generates data with the effect you specify, run your exact planned analysis on it, repeat a few thousand times, and count the proportion of runs in which you detected the effect.

library(tidyverse) # STEP 1 — generate one dataset with a KNOWN effect sim_one <- function(n_per, d, attrition = 0.15) { dat <- tibble( id = 1:(2 * n_per), condition = rep(c("control", "treat"), each = n_per), pre = rnorm(2 * n_per, 20, 6), post = 0.65 * pre + if_else(condition == "treat", -d * 6, 0) + rnorm(2 * n_per, 7, 4.5) ) dat |> mutate(post = if_else(runif(n()) < attrition, NA_real_, post)) } # STEP 2 — run the EXACT planned analysis and return the decision analyse <- function(dat) { m <- lm(post ~ condition + pre, data = dat) ci <- confint(m)["conditiontreat", ] c(p = summary(m)$coefficients["conditiontreat", 4], est = coef(m)["conditiontreat"], lo = ci[1], hi = ci[2]) } # STEP 3/4 — repeat and count set.seed(2026) power_at <- function(n_per, d, reps = 1000) { out <- t(replicate(reps, analyse(sim_one(n_per, d)))) c(n_per = n_per, d = d, power = mean(out[, "p"] < .05), mean_ci_width = mean(out[, "hi"] - out[, "lo"])) } map_dfr(c(40, 60, 80, 100, 120), power_at, d = 0.40)#> # A tibble: 5 × 4 #> n_per d power mean_ci_width #> <dbl> <dbl> <dbl> <dbl> #> 1 40 0.4 0.512 3.84 #> 2 60 0.4 0.681 3.13 #> 3 80 0.4 0.798 2.71 #> 4 100 0.4 0.874 2.42 #> 5 120 0.4 0.921 2.21 #> ← 80 per group for 80% power WITH 15% attrition; the analytic answer #> of 99 assumed complete data, so simulation is both more and less #> demanding depending on what you model.
Why this replaces almost everything else
Whatever you can analyse, you can simulate — including things no formula covers: attrition, a floor effect, an ordinal outcome, a covariate that only correlates .3, a moderated mediation, a preregistered exclusion rule. And the generating code is reusable: the same sim_one() becomes the analysis-script test on day 52, the precision-based design on the next block, and the simulation study of day 84. One function, four purposes.
# Power is not the only target. Precision often matters more. # "What n gives me a CI no wider than 0.3 SDs?" map_dfr(c(80, 120, 160, 200), power_at, d = 0.40) |> mutate(width_in_sd = mean_ci_width / 6) # Use the parallel package for speed; simulations are embarrassingly parallel library(furrr); plan(multisession, workers = 4) future_map_dfr(c(40, 60, 80, 100, 120), power_at, d = 0.40, .options = furrr_options(seed = 2026)) # Always check your generator: does it produce data that look real? sim_one(80, 0.4) |> ggplot(aes(condition, post)) + geom_jitter(width = 0.1, alpha = 0.3) + stat_summary(fun.data = mean_cl_normal)
Do this now · 30 minutes
Write sim_one() and analyse() for your own planned study, including your attrition estimate and exclusion rules. Produce the power curve across five sample sizes. This is the core of your preregistration and it took under an hour.
Day 80 60 minutes · participants or trials?

simr: the design trade-off only simulation can answer

In a multilevel design you can add participants or add observations per participant, and they are not interchangeable: more clusters buys power for between-person effects, more observations per cluster buys power for within-person effects, and the exchange rate depends on the ICC. simr answers it by simulating from a fitted or hand-specified lmer model.

library(lme4); library(simr) # Option A: start from a pilot or published model, then set the effect you # care about rather than trusting the pilot's estimate. m <- lmer(mood ~ stress_cw + (1 + stress_cw | id), data = pilot) fixef(m)["stress_cw"] <- -0.20 # the SMALLEST effect of interest powerSim(m, nsim = 500, test = fixed("stress_cw", "t"))#> Power for predictor 'stress_cw', (95% confidence interval): #> 68.40% (64.21, 72.42) #> Test: t-test #> Based on 500 simulations, (0 warnings, 0 errors) #> alpha = 0.05, nrow = 1410 # The two ways to grow the design, compared more_people <- extend(m, along = "id", n = 150) # 94 → 150 participants more_beeps <- extend(m, within = "id", n = 60) # 30 → 60 beeps each powerCurve(more_people, along = "id", breaks = c(60, 90, 120, 150), nsim = 200) |> plot() powerCurve(more_beeps, within = "id", breaks = c(20, 30, 45, 60), nsim = 200) |> plot()#> Power for predictor 'stress_cw' (more participants): #> 60: 52.0% | 90: 67.5% | 120: 78.5% | 150: 86.0% #> #> Power for predictor 'stress_cw' (more beeps per person): #> 20: 58.5% | 30: 68.0% | 45: 79.5% | 60: 86.5% #> ← for this WITHIN-person effect, 30 extra beeps from existing participants #> buys as much as 56 extra participants — and costs far less to collect. # No pilot data? Build the model from scratch with makeLmer(). subj <- data.frame(id = factor(1:100)) design <- expand.grid(id = subj$id, trial = 1:40) design$cond <- rep(c("a", "b"), each = 20) m_spec <- makeLmer(rt ~ cond + (1 + cond | id), fixef = c(650, -25), # intercept, effect in ms VarCorr = matrix(c(2500, -150, -150, 400), 2, 2), sigma = 80, data = design) powerSim(m_spec, nsim = 500, test = fixed("condb", "t")) # Effect sizes in a mixed model are variance-relative: you must specify the # random-effect SDs and residual SD too. Get them from a published model, # a pilot, or a defensible range — and report which.#> Power for predictor 'condb', (95% confidence interval): #> 84.20% (80.81, 87.21) #> Based on 500 simulations, (12 warnings, 0 errors) #> ← warnings are usually singular fits in individual simulated datasets; #> check they are not the majority
Two practical warnings
It is slow. 500 simulations of a moderate mixed model can take an hour; develop with nsim = 50, then run the real thing once with more. The answer is only as good as the variance components you assume. Report them explicitly in the preregistration — "assuming a random-intercept SD of 0.85 and residual SD of 0.95, as observed in [source]" — and give power at a pessimistic assumption too. For SEM the equivalent tool is simsem.
Do this now · 30 minutes
Build a makeLmer() specification for your design, using variance components from a published model in your area, and produce both power curves — clusters and observations per cluster. The crossing point is a design decision worth a paragraph in your proposal.
Day 81 50 minutes · when n is already fixed

Sensitivity analysis: what effect could this study have detected?

Secondary data, an existing cohort, a recruitment ceiling — often n is not yours to choose. The right analysis is then inverted: hold n and power fixed, solve for the effect size. That number is your study's detection threshold, and reporting it is far more informative than a post-hoc power calculation, which is a mathematical restatement of your p-value and tells nobody anything.

library(pwr) # n is fixed at 64 per group. What could we have detected at 80% power? pwr.t.test(n = 64, power = 0.80, sig.level = 0.05, type = "two.sample") # And the effect detectable at more realistic power levels sapply(c(0.50, 0.80, 0.95), \(pw) pwr.t.test(n = 64, power = pw, type = "two.sample")$d) |> round(2)#> Two-sample t test power calculation #> n = 64 #> d = 0.4986 #> power = 0.8 #> ← this study could reliably detect d ≈ 0.50 and nothing smaller. If the #> field's meta-analytic estimate is 0.30, the design was never adequate. #> #> [1] 0.34 0.50 0.65
Never report post-hoc power
"Observed power" computed from your own observed effect size is a deterministic function of your p-value — p = .05 always gives about 50% power — so it adds no information and creates a circular argument ("the null result is explained by low power, which we know because the result was null"). Journals and reviewers increasingly flag it. The two legitimate post-data analyses are sensitivity analysis (this day) and equivalence testing (next block).
# Claiming a null result properly: equivalence testing. library(TOSTER) # "Any difference smaller than d = 0.35 is not meaningful to us" tsum_TOST(m1 = 18.1, m2 = 17.7, sd1 = 5.4, sd2 = 5.2, n1 = 64, n2 = 64, low_eqbound = -0.35, high_eqbound = 0.35, eqbound_type = "SMD") # The Bayesian route to the same claim (volume 7) # BayesFactor::ttestBF(...) with BF10 < 1/3, or a ROPE from bayestestR#> Welch Modified Two-Sample t-Test #> t = 0.426, df = 125.4, p = 0.671 ← the ordinary NHST result #> #> TOST Results: #> t df p.value #> t-test 0.426 125.40 0.671 #> TOST Lower 2.404 125.40 0.009 #> TOST Upper -1.552 125.40 0.062 ← upper bound NOT rejected #> #> Equivalence Test: don't reject null equivalence hypothesis #> ← we can exclude effects more negative than -0.35 but not more positive #> than +0.35. Honest conclusion: inconclusive, not "no effect".
SituationThe right analysis
Designing a studyA priori power for your smallest effect of interest (days 78–80)
n fixed by circumstance, before analysisSensitivity analysis: the detectable effect at 80% power
A null result you want to interpretEquivalence test (TOST) against a preregistered bound, or a Bayes factor
After a significant resultNothing. Report the estimate and its interval
A reviewer asks for post-hoc powerPolitely supply the sensitivity analysis instead, with a one-line explanation
Do this now · 20 minutes
For your existing or planned dataset, compute the effect detectable at 50%, 80% and 95% power, and compare with your field's meta-analytic estimate. Then define the equivalence bound you would defend, and write down where it came from.
Day 82 50 minutes · stopping early, legitimately

Sequential analysis: looking at your data without breaking it

Peeking at accumulating data and stopping when p drops below .05 inflates the false-positive rate badly — to around 20% with a handful of looks. But planned sequential designs let you do almost exactly that, legitimately, by spending your alpha across the interim analyses. For long, expensive data collection this is the most valuable design technique in this volume.

# The problem, quantified: unplanned peeking set.seed(2026) peek <- function(max_n = 200, looks = seq(40, 200, 40)) { x <- rnorm(max_n); y <- rnorm(max_n) # no true effect at all any(sapply(looks, \(n) t.test(x[1:n], y[1:n])$p.value < .05)) } mean(replicate(2000, peek()))#> [1] 0.1435 #> ← 14% false positives at a nominal 5%, with only five looks. library(rpact) # Plan the looks and the alpha spending BEFORE collecting design <- getDesignGroupSequential( informationRates = c(0.4, 0.7, 1.0), # three looks typeOfDesign = "asOF", # O'Brien-Fleming alpha spending alpha = 0.05, beta = 0.20, sided = 2) design summary(design) # Sample size at each stage for the effect you care about getSampleSizeMeans(design, alternative = 0.4, stDev = 1)#> Stage 1 2 3 #> Information rate 0.400 0.700 1.000 #> Efficacy boundary (z) 3.475 2.454 2.004 #> Cumulative alpha spent 0.0005 0.0142 0.0500 #> #> Maximum number of subjects: 208 #> Expected number of subjects under H1: 158.4 #> ← on average you recruit 50 fewer participants than the fixed design, #> and the false-positive rate is exactly 5%. # The Bayesian alternative: stop when evidence is sufficient. No alpha to spend. library(BayesFactor) sequential_bf <- function(x, y, from = 20) { sapply(from:length(x), \(n) extractBF(ttestBF(x[1:n], y[1:n]))$bf) } # Stop when BF10 > 10 (evidence for H1) or BF10 < 1/10 (evidence for H0). # Optional stopping does not bias a Bayes factor — but DO preregister the # rule, the maximum n, and the prior, or the analysis is unreviewable.
DesignBuys youCosts
Fixed-nSimplicity; no special analysisYou collect the full sample even when the answer is obvious by half
Group-sequential (rpact)Early stopping for efficacy or futility; exact error controlMust be planned in advance; slightly larger maximum n
Sequential Bayes factorStop when evidence suffices, in either directionPrior-dependent; needs a preregistered rule and cap
Unplanned peekingNothingA false-positive rate you cannot report
Do this now · 20 minutes
If your data collection runs over months, design a three-look group-sequential plan with rpact and note the expected saving in participants. If not, run the peeking simulation and keep the number — it is the clearest argument against "let me just check whether it is significant yet".
Day 83 55 minutes · robustness as a result

Specification curves: every defensible analysis at once

Most analyses involve arbitrary choices — which outliers to exclude, whether to control for age, which of three depression measures, log or raw. Each is defensible, and reporting only one hides how much the conclusion depends on it. A multiverse analysis runs all defensible combinations and shows the distribution of results, turning an invisible degree of freedom into a reportable finding.

library(specr) specs <- setup( data = d, y = c("bdi_total", "bdi_log", "phq9"), # 3 outcomes x = c("stress", "stress_winsor"), # 2 predictors model = c("lm"), controls = c("age", "gender", "ses"), # all subsets subsets = list(complete_only = c(TRUE, FALSE)) # 2 exclusion rules ) results <- specr(specs) # 3*2*8*2 = 96 models summary(results, type = "curve") plot(results, type = "curve")#> Median effect: 0.312 Range: 0.104 to 0.482 #> Proportion of specifications with p < .05: 0.885 (85 of 96) #> Proportion positive: 1.00 #> #> Effects by decision: #> outcome bdi_total 0.381 | bdi_log 0.298 | phq9 0.264 #> subset complete_only 0.341 | all 0.294 #> ← the sign never flips and 85 of 96 specifications reach significance: #> a genuinely robust association, with magnitude depending mostly on #> which outcome measure you pick.
How to report a multiverse honestly
Report four things: the median effect across specifications, the range, the proportion of specifications supporting your conclusion, and which decisions move the estimate most (the lower panel of the curve). Then name your preregistered primary specification and show where it sits in the distribution. The failure mode is using a multiverse to find one supportive specification — the analysis exists to prevent exactly that, so present the whole curve or none of it.
# Which decisions matter? Decompose the variance across specifications. plot(results, type = "variance") # A permutation test for the curve as a whole: is this pattern of results # more extreme than chance across the whole multiverse? # (specr does not do this directly; shuffle the predictor and re-run) null_curves <- replicate(200, { d_perm <- d |> mutate(stress = sample(stress)) median(specr(setup(data = d_perm, y = "bdi_total", x = "stress", model = "lm", controls = c("age", "gender")))$estimate) }) mean(abs(null_curves) >= abs(median(results$estimate)))#> Variance decomposition: #> outcome 41.2% #> controls 18.4% #> subset 8.1% #> predictor 3.2% #> residual 29.1% #> [1] 0.005 ← the observed curve is far more extreme than permuted curves
Decision typeInclude in the multiverse?
Equally defensible outcome measuresYes — this is the clearest case
Plausible covariate setsYes, but only covariates that pass day 33's causal test
Different exclusion rulesYes; each must be one you would have defended in advance
TransformationsYes, where more than one is standard in the field
Specifications you know are wrongNo. A multiverse is the set of defensible analyses, not all possible ones
Different hypothesesNo. That is not robustness, that is fishing
Do this now · 25 minutes
List every arbitrary decision in your own analysis, mark which are genuinely defensible alternatives, and build the specification curve. Report the median, the range and the proportion supporting your conclusion — then find your preregistered specification on the curve.
Day 84 60 minutes · portfolio artefact seven

A simulation study: methodological maturity, in one afternoon

A simulation study demonstrates a property of a method — bias, coverage, power, robustness to violation — by generating data where the truth is known. It is the single most transferable skill in this course, it makes a compelling job-talk slide, and it reads as methodological seriousness that very few applicants demonstrate. The structure is always the same four pieces.

1 · design grid
Every combination of conditions you want to vary: n, effect size, degree of violation. expand_grid().
2 · generator
A function that produces one dataset from one row of the grid, with the truth built in.
3 · estimators
Two or more analyses applied to the same dataset, so the comparison is fair.
4 · summaries
Bias, RMSE, coverage of the 95% interval, power, type-I error. Computed over replications.
library(tidyverse) # Question: how much does ignoring clustering damage inference, and does # cluster-robust SE fix it as well as a mixed model? grid <- expand_grid(n_cluster = c(10, 30, 60), icc = c(0.05, 0.20, 0.40), rep = 1:500) gen <- function(n_cluster, icc, m = 10, beta = 0.3) { sd_u <- sqrt(icc); sd_e <- sqrt(1 - icc) tibble(cluster = rep(1:n_cluster, each = m), u = rep(rnorm(n_cluster, 0, sd_u), each = m), x = rnorm(n_cluster * m), y = beta * x + u + rnorm(n_cluster * m, 0, sd_e)) } estimate <- function(dat) { lm_fit <- lm(y ~ x, data = dat) ci_lm <- confint(lm_fit)["x", ] ml_fit <- lme4::lmer(y ~ x + (1 | cluster), data = dat) ci_ml <- confint(ml_fit, method = "Wald", parm = "x") tibble(method = c("lm", "lmer"), est = c(coef(lm_fit)["x"], lme4::fixef(ml_fit)["x"]), lo = c(ci_lm[1], ci_ml[1]), hi = c(ci_lm[2], ci_ml[2])) } # Run it (start with rep = 1:20 while debugging) set.seed(2026) out <- grid |> mutate(fit = pmap(list(n_cluster, icc), gen) |> map(estimate)) |> unnest(fit) # The four summaries that make it a simulation STUDY rather than a demo summ <- out |> group_by(n_cluster, icc, method) |> summarise(bias = mean(est) - 0.3, rmse = sqrt(mean((est - 0.3)^2)), coverage = mean(lo <= 0.3 & hi >= 0.3), ci_width = mean(hi - lo), .groups = "drop") summ |> filter(icc == 0.40) ggplot(summ, aes(factor(n_cluster), coverage, colour = method, group = method)) + geom_hline(yintercept = 0.95, linetype = "dashed") + geom_line() + geom_point() + facet_wrap(~ paste("ICC =", icc)) + scale_y_continuous(limits = c(0.5, 1)) + labs(x = "Number of clusters", y = "Coverage of the 95% CI") + theme_paper()#> # A tibble: 6 × 7 #> n_cluster icc method bias rmse coverage ci_width #> 10 0.4 lm -0.002 0.081 0.631 0.242 #> 10 0.4 lmer 0.001 0.079 0.918 0.352 #> 30 0.4 lm -0.001 0.046 0.642 0.139 #> 30 0.4 lmer 0.000 0.045 0.941 0.192 #> ← both are UNBIASED; only the intervals differ. lm's coverage is 64% #> instead of 95% — the entire damage is to uncertainty, not to the estimate.
Reporting a simulation study
Five sections: the question, the data-generating model written out in equations, the design grid, the estimators compared, and the performance measures (bias, RMSE, coverage, power, type-I error rate) with Monte Carlo standard errors. State the number of replications and the seed, and deposit the code. Ten pages of this is a publishable methods note; two figures of it is the strongest slide in a PhD interview.
Do this now · 30 minutes
Build a complete simulation study on a question that matters to your own design — one generator, two estimators, one violated assumption, 500 replications — and produce the coverage figure. Deposit the script. That is portfolio artefact seven, finished.
Reference

Which tool for which design

Analytic where a formula exists, simulation everywhere else.

DesignTool
Two groups, one outcomepwr.t.test()
One-way ANOVApwr.anova.test()
Correlationpwr.r.test()
Multiple regression, added predictorspwr.f2.test()
Repeated measures / mixed ANOVAWebPower::wp.rmanova()
MediationWebPower::wp.mediation(), or simulate
Logistic regressionWebPower::wp.logistic()
Multilevel / mixed modelsimr::powerSim(), powerCurve()
SEM or CFAsimsem, or a Monte Carlo study (day 84)
Anything with attrition, exclusions or a custom analysisSimulation, day 79
Meta-analysisdmetar::power.analysis()
n already fixedSensitivity analysis, day 81
Reference

What a preregistered power analysis contains

The effect size, and its source
"d = 0.35, the bias-adjusted pooled estimate from our meta-analysis (day 76)" — never "from our pilot" without justification.
Why that effect is the smallest of interest
A clinical threshold, a published benchmark, or an explicit judgement. One sentence.
The exact planned analysis
The power analysis must match the model you will actually fit, covariates and all.
Assumptions you had to make
Variance components, ICC, attrition rate, correlation between waves — each with a source.
The result, with a curve
Required n at 80% and 95% power, plus the power curve as a figure.
The stopping rule
Fixed n, or the sequential plan with its alpha spending (day 82).
Code
The simulation script, deposited. This is what makes it checkable.
Reference

Install these once

install.packages(c( "pwr", # analytic power for the classic designs "WebPower", # repeated measures, mediation, logistic, SEM "simr", # simulation power for lme4 models "simsem", # Monte Carlo power for lavaan models "rpact", # group-sequential designs and alpha spending "TOSTER", # equivalence testing "specr", # specification curve / multiverse analysis "furrr", # parallel simulation "dmetar" # power analysis for meta-analysis ))
Checkpoint

Six questions before volume 9

{{ quizCounter }}
{{ quizScore }}
{{ quizQ }}
{{ quizFb }}

Next: volume 9

The portfolio volume. Twelve days building the artefacts a PhD application is judged on: a Quarto manuscript whose numbers come from the model, APA tables, papaja, Git and GitHub, renv, targets pipelines, your own functions, your own package, and OSF preregistration.

Start day 85 →