Describe the design. The model follows.
Choosing a test by searching for the name you half-remember is how the wrong analysis gets run. The reliable route is mechanical: state what one row is, what the outcome is made of, whether rows are independent, and what claim you want to make. Four answers determine the model almost completely — and where they leave a genuine choice, this page says so rather than pretending otherwise.
Decide in this order, always
The order matters because each answer constrains the next. Reversing it — starting from a test and asking whether your data fit — is how ANOVAs get run on binary outcomes.
lm/glm and the mixed-model family.check_model(). Assumptions are about residuals and are checked after fitting — they do not choose the model. Day 28.By outcome type
| Outcome | Independent rows | Clustered or repeated rows |
|---|---|---|
| Continuous (scale score, RT, minutes) | lm() | lmer(), or brm() for hard random structures |
| Binary (relapse, correct, endorsed) | glm(family = binomial) | glmer(family = binomial) |
| Count (episodes, errors, sessions) | glm(family = poisson), then check overdispersion | glmmTMB(family = nbinom2) |
| Single ordinal item (one Likert question) | ordinal::clm() | ordinal::clmm() or brm(family = cumulative()) |
| Proportion or percentage bounded 0–1 | glm(family = binomial) with weights, or beta regression | glmmTMB(family = beta_family()) |
| A construct measured by several items | Score it (day 14) and use lm, or model it latently with lavaan::sem() | Multilevel SEM, or score it and use lmer() |
| Time to an event | Survival analysis — survival::coxph(). Beyond this course; the framing is the same | Frailty models, coxme |
By data structure
| What your data look like | Random structure | Note |
|---|---|---|
| One row per participant | None | lm/glm. The only case where volume 4 stands alone |
| Two occasions per participant | (1 | id), or analyse change directly | For a randomised trial, lm(post ~ condition + pre) is more powerful than a change score |
| Three or more occasions, balanced | (1 + time | id) | aov_ez(within =) also works; the mixed model tolerates missing occasions |
| Many unequally spaced occasions (ESM) | (1 + x | id), possibly + (1 | id:day) | Split predictors into within and between components. Day 44 |
| Trials within participants, all seeing the same items | (1 + cond | subject) + (1 | item) | Crossed, not nested. Treating items as fixed inflates error. Day 47 |
| Pupils in classes in schools | (1 | school/class) | Genuinely nested; check with a cross-tabulation |
| Patients within therapists | (1 | therapist) | Therapist effects are real and usually ignored in trials |
| Dyads — couples, parent and child | (1 | dyad) with role as a fixed effect | Distinguishable dyads need a different parameterisation from indistinguishable ones |
| Multi-site trial | (1 | site), or site as fixed if only a few | Fewer than about five sites: fixed. Day 42 |
By the claim you want to make
| The sentence you want to write | The call | Then |
|---|---|---|
| "These two groups differ" | lm(y ~ group) or t.test() | Estimate, CI, Cohen's d with CI |
| "These three or more groups differ" | aov_ez(), or lm() with a factor | emmeans contrasts — the omnibus F is not a finding |
| "Treatment beat control, adjusting for baseline" | lm(post ~ condition + pre) | Adjusted means with emmeans |
| "These two variables are associated" | cor.test() or lm() | The interval, and the reliabilities of both measures |
| "X predicts Y, controlling for Z" | lm(y ~ x + z) | Justify every covariate causally first. Day 33 |
| "The effect depends on a third variable" | lm(y ~ x * mod), centred | Simple slopes and a figure; never read the product term alone |
| "X works through M" | lavaan::sem() with bootstrapped ab | State the causal assumptions and the temporal design |
| "People changed over time" | lmer(y ~ time + (1 + time | id)) | Code time deliberately; report slope SD as well as the mean slope |
| "The treatment changed the rate of change" | lmer(y ~ time * condition + (1 + time | id)) | emtrends() for the slope difference |
| "X and Y influence each other over waves" | RI-CLPM in lavaan | Report the interval length; effects depend on it. Day 65 |
| "My scale measures one thing" | fa.parallel(), then cfa() | All four fit indices plus omega. Days 57–61 |
| "These groups can be compared on this scale" | Invariance sequence in lavaan | ΔCFI at each step before any mean comparison. Day 63 |
| "There is no meaningful effect" | TOSTER equivalence test, or a Bayes factor | The bound must be preregistered. Days 72, 81 |
| "The literature as a whole says…" | metafor::rma() | tau and the prediction interval, not just the pooled estimate |
| "I can predict this for new people" | tidymodels workflow | Cross-validated metrics, calibration, no causal language |
Twelve mistakes a reviewer will catch
Each of these is common in published psychology, each is checkable in a minute, and each has a day in this course.
| The mistake | What to do instead | Day |
|---|---|---|
| A normality test chooses the analysis | Judge residuals visually; choose from the design | 28 |
| Reporting only p-values | Estimate, interval, effect size — then p | 25 |
| The omnibus F reported as the finding | Planned contrasts or marginal means | 29, 30 |
| Significant in one group, not the other, called a difference | Test the interaction | 31 |
Clustered rows analysed with lm | Mixed model, or cluster-robust SEs at minimum | 41 |
| Stimuli treated as fixed effects | Crossed random effects for subject and item | 47 |
| Raw predictors in a multilevel model | Split into within- and between-person components | 44 |
| Uncentred predictors in a moderation | Centre, then interpret simple slopes | 38 |
| PCA called factor analysis | EFA or CFA for construct claims | 59 |
| Group means compared without invariance testing | Configural, metric, scalar first | 63 |
| Modification indices followed mechanically | One justified change, both models reported | 62 |
| In-sample R² reported as predictive accuracy | Cross-validated metrics on held-out data | 99 |
What every result needs, without exception
If two models are defensible
Report both. A result that survives a defensible alternative specification is stronger evidence than a result from the single analysis you preferred, and a result that does not survive is something your reader is entitled to know. Where the choices multiply — three outcomes, two exclusion rules, four covariate sets — that is a specification curve rather than a dilemma. Day 83.