Volume 4 · The t and F family, end to end

Every group comparison you will ever run is the same ratio with different bookkeeping.

Sixteen days through the tests that carry most of published psychology. Each one is built the same way — an effect divided by an estimate of noise, located in a null distribution — and each one is taught here with the arithmetic examiners ask for, the assumption that actually matters, the effect size that must accompany it, and the reporting sentence.

16
days of 45–55 min
12
tests, with their effect sizes
1
source table, by hand
8
assumption decisions
Day 3550 minutes · the first inferential test

The z-test and the one-sample t-test

The simplest inferential question: is my sample mean different from some stated value? A clinic's mean anxiety score against the published norm, a class's mean against the national average, a reaction time against a theoretical baseline. Everything more complicated is an elaboration of this.

z = (M − μ) / (σ/√n) t = (M − μ) / (s/√n)how many standard errors your mean sits from the value you are testing againstthe effect: observed mean minus hypothesised valuestandard error — known σ gives z, estimated s gives tn − 1 for the t versionClinic sample: M = 23.4, s = 7.2, n = 25 Published norm: μ = 20.0 SE = 7.2 / √25 = 1.44 t = (23.4 − 20.0) / 1.44 = 2.36 df = 24, two-tailed critical t = 2.064 2.36 > 2.064, so reject H₀; p = .027 Effect size: d = (23.4 − 20.0) / 7.2 = 0.47 Report: t(24) = 2.36, p = .027, d = 0.47, 95% CI [0.05, 0.88]

Use z only when σ is genuinely known from a large normative sample — which in practice means almost never. Otherwise s estimates σ, the extra uncertainty fattens the tails, and the t distribution is the correct reference. At n above about 100 the two give nearly identical answers, which is why large-sample work is often described in z terms.

Assumptions, in order
Independence of observations (design question, non-negotiable). Approximately normal sampling distribution of the mean — the CLT handles this at moderate n unless the data are severely skewed. Interval or ratio measurement. That is the entire list, and the first item is the only one that routinely destroys a study.
The mistake this topic produces
Comparing your sample to a norm that was collected on a different population, in a different decade, or with a different version of the instrument. The statistics will run perfectly; the comparison is meaningless.
Do this now · 15 minutes
Find a published norm for a measure you use. Run a one-sample t comparing your sample to it by hand, compute d, and write the full APA sentence. Then write one sentence on whether that norm group is an appropriate comparison for your sample.
Day 3655 minutes · Welch by default

The independent-samples t-test

Two unrelated groups, one continuous outcome. The signal is the difference between means; the noise is the standard error of that difference, which — as day 14 warned — is built from both groups' variances added together.

t = (M₁ − M₂) / √(s²p (1/n₁ + 1/n₂))difference between means, over the standard error of that differencepooled variance: ((n₁−1)s₁² + (n₂−1)s₂²) / (n₁+n₂−2)n₁ + n₂ − 2 for the pooled versiona fractional df computed from the two variances separately — do not be alarmed by df = 43.7Group 1: M = 18.2, s = 6.1, n = 30 Group 2: M = 22.9, s = 6.5, n = 30 s²p = ((29 × 37.21) + (29 × 42.25)) / 58 = 39.73 SE = √(39.73 × (1/30 + 1/30)) = √2.649 = 1.63 t = (18.2 − 22.9) / 1.63 = −2.88 df = 58, p = .006 d = −4.7 / 6.30 = −0.75 Report: t(58) = −2.88, p = .006, d = −0.75, 95% CI [−1.27, −0.22]
Why Welch, and why by default
The pooled test assumes equal population variances. When variances and group sizes are both unequal, its error rate goes badly wrong — too liberal when the smaller group has the larger variance, too conservative in reverse. Welch's version estimates the standard error without pooling and adjusts df downward. It performs essentially as well as the pooled test when variances are equal, and far better when they are not. Simulation work over three decades supports using it unconditionally; R's t.test() already defaults to it. The traditional 'run Levene's test first, then choose' procedure is a two-stage decision that inflates error rates and should be retired — though you should recognise it in an exam.

Note the identity from day 21: with two groups, t² equals the F from a one-way ANOVA on the same data. They are one test. Which you report is a matter of convention and of how many groups you have.

The mistake this topic produces
Running an independent-samples t on paired data — the same people before and after. It discards the pairing, inflates the error term with between-person variance, and throws away most of your power. Check whether each row of your data is one person or two.
Do this now · 15 minutes
Run both pooled and Welch t-tests on two groups of your own with clearly unequal ns or variances. Compare df, p and the CI. Write which you would report and why — that one sentence is a defensible methods choice.
Day 3745 minutes · difference scores

The paired-samples t-test

When the two measurements come from the same participant — before and after, two conditions, matched pairs — you can subtract within person first. That single move removes all stable between-person variability from the error term, which is why within-subject designs are so much more powerful.

t = M_D / (s_D / √n)the mean of the difference scores, over its standard errormean of each person's difference scoreSD of those difference scores — not of the raw scoresnumber of pairs, so df = n − 1M_D / s_D, the effect size for a paired designTen participants, pre minus post differences: 4, 6, 2, 8, 5, 1, 7, 3, 6, 4 M_D = 46/10 = 4.6 s_D = 2.22 SE = 2.22 / √10 = 0.70 t = 4.6 / 0.70 = 6.57, df = 9, p < .001 d_z = 4.6 / 2.22 = 2.07 (large, but note d_z and between-subject d are not on the same scale — say which one you report)

The power advantage depends entirely on the correlation between the two measurements. When pre and post correlate at .80, the SD of the differences is far smaller than either raw SD and the design is dramatically more efficient. When the correlation is near zero, pairing buys you nothing and costs you degrees of freedom.

Which effect size for a within-subject design
Three are in circulation and they are not interchangeable. d_z divides by the SD of the differences and is what G*Power wants for a paired test. d_av divides by the average of the two raw SDs and is comparable with between-subject d. d_rm corrects for the correlation. Meta-analysts usually want d_av. Whichever you report, name it — an unnamed within-subject d is ambiguous by a factor that can exceed two.
The mistake this topic produces
Reporting a paired d_z alongside between-subject ds from other studies as if they were on one scale. Because d_z divides by the SD of the differences, a high pre-post correlation makes it enormous without the raw effect being any bigger.
Do this now · 15 minutes
For a within-subject comparison of yours, compute the correlation between the two measurements, then the paired t and d_z. Then compute what an independent-samples t would have given on the same numbers. The difference in p is the value of your design.
Day 3845 minutes · magnitude and precision

Effect sizes and intervals for mean differences

A t-test gives you a decision. The three quantities that make it a finding are the raw difference, the standardised difference, and an interval around both. Journals now expect all of them, and entrance papers increasingly ask you to compute the second.

QuantityFormulaWhat it is for
Raw differenceM₁ − M₂The number a clinician can act on: points on a scale, milliseconds, kilograms.
Cohen's d(M₁ − M₂) / s_pooledComparable across measures and studies; the currency of meta-analysis.
Hedges' gd × (1 − 3/(4df − 1))Small-sample correction. Below about n = 20 per group, d is biased upward and g is preferred.
Glass's delta(M₁ − M₂) / s_controlWhen the intervention plausibly changes the variance, standardise on the control group only.
CI on the difference(M₁ − M₂) ± t_crit × SE_diffThe most informative single result in a two-group study.
Common language ESP(a random case from group 1 exceeds one from group 2)For communicating to non-statisticians: 'a randomly chosen treated patient improves more than an untreated one 70% of the time'.
Interpreting d without the benchmark crutch
d = 0.50 means the group means differ by half a standard deviation, that about 69 per cent of the treatment group exceeds the control mean, and that the distributions overlap by roughly 80 per cent. Those three translations are more useful than the word 'medium', and each is directly derivable: the percentage superior is the normal area above −d.
The mistake this topic produces
Computing d using the SD of the difference scores in a between-subjects design, or the pooled raw SD in a within-subjects design. The denominator defines the effect size; using the wrong one can change the number by a factor of two or more.
Do this now · 15 minutes
Recompute the effect sizes for two of your existing analyses, with intervals. Then write each result as a single sentence containing the raw difference, the standardised difference and the interval — and check that a non-statistician could act on it.
Day 3955 minutes · why not many t-tests

One-way ANOVA: partitioning variance

With three or more groups, running all pairwise t-tests inflates the familywise error rate (day 31). ANOVA answers a single omnibus question instead — do these means differ anywhere? — with one test at one α. The way it does so is the most elegant idea in classical statistics: split the total variability into the part explained by group membership and the part left over.

Total variability, measured as a sum of squares around the grand mean, splits exactly into between-groups SS (how far each group mean sits from the grand mean) and within-groups SS (how far each score sits from its own group mean). Divide each by its df to get mean squares, which are variance estimates, and take their ratio. Under H₀ both estimate error variance alone, so F hovers near 1. When group means genuinely differ, the numerator grows.

Push difference between means up and watch the between bar and F grow together. Now, holding that difference fixed, push within-group SD up: the same real effect becomes undetectable. That is the whole argument for careful measurement — reducing noise is mathematically equivalent to increasing the effect, and usually cheaper than recruiting.

What the omnibus F does not tell you
A significant F says at least one contrast among the means is non-zero. It does not say which, and with four groups there are six pairwise comparisons and many more complex contrasts. That is what days 41 and 42 are for. A non-significant omnibus F, meanwhile, does not license the claim that all means are equal — the same estimation logic from day 23 applies.
The mistake this topic produces
Running an ANOVA on two groups and reporting it as though it were more sophisticated than a t-test. It is the identical test: F = t². Use whichever the reader expects, but do not imply extra rigour.
Do this now · 15 minutes
In the simulator, find a combination of separation and within-SD that produces F just above and just below significance at n = 20 per group. Write down both η² values. Notice how similar the effect sizes can be on either side of the threshold — that is day 26's lesson, in pictures.
Day 4055 minutes · compute it by hand once

The ANOVA source table, computed by hand

Filling in a partially completed source table is the single most common computational question in psychology entrance papers, and it is entirely mechanical once you know that the rows add up. Do it by hand once and you will never be troubled by it again.

F = MS_between / MS_within, MS = SS / dfexplained variance per degree of freedom, divided by unexplained variance per degree of freedomΣ n_j (M_j − M_grand)²Σ Σ (X − M_j)², the sum of squared deviations inside each groupk − 1, where k is the number of groupsN − k, total observations minus number of groupsSS_between / SS_total — the proportion of variance explainedThree groups of 5. Group means 12, 15, 21; grand mean 16. SS_within computed from the raw scores = 120. SS_between = 5(12−16)² + 5(15−16)² + 5(21−16)² = 5(16) + 5(1) + 5(25) = 210 SS_total = 210 + 120 = 330 df_between = 3 − 1 = 2 df_within = 15 − 3 = 12 df_total = 14 (check: 2 + 12 = 14 ✓) MS_between = 210 / 2 = 105.0 MS_within = 120 / 12 = 10.0 F(2, 12) = 105 / 10 = 10.5, p = .002 η² = 210 / 330 = .636 ω² = (210 − 2(10)) / (330 + 10) = .559
SourceSSdfMSF
Between groups2102105.010.5
Within groups (error)1201210.0—
Total33014——
How to fill a partially blank table
Two identities do all the work: SS_total = SS_between + SS_within, and df_total = df_between + df_within. Given a table with any two of three SS values and any two of three df values, everything else follows, and F = MS_b/MS_w closes it. From df_between you recover k = df_b + 1; from df_total you recover N = df_total + 1. Exam questions are constructed so this always works — if it does not, you have mis-transcribed a number.
The mistake this topic produces
Reporting η² when the design has more than one factor and the software gave you partial η². They differ, sometimes substantially, because partial η² removes other factors from the denominator. Check which your output produced before you write it down.
Do this now · 15 minutes
Construct a source table by hand from a three-group dataset of your own, then confirm every cell against software. Then delete three cells at random and reconstruct them from the identities above — that is the exam question, set by you.
Day 4150 minutes · which pairs differ

Post hoc tests after a significant F

The omnibus F says something differs. Post hoc tests say what, while keeping the familywise error rate at the level you claimed. They are called post hoc because they are chosen after seeing that something is there — which is precisely why they carry a multiplicity penalty that planned contrasts do not.

TestControlsUse when
Tukey HSDFamilywise error across all pairwise comparisonsThe default for all-pairs comparisons with roughly equal n and equal variances. Good power for its protection.
BonferroniFWER, conservativelyA small number of specific comparisons. Simple to defend, loses power as k grows.
HolmFWER, sequentiallyStrictly better than Bonferroni. Use it wherever Bonferroni is expected.
SchefféFWER over ALL possible contrasts, not just pairwiseVery conservative; justified when you want the freedom to test complex contrasts suggested by the data.
Games-HowellFWER without assuming equal variancesThe Welch of post hoc tests. Unequal variances or unequal n.
DunnettFWER for comparisons against one controlSeveral treatments versus one control and nothing else — more powerful than Tukey because it tests fewer contrasts.
Fisher's LSDNothing beyond the omnibus testOnly defensible with exactly three groups after a significant F. Otherwise it is uncorrected testing with a respectable name.
Reporting them properly
Give the omnibus F with its df, p and η², then the specific comparisons with adjusted p values, mean differences and their confidence intervals. A results paragraph that reports only 'Tukey showed groups A and C differed (p = .03)' is missing the magnitude — which for a reader is the whole point of having run the study.
The mistake this topic produces
Running post hoc tests after a non-significant omnibus F and reporting whatever survives. With a protected-test logic, the omnibus result is the gate. If you had specific comparisons in mind, they should have been planned contrasts, and then the omnibus F is not required at all.
Do this now · 15 minutes
For a three- or four-group comparison, run Tukey and Bonferroni on the same data and compare the adjusted p values and intervals. Note which comparisons change status, and write one sentence on why Tukey is usually more powerful here.
Day 4250 minutes · a priori beats post hoc

Planned contrasts and trend analysis

If your hypothesis is more specific than 'something differs', you should not be running an omnibus test and then fishing. A planned contrast tests a specific, theory-derived comparison — two treatments combined against a control, or a linear increase across doses — with more power and a cleaner interpretation.

A contrast is a set of weights, one per group, that sum to zero. Weights of (1, −1, 0) compare groups 1 and 2 and ignore group 3. Weights of (1, 1, −2) compare the average of the first two against the third. Two contrasts are orthogonal — statistically independent — when the products of their corresponding weights sum to zero, and with k groups you can construct k − 1 mutually orthogonal contrasts that partition the between-groups SS completely.

L = Σ c_j M_j F = L² / (MS_within × Σ (c_j²/n_j))weighted combination of means, tested against error variancethe weights, summing to zeroeach group meanthe contrast value — the quantity you are testing1 for the numerator, always; error df from the ANOVAMeans: control 12, drug A 15, drug B 21; n = 5 each, MS_within = 10 Contrast: both drugs vs control, weights (−2, 1, 1) L = (−2)(12) + (1)(15) + (1)(21) = 12 Σ(c²/n) = (4/5) + (1/5) + (1/5) = 1.2 F(1, 12) = 144 / (10 × 1.2) = 12.0, p = .005 One test, one df, full power, and it answers the question the study was designed to ask.
Trend analysis, when the factor is ordered
With a quantitative factor — dose, practice sessions, age bands — the interesting question is usually the shape, not which pair differs. Orthogonal polynomial contrasts test this directly: linear weights (−1, 0, 1) for three levels, quadratic (1, −2, 1), and so on. A significant linear trend with a non-significant quadratic says the relationship is essentially straight; a significant quadratic says it bends, which is often the theoretically interesting result (think Yerkes-Dodson).
The mistake this topic produces
Calling a contrast 'planned' after having looked at the means. Planned means specified before the data existed, ideally in a preregistration; a contrast chosen after seeing which means are furthest apart is a post hoc test and needs the corresponding correction.
Do this now · 15 minutes
Write out the contrast weights for the specific hypothesis in a multi-group study of yours. Check that they sum to zero, and check orthogonality if you have more than one. Then run it and compare the p-value to the omnibus F — the contrast will usually be far more powerful.
Day 4355 minutes · two factors at once

Factorial ANOVA: two factors, three questions

A factorial design crosses two or more factors so that every combination of levels appears. A 2 × 3 design with therapy type (CBT, control) and severity (mild, moderate, severe) has six cells, and it answers three questions for the price of one study: the main effect of therapy, the main effect of severity, and their interaction.

The partitioning logic extends exactly as you would hope. Total SS splits into SS_A, SS_B, SS_AxB and SS_error, each with its own df, each mean square divided by MS_error to give its own F. The efficiency gain is real: testing two factors in one study takes far fewer participants than two separate studies, and only a factorial design can detect an interaction at all.

SourcedfWhat it tests
Factor Aa − 1Differences among A's marginal means, averaging over B.
Factor Bb − 1Differences among B's marginal means, averaging over A.
A × B(a − 1)(b − 1)Whether A's effect differs across levels of B. The reason to run a factorial design.
ErrorN − abWithin-cell variability — the noise term for all three F tests in a fully between-subjects design.
TotalN − 1Check: the four df above must sum to this.
Type I, II and III sums of squares
With unequal cell sizes the factors become correlated and the SS no longer partition uniquely, so the order of entry matters. Type I is sequential — each term adjusted only for those before it. Type III adjusts every term for all others and is what SPSS reports by default and what most psychology journals expect. Type II is more powerful when there is no interaction. Report which you used when cells are unbalanced; with equal n all three give identical answers and the question is moot.
The mistake this topic produces
Interpreting a main effect in the presence of a strong interaction. If therapy helps the mildly affected and harms the severe, the 'main effect of therapy' averages those into a number that describes nobody. Look at the interaction first — always.
Do this now · 15 minutes
Lay out the source table skeleton for a 2 × 4 factorial with 10 participants per cell: every source, its df, and the total. Confirm the df sum correctly. Then write the three hypotheses the design tests, in words.
Day 4455 minutes · the interesting result

Interactions, simple effects, and how to read a plot

An interaction means the effect of one factor depends on the level of another. It is usually the most theoretically interesting result in a factorial design, because moderation — 'it works, but only for these people' — is what most psychological theory actually predicts.

Ordinal interaction
Lines diverge, order preserved
One group benefits more, but the direction is the same for everyone. Main effects remain interpretable, with care.
Disordinal (crossover)
Lines cross
The effect reverses across levels. Main effects are actively misleading here and should not be interpreted.
Simple effects
The follow-up
Test the effect of A separately at each level of B. This is what you report after a significant interaction, not pairwise comparisons of all six cells.
Simple contrasts
Focused follow-up
A specific contrast within one level of the other factor — more powerful and more interpretable than running everything.

Reading the plot: parallel lines mean no interaction; non-parallel lines suggest one; crossing lines suggest a disordinal interaction. But eyeballing is not testing — a small departure from parallel in a low-powered design is noise, and interactions typically need roughly four times the sample size to detect than main effects of the same raw magnitude.

The interaction fallacy, worth knowing by name
Showing that an effect is significant in group A (p = .03) and not significant in group B (p = .21) does not establish that the effect differs between groups. The difference between significant and non-significant is not itself significant. To claim moderation you must test the interaction term directly — and papers that make this error are common enough that spotting it is a genuinely useful reviewing skill.
The mistake this topic produces
Following up a significant interaction with all pairwise cell comparisons. In a 2 × 3 design that is fifteen tests answering no clear question. Simple effects — the effect of factor A at each level of B — answer the question the interaction raised, in two or three tests.
Do this now · 15 minutes
Plot the cell means of a factorial dataset with error bars, classify the interaction as ordinal or disordinal, and run the simple effects. Then write the results paragraph in the correct order: interaction first, simple effects second, main effects last or not at all.
Day 4555 minutes · subject variance removed

Repeated-measures ANOVA and sphericity

When every participant experiences every condition, the same logic as the paired t-test scales up. Total variability now splits three ways: between-subjects (stable individual differences), the effect of condition, and the residual condition-by-subject interaction which becomes the error term. Because between-person variability is pulled out of the error, the design is substantially more powerful than its between-subjects equivalent.

SourcedfNote
Between subjectsn − 1Removed from error — this is where the power gain comes from. Not usually tested.
Within: conditionk − 1The effect of interest.
Within: error (condition × subject)(n − 1)(k − 1)The denominator for F. Typically much smaller than a between-subjects error term.
Totalnk − 1Check the df add up.
Sphericity, in plain language
Repeated-measures ANOVA assumes that the variances of the differences between every pair of conditions are equal. That is sphericity — a stronger and less intuitive assumption than homogeneity of variance. It is frequently violated when conditions are ordered in time, because adjacent measurements correlate more than distant ones. Mauchly's test detects it, badly: underpowered in small samples, oversensitive in large ones. The practical protocol is to apply a correction routinely. Greenhouse-Geisser multiplies both df by epsilon (conservative, recommended when epsilon is below .75); Huynh-Feldt is less conservative above that. Report the corrected df, which is why you will see F(1.62, 48.7) in published papers — not a typo.

With only two conditions sphericity cannot be violated (there is only one difference), which is why a paired t-test never needs it. The modern alternative for anything more complex is a mixed-effects model (volume 9), which handles the covariance structure explicitly and copes with missing data instead of deleting the whole participant.

The mistake this topic produces
Losing a participant's entire data because one condition is missing. Classical repeated-measures ANOVA is listwise: a single missing cell deletes the person. In a design with four sessions and typical attrition this can cost you a third of your sample — and it is the strongest practical argument for mixed models.
Do this now · 15 minutes
Run a repeated-measures ANOVA on data of yours, note Mauchly's result and epsilon, and report both corrected and uncorrected F. Then count how many participants were dropped for incomplete data, and write down what a mixed model would have retained.
Day 4650 minutes · between × within

Mixed designs: one between factor, one within

The commonest design in intervention research: participants are randomised to treatment or control (between) and measured before and after (within). It is powerful, it is efficient, and its output confuses people because it has two different error terms.

SourceError term usedQuestion
Group (between)Between-subjects errorDo the groups differ on average across time points? Usually not the question of interest.
Time (within)Within-subjects errorDid everyone change over time? Also usually not the question.
Group × TimeWithin-subjects errorDid the groups change differently? This is the treatment effect in a pre-post RCT.
Between error—Individual differences among people, pooled across cells.
Within error—Time × subject variability; typically smaller, which is why within effects are better powered.
The interaction is the result
In a pre-post randomised design, the treatment effect is the group × time interaction. A significant main effect of time only says both groups changed — which placebo, regression to the mean and maturation all predict. A significant main effect of group in a randomised trial is usually a baseline imbalance, which randomisation was supposed to prevent. Report the interaction, then the simple effect of time within each group, then the between-group difference at follow-up with its CI.

A word on the alternative. For pre-post data, ANCOVA on the post-test score with the pre-test as covariate is generally more powerful than the mixed ANOVA on change, and it is what most methodologists now recommend for randomised designs. Day 47 explains why — and why the same recommendation is dangerous in non-randomised ones.

The mistake this topic produces
Reporting the main effect of time in an RCT as evidence the intervention worked. Both groups improving is not a treatment effect. Only the differential change — the interaction — speaks to the intervention.
Do this now · 15 minutes
Analyse a pre-post two-group dataset three ways: mixed ANOVA, t-test on change scores, and ANCOVA on post with pre as covariate. Compare the p-values and effect sizes, and write which you would report in a randomised design and why.
Day 4755 minutes · adjustment has conditions

ANCOVA, and what you must never control for

Analysis of covariance removes variance associated with a continuous covariate from the error term before testing group differences. Done properly, it increases power substantially and adjusts for chance baseline differences. Done improperly, it manufactures effects that do not exist.

Mechanically, ANCOVA is a regression with both a categorical predictor and a continuous one. The group means are adjusted to what they would be if every group had the same covariate mean, and the error term shrinks by however much the covariate explains. With a covariate correlating .60 with the outcome, error variance drops by 36 per cent — a free increase in power.

01
The covariate must be measured before the manipulation
Or at least be unaffected by it. Controlling for something the treatment changed removes part of the treatment effect itself — mediation analysis, badly done.
02
Homogeneity of regression slopes
The covariate-outcome relationship must be similar in each group. Test the covariate × group interaction: if it is significant, ANCOVA's adjustment is misleading and you should be modelling the interaction instead.
03
Covariate measured reliably
Unreliable covariates under-adjust, leaving residual confounding that looks like a group effect. This is a serious problem with single-item or noisy baseline measures.
04
Linear relationship with the outcome
Check the scatter within each group before trusting the adjustment.
05
Groups must be randomly assigned for causal adjustment
In randomised designs ANCOVA adjusts for chance imbalance and gains power. In non-randomised designs it cannot fix systematic pre-existing differences, whatever the output suggests.
Lord's paradox, which every viva loves
Take two naturally different groups measured twice. Analysing change scores and analysing post-scores adjusted for pre-scores can give opposite answers from the same data — one showing no difference, the other a large one. Neither is a computational error; they answer different questions. Change scores ask 'did these groups change differently?'; ANCOVA asks 'for two people with the same baseline, does group predict the outcome?'. With randomisation the two converge. Without it, no statistical adjustment settles which is causally correct — that requires a causal model (day 112).
The mistake this topic produces
Controlling for a variable that lies on the causal pathway between your predictor and outcome. Controlling for socio-economic status when studying the effect of an intervention that changes socio-economic status removes the very effect you are trying to measure. This is one of the most common subtle errors in applied research.
Do this now · 15 minutes
For a planned analysis, list every covariate you intend to include and justify each in one sentence: what confound does it remove, and is it certain the treatment could not have affected it? Delete any you cannot justify — each unjustified covariate is a researcher degree of freedom.
Day 4850 minutes · several outcomes at once

MANOVA: several dependent variables at once

When a study has several conceptually related outcomes — the five subscales of a wellbeing measure, say — running five separate ANOVAs inflates the familywise error rate and ignores the correlations among the outcomes. MANOVA tests them jointly: does the group membership predict the combination of dependent variables?

It works by forming the linear combination of the DVs that maximally separates the groups, and testing that. Four multivariate test statistics appear in output — Wilks' lambda (most commonly reported), Pillai's trace (most robust to assumption violations), Hotelling's trace and Roy's largest root. Report Pillai's when assumptions are shaky; Wilks' otherwise, and mention which.

When MANOVA earns its complexity — and when it does not
It earns it when the DVs are moderately correlated, conceptually part of one construct, and you have a genuine multivariate hypothesis. It does not when the DVs are unrelated outcomes you happened to measure — then it answers a question nobody asked and you should simply correct for multiplicity across separate univariate tests. And the widespread practice of running MANOVA purely as a 'protection' gate before univariate ANOVAs has been criticised for decades: a significant multivariate test does not control the error rate of the follow-ups.

Assumptions are the multivariate versions of the familiar ones: multivariate normality, homogeneity of covariance matrices (Box's M, which is notoriously oversensitive), independence, and no severe multicollinearity among the DVs. When the DVs correlate above about .90 you are measuring the same thing twice and should combine them instead.

The mistake this topic produces
Following a significant MANOVA with uncorrected univariate ANOVAs and treating the multivariate test as having licensed them. Use discriminant function analysis (day 75) to understand what drove the multivariate effect, or correct the univariate follow-ups properly.
Do this now · 15 minutes
For a study with multiple outcomes, write down the correlation matrix of the DVs and decide honestly whether a multivariate hypothesis exists. If it does not, write the multiplicity-correction plan you will use instead.
Day 4955 minutes · frequencies, not means

Chi-square: goodness of fit and independence

When the outcome is a count of cases in categories rather than a score, the comparison is between observed frequencies and the frequencies a hypothesis predicts. Both chi-square tests use the same statistic and differ only in where the expected frequencies come from.

χ² = Σ (O − E)² / Esquared discrepancy between observed and expected, relative to how much was expectedobserved frequency in a cell — a count, never a percentage or a meanexpected frequency: total/k for goodness of fit; (row total × column total)/N for independencek − 1 for goodness of fit; (r − 1)(c − 1) for independenceIndependence: therapy type × improved? Improved Not Row total CBT 45 15 60 Control 30 30 60 Col total 75 45 120 E(CBT, improved) = (60 × 75) / 120 = 37.5 E(CBT, not) = (60 × 45) / 120 = 22.5 E(Control, improved) = 37.5, E(Control, not) = 22.5 χ² = (45−37.5)²/37.5 + (15−22.5)²/22.5 + (30−37.5)²/37.5 + (30−22.5)²/22.5 = 1.5 + 2.5 + 1.5 + 2.5 = 8.0 df = (2−1)(2−1) = 1, critical χ² = 3.84, p = .005 phi = √(8/120) = .26 — a small-to-moderate association
RequirementRuleIf violated
Expected frequenciesAll E ≥ 5, or at least 80% of cells with none below 1Use Fisher's exact test for 2 × 2, or collapse categories with a substantive justification.
Independence of observationsEach case contributes to exactly one cellUse McNemar's test for paired nominal data (before/after in the same people).
Raw counts onlyNever run it on percentages or meansThe test is meaningless; recompute from the raw frequencies.
Effect sizephi for 2 × 2, Cramér's V for larger tablesχ² grows with N, so it is not an effect size — a huge χ² in a huge sample can be a trivial association.
Yates' continuity correction
Applied to 2 × 2 tables to compensate for approximating a discrete distribution with a continuous one. It is conservative — some say excessively — and modern practice is to use Fisher's exact test for small tables and the uncorrected χ² otherwise. Know it exists, because syllabi still ask.
The mistake this topic produces
Running a chi-square when the same participants appear in more than one cell — for example, counting each of three responses per person. The independence requirement is violated and the test is invalid. McNemar or a multilevel logistic model is the answer.
Do this now · 15 minutes
Build a contingency table from your own categorical data, compute expected frequencies and χ² by hand, and check every E against the ≥ 5 rule. Then compute Cramér's V and write the result as a sentence that includes the association's size, not just its significance.
Day 5055 minutes · ranks answer differently

Nonparametric alternatives, and what they really test

Rank-based tests replace scores with their ranks and work with those. They require no distributional assumption about the population, which makes them useful for small samples, ordinal outcomes and severely skewed data. They are not, however, assumption-free — and they do not test the same hypothesis.

Parametric testRank alternativeNotes
Independent tMann-Whitney U (Wilcoxon rank-sum)Tests whether one group's scores tend to exceed the other's. Effect size: rank-biserial r, or the common language A statistic.
Paired tWilcoxon signed-rankUses both the sign and the magnitude of the ranked differences. The sign test uses only direction and is weaker still.
One-way ANOVAKruskal-Wallis HDistributed as χ² with k − 1 df. Follow up with Dunn's test, not with multiple Mann-Whitneys.
Repeated-measures ANOVAFriedman testRanks within each participant across conditions. Follow up with Wilcoxon plus correction.
Pearson rSpearman rho, Kendall tauMonotone rather than linear association. Kendall behaves better with small n and ties.
Chi-square, small samplesFisher's exact testExact probabilities from the hypergeometric distribution; no expected-frequency requirement.
What Mann-Whitney actually tests
It is routinely described as 'comparing medians'. That is only true under the additional assumption that the two distributions have the same shape and spread, differing only in location. Without that assumption it tests stochastic dominance: the probability that a randomly chosen case from group 1 exceeds one from group 2. Two distributions with identical medians but different shapes can produce a highly significant U. Say what you are testing.

And the standard reason to reach for these tests is weaker than it looks. With moderate n the CLT protects the t-test well against non-normality; where rank tests genuinely win is with ordinal outcomes, tiny samples, or extreme outliers that a mean cannot survive. For heavy-tailed continuous data, modern robust methods — trimmed means with bootstrap intervals, day 110 — usually outperform rank tests while keeping an interpretable effect size.

The mistake this topic produces
Choosing a nonparametric test because Shapiro-Wilk was significant at n = 200. At that sample size the t-test is robust and the rank test answers a different question with less interpretable output. Base the choice on the shape and the measurement level, not on a normality test.
Do this now · 15 minutes
Run both the parametric and the rank-based version of one of your comparisons. Compare p-values and, more importantly, write down what each test's null hypothesis actually says. Decide which matches your research question.
Reference

Which test, for which design

The volume in one table. Take your design across the row; the parametric test, its rank-based alternative and the required effect size follow.

DesignParametricNonparametricEffect size
One sample vs a known valueOne-sample tWilcoxon signed-rankCohen's d
Two independent groupsWelch's tMann-Whitney UCohen's d / Hedges' g
Two related measuresPaired tWilcoxon signed-rankd_z or d_av
Three or more independent groupsOne-way ANOVAKruskal-Wallisη², ω²
Three or more related measuresRepeated-measures ANOVAFriedmanpartial η², generalised η²
Two crossed factorsFactorial ANOVAAligned rank transformpartial η² per effect
Between × withinMixed ANOVAART or mixed modelpartial η² for the interaction
Groups plus a baseline covariateANCOVARank ANCOVA (rare)partial η², adjusted mean difference
Several related outcomesMANOVA—Multivariate η², Pillai's
Two categorical variablesChi-square independenceFisher's exactphi, Cramér's V
One categorical variable vs expectedChi-square goodness of fitExact multinomialw (Cohen's)
Paired categorical (before/after)McNemarExact McNemarOdds ratio
Reference

Reporting sentences for every test in this volume

Fill the numbers, keep the structure. Every one includes df, exact p, an effect size and an interval — which is what current reporting standards require.

TestSentence
Independent tThe treatment group improved more (M = 18.2, SD = 6.1) than controls (M = 22.9, SD = 6.5), t(58) = 2.88, p = .006, d = 0.75, 95% CI [0.22, 1.27].
Paired tSymptoms fell from pre-test (M = 24.1, SD = 5.8) to post-test (M = 19.5, SD = 6.2), t(29) = 6.57, p < .001, d_z = 1.20, 95% CI of the difference [3.2, 6.0].
One-way ANOVAThe three conditions differed, F(2, 12) = 10.50, p = .002, η² = .64, 90% CI [.24, .76].
Post hocTukey-corrected comparisons showed group C exceeded group A by 9.0 points, 95% CI [3.4, 14.6], p = .003; A and B did not differ, p = .41.
FactorialThe therapy × severity interaction was significant, F(2, 84) = 5.12, p = .008, partial η² = .11; simple-effects analysis showed therapy helped mild cases, F(1, 84) = 12.4, p < .001, but not severe cases, F(1, 84) = 0.41, p = .52.
Repeated measuresMauchly's test indicated a violation of sphericity, so Greenhouse-Geisser corrected values are reported, F(1.62, 46.9) = 8.31, p = .002, partial η² = .22.
ANCOVAControlling for baseline severity, the adjusted post-test means differed, F(1, 57) = 9.14, p = .004, partial η² = .14; adjusted difference 3.8 points, 95% CI [1.3, 6.3].
Chi-squareImprovement was associated with therapy type, χ²(1, N = 120) = 8.00, p = .005, phi = .26.
Mann-WhitneyScores were higher in the treatment group (Mdn = 21) than in controls (Mdn = 17), U = 254, z = −2.61, p = .009, rank-biserial r = .38.
Checkpoint

Ten questions before you move on

A mix of computation, choice-of-test and interpretation — the three things group-comparison questions are built from. Answer without looking back.

{{ quizCounter }}
{{ quizScore }}
{{ quizQ }}
{{ quizFb }}

Next: volume 5, Association

Fourteen days on correlation and the general linear model: r and its many disguises, partial correlation, regression and its diagnostics, multiple predictors, moderation, mediation, logistic and count models — and the moment you realise every test in volume 4 was a regression in costume.

Start day 51 →