Volume 5 · Correlation, regression, prediction

One framework underneath everything: a line, and how wrong it is.

Fourteen days from covariance to logistic regression. Along the way the realisation that reorganises the whole subject: t-tests, ANOVA and ANCOVA are all special cases of regression, and once you can read a regression table you can read most of the empirical literature in your field.

14
days of 45–55 min
9
kinds of correlation
5
regression diagnostics
1
general linear model
Day 5150 minutes · from cross-products to r

Covariance and the Pearson correlation

Two variables covary when people above the mean on one tend to be above the mean on the other. Covariance measures that directly: multiply each person's two deviations together and average. Positive products dominate when the association is positive; they cancel when there is none.

The trouble with covariance is that its size depends on the units — covariance between height in millimetres and weight in grams is a huge number that means nothing. Dividing by both standard deviations strips the units out and bounds the result between −1 and +1. That is the Pearson correlation: a standardised covariance.

cov = Σ(X − M_x)(Y − M_y) / (n − 1) r = cov / (s_x s_y)average co-deviation, divided by the two standard deviations to make it unit-freea cross-product: positive when both deviations share a signsame df logic as variance — two means were estimatedbounded −1 to +1; zero means no LINEAR associationn − 2, because the line costs a slope and an interceptFive participants, hours slept (X) and mood (Y): X: 5, 6, 7, 8, 9 M_x = 7, s_x = 1.58 Y: 3, 5, 4, 8, 10 M_y = 6, s_y = 2.92 (X−M)(Y−M): (−2)(−3)=6, (−1)(−1)=1, (0)(−2)=0, (1)(2)=2, (2)(4)=8 Σ = 17 cov = 17 / 4 = 4.25 r = 4.25 / (1.58 × 2.92) = 0.92 Test: t = r√((n−2)/(1−r²)) = .92 × √(3/0.1536) = 4.07 df = 3, p = .027

Spend five minutes with the simulator before moving on. Set r to .30 — the size of a typical published effect in social and personality psychology — and look at the cloud. Most people are surprised by how weak a 'moderate' correlation looks. Then press hide r — guess it a dozen times. Calibrating your eye against real scatter is worth more than memorising the benchmarks.

The mistake this topic produces
Reading r as a percentage. An r of .50 is not 'half' of anything; it is 25 per cent of shared variance. And r is not a linear scale — the difference between .20 and .30 is not the same amount of association as between .80 and .90.
Do this now · 15 minutes
Compute r by hand for five or six pairs of your own scores, laying out the cross-product table. Then check against software and against the scatterplot. If the plot shows a curve, note that your r is understating a real relationship.
Day 5250 minutes · the four distortions

What r does not mean: range, reliability, outliers, cause

The correlation coefficient is the most reported and most over-read statistic in psychology. Four things distort it systematically, and all four are examinable.

Restriction of range
Truncated variability shrinks r
Correlate SAT scores with university grades among admitted students only and the correlation looks weak — because everyone admitted scored highly. Selection on either variable attenuates r, sometimes dramatically.
Attenuation by unreliability
Noise caps the correlation
The maximum observable correlation between two measures is √(r_xx × r_yy). With two scales at .70 reliability, r cannot exceed .70 even if the constructs correlate perfectly.
Outliers and leverage
One point can do anything
A single extreme case can create an r of .60 in uncorrelated data, or hide a real relationship. Always look at the scatterplot — Anscombe's quartet is the classic demonstration.
Non-linearity
r measures straight lines only
Yerkes-Dodson arousal and performance is a strong relationship with an r near zero. A correlation of zero means no linear trend, not no relationship.
r_true = r_observed / √(r_xx × r_yy)what the correlation would be if both measures were perfectly reliablereliability coefficients of the two measuresthe disattenuated estimate — an upper bound, to be reported as suchObserved r = .40 Reliabilities: .75 and .80 r_true = .40 / √(.75 × .80) = .40 / .775 = .52 So a third of the 'missing' association was measurement error. Report both values, and never the corrected one alone — it can exceed 1.0 when reliabilities are poorly estimated.
Correlation and causation, stated usefully
Everyone can recite that correlation does not imply causation. The useful version names the four alternatives to X causing Y: Y causes X (reverse causation); a third variable Z causes both (confounding); selection into the sample created the association (collider bias — day 112); or it is sampling noise. When you write a discussion section, address whichever of these four is most plausible for your data. That paragraph is what distinguishes a careful thesis from a careless one.
The mistake this topic produces
Reporting a correlation matrix of twenty variables and interpreting whichever coefficients reached .05. With 190 correlations, about ten will be 'significant' by chance alone. Correct, or label the matrix as exploratory and interpret the pattern rather than individual cells.
Do this now · 15 minutes
For your main correlation, compute r², compute the disattenuated estimate using your scales' reliabilities, and check the scatterplot for outliers and curvature. Then write the sentence you would publish, including r² and the caveat that matters most for your data.
Day 5345 minutes · match the coefficient to the scales

Spearman, Kendall, point-biserial and the rest

Pearson r assumes two continuous, roughly linearly related variables. When your measurement levels differ, the coefficient changes name — but nearly all of them are mathematically Pearson's r applied to transformed or coded data, which is why the interpretation carries across.

CoefficientVariablesNotes
Pearson rTwo continuousThe default. Linear association only.
Spearman rhoTwo ordinal, or non-linear monotoneLiterally Pearson r computed on ranks. Robust to outliers and to monotone non-linearity.
Kendall tauTwo ordinal, small n or many tiesBased on concordant vs discordant pairs. More interpretable as a probability, and better behaved than rho with small samples.
Point-biserialOne genuine dichotomy, one continuousPearson r with the dichotomy coded 0/1. Directly related to the independent t-test: they test the same thing.
BiserialOne artificial dichotomy (a split continuum), one continuousEstimates what r would have been without the dichotomisation. Larger than the point-biserial.
PhiTwo genuine dichotomiesPearson r on two 0/1 variables; equals √(χ²/N) for a 2 × 2 table.
TetrachoricTwo artificial dichotomiesEstimates the underlying continuous correlation. Used in IRT and in factor-analysing binary items.
PolychoricTwo ordinal items assumed to have continuous latentsThe correct input for factor-analysing Likert items — Pearson r on 5-point items underestimates loadings.
Eta (η)One nominal with k levels, one continuousNon-linear association; η² is the ANOVA effect size from day 40.
The dichotomisation warning
Splitting a continuous variable at the median to make groups throws away information, reduces power by roughly the equivalent of discarding a third of your sample, and can create spurious interactions. It remains common because group comparisons feel simpler to describe. Keep the variable continuous and use regression; if you must present groups, do it in a figure while analysing the continuum.
The mistake this topic produces
Factor-analysing five-point Likert items using Pearson correlations and then wondering why the loadings are modest and an extra 'difficulty factor' appears. Polychoric correlations are the correct input, and most modern software offers them.
Do this now · 15 minutes
Identify a pair of variables in your data whose measurement levels do not fit Pearson. Compute the appropriate coefficient, and compute Pearson as well. Note the difference — and write which you would report and why.
Day 5450 minutes · holding things constant

Partial and semipartial correlation

Two variables can correlate because they share a cause. Partial correlation asks what remains of their association once a third variable is statistically removed from both. It is the arithmetic behind the phrase 'controlling for', and it is the conceptual bridge to multiple regression.

r_xy.z = (r_xy − r_xz·r_yz) / √((1 − r²_xz)(1 − r²_yz))the correlation between X and Y after removing Z's influence from boththe original zero-order correlationeach variable's correlation with the control variablethe partial: read the dot as 'controlling for'Ice-cream sales and drownings: r_xy = .70 Both correlate with temperature: r_xz = .80, r_yz = .75 r_xy.z = (.70 − (.80 × .75)) / √((1 − .64)(1 − .5625)) = (.70 − .60) / √(.36 × .4375) = .10 / .397 = .25 Most of the association was temperature. Semipartial (removing Z from X only): sr = (.70 − .60) / √(1 − .5625) = .15
Partial versus semipartial, and why regression uses the second
A partial correlation removes Z from both X and Y. A semipartial (or part) correlation removes Z from X only, leaving Y intact. The squared semipartial is the unique proportion of the total variance in Y explained by X over and above the other predictors — which is exactly what ΔR² reports when you add a predictor to a regression (day 58). That is why semipartials, not partials, are the natural regression quantity.

A caution that day 112 will formalise. Statistical control is not experimental control. It adjusts for the variable as you measured it, so an unreliable or coarsely measured control leaves residual confounding behind — and controlling for the wrong variable, such as a mediator or a collider, actively introduces bias rather than removing it.

The mistake this topic produces
Interpreting a partial correlation as though it established what would happen if the control variable were physically held constant. It shows the association among people who happen to be similar on Z, in your sample, as measured. That is considerably weaker than an intervention.
Do this now · 15 minutes
Compute a partial correlation by hand for a triple of your variables using the formula above, then reproduce it in software. Write one sentence interpreting the change from the zero-order r, naming what the control variable plausibly represents.
Day 5555 minutes · least squares

Simple linear regression

Correlation says how strongly two variables move together. Regression says by how much: for every unit increase in X, how many units does Y change? That makes it a prediction tool and, more importantly, the general framework into which every design in volume 4 fits.

Ŷ = a + bX b = r(s_y / s_x) a = M_y − b·M_xpredicted score equals intercept plus slope times predictorunstandardised slope: the change in Y per one-unit change in X, in real unitsintercept: predicted Y when X = 0 — meaningful only if X = 0 is meaningfulstandardised slope; in simple regression β = r exactlyY − Ŷ: what the line got wrong for that personSleep hours (X) predicting mood (Y) r = .92, s_x = 1.58, s_y = 2.92, M_x = 7, M_y = 6 b = .92 × (2.92 / 1.58) = 1.70 a = 6 − (1.70 × 7) = −5.90 Ŷ = −5.90 + 1.70X Each extra hour of sleep predicts 1.70 more mood points. At X = 8: Ŷ = −5.90 + 13.60 = 7.70 r² = .846 — the model accounts for 85% of variance in mood. Note the intercept: mood at zero hours of sleep is negative, which is nonsense. That is extrapolation, not a defect of the fit.

Least squares chooses the a and b that make the sum of squared residuals as small as possible. That is a specific decision with consequences: squaring makes large errors dominate, so a single extreme case can pull the line noticeably. It is also why regression and ANOVA share the same partitioning logic — regression's SS_regression and SS_residual are volume 4's between and within, wearing different labels.

b or β — which to report
Report the unstandardised b with its confidence interval as the primary result: it is in real units and can be acted on ('each session reduced symptoms by 1.7 points'). Report β when you want to compare predictors measured on different scales within the same model. Never compare β across different samples or models — β depends on the sample's variances, so it changes when the sample does, even if the underlying relationship does not.
The mistake this topic produces
Extrapolating beyond the range of X in your data. The line continues forever mathematically; your evidence does not. Predicting from a model fitted on adults aged 18–35 what will happen at age 70 is not statistics, it is optimism.
Do this now · 15 minutes
Fit a simple regression on your own data by hand using the formulas above — b from r and the two SDs, then a. Confirm against software, then write the interpretation of b as a sentence a non-statistician would understand.
Day 5655 minutes · read the residuals

Diagnostics: the four plots that check everything

Regression assumptions are about residuals, not about your raw variables — a distinction worth repeating because it saves students from unnecessary transformations. Four plots check almost everything that matters, and reading them is faster than running tests.

AssumptionHow to checkWhat a violation does
LinearityResiduals vs fitted values — look for curvatureThe model answers the wrong question; the slope is an average over a curve that fits nobody.
HomoscedasticityThe same plot — look for a funnelCoefficients stay unbiased, but standard errors and p-values are wrong. Use robust (HC3) standard errors.
Normality of residualsQ-Q plot of residualsMild violations matter little at moderate n; severe ones affect intervals and prediction.
IndependenceDesign, plus Durbin-Watson for ordered dataCatastrophic, as always. Clustered or time-ordered data need a different model.
No influential outliersCook's distance, leverage, standardised residualsA single case can change your slope, your significance and your conclusion.
Leverage, discrepancy, influence
Three different things, often confused. Leverage is being extreme on the predictors — an unusual X. Discrepancy is a large residual — an unusual Y given X. Influence is the product: a case that is both, and therefore actually moves the line. Cook's distance measures influence; values above 1 (or above 4/n as a sensitive screen) deserve inspection. An extreme case with low leverage barely matters; a moderate case at the far end of X can matter enormously.

What to do when you find one: never delete silently. Check for data error first. If the case is real, report the analysis with and without it. If the conclusion depends on one participant out of 120, that is the finding, and it belongs in the results rather than being quietly removed.

The mistake this topic produces
Checking normality of the predictor or the outcome instead of the residuals. Regression makes no assumption at all about the distribution of X, and the assumption about Y is conditional on X. A heavily skewed predictor is not in itself a problem.
Do this now · 15 minutes
Produce the four diagnostic plots for a regression of yours. Write one sentence per plot on what you see. Then compute Cook's distance and identify your most influential case — and find out who that participant is and why they are unusual.
Day 5755 minutes · several predictors

Multiple regression: partial slopes and R²

Add predictors and the equation extends: Ŷ = a + b₁X₁ + b₂X₂ + ... Each b is now a partial slope — the change in Y per unit of that predictor holding all the others constant. That phrase is doing enormous work, and misunderstanding it is the source of most misreadings of regression tables.

Ŷ = a + b₁X₁ + b₂X₂ + … + b_k X_keach slope is that predictor's unique contribution, with the others held fixedproportion of variance in Y explained by all predictors togetherR² penalised for the number of predictors — always report this one tootests the whole model: F(k, n − k − 1)tests whether that predictor adds anything beyond the othersPredicting exam score from study hours and anxiety, n = 100 b_hours = 2.41, SE = 0.55, t = 4.38, p < .001, β = .41 b_anxiety = −0.87, SE = 0.31, t = −2.81, p = .006, β = −.26 R² = .31, adjusted R² = .29 F(2, 97) = 21.8, p < .001 Each extra study hour predicts 2.41 more marks for students with the SAME anxiety level. Together the two predictors account for 31% of variance in exam score.
The two F tests and the t tests
The omnibus F asks whether the set of predictors together explains variance. Each t asks whether one predictor adds anything unique beyond all the others. They can disagree, and the disagreement is informative: a significant F with no significant individual predictors is the classic signature of multicollinearity (day 58) — the predictors jointly explain a lot, but overlap so heavily that none has a distinguishable unique contribution.

The unifying realisation of this volume: ANOVA is multiple regression with dummy-coded predictors, the t-test is regression with one dummy, ANCOVA is regression with a dummy and a continuous predictor, and multiple regression is all of them at once. The general linear model is one model. Once you see that, volume 4's menu of named tests becomes a single framework with different inputs.

The mistake this topic produces
Interpreting a non-significant predictor as unimportant. It means that predictor adds nothing beyond the others in this model. Remove a correlated predictor and it may become highly significant. Significance in a regression table is always conditional on the rest of the model.
Do this now · 15 minutes
Fit a two-predictor regression and write out the interpretation of each b, each β, R² and adjusted R² in full sentences. Then remove one predictor and refit. Note how the surviving coefficient changes — and explain why in terms of the correlation between the predictors.
Day 5855 minutes · order and overlap

Hierarchical entry, stepwise, and multicollinearity

How predictors enter a model is a substantive decision, not a technical one, and reviewers judge it. There are three approaches, and only two are defensible.

Simultaneous
All predictors at once
The default. Appropriate when you have no theoretical ordering and want each predictor's unique contribution.
Hierarchical
Blocks in a theory-driven order
Enter covariates first, then the predictor of interest, and test ΔR² at each step. This is how you show a variable predicts above and beyond established ones — the standard design for incremental validity.
Stepwise
Software chooses by p-value
Capitalises on chance, produces unstable models that differ with tiny data changes, and yields invalid p-values and inflated R². Widely criticised for forty years and still widely used. Avoid, and be ready to say why in a viva.
Best subsets / LASSO
Algorithmic but principled
If the goal is genuinely prediction rather than explanation, use regularisation with cross-validation (day 64) — not stepwise.
F_change = (ΔR² / m) / ((1 − R²_full) / (n − k − 1))how much new variance this block adds, relative to what is still unexplainedR² with the block minus R² without itnumber of predictors added in this blocktotal predictors in the full modelequals the squared semipartial correlation when the block is a single predictorStep 1: age and gender R² = .12 Step 2: add self-efficacy R² = .27 ΔR² = .15 with m = 1, n = 120, k = 3 F = (.15 / 1) / ((1 − .27) / 116) = .15 / .00629 = 23.8 df = (1, 116), p < .001 Self-efficacy explains an additional 15% of variance beyond demographics.
Multicollinearity: when predictors overlap too much
Highly correlated predictors make it impossible to separate their unique contributions. Symptoms: a significant overall F with no significant individual predictors; coefficients that flip sign when another predictor is added; standard errors that balloon. Diagnose with the variance inflation factor — VIF above 5 warrants attention, above 10 is generally treated as serious, though these are conventions. Remedies: drop one of the pair, combine them into a composite, centre interaction terms (day 60), or use regularisation. Note that multicollinearity does not bias the coefficients; it makes them unstable and imprecise.
The mistake this topic produces
Running stepwise regression and reporting the resulting p-values as if they were valid. They are not — the selection process has already used the data, so the reported significance is badly optimistic and the model rarely replicates.
Do this now · 15 minutes
Run a hierarchical regression with a theoretically motivated order: controls first, focal predictor second. Report ΔR² and its F test at each step, and check VIFs. Then write the incremental-validity sentence that such a model exists to produce.
Day 5950 minutes · categories in a line

Dummy coding: how categories enter a regression

Regression needs numbers, and categories are not numbers. Coding schemes translate a k-level factor into k − 1 numeric predictors — which is where the k − 1 degrees of freedom in ANOVA came from. Understanding the coding is what lets you read the coefficients.

SchemeHowThe intercept and slopes mean
Dummy (indicator)One 0/1 variable per non-reference groupIntercept = mean of the reference group. Each b = that group's difference from the reference. The default everywhere.
Effect (deviation)Codes of 1, 0, −1Intercept = grand mean. Each b = that group's deviation from the grand mean. This is what classical ANOVA uses internally.
ContrastCustom orthogonal weightsEach b tests a specific planned comparison (day 42). Most informative when you have real hypotheses.
HelmertEach level against the mean of subsequent levelsUseful for ordered categories or nested comparisons.
Ŷ = a + b₁D₁ + b₂D₂with dummy coding, the intercept is the reference group and each slope is a difference from it1 if group 2, else 01 if group 3, else 0the group coded 0 on both — chosen by you, and it should be the theoretically meaningful baselineControl M = 12, Drug A M = 15, Drug B M = 21 Control as reference: a = 12.0 (control mean) b₁ = 3.0 (drug A − control) b₂ = 9.0 (drug B − control) Each t tests that difference against the control. Switch the reference to drug A and the coefficients change, though the model fit, R² and F stay identical.
Why this matters beyond bookkeeping
Because the reference group is a choice, the individual coefficients and their p-values depend on a decision you made. The overall model — R², the omnibus F — does not. If a reviewer asks why one comparison is significant and another is not, the answer often lies in the reference category rather than in the data. Choose the baseline that matches your research question, and say that you did.
The mistake this topic produces
Entering a nominal variable coded 1, 2, 3 as a single numeric predictor. That imposes equal spacing and a linear order on categories that have neither, and the resulting coefficient is meaningless. Always create k − 1 coded variables, which every statistics package will do if you declare the variable as a factor.
Do this now · 15 minutes
Recreate one of your one-way ANOVAs as a regression with dummy codes. Confirm that F, R² and η² match exactly. Then change the reference group and note which coefficients change and which do not.
Day 6055 minutes · it depends

Moderation: interactions in regression

A moderator changes the strength or direction of the relationship between a predictor and an outcome. Statistically, it is an interaction term: the product of the two predictors, entered after both main effects. If the product term is significant, the slope of X on Y depends on the level of Z.

Ŷ = a + b₁X + b₂Z + b₃(X × Z)the effect of X on Y is (b₁ + b₃Z) — a slope that changes with Zthe interaction: how much X's slope changes per unit of Zthe effect of X when Z = 0 — which is why centring mattersb₁ + b₃Z, evaluated at chosen values of ZStress (X) predicting depression (Y), social support (Z) as moderator. Both predictors mean-centred. b₁ (stress) = 0.62, p < .001 b₂ (support) = −0.34, p = .003 b₃ (interaction) = −0.21, p = .012 Simple slopes at ±1 SD of support (SD = 1.0): Low support (Z = −1): 0.62 + (−0.21)(−1) = 0.83 High support (Z = +1): 0.62 + (−0.21)(+1) = 0.41 Stress predicts depression twice as steeply for people with low support. That is the buffering hypothesis, tested.
Centre the predictors, and know why
Subtracting the mean from X and Z before forming the product does two things. It makes b₁ and b₂ interpretable — they become the effect of each predictor at the average level of the other, rather than at a zero that may not exist. And it dramatically reduces the correlation between the product term and its components, removing one source of multicollinearity. It does not change b₃, its p-value or the model fit. Centring is interpretive hygiene, not a statistical fix.

Follow up a significant interaction the same way as in factorial ANOVA: with simple slopes at meaningful values of the moderator (conventionally the mean and ±1 SD), and a plot. The Johnson-Neyman technique goes further, identifying the exact range of the moderator over which the effect of X is statistically significant — more informative than three arbitrary points, and increasingly expected.

One practical warning that governs design: interactions require substantially more power than main effects. Detecting a moderation of realistic size often needs four times the sample of the corresponding main effect, which is why the literature is full of underpowered and unreplicable moderation findings.

The mistake this topic produces
Interpreting b₁ as 'the main effect of X' when the interaction is in the model and the predictors are uncentred. With raw scores, b₁ is the effect of X when Z equals exactly zero — often an impossible value such as an IQ of 0 — which makes it uninterpretable rather than merely awkward.
Do this now · 15 minutes
Test a moderation hypothesis of your own: centre both predictors, create the product, enter main effects then the interaction, and report ΔR² for the product term. If significant, compute and plot simple slopes at ±1 SD.
Day 6155 minutes · through what mechanism

Mediation: testing a mechanism

A mediator is the mechanism through which a predictor affects an outcome: therapy reduces depression by increasing behavioural activation. Where moderation asks when, mediation asks how. It is the most popular and most over-claimed analysis in applied psychology.

PathMeaningNotation
cTotal effect of X on Y, ignoring the mediatorY = a + cX
aEffect of X on the mediatorM = a₀ + aX
bEffect of M on Y, controlling for XY = a₁ + c′X + bM
c′Direct effect of X on Y after accounting for MWhat is left over
a × bThe indirect effect — the quantity of interestc = c′ + ab in linear models
The exam answer and the modern answer
The exam answer is Baron and Kenny's four steps: show X predicts Y, X predicts M, M predicts Y controlling for X, and the X–Y relationship shrinks (full mediation if to non-significance, partial if it merely weakens). Know it — it is still examined. The modern answer is that the causal-steps approach has low power, that step one is unnecessary (an indirect effect can exist with no total effect when paths offset), that the Sobel test assumes a normal sampling distribution the product ab does not have, and that the correct test is a bootstrap confidence interval on ab with 5,000 or more resamples. If it excludes zero, the indirect effect is supported. Report both the indirect effect with its bootstrap CI and the completely standardised effect size.

The caution that matters more than the statistics. Mediation is a causal claim, and running it on cross-sectional data does not make it one. With X, M and Y all measured at one time point, the data are equally consistent with several different causal orderings — including M causing X. Establishing mediation properly needs temporal separation at minimum, and ideally experimental manipulation of the mediator.

The mistake this topic produces
Describing 'full mediation' because c′ became non-significant. Non-significance is not absence (day 25), and c′ dropping below the threshold often reflects the loss of power from adding a correlated predictor. Report the size of the indirect effect and its interval, and drop the full-versus-partial language entirely.
Do this now · 15 minutes
Specify a mediation model from your own theory, drawing the path diagram and labelling a, b, c and c′. Write down what temporal ordering would be needed for the causal claim to be credible — and note honestly whether your design has it.
Day 6255 minutes · binary outcomes

Logistic regression: when the outcome is yes or no

Relapsed or not, passed or failed, diagnosed or not. Ordinary regression on a binary outcome predicts impossible values below 0 and above 1, and its residuals violate every assumption at once. Logistic regression models the log-odds instead, which is unbounded and behaves properly.

ln(p / (1 − p)) = a + b₁X₁ + b₂X₂predict the log-odds of the event; exponentiate a coefficient to get an odds ratioodds: probability of the event over probability of no eventthe logit — linear in the predictors, which is the whole trickthe odds ratio: multiplicative change in odds per unit of Xno effect; above 1 increases the odds, below 1 decreases themPredicting relapse from baseline severity b = 0.41, SE = 0.15, Wald z = 2.73, p = .006 OR = e^0.41 = 1.51, 95% CI [1.12, 2.02] Each one-point increase in severity multiplies the odds of relapse by 1.51 — a 51% increase in ODDS, not in risk. At 10% baseline risk: odds .111 → .168 → probability 14.4% So a 51% rise in odds is a 4.4 percentage-point rise in risk.
Odds ratios are not risk ratios
When the outcome is rare, the two are close. When it is common, the odds ratio badly exaggerates the risk difference — an OR of 2 with a 40 per cent baseline risk corresponds to a risk ratio of only about 1.4. Clinical and media misreporting of this is endemic. Report absolute risks or predicted probabilities alongside the OR, because those are what a decision-maker needs.

Other departures from ordinary regression: estimation is by maximum likelihood rather than least squares, so there is no true R² — the pseudo-R² measures (Cox and Snell, Nagelkerke, McFadden) are analogues, not proportions of variance, and should be labelled as such. Model comparison uses the likelihood ratio test or AIC. Classification tables and ROC curves with the area under the curve describe predictive accuracy, and the AUC has a nice interpretation: the probability that a randomly chosen case scores higher than a randomly chosen non-case.

The mistake this topic produces
Reporting the raw logit coefficient b as if it were interpretable to readers. Almost nobody thinks in log-odds. Exponentiate to an odds ratio, give its confidence interval, and add a predicted probability at a couple of meaningful predictor values.
Do this now · 15 minutes
Fit a logistic regression on a binary outcome of yours. Report the OR with its CI, then compute the predicted probability of the event at the 25th, 50th and 75th percentile of your key predictor. That three-number table communicates more than the coefficients do.
Day 6350 minutes · beyond binary

Ordinal, count and other generalized linear models

Logistic regression is one member of a family. A generalized linear model has three parts: a distribution for the outcome, a link function connecting the linear predictor to the mean of that distribution, and the linear predictor itself. Choose the distribution to match your outcome, and the rest follows.

OutcomeModelLinkReports
Continuous, symmetricLinear regressionidentityb, β, R²
BinaryLogisticlogitOdds ratios, AUC
Ordinal (Likert, stage, severity band)Proportional odds (ordinal logistic)cumulative logitCumulative odds ratios
Nominal with 3+ unordered levelsMultinomial logisticgeneralised logitORs per category vs a reference
Counts, variance ≈ meanPoisson regressionlogIncidence rate ratios
Counts, overdispersedNegative binomiallogIRRs plus a dispersion parameter
Counts with excess zerosZero-inflated or hurdle modellog + logitTwo sets of coefficients — one for the zeros, one for the counts
Proportions bounded 0–1Beta regressionlogitCoefficients on the proportion scale
The proportional odds assumption
Ordinal logistic regression assumes the effect of each predictor is the same across every cut-point of the outcome — the odds of scoring above 1 versus above 2 versus above 3 all shift by the same factor. Test it (the Brant test, or by comparing against a model that relaxes it). When it fails, either use a partial proportional odds model or a multinomial model, and say which. This assumption is the ordinal analogue of homogeneity of regression slopes in ANCOVA, and it is skipped far more often than it is checked.

Why this matters for the exam and for your thesis: treating a five-point Likert outcome as continuous is widespread and usually acceptable for a scale mean, but modelling a single ordinal item with linear regression can produce impossible predictions and distorted effects — especially when responses pile up at the ends. Knowing that the ordinal model exists, and when it is needed, is increasingly what separates competent applied work.

The mistake this topic produces
Fitting Poisson to overdispersed counts because the outcome is a count. Check the ratio of residual deviance to df; when it is well above 1, negative binomial is the answer and the Poisson standard errors were too small — making everything look more significant than it is.
Do this now · 15 minutes
Identify the correct GLM for each outcome in your own research, using the table above. For any that you have previously analysed with ordinary regression, note what the consequence of the mismatch was likely to be.
Day 6450 minutes · two different goals

Explanation versus prediction

Psychology's statistical training is built almost entirely around explanation: which variables have non-zero coefficients, and why. Prediction is a different goal with a different scorecard, and confusing them produces models that fit beautifully and forecast nothing.

Explanatory goal
Which variables matter, and how much
Judged by coefficient interpretability, theory, and whether confounds were addressed. Significance testing belongs here.
Predictive goal
How accurately can I forecast a new case
Judged by out-of-sample error. A predictor with no theoretical meaning is welcome if it improves accuracy.
Overfitting
Modelling your sample's noise
R² always rises when you add predictors, even random ones. With k predictors and small n, a model can fit your data almost perfectly and predict new data no better than the mean.
Shrinkage
The honest correction
Adjusted R² penalises for k. Cross-validated R² is better still: it estimates what will happen on data the model has never seen.
R²_adj = 1 − [(1 − R²)(n − 1) / (n − k − 1)]R² penalised for how many predictors you spent to get itnumber of predictorssample size — the penalty bites hardest when n is small relative to kR² = .40, n = 50, k = 8 R²_adj = 1 − [(1 − .40)(49) / (41)] = 1 − [(.60 × 49) / 41] = 1 − .717 = .283 Eleven points of R² were bought with degrees of freedom. With n = 500 and the same k, adjusted R² would be .390 — the penalty is negligible when you have data to spare.
Cross-validation in one paragraph
Split the data into k folds, usually five or ten. Fit the model on k − 1 folds and measure error on the held-out fold. Rotate, and average. The result estimates how the model performs on data it has not seen, which is the only honest measure of prediction. Doing this once on your own data is the fastest cure for over-confidence in R² — a model with R² = .45 in-sample routinely drops to .25 cross-validated. Day 111 returns to this with machine learning.

You now have the general linear model in full. Volume 6 relaxes the assumption that your variables are the things you care about — introducing latent variables, measured imperfectly by several indicators at once, which is the conceptual world psychometrics and most theory-driven psychology actually inhabits.

The mistake this topic produces
Reporting in-sample R² as evidence that your model 'predicts' the outcome. It quantifies fit to the data you already have. Prediction is a claim about future cases, and only out-of-sample validation supports it.
Do this now · 15 minutes
Take your fullest regression model and cross-validate it with five folds. Compare in-sample R² with the cross-validated estimate. Write down both numbers, and decide which one you would be willing to defend in a viva.
Reference

Everything in volume 4 was a regression

The single most useful table in this course. Every named test you learned as a separate procedure is one regression model with particular predictors.

Named testEquivalent regressionIdentity
One-sample tIntercept-only model, testing a = μ₀t is identical
Independent tY on one dummy predictort identical; t² = F
One-way ANOVA (k groups)Y on k − 1 dummy predictorsF and R² = η² identical
Factorial ANOVAY on dummies plus their productsInteraction term = product term
ANCOVAY on dummies plus a continuous covariateAssumes no dummy × covariate interaction
Pearson correlationY on standardised Xβ = r; t identical
Point-biserial correlationY on one dummySame test as the independent t
Repeated-measures ANOVAMixed model with random intercepts per subjectIdentical when the design is balanced and complete
Chi-square independenceLog-linear or logistic modelIdentical for 2 × 2 tables
Reference

The volume 5 formula sheet

QuantityFormulaNote
CovarianceΣ(X − M_x)(Y − M_y) / (n − 1)Unit-dependent; standardise it to get r.
Pearson rcov / (s_x s_y)Linear association only; bounded ±1.
Test of rt = r√((n − 2)/(1 − r²)), df = n − 2Two parameters estimated for the line.
Fisher z transformz = 0.5·ln((1 + r)/(1 − r)), SE = 1/√(n − 3)Required for CIs on r and for averaging correlations.
Coefficient of determinationr²Proportion of shared variance.
Partial correlation(r_xy − r_xz r_yz) / √((1 − r²_xz)(1 − r²_yz))Removes Z from both variables.
Regression slopeb = r(s_y / s_x)β = r in simple regression.
Intercepta = M_y − b·M_xThe line passes through (M_x, M_y) always.
Standard error of estimates_e = √(SS_residual / (n − 2))Average size of the prediction error.
Adjusted R²1 − [(1 − R²)(n − 1)/(n − k − 1)]Always report alongside R².
F for the modelF = (R²/k) / ((1 − R²)/(n − k − 1))Tests all predictors jointly.
ΔR² testF = (ΔR²/m) / ((1 − R²_full)/(n − k − 1))The hierarchical-entry test.
Attenuation correctionr / √(r_xx r_yy)An upper bound; report the observed r too.
Odds ratioe^b from logistic regressionNot a risk ratio when the outcome is common.
Checkpoint

Ten questions before you move on

Correlation and regression questions dominate the statistics section of most entrance papers. These are the forms they take.

{{ quizCounter }}
{{ quizScore }}
{{ quizQ }}
{{ quizFb }}

Next: volume 6, Multivariate

Twelve days on many variables at once and on latent structure: PCA, exploratory and confirmatory factor analysis, rotation, SEM and its fit indices, measurement invariance, cluster and discriminant analysis — the toolkit psychometric and individual-differences research runs on.

Start day 65 →