Every p-value you will ever read is an answer to a question about repeated sampling.
Eleven days on the machinery underneath inference. Probability rules and Bayes' theorem, the distributions psychology actually uses, and then the sampling distribution — the idea that separates people who can use statistics from people who can only run them. If volume 3 ever feels like arbitrary ritual, the missing piece is in here.
11
days of 40–55 min
5
distributions, derived not asserted
2
simulators to break
1
idea that unlocks all inference
Day 1245 minutes · the rules themselves
Probability rules, and the conditional that matters
Probability is a number between 0 and 1 expressing how often something happens in the long run. Three rules carry almost all of applied statistics, and the third is where psychology students lose marks.
Addition
P(A or B)
P(A) + P(B) − P(A and B). Subtract the overlap, or you count it twice. For mutually exclusive events the overlap is zero.
Multiplication
P(A and B)
P(A) × P(B|A). Only reduces to P(A) × P(B) when the events are independent — which is exactly the assumption every test makes about your participants.
Complement
P(not A)
1 − P(A). Quietly the most useful: 'at least one' problems are almost always easier solved as 1 − P(none).
Conditional
P(A|B)
P(A and B) / P(B). The probability of A in the restricted world where B has already happened. Not symmetric: P(A|B) ≠ P(B|A).
The asymmetry that underlies the entire p-value confusion
P(positive test | disease) is high for any decent test. P(disease | positive test) can still be tiny if the disease is rare. These are different quantities and reversing them is called the conditional probability fallacy. Hold on to it: a p-value is P(data this extreme | H₀ true), and the near-universal misreading treats it as P(H₀ true | data). Same error, different costume. Day 26.
P(A | B) = P(A and B) / P(B)of the times B happened, how often did A happen toojoint probability — both occurthe new, restricted universe you are dividing bywhen P(A|B) = P(A): knowing B tells you nothing about AOf 200 clients, 60 have an anxiety diagnosis; 40 of those 60 also report insomnia.
Of the whole sample, 90 report insomnia.
P(insomnia | anxiety) = 40 / 60 = .67
P(anxiety | insomnia) = 40 / 90 = .44
Same 40 people. Two very different numbers.
The denominator is the whole question.
The mistake this topic produces
Multiplying probabilities of events that are not independent. Two measurements from the same participant, two trials by the same rat, two children in the same classroom — these share variance, and treating them as independent inflates your effective sample size and your Type I error rate. That is the whole motivation for multilevel models in volume 9.
Do this now · 15 minutes
Build a 2 × 2 contingency table from any categorical pair in your own data. From it compute both conditional probabilities and the joint probability. Write in one sentence why the two conditionals differ, in the language of your actual variables.
Day 1350 minutes · base rates bite
Bayes' theorem, and why a positive test often is not
Bayes' theorem is the formal machine for reversing a conditional probability. It is examined as a formula and it is useful as a habit of thought: evidence updates a prior belief, and how much it updates depends on how rare the thing was to begin with.
P(H | E) = P(E | H) × P(H) / P(E)posterior = likelihood × prior, scaled so the probabilities add to oneprior — how common the hypothesis or condition is before the evidencelikelihood — how probable this evidence is if the hypothesis holds (test sensitivity)total probability of the evidence, across all hypothesesposterior — what you actually want to knowScreening for a disorder with prevalence 1% in 10,000 people.
Test sensitivity 90%, false positive rate 5%.
True positives = 100 × .90 = 90
False positives = 9,900 × .05 = 495
Total positives = 585
P(disorder | positive) = 90 / 585 = .154
A 90%-accurate test, a positive result — and an 85% chance
the person does not have the disorder. The base rate did that.
Work it as frequencies, not fractions: imagine 10,000 people and count. Every published demonstration that clinicians, doctors and judges misjudge these problems also shows that the frequency format fixes most of the error. Use it in exams and use it when you explain a screening result to a client.
Where this returns
In psychometrics as the logic behind predictive values and base-rate problems in clinical decision-making (volume 7). In volume 9 as the foundation of Bayesian estimation — where the prior is your state of knowledge before the study and the posterior is your state of knowledge after it. And immediately in your reading of the literature: if only 10 per cent of the hypotheses in a field are true, a literature of p < .05 findings contains far more false positives than the 5 per cent people imagine.
The mistake this topic produces
Ignoring the base rate because the test accuracy sounds impressive. The examinable version is the screening problem above; the professional version is a clinician over-diagnosing a rare condition because an instrument flagged it. Sensitivity and specificity are properties of the test; predictive value depends on the population.
Do this now · 15 minutes
Take a real cutoff-based instrument you know — a depression screener, a cognitive screen — and its published sensitivity and specificity. Compute the positive predictive value at a prevalence of 20 per cent and again at 2 per cent. Write one sentence about what that difference means for screening in a general population.
Day 1445 minutes · expectation and variance
Random variables, expected value, and the algebra of variance
A random variable is a variable whose value is determined by a chance process — the score of a randomly chosen participant, the mean of a randomly drawn sample. Everything inferential treats your statistics as random variables and asks what their distributions look like.
E(X) = Σ x·P(x) Var(X) = Σ (x − μ)²·P(x)the long-run average, and the long-run average squared distance from itexpected value — the mean you would get over infinite repetitionsprobability of each possible valuevariance of the random variable; its square root is the SDA gamble: win 10 with probability .2, win 0 with probability .8
E(X) = (10 × .2) + (0 × .8) = 2.0
Var(X) = (10 − 2)² × .2 + (0 − 2)² × .8
= 12.8 + 3.2 = 16.0
SD = 4.0
The expected value is 2 — a value that never actually occurs.
Rule
Statement
Why you need it
Adding a constant
E(X + c) = E(X) + c; Var(X + c) = Var(X)
Shifting every score leaves spread untouched. This is why centring predictors changes intercepts but not slopes (day 60).
Multiplying by a constant
E(cX) = c·E(X); Var(cX) = c²·Var(X)
Variance scales with the square — the reason SD, not variance, is in the original units.
Sum of independent variables
Var(X + Y) = Var(X) + Var(Y)
Variances add when independent. This is the engine behind the standard error formula on day 20.
Difference of independent variables
Var(X − Y) = Var(X) + Var(Y)
They still add. A surprise every year in exams — and the reason a difference score is noisier than either score.
Why the difference rule matters clinically
A change score (post minus pre) inherits the error variance of both measurements. If each has a reliability of .80, the difference score can have a reliability below .50 — which is why 'reliable change' needs its own index (day 81), and why analysing post-scores with pre-scores as a covariate is usually the more powerful design.
The mistake this topic produces
Assuming that because expected value is 'the average', it must be a value the variable can actually take. Expected family size 2.3, expected value of a gamble 2 when only 0 and 10 are possible — the expectation is a long-run balance point, not a prediction about any single case.
Do this now · 15 minutes
Write out the probability distribution of one discrete variable in your data — the number of errors per participant, say. Compute E(X) and Var(X) from the frequencies by hand, and confirm against the sample mean and variance from software.
Day 1545 minutes · n trials, two outcomes
The binomial distribution
The binomial describes the number of successes in n independent trials with constant probability p. It is the model behind the sign test, behind accuracy data in forced-choice tasks, behind pass/fail item responses, and behind any question of the form 'how likely is this many hits by chance'.
P(k) = nCk · p^k · (1 − p)^(n − k)ways it could happen × probability of one such waynumber of independent trialsnumber of successes you are asking aboutprobability of success on any single trialcombinations: n! / (k!(n − k)!) — how many orderings give k successesA participant guesses on 10 two-alternative trials (p = .5).
Probability of getting exactly 8 correct:
10C8 = 10! / (8!·2!) = 45
P(8) = 45 × .5^8 × .5^2 = 45 × .00390625 = .0439
Probability of 8 OR MORE (the one-tailed test):
P(8) + P(9) + P(10) = .0439 + .0098 + .0010 = .0547
Just above .05 — so eight out of ten is not quite
evidence of above-chance performance.
Mean of a binomial is np, variance is np(1 − p). Both are worth memorising: they let you sanity-check accuracy data instantly. Fifty trials at chance gives an expected 25 correct with an SD of 3.54, so a participant scoring 33 sits more than two SDs above chance.
The normal approximation, and its condition
When np and n(1 − p) both exceed about 10, the binomial is well approximated by a normal distribution with mean np and SD √(np(1−p)). That is why proportions can be tested with z, and why the whole apparatus of the chi-square test for frequencies works. Below that, use exact binomial probabilities — which is precisely when exam questions ask for the formula.
The mistake this topic produces
Treating trials from one participant as independent when performance drifts with fatigue, learning or strategy change. The binomial assumes constant p across trials; systematic drift violates it and makes the resulting p-value optimistic.
Do this now · 15 minutes
For a forced-choice task in your field, compute the number correct needed to beat chance at p < .05 one-tailed, for n = 20 and n = 50 trials. Note how the required proportion falls as n rises — that is statistical power arriving early (day 30).
Day 1640 minutes · counts and rare events
Poisson and the other discrete models
When your outcome is a count in a fixed window — seizures per month, aggressive incidents per week, lapses of attention per session, words recalled — the Poisson distribution is usually a better model than the normal. It has one parameter, λ, which is both its mean and its variance.
Distribution
Models
Psychology example
Binomial
Successes in a fixed number of trials
Correct responses out of 40 items; number of participants relapsing out of 60.
Poisson
Counts in a fixed interval, mean = variance
Self-harm episodes per month; eye fixations per trial; errors per page.
Negative binomial
Counts with variance greater than the mean
Almost all real behavioural counts — most people have few episodes, a handful have many.
Geometric
Trials until the first success
Number of sessions until first symptom-free week.
Hypergeometric
Sampling without replacement
Fisher's exact test for small contingency tables.
Overdispersion, the thing that catches people
Poisson insists variance equals the mean. Real behavioural counts are almost always more variable than that, because people differ — a phenomenon called overdispersion. Fitting Poisson to overdispersed data produces standard errors that are too small and p-values that are too impressive. Check the ratio of residual deviance to df; if it is well above 1, move to negative binomial or a mixed model (day 104).
Why this matters even at foundation level: the reflexive alternative is to take the mean count per participant and run a t-test. With counts that are mostly zeros and occasionally large, that analysis has poor power and misstates the uncertainty — and 'we log-transformed the counts' is now an answer a reviewer will push back on.
The mistake this topic produces
Analysing zero-inflated counts — where a large subgroup can only score zero, such as drinks consumed among non-drinkers — with any ordinary count model. Those data hold two processes at once, and there are zero-inflated models built for exactly this situation.
Do this now · 15 minutes
Take any count variable you have. Compute its mean and variance. If the variance is more than roughly twice the mean, write down that you have overdispersion and name the model you would fit instead. Keep the note for day 63.
Day 1745 minutes · a model, not a law
The normal distribution as a model of the world
Why does the bell curve keep appearing? Because when many small independent influences add together, their sum tends toward normality regardless of the distribution of each influence. Height is the sum of many genetic and environmental contributions; so, arguably, is a trait score built from many items. This is a mathematical consequence, not a mystical one, and it is also the reason the normal curve is a good model for measurement error.
Two properties make it so tractable. It is completely defined by μ and σ — tell me those two numbers and I can reproduce the whole distribution. And it is closed under addition: sums and means of normal variables are themselves normal, exactly, which is what makes the mathematics of t, F and regression work out in closed form.
Where psychology is emphatically not normal
Reaction times (right-skewed, always). Symptom counts in a general population (mass at zero). Income, social network size, publication counts (heavy-tailed, sometimes power-law). Likert responses (discrete, bounded, often piled at the ends). Clinical samples on clinical measures (selected to be extreme, hence truncated). Assuming normality in these cases is not a small technical slip — it changes which participants dominate your estimate.
The practical resolution is to separate three different normality questions, which textbooks often blur: is the population normal (usually unknowable and usually not), is the sample normal (checkable, and rarely essential), and is the sampling distribution of your statistic normal (what your test actually needs, and what the CLT usually delivers). Day 20 makes that third one visible.
The mistake this topic produces
Describing a distribution as 'normally distributed' on the basis of a non-significant Shapiro-Wilk test in a small sample. Failing to reject normality is not evidence of normality — especially at n = 25, where the test has almost no power to detect anything.
Do this now · 15 minutes
Plot a Q-Q plot for two of your variables. For each, describe where the points depart from the diagonal and what that departure means substantively — heavy right tail, ceiling, discreteness. A Q-Q plot read properly replaces three normality tests.
Day 1850 minutes · who is in your data
Sampling: what your data can and cannot speak for
Every inferential statistic assumes your sample was drawn randomly from the population you want to generalise to. Almost no psychology sample is. Understanding the gap between the assumption and the reality is what separates a defensible limitations section from an embarrassing one.
Method
How it works
Cost or risk
Simple random
Every member has an equal, independent chance.
The theoretical ideal; requires a complete sampling frame you rarely have.
Systematic
Every kth case from a list, random start.
Fast; fails badly if the list has a periodicity matching k.
Stratified
Random sampling within predefined strata.
More precise than simple random when strata differ; needs known strata proportions.
Cluster
Sample whole groups — schools, clinics — then all or some within.
Practical and cheap, but cases within a cluster are correlated: the effective N is smaller than the raw N (day 101).
Quota
Fill category targets non-randomly.
Looks representative on the quota variables, arbitrary on everything else.
Convenience
Whoever is available.
The actual method in most psychology. Generalisation becomes an argument, not a calculation.
Snowball
Participants recruit participants.
The only route into hidden populations; produces networked, non-independent samples.
The WEIRD problem, stated as a statistical one
Samples drawn from Western, Educated, Industrialised, Rich, Democratic populations — and especially from undergraduate participant pools — are not merely non-random. On many measured dimensions they are outliers relative to the global population. Statistically, this is a restriction-of-range and selection problem: it biases means, it attenuates correlations, and it silently narrows the population your parameter estimates refer to.
Two questions to answer explicitly in any write-up: what population would this sampling procedure legitimately support inference to, and who is systematically absent. Answering them well is worth more than any robustness check.
The mistake this topic produces
Writing 'a random sample of 120 undergraduates' when what happened was a convenience sample of whoever signed up. Random assignment to conditions — which licenses causal claims — is a different thing from random sampling from a population, which licenses generalisation. Most psychology has the first and not the second.
Do this now · 15 minutes
Write the sampling paragraph for your own study: the procedure, the population it can support inference to, and who is absent. Then write one sentence naming the direction you expect each absence to bias your estimates.
Day 1955 minutes · the central object
The sampling distribution: the idea everything rests on
Here is the move that makes inference possible. Your sample mean is one value. But imagine drawing another sample of the same size from the same population — you would get a slightly different mean. And another, and another. The distribution of all those hypothetical means is the sampling distribution of the mean, and it is what every test, interval and p-value is a statement about.
It is crucial and slippery because it is entirely imaginary. You have one sample. You never observe the sampling distribution. But its properties are derivable mathematically, and that derivation is what lets you say how far your one mean might plausibly be from μ.
Work the simulator properly, it is the day's real content. Set the population to skewed and n to 5, draw 2000 samples, and note the shape of the means. Now raise n to 30 and reset. Two things change, and they change independently: the distribution of means becomes symmetric, and it becomes narrower. The first is the central limit theorem; the second is the standard error. Confusing them costs marks constantly.
Three distributions, never to be merged
The population distribution: all the scores of everyone, shape unknown, described by μ and σ. The sample distribution: the scores you actually collected, described by M and s, and roughly resembling the population. The sampling distribution: not scores at all, but a distribution of statistics over hypothetical repeated samples, centred on μ with spread σ/√n. Exams test this three-way distinction directly, and most inferential confusion downstream is a collapse of the third into one of the first two.
The mistake this topic produces
Believing the CLT says your data become normal as n increases. It says nothing whatsoever about your data. A sample of 10,000 reaction times is exactly as skewed as a sample of 50; it is the distribution of their means over repeated samples that goes normal.
Do this now · 15 minutes
In the simulator, find the smallest n at which the bimodal population produces a sampling distribution you would be comfortable calling normal. Then do the same for the skewed population. Write both numbers down — you now have your own answer to 'how big does n need to be', which is better than the textbook's flat 30.
Day 2055 minutes · sigma over root n
Standard error and the central limit theorem
The standard deviation of the sampling distribution has its own name — the standard error — because it plays a different role from the SD of your data. The SD says how much individual people differ. The standard error says how much your estimate would differ from sample to sample. Confusing them is the most consequential notational error in applied statistics.
SE = σ / √n (estimated as s / √n)sample-to-sample variability of the mean shrinks with the square root of sample sizespread of individual scores in the population or samplesample size — under a square root, which is the whole storystandard deviation of the sampling distribution of MA depression scale: s = 12, n = 36
SE = 12 / √36 = 12 / 6 = 2.0
So a 95% CI on the mean is roughly M ± 1.96 × 2.0 = M ± 3.9
To halve that SE to 1.0 you need n = 144 — four times the data
for twice the precision. That square root governs every
sample-size decision you will ever make.
The central limit theorem is the second half. It states that as n increases, the sampling distribution of the mean approaches normality regardless of the population's shape. Together, the CLT and the SE formula tell you the shape and the spread of the sampling distribution without ever observing it — and that is what makes a t-test possible from a single sample.
How large is large enough
The textbook answer is n ≥ 30, and it is a crude one. The honest answer depends on how non-normal the population is: for mild skew, n = 15 is plenty; for strongly skewed data such as reaction times or income, n = 100 may not be; for heavy-tailed distributions with extreme outliers, convergence can be painfully slow. You verified this yourself in yesterday's simulator, which is why it came first.
One practical consequence worth internalising: because SE depends on both s and n, you can buy precision two ways. Collect more participants, or measure more precisely. Halving your measurement error has the same effect on SE as quadrupling your sample — and is often far cheaper. That is the statistical case for careful methodology.
The mistake this topic produces
Plotting error bars without saying what they are. SD bars, SE bars and 95% CI bars have very different widths and completely different meanings, and a figure with unlabelled bars cannot be interpreted. SE bars are roughly half the width of 95% CI bars, which is why they are so popular and so often misleading.
Do this now · 15 minutes
For each key variable in your data, compute the SD and the SE. Write both in a sentence describing your sample, and make explicit what each one tells the reader. Then compute what n you would need to halve the SE — and decide whether that study is feasible.
Day 2150 minutes · three derived distributions
Where t, chi-square and F come from
These three distributions are not arbitrary tables at the back of a textbook. Each arises from a specific operation on normal variables, and knowing the origin tells you what each test is really asking.
t
Normal with σ estimated
When you must estimate σ from the same small sample, the extra uncertainty fattens the tails. As df grows, t converges on z — at df = 100 they are practically identical.
Chi-square
Sum of squared normals
Add k squared z-scores and you get χ² with k df. Because it is a sum of squares, it is non-negative and right-skewed. It tests frequencies and variances.
F
Ratio of two variance estimates
Divide one chi-square by another, each over its df. Under H₀ both estimate the same error variance, so F hovers near 1. Large F means the numerator found something.
All three
Shape depends on df
Degrees of freedom are not bureaucracy — they change the distribution, and therefore the critical value, and therefore your decision.
Switch between the tabs above with the same α. Notice that the t critical value at df = 5 is far out at about 2.57 while at df = 100 it is 1.98, and that χ² and F have no left tail to reject in — they are one-tailed by construction, because a variance ratio cannot be negative. Students routinely ask why ANOVA has no two-tailed option; this picture is the answer.
The relationships worth memorising
t² = F when the numerator has 1 df — an independent-samples t-test and a two-group one-way ANOVA are literally the same test, and t = 2.5 corresponds to F = 6.25. z² = χ² with 1 df. And as df goes to infinity, t becomes z and F becomes χ²/df. Exams love these identities because they reveal whether you understand the family or merely memorised the tables.
The mistake this topic produces
Using a z critical value of 1.96 when σ is estimated from a small sample. With n = 10 the correct two-tailed critical value is 2.26, and using 1.96 makes your test more liberal than you claim — the exact error the t distribution was invented to fix.
Do this now · 15 minutes
For df = 5, 10, 30 and 100, look up the two-tailed .05 critical value of t and plot the four numbers against df. Then write in one sentence what the curve you have drawn means for planning a small-sample study.
Day 2245 minutes · not bureaucracy
Degrees of freedom, explained so it stays explained
Degrees of freedom are the number of values in a calculation that are free to vary once the constraints are imposed. The classic demonstration: if four numbers must average 10, you may choose any three freely, but the fourth is then fixed. Three degrees of freedom, one constraint.
Every parameter you estimate from the data imposes one such constraint. That is the whole rule, and it generates every df formula you will meet.
Test
df
Which constraints
One-sample t
n − 1
One mean estimated from the data.
Independent t
n₁ + n₂ − 2
One mean estimated in each group.
Paired t
n pairs − 1
One mean of the difference scores.
One-way ANOVA
between k − 1, within N − k
k group means; total df = N − 1 always splits exactly.
Factorial interaction
(a − 1)(b − 1)
Cells left free after both sets of main-effect means are fixed.
Chi-square goodness of fit
k − 1
Categories minus the fixed total.
Chi-square independence
(r − 1)(c − 1)
Row and column totals are fixed by the margins.
Pearson r significance
n − 2
Two parameters estimated for the line: slope and intercept.
Multiple regression
n − k − 1
k slopes plus an intercept.
Why it changes your answer
df determines which distribution your test statistic is compared against, and small-df distributions have fatter tails and more demanding critical values. Getting df wrong by one or two in a small study can flip a decision. In ANOVA it is also a free error check: between-df plus within-df must equal total df, and if it does not, you have miscounted a group or a case.
One conceptual payoff. Degrees of freedom are what you spend to learn things from data. Every parameter costs one; every df spent leaves less information to estimate error. That is why a model with many predictors and few participants can fit brilliantly and predict nothing — you spent all your information on the fit (day 64).
The mistake this topic produces
Reporting a test without its df. 'A significant difference was found (t = 2.31, p = .02)' is incomplete: t(18) = 2.31 and t(180) = 2.31 are very different findings, and the df is how the reader recovers your sample size and checks your arithmetic.
Do this now · 15 minutes
Take three analyses you have run or plan to run, and write out the df for each from first principles — how many observations, how many parameters estimated, what remains. Then check against software. Any mismatch is telling you something about your design you had not noticed.
Reference
The volume 2 formula sheet
The computational content of this volume, with the condition each formula depends on. Conditions are examined as often as formulas.
Quantity
Formula
Condition or caution
Addition rule
P(A or B) = P(A) + P(B) − P(A and B)
Subtract the overlap unless the events are mutually exclusive.
Multiplication rule
P(A and B) = P(A) × P(B|A)
Simplifies to P(A)P(B) only under independence.
Conditional
P(A|B) = P(A and B) / P(B)
Not symmetric. Reversing it is the conditional probability fallacy.
Bayes
P(H|E) = P(E|H)P(H) / P(E)
Work it in frequencies out of 10,000 to avoid arithmetic slips.
Binomial
P(k) = nCk p^k (1−p)^(n−k)
Independent trials, constant p.
Binomial mean / variance
μ = np, σ² = np(1 − p)
Normal approximation valid when np and n(1−p) both exceed ~10.
Poisson
P(k) = λ^k e^(−λ) / k!
Assumes variance equals the mean — check for overdispersion.
Standard error of M
SE = σ/√n, estimated s/√n
Precision improves with the square root of n, not with n.
SE of a proportion
SE = √(p(1 − p)/n)
Maximum at p = .5; narrower near 0 or 1.
z for a sample mean
z = (M − μ) / (σ/√n)
Uses σ. With s instead, it is a t statistic with n − 1 df.
Reference
The three distributions, side by side
The distinction that resolves most of the confusion in this volume — worth rereading whenever a later volume stops making sense.
Population
Sample
Sampling distribution
What it contains
All scores of everyone
The scores you collected
Statistics from infinitely many hypothetical samples
Do you observe it
Essentially never
Yes — it is your dataset
Never; it is derived mathematically
Centre
μ
M
μ (the mean of means is the population mean)
Spread
σ
s
σ/√n, the standard error
Shape
Whatever it is — often skewed
Roughly mirrors the population
Approaches normal as n grows (CLT)
What it is for
The target of inference
The evidence
The bridge between the two
Checkpoint
Nine questions before you move on
These are the questions whose absence shows up three volumes later as 'I can run it but I cannot explain it'. Answer from memory first.
{{ quizCounter }}
{{ quizScore }}
{{ quizQ }}
{{ quizFb }}
Next: volume 3, Inference
Twelve days turning the sampling distribution into decisions: confidence intervals, the logic of null hypothesis testing, what a p-value is and is not, error rates, effect sizes, power, and the honest alternatives to the ritual.