Volume 3 · Deciding under uncertainty

The ritual, dismantled — then rebuilt so you can defend every step.

Twelve days on the reasoning that every test in volumes 4 and 5 is an instance of. What an interval means, what the null hypothesis is for, what a p-value can and cannot say, what α and β cost each other, how big an effect has to be to matter, and how many participants that requires. This is the volume that decides whether you are a statistician or a button-pusher.

12
days of 40–55 min
5
p-value misreadings, named
3
simulators
1
power calculation before every study
Day 2345 minutes · estimate first

Estimation comes before testing

Inference has two jobs, and the teaching order of most curricula gets them backwards. The first job is estimation: what is the size of the thing, and how precisely do we know it? The second is testing: can we rule out a specific value, usually zero? Estimation is the more informative, and testing is a special case of it.

A point estimate is a single best guess — the sample mean for μ, the sample correlation for ρ. Good estimators are unbiased (right on average over repeated samples), efficient (small sampling variance), and consistent (converging on the parameter as n grows). The sample mean is all three, which is why it is everywhere.

But a point estimate alone is dishonest, because it conceals how much it would have moved with a different sample. An interval estimate attaches that uncertainty. Reporting M = 14.2 tells the reader the answer; reporting M = 14.2, 95% CI [11.9, 16.5] tells them how much to trust it, and whether the study was even capable of answering the question.

The question that reframes your results section
Instead of asking 'is there an effect?', ask 'how big is it, and how precisely have I pinned it down?'. That reframing does three things: it forces an effect size, it makes a wide interval visible as low information rather than a null finding, and it removes the false binary that makes a p of .049 feel categorically different from .051. Every modern reporting standard — APA 7, the Cochrane guidance, most journals' statistics policies — is pushing in this direction.
The mistake this topic produces
Treating a non-significant result as evidence of no effect when the interval is enormous. 'No difference was found, t(18) = 1.2, p = .25' with a CI running from −4 to +14 points is not a null result. It is a study that could not tell the difference between a trivial effect and a large one, which is a statement about the design, not about the world.
Do this now · 15 minutes
Take one finding you already have and rewrite the sentence three ways: as a significance statement, as a point estimate with a CI, and as an effect size with its interval. Decide which sentence a reader would learn most from — and then write your results that way from now on.
Day 2455 minutes · the procedure, not your interval

Confidence intervals, and the four things they do not mean

A 95% confidence interval is built so that, over infinitely many repetitions of the study, 95 per cent of the intervals constructed this way would contain the true parameter. The confidence attaches to the procedure. Your particular interval either contains μ or does not, and you will never know which.

CI = estimate ± (critical value × standard error)your best guess, plus and minus a margin of uncertaintyM, or a mean difference, or b, or r after transformationt with n − 1 df for a mean; 1.96 only when σ is known or n is larges/√n for a mean — the same standard error from day 20M = 24.6, s = 8.0, n = 25 SE = 8.0 / √25 = 1.60 t(24) at 95% two-tailed = 2.064 CI = 24.6 ± (2.064 × 1.60) = 24.6 ± 3.30 = [21.30, 27.90] Width is 6.6 points. Whether that is precise enough is a substantive question, not a statistical one.

Press run 500. Roughly five per cent of the intervals will be drawn in terracotta because they missed μ entirely — and notice that they look exactly like the others from the inside. That is the whole lesson: nothing about your single interval tells you whether it is one of the misses. Now drop n to 8 and watch the intervals widen and wobble.

The claimVerdictWhy
There is a 95% probability μ lies in my intervalWrongμ is a fixed constant; it is either in or out. The probability is a property of the long-run procedure.
95% of my sample lies in this intervalWrongThat would be a tolerance interval on scores. A CI is about the estimate, and it is roughly √n times narrower.
If I repeated the study, 95% of new means would fall in this intervalWrongRoughly 83% would. The prediction interval for a future mean is wider.
Any value inside the interval is equally plausibleWrongValues near the point estimate are considerably more consistent with the data than values near the limits.
Values outside the interval would be rejected at α = .05RightThe duality of intervals and tests — and the only way to read a CI as a test.
The practical reading, which is allowed
Look at both ends and ask what each would mean substantively. An intervention with a CI of [0.5, 14.0] points on a depression scale is compatible with a clinically worthless effect and with a large one — the study did not settle anything. A CI of [4.1, 5.9] is informative even if you never mention a p-value. Reading both limits is the single best habit in this volume.
The mistake this topic produces
Concluding two groups do not differ because their individual confidence intervals overlap. Overlapping CIs can still correspond to a significant difference — the test concerns the CI of the difference, which is what you should compute and plot. Non-overlap does imply significance; overlap implies nothing.
Do this now · 15 minutes
Compute a 95% CI by hand for one mean in your data, then a 90% and a 99%. Note the trade-off: more confidence, wider interval, less precision. Then compute the CI on a difference between two groups and compare it with the two separate CIs.
Day 2550 minutes · the chain of reasoning

The logic of null hypothesis testing

Null hypothesis significance testing is a proof by contradiction run on probabilities. You assume the thing you doubt, work out what data that assumption predicts, and see whether what you observed is surprising under it. Nothing more mystical is going on, and being able to state the chain in five steps answers most exam questions on the topic.

01
State H₀ and H₁
H₀ is a precise claim — usually μ₁ = μ₂ or ρ = 0 — because you can only compute predictions from a specific value. H₁ is everything else, or everything on one side.
02
Choose α before you look
The rate of false positives you will tolerate over the long run. Conventionally .05, which is a convention and not a law of nature.
03
Compute a test statistic
t, F, χ² — each one is a signal-to-noise ratio: the effect you saw, divided by how much it would vary by chance.
04
Find its position in the null distribution
The p-value: the probability of a statistic at least this extreme if H₀ were exactly true.
05
Decide, and state the effect size
p ≤ α: reject H₀. p > α: retain it — never 'accept' it. Then report the effect size and interval, because the decision alone is not a finding.
Why we never accept the null
Absence of evidence is not evidence of absence. A non-significant result can mean there is no effect, or that there is an effect and your study was too small, too noisy or too brief to see it — and the test cannot distinguish these. If you genuinely need to argue for no effect, you need an equivalence test or a Bayes factor (day 34), both of which can provide evidence for the null. The word 'accept' in a results section is a red flag to every examiner.

The deeper criticism, which you should be able to voice: H₀ as usually stated — that the effect is exactly zero to infinite decimal places — is almost never plausible in psychology. Any two groups differ by something. With a large enough n, everything is significant. This is why the field's centre of gravity has moved toward estimation, effect sizes and equivalence testing, and why you are being taught the logic rather than the ritual.

The mistake this topic produces
Writing that the results 'proved' or 'disproved' the hypothesis. NHST is probabilistic: it licenses 'the data are unlikely under H₀' and nothing stronger. Proof belongs to mathematics.
Do this now · 15 minutes
For a study of your own, write H₀ and H₁ in both words and symbols, state α, and name the test statistic you would compute. Then write the two sentences you would publish — one for a significant result and one for a non-significant one — with no overclaiming in either.
Day 2650 minutes · five misreadings

What a p-value is, and the five things it is not

A p-value is the probability of obtaining a result at least as extreme as the one observed, assuming the null hypothesis is exactly true. Every part of that sentence is doing work: 'at least as extreme' (not 'exactly this'), and the conditional 'assuming H₀ is true' (which is the assumption, not the conclusion).

MisreadingWhat it confusesThe correction
p is the probability that H₀ is trueP(data|H₀) with P(H₀|data)The conditional-probability fallacy from day 12. To get P(H₀|data) you need a prior and Bayes' theorem — day 106.
1 − p is the probability H₁ is trueThe same reversal, restatedp says nothing about the probability of any hypothesis.
p is the probability the result was due to chanceA colloquial version of the same errorThe p-value is computed assuming chance alone; it cannot then tell you how likely chance is.
A smaller p means a bigger effectSignificance with magnitudep depends on effect size AND sample size. A trivial effect at n = 5,000 gives a tiny p; a large effect at n = 12 may not reach .05.
p > .05 means there is no effectRetaining with acceptingIt means the data are not surprising under H₀ — often because the study lacked power.
p = .06 is a trend toward significanceA threshold with a continuumEither use the threshold you preregistered, or report the estimate and interval and drop the threshold language entirely.
The ASA statement, in one line each
p-values can indicate how incompatible data are with a specified model. They do not measure the probability that the hypothesis is true, nor the size or importance of an effect. Decisions should not be based on whether p crosses a threshold. Proper inference requires full reporting and transparency. And a p-value without an effect size and an interval is not a complete result. Those five sentences are quotable in any answer about the limitations of significance testing.

The constructive reading: a p-value is a rough index of surprise under a specific model. Small p means the model — which includes H₀ and all your auxiliary assumptions about independence, distribution and measurement — sits poorly with the data. Notice that the null is not the only thing that can be wrong. A tiny p can be produced by a violated independence assumption just as well as by a real effect.

The mistake this topic produces
Reporting p = .000. No p-value is ever exactly zero; software has simply rounded. Write p < .001. It is a small thing that examiners and copy editors both catch.
Do this now · 15 minutes
Find three sentences in published papers in your field that interpret a p-value. Classify each as correct or as one of the misreadings above. You will find the misreadings faster than you expect, which is itself the lesson.
Day 2750 minutes · two ways to be wrong

Type I and Type II error, and the trade-off between them

Any decision rule under uncertainty can be wrong in two directions. You can claim an effect that is not there — a Type I error, false positive, rate α. Or you can miss an effect that is there — a Type II error, false negative, rate β. Power is 1 − β: the probability of catching a real effect of a given size.

H₀ is actually trueH₀ is actually false
You reject H₀Type I error — probability α. You publish something that is not real.Correct decision — probability 1 − β, called power.
You retain H₀Correct decision — probability 1 − α.Type II error — probability β. A real effect goes unreported.

The trade-off is structural: with everything else fixed, lowering α to reduce false positives moves the decision threshold outward, which increases β and reduces power. You cannot minimise both by choosing a number. The only way to reduce both simultaneously is to get more information — more participants, less measurement error, a stronger manipulation, a within-subjects design.

Which error is worse depends on the decision, not on statistics
Screening for a treatable illness: a false negative may cost a life, so you tolerate more false positives. Approving a drug with serious side effects: a false positive harms many people, so α is set brutally low. Exploratory research intended to generate hypotheses for a larger study: false negatives close off lines of enquiry, so a liberal α may be defensible if labelled. The .05 convention is a default for none of these situations in particular — Fisher himself described it as merely convenient.

There is a third error worth knowing for vivas. A Type III error is getting the right answer to the wrong question — or, in its common form, correctly rejecting the null but misidentifying the direction or the source of the effect. A significant interaction misinterpreted as a main effect is a Type III error, and no amount of statistical rigour protects against it.

The mistake this topic produces
Believing α is the proportion of published significant findings that are false. It is not: α is the false-positive rate among tests where the null is true. The proportion of your significant findings that are false — the false discovery rate — also depends on how many of the hypotheses you test are true and on your power. In a low-power field testing mostly false hypotheses, most significant results are wrong even with α = .05.
Do this now · 15 minutes
For your own research area, write down the practical consequence of a Type I error and of a Type II error — who is harmed, and how. Then decide whether .05 is the right α for your next study, and write one sentence justifying whatever you choose.
Day 2840 minutes · direction costs credibility

One-tailed and two-tailed tests

A two-tailed test splits α between both ends of the null distribution: you will reject if the effect is surprisingly large in either direction. A one-tailed test puts the whole of α in one tail, which makes it easier to reach significance in that direction and impossible in the other.

The arithmetic consequence is direct. At α = .05, the two-tailed z critical value is 1.96; one-tailed it is 1.645. So a one-tailed test buys you power — roughly the power you would gain from adding a fair number of participants — in exchange for a commitment: if the effect runs the other way, however strongly, you must report a null result.

When one-tailed is legitimate
Three conditions, all of them: a strong theoretical or empirical basis for the direction; a decision rule that genuinely treats a reverse effect as equivalent to no effect (rare — in clinical work a treatment that harms is not the same as one that does nothing); and the direction stated before the data were seen, ideally in a preregistration. Absent all three, use two-tailed. Choosing one-tailed after seeing a p of .08 is p-hacking with extra steps, and reviewers can spot it from the critical value.

Practical note for reading the literature: many older papers report one-tailed tests without saying so. If a paper's t and df imply a two-tailed p of .09 but reports p < .05, that is what happened. Being able to reconstruct a p from t and df is a genuinely useful reviewing skill.

The mistake this topic produces
Running two-tailed, finding p = .07, then reporting the one-tailed p = .035 as if the hypothesis had been directional all along. This is one of the most common forms of undisclosed flexibility, and it inflates your true Type I rate well beyond the stated α.
Do this now · 15 minutes
Look at your own hypotheses. For each, decide honestly whether you would report a strong effect in the opposite direction as a null finding. If you would not — and you usually would not — the test is two-tailed. Write that down now, before you have data.
Day 2955 minutes · how big, not whether

Effect sizes: the number that actually answers the question

A test tells you whether an effect is distinguishable from zero. An effect size tells you how big it is — which is the question anyone outside statistics was actually asking. Every journal that matters now requires one, every meta-analysis is built from them, and every power calculation needs one as input.

d = (M₁ − M₂) / s_pooledthe gap between two means, measured in standard deviationsthe raw difference, in the units you measured√[((n₁−1)s₁² + (n₂−1)s₂²) / (n₁+n₂−2)] — the common within-group SDunitless, so comparable across measures and studiesd corrected for small-sample bias; use it below about n = 20 per groupTreatment M = 18.2, SD = 6.1, n = 30 Control M = 22.9, SD = 6.5, n = 30 s_pooled = √[((29 × 37.21) + (29 × 42.25)) / 58] = 6.30 d = (18.2 − 22.9) / 6.30 = −0.75 A medium-to-large effect: the groups differ by three quarters of a standard deviation. Roughly 77% of treated participants score below the control mean.
Effect sizeUsed withSmall / medium / largeNote
Cohen's d, Hedges' gTwo means.20 / .50 / .80Benchmarks are conventions, not laws; compare to your own literature.
Pearson rTwo continuous variables.10 / .30 / .50r² is the shared variance and is usually the more sobering number.
η² and partial η²ANOVA.01 / .06 / .14Partial η² is inflated in factorial designs and cannot be compared across studies with different designs.
ω²ANOVAas η²Less biased than η²; preferred when available.
Cohen's fANOVA power.10 / .25 / .40The input G*Power wants; f = √(η²/(1 − η²)).
Odds ratio, risk ratioBinary outcomescontext-specificOR overstates risk when the outcome is common; report absolute risk too.
Cramér's VChi-square.10 / .30 / .50 (df-dependent)Phi for 2 × 2 tables.
R²Regression.02 / .13 / .26In-sample R² is optimistic; adjusted or cross-validated is honest.
The benchmarks are a last resort
Cohen offered his small/medium/large values as rough guidance in the absence of anything better, and explicitly regretted how mechanically they were adopted. A d of .20 for a cheap public-health intervention delivered to millions is enormous. A d of .80 for an intensive year-long therapy compared to waiting list may be unimpressive. Always give the raw effect as well — points on the scale, milliseconds, percentage of patients recovered — because that is the number a clinician or policymaker can act on.
The mistake this topic produces
Reporting partial η² from a factorial design and comparing it with an η² from a different study. Partial η² removes other factors' variance from the denominator, so it grows as you add factors — it is not comparable across designs, and inflating it accidentally is common.
Do this now · 15 minutes
Compute Cohen's d and its 95% CI for a comparison you have made. Then translate it into plain English two ways: as a percentage overlap between groups, and as the raw difference in the units of your measure. Decide which sentence you would put in an abstract.
Day 3055 minutes · before, not after

Power analysis: deciding n before you collect anything

Power is the probability of detecting an effect of a specified size if it truly exists. It is determined by four quantities, and fixing any three determines the fourth: effect size, sample size, α, and power itself. That is the whole of power analysis, and the simulator below is the whole of it made visible.

Set d = .50 and n = 30 per group: power lands near 48 per cent — a coin flip, for one of the most common designs in psychology. Set d = .30, the more realistic value for many social and clinical effects, and power collapses toward 20 per cent. Then switch to the power curve and look at the n you would need. That gap between typical practice and adequate design is the mechanical explanation for the replication crisis.

Kind of power analysisWhat it fixesVerdict
A prioriEffect size, α, desired power → gives required nThe only one that should drive a design. Do it before recruiting.
Sensitivityn, α, power → gives the smallest detectable effectExcellent for an existing dataset or a fixed clinical sample: 'this study could only detect d ≥ .62'.
CompromiseRatio of β to α, n, effect → balances the two errorsUseful where a Type II error is genuinely as costly as a Type I.
Post hoc / observedComputes power from the effect you actually gotUninformative and discouraged. It is a monotone function of your p-value and tells you nothing new.
Where the effect size for the calculation comes from
Four defensible sources, in descending order: a meta-analysis in your area; a well-powered previous study using the same measures; a pilot study, with the caveat that pilot effect sizes are noisy and usually overestimates; or a smallest effect size of interest — the smallest effect that would change practice or theory, which is often the most honest input. Never use the effect size from a single small significant study: published small studies overestimate effects precisely because they had to be large to get published.

One structural point about underpowered research that is worth carrying into every design meeting. Low power does not merely mean you might miss things. It also means that any effect you do detect must have been unusually large in your sample to clear the threshold — so the published estimate is inflated, sometimes by a factor of two. Low power corrupts the literature in both directions at once.

The mistake this topic produces
Running a post hoc power analysis after a null result to argue the study was adequately powered. Observed power is computed from the observed effect, so a non-significant result always yields low observed power. It is circular, and reviewers who know this will say so. Report a sensitivity analysis instead.
Do this now · 15 minutes
Run an a priori power analysis for your next study using the simulator or G*Power. Write down the four numbers — effect size and where it came from, α, target power, resulting n. If the required n is infeasible, write the sentence that says so honestly and name what you will change: the design, the measure, or the question.
Day 3150 minutes · the more tests, the more luck

Multiple comparisons and the inflation of error

Run one test at α = .05 and the chance of a false positive is 5 per cent. Run twenty independent tests with no real effects anywhere and the chance that at least one comes out significant is 1 − .95²⁰ = 64 per cent. Nothing has gone wrong in any individual test; the problem is the family of tests taken together.

FWER = 1 − (1 − α)^kthe probability of at least one false positive across k independent teststhe per-comparison error rate, usually .05number of tests in the familyprobability of at least one Type I error somewhere in the familyk = 3 → 1 − .95³ = .143 k = 10 → 1 − .95¹⁰ = .401 k = 20 → 1 − .95²⁰ = .642 k = 100 → 1 − .95¹⁰⁰ = .994 Bonferroni protects by using α/k per test: with k = 10, each test is judged at .005.
CorrectionHowWhen to use it
BonferroniDivide α by the number of tests.Simple, exact, conservative. Fine for a handful of planned comparisons; brutal for many.
Holm-BonferroniSequential: compare the smallest p to α/k, next to α/(k−1), and so on.Uniformly more powerful than Bonferroni with the same protection. There is no good reason to prefer plain Bonferroni.
Tukey HSDExact control across all pairwise comparisons after ANOVA.The default for all-pairs post hoc comparisons with equal n.
Šidák1 − (1 − α)^(1/k).Slightly less conservative than Bonferroni when tests are independent.
False discovery rate (BH)Controls the expected proportion of false positives among rejected tests.The standard for large-scale testing — imaging, genetics, item-level analyses. Far more powerful than FWER methods at k in the hundreds.
No correctionReport everything, label it exploratory.Defensible when the analyses are genuinely exploratory and reported as such — undeclared is what is indefensible.
When correction is not required
A small number of preregistered, theoretically motivated comparisons — planned contrasts — do not require the same protection as an omnibus fishing expedition, and imposing Bonferroni on three hypotheses stated in advance simply wastes power. The real question is not 'how many tests did I run' but 'how much undisclosed flexibility was there'. Preregistration answers that question better than any correction.
The mistake this topic produces
Correcting for the tests you report while ignoring the ones you ran and dropped. The family is defined by the analyses you conducted, not the ones that survived. This is why preregistering the analysis plan is the only genuine solution to multiplicity.
Do this now · 15 minutes
Count the number of statistical tests in the last paper you read in your field — including every correlation in a matrix and every post hoc. Compute the familywise error rate at α = .05 with no correction. Then check what the paper did about it.
Day 3250 minutes · ranked by consequence

Assumptions, ranked by how much the violation actually costs

Every test makes assumptions, and textbooks list them as equally important. They are not. Ranked by the damage a violation does, the order is roughly: independence, then correct model specification, then homogeneity of variance with unequal n, then normality. Most students spend their anxiety in exactly the reverse order.

01
Independence of observations — catastrophic
Two responses from the same person, pupils within a classroom, trials within a participant. Violating this inflates Type I error dramatically and no robustness saves you. The fix is a design-appropriate model: repeated-measures, or a mixed model (volume 9).
02
Correct specification — severe
A curved relationship fitted with a straight line, an omitted confounder, an ignored interaction. The parameters answer a question you did not ask. Plot your residuals.
03
Homogeneity of variance — moderate, conditional
With equal group sizes, ANOVA and t are fairly robust. With unequal n and unequal variances, error rates go badly wrong — which is why Welch's t is now the recommended default (day 36).
04
Normality of residuals — mild at moderate n
The CLT protects means. It matters most in small samples, and for tests of variances rather than means. Check with a Q-Q plot, not a significance test.
05
Linearity and additivity — depends entirely on the data
Check with a scatter or partial residual plot. In psychology, genuinely non-linear relationships — Yerkes-Dodson, dose-response, practice curves — are common and often theoretically interesting.
The order to check them in
Independence first, and it is a question about your design, not your data — you answer it by describing how the observations were collected. Then plot the residuals: one plot of residuals against fitted values catches non-linearity and heteroscedasticity at once, and a Q-Q plot of residuals catches non-normality. Two plots settle most assumption checking. Levene's, Shapiro-Wilk and their relatives are supplements, and are unreliable at exactly the sample sizes where they matter.
The mistake this topic produces
Testing assumptions with hypothesis tests and treating a non-significant result as clearance. At small n these tests have no power to detect real violations; at large n they flag trivial ones. In both cases the significance of an assumption test is a poor guide to whether your analysis is in trouble.
Do this now · 15 minutes
For your main planned analysis, list its assumptions, rank them by consequence as above, and write next to each how you will check it and what you will do if it fails. Doing this before you collect data is preregistration; doing it afterwards is when the temptation starts.
Day 3345 minutes · no villains required

p-hacking, forking paths, and how honest people produce false findings

Researcher degrees of freedom are the many small, defensible-looking decisions made during analysis: which outliers to exclude, whether to transform, which covariates to include, when to stop collecting, which of three measures to treat as primary. Each choice is arguable. Made after seeing the data, in the direction that helps, they collectively drive the false-positive rate far above .05 — simulations put it above 60 per cent with only a handful of such freedoms.

Optional stopping
Testing as you go
Checking p after every ten participants and stopping when it dips below .05 gives a false-positive rate near 25 per cent. Sequential designs with corrected boundaries make this legitimate — ad hoc peeking does not.
Outlier flexibility
Choosing the rule afterwards
Trying 2 SD, then 2.5, then 3, and reporting the one that worked. Set the rule in advance and report sensitivity to it.
Covariate shopping
Adding controls until it works
Each covariate is a new analysis. Report the model you planned; anything else is exploratory.
HARKing
Hypothesising after results are known
Presenting a post hoc explanation as an a priori prediction. It makes an exploratory finding look confirmatory and removes the reader's ability to calibrate.
Selective outcome reporting
Five measures, one reported
The family for multiplicity is everything you measured, not everything you printed.
Garden of forking paths
No conscious cheating at all
Even a researcher who runs exactly one analysis has often chosen it in light of the data. The flexibility does not need to be exercised to have inflated error — it only needs to have been available.
The fixes, in order of effectiveness
Preregistration: state hypotheses, sample size, exclusions and the analysis plan before collecting, with a timestamp. Registered reports: peer review of the plan before results exist, so publication does not depend on the outcome. Multiverse or specification-curve analysis: run all defensible analyses and show the distribution of results. Open data and code, so others can run the paths you did not. And in your own write-up, one honest sentence: which analyses were planned and which were exploratory. That sentence costs nothing and is worth more than most robustness checks.
The mistake this topic produces
Believing that because you did not intend to bias anything, you did not. The entire literature on this shows that motivated flexibility operates without any conscious dishonesty — which is why procedural solutions such as preregistration work better than resolutions to be careful.
Do this now · 15 minutes
Take an analysis you have already run and list every decision point where you could reasonably have chosen otherwise. Count them, and compute 2 to that power — the number of analyses in your garden. Then rerun the analysis under three of those alternative paths and see whether your conclusion survives.
Day 3450 minutes · the honest alternatives

Beyond the p-value: estimation, equivalence, and Bayes

Nothing in the previous eleven days argues against inference. It argues against a single dichotomous decision standing in for a result. Three alternatives are now standard enough to be expected in good work, and all three are examinable as 'current developments'.

Estimation-first reporting
The new-statistics approach
Report the effect size with its confidence interval as the primary result; p becomes a footnote. Meta-analysis then becomes possible because your paper contributes an estimate rather than a verdict.
Equivalence testing (TOST)
Evidence for the absence of an effect
Define the smallest effect of interest, then run two one-sided tests against those bounds. A significant TOST licenses the claim that the effect is smaller than that — which a non-significant t-test never can.
Bayes factors
Relative evidence between models
BF₁₀ quantifies how much more likely the data are under H₁ than under H₀. Unlike p, it can support the null, and it permits honest sequential testing. Volume 9, days 105–106.
Multiverse analysis
All defensible paths at once
Report the distribution of results across every reasonable analytic choice. Robust findings survive most paths; fragile ones show it immediately.
What a reviewer or examiner wants to see in 2026
An a priori justification for the sample size. Effect sizes with intervals for every reported test. A statement distinguishing confirmatory from exploratory analyses. Assumptions checked and reported, with sensitivity analyses where a decision was arguable. Data and analysis code available unless there is an ethical reason not to. None of that is exotic any more — it is the baseline, and a thesis that has it is markedly easier to defend.

You now have the reasoning that the rest of the course applies. Volumes 4 and 5 are, in a real sense, the same twelve days repeated with different test statistics: state the model, check what it assumes, compute the signal-to-noise ratio, locate it in the appropriate null distribution, report the estimate and its uncertainty, and name the mistake you avoided.

The mistake this topic produces
Adopting the language of the new statistics while keeping the old logic — reporting a CI and then interpreting it purely by whether it excludes zero. That is a p-value in different clothing. Read both limits and say what each would mean substantively.
Do this now · 15 minutes
Rewrite the results paragraph of something you have written, estimation-first: effect size and CI leading, p in parentheses, one sentence on what the interval's lower and upper limits would each mean in practice. Compare the two versions and notice which one a non-statistician can act on.
Reference

The volume 3 reference sheet

The quantities and definitions from this volume that are examined directly, with the precise wording that earns the mark.

TermPrecise statementCommon wrong answer
p-valueP(a result at least this extreme | H₀ true)The probability that H₀ is true.
αThe long-run Type I error rate you accept, chosen before analysisThe probability the result is due to chance.
βThe probability of failing to reject a false H₀The standardised regression coefficient (also β — context decides).
Power1 − β: the probability of detecting a true effect of a stated sizeA property of a study alone, without reference to an effect size.
Confidence intervalA range from a procedure that captures the parameter in 95% of repetitionsA 95% probability that μ is inside this interval.
Effect sizeA standardised or raw measure of magnitude, independent of nSomething a small p-value implies.
Type I errorRejecting a true nullAny incorrect conclusion.
Type II errorRetaining a false nullA mistake in the calculation.
FWER1 − (1 − α)^k across k independent testsAlways requires Bonferroni correction.
Equivalence testTwo one-sided tests against bounds of the smallest effect of interestA non-significant t-test.
Reference

Reporting templates you can paste

APA-conformant sentences with the slots marked. Fill the numbers, keep the structure — these are the patterns that stop a results section from being rewritten by a reviewer.

ResultTemplate
Independent tParticipants in the treatment group scored lower (M = 18.2, SD = 6.1) than controls (M = 22.9, SD = 6.5), t(58) = 2.89, p = .005, d = 0.75, 95% CI [0.22, 1.27].
Non-significant, with sensitivityThe groups did not differ reliably, t(38) = 1.12, p = .27, d = 0.35, 95% CI [−0.28, 0.98]. With n = 20 per group the study had 80% power to detect d ≥ 0.91, so smaller effects cannot be ruled out.
Confidence interval firstThe intervention reduced symptoms by 4.7 points, 95% CI [1.5, 7.9], on a scale where 3 points is the accepted minimal clinically important difference.
EquivalenceThe difference fell within the equivalence bounds of ±0.4 SD, TOST p = .018, supporting practical equivalence.
Power statement in the methodA priori power analysis (G*Power 3.1) indicated that 64 participants per group were required to detect d = 0.50 with 80% power at α = .05, two-tailed.
Confirmatory vs exploratoryHypotheses 1 and 2 and the analyses reported in Table 2 were preregistered (osf.io/xxxxx). The moderation analyses in Table 3 were exploratory and are reported as such.
Checkpoint

Ten questions before you move on

Several of these are the exact wordings that appear in NET papers and in viva questions. If you score below seven, reread before volume 4 — everything there assumes this reasoning is automatic.

{{ quizCounter }}
{{ quizScore }}
{{ quizQ }}
{{ quizFb }}

Next: volume 4, Comparing groups

Sixteen days on the t and F family, worked end to end: every t-test, the full ANOVA source table by hand, post hoc tests and contrasts, factorial designs and interactions, repeated measures and sphericity, ANCOVA, MANOVA, chi-square and the rank-based alternatives.

Start day 35 →