Before you can infer anything, you have to describe it honestly.
Eleven days on the part everyone treats as revision and then gets wrong under exam conditions: what a mean actually assumes, why variance divides by n minus one, what a z-score buys you, and what a normal curve does and does not promise. Nothing here is inferential. Every single thing here is load-bearing for the ninety-eight days that follow.
11
days of 45–50 min
6
formulas worked by hand
4
scales of measurement
1
results table you can defend
Day 145 minutes · no formulas yet
What statistics is actually for
Statistics exists because of one gap: you can never measure everyone you want to talk about. You want to say something about adolescents with social anxiety, or about working memory in healthy adults, or about how a drug affects rats. You measure forty of them. Every technique in the next 111 days is machinery for crossing that gap responsibly — and for stating honestly how wide the gap still is.
The group you want to conclude about is the population. The group you actually measured is the sample. A number describing a population is a parameter — μ, σ, ρ, π, written in Greek. A number computed from your sample is a statistic — M, s, r, p, written in Latin. You almost never see a parameter. You compute statistics and use them to estimate parameters, and the entire discipline is the study of how wrong that estimate is likely to be.
Descriptive
Summarising what you have
Mean, SD, correlation, a histogram. Claims stop at the edge of your dataset. No probability is involved and nothing can be wrong in the inferential sense.
Inferential
Reasoning beyond what you have
Tests, intervals, models. Every one carries a probability statement about how sample results relate to a population you did not measure.
Exploratory
Finding what might be there
Legitimate and necessary — as long as it is labelled. Exploration that reports itself as confirmation is the single biggest reason published findings fail to replicate.
Predictive
Forecasting a new case
Judged on out-of-sample accuracy, not p-values. A different goal with a different scorecard; volume 5 draws the line properly.
The distinction exams test constantly
Parameter vs statistic, and population vs sample, are examined as vocabulary and then punished as reasoning. When a question says the mean IQ of the sample was 104, that is a statistic — 104 is known exactly and has no standard error. When it says the population mean is 100, that is a parameter, assumed or defined. Standard error, confidence intervals and p-values only exist because statistics vary from sample to sample while parameters sit still.
One more distinction to fix now, because it decides which half of this course you are in. A variable is anything that takes different values across cases: depression score, condition, reaction time, handedness. A constant does not vary, and cannot correlate with anything. If your manipulation produced no variability — everyone scored the ceiling on your task — no statistical technique in existence can rescue the study. Most 'my analysis failed' problems are design problems arriving late.
The mistake this topic produces
Writing 'the sample mean is 104, therefore the population mean is 104'. The sample mean is an unbiased estimate of μ, which means it is right on average across infinitely many samples — not right in yours. Saying the second sentence in a viva invites twenty minutes of questions you will not enjoy.
Do this now · 15 minutes
Write down, in one sentence each, the population, the sample, one parameter you care about and the statistic you would compute to estimate it — for a study you have actually run or want to run. Then write what would make your sample a bad stand-in for that population. Keep this page; you will reuse it on day 18.
Day 245 minutes · nominal to ratio
Scales of measurement, and what each one permits
Stevens' four scales are the most reliably examined content in any psychology statistics paper, and the most commonly reduced to memorised labels. The point of the taxonomy is not classification. It is permission: the scale of a variable determines which arithmetic operations mean anything, and therefore which statistics are legitimate.
Scale
What the numbers do
Legitimate
Psychology examples
Nominal
Name categories. Numbers are labels only; order is meaningless.
Socio-economic class, rank in a race, stage of illness, a single Likert item (arguably).
Interval
Equal intervals, arbitrary zero. Differences are meaningful, ratios are not.
Mean, SD, Pearson r, t, ANOVA, regression.
Temperature in °C, calendar year, IQ and most test scores by convention.
Ratio
Equal intervals and a true zero. Ratios are meaningful.
Everything above, plus the coefficient of variation and ratio statements.
Reaction time, number of errors, income, time on task, cortisol concentration.
The interval-vs-ratio line matters less in practice than the textbook implies — both permit the same tests. The nominal-vs-ordinal-vs-interval boundary matters enormously. An IQ of 130 is not twice as intelligent as 65, because IQ has no true zero; that is an interval-scale statement and it is simply meaningless. But the mean of IQ scores is fine, because intervals are treated as equal.
The Likert argument, settled enough to answer
A single Likert item is ordinal: nobody can defend the claim that the psychological distance from strongly disagree to disagree equals that from neutral to agree. A Likert scale — the sum or mean of many such items — is treated as interval, and simulation work over four decades shows parametric tests on such sums behave well with five or more response options and roughly symmetric distributions. In an exam: item ordinal, summed scale interval. In your thesis: say which you did and why.
Two further distinctions live alongside the scales and are examined with them. Discrete variables take separate values with nothing in between (number of children, errors committed); continuous variables can in principle take any value in a range (time, weight). And a variable's role in a design — independent, dependent, confounding, mediating, moderating — is a separate question from its scale. Reaction time is ratio whether it is your outcome or your predictor.
The mistake this topic produces
Computing a mean of a nominal variable. It happens invisibly: code male = 1, female = 2, and software will happily report a 'mean gender of 1.4'. It is meaningless — though note that a mean of a 0/1 dummy variable IS meaningful: it is the proportion of ones, which is why dummy coding works in regression. Know which of the two you are doing.
Do this now · 15 minutes
List every variable in a dataset you have, and classify each as nominal, ordinal, interval or ratio, and as discrete or continuous. Flag any where you are unsure — those are exactly the ones where your analysis choice will later be questioned. For each, write the one descriptive statistic that is legitimate.
Day 345 minutes · look before you compute
Frequency distributions: looking before you compute
A frequency distribution is just a count of how often each value occurs. It is also the last point at which your data are still fully visible. Every statistic after this is a compression — the mean throws away everything except the centre — and the compression can only be trusted if you have seen what is being compressed.
For a small number of categories, a simple frequency table does it: value, frequency f, relative frequency f/N, cumulative percentage. For continuous data you group into class intervals, and the number of intervals is a real decision: too few and you hide structure, too many and you show noise. Ten to twenty bins is the usual working range, and the honest move is to try three bin widths and report whether your conclusion changes.
Histogram
Continuous data, one variable
Bars touch, because the underlying scale is continuous. The shape — one hump, two humps, a long tail — decides half your later analysis choices.
Bar chart
Categorical data
Bars separated, because categories are not adjacent points on a scale. Never use one to show group means without also showing the spread.
Stem-and-leaf
Small samples, by hand
Keeps the actual digits visible while showing shape. Exam favourite because it can be read and reconstructed without software.
Boxplot
Comparing groups
Median, quartiles, whiskers, outliers. Best single plot for 'do these groups differ and how much do they overlap'.
What you are looking for, in order
One: how many modes — a second hump usually means a mixed population you have not accounted for. Two: skew — a long right tail is the default for reaction times, income, symptom counts, and it changes which centre is honest. Three: gaps and impossible values — a 999 age, a negative latency, a straight line of identical responses. Four: the range against the possible range — if your depression scale runs 0–63 and nobody scored above 12, you have a floor effect, and no test will find an effect where there is no variance.
This is not a preliminary. It is the analysis step where the most consequential decisions are made, and it is the step most often skipped in favour of running the test that was planned before the data existed.
The mistake this topic produces
Reporting a mean and SD for a bimodal or heavily skewed distribution as if they described it. A mean of 50 from a distribution with humps at 20 and 80 describes nobody in the dataset. Plot first, and if the plot and the summary disagree, the plot is right.
Do this now · 15 minutes
Take any continuous variable you have and make three histograms with different bin widths — narrow, medium, wide. Write one sentence on what each reveals or hides. Then make a boxplot by group, and note whether any conclusion you already believed survives.
Day 445 minutes · three centres
Central tendency, and when each measure lies
Three ways to say where a distribution sits, each answering a subtly different question. The mean is the balance point: the value that makes the deviations sum to exactly zero. The median is the positional middle: half above, half below. The mode is the most frequent value, and the only centre available for nominal data.
M = ΣX / nadd everything up, divide by how many things you addedsum of all scores — capital sigma means 'add these'number of scores in the samplethe sample mean; μ is its population counterpartScores: 4, 7, 7, 9, 13
ΣX = 4 + 7 + 7 + 9 + 13 = 40
n = 5
M = 40 / 5 = 8.0
Check: deviations from the mean are −4, −1, −1, +1, +5
They sum to 0 — always. That property IS the mean.
The mean uses every value, which is its strength and its weakness: it is the most efficient estimator when the distribution is roughly symmetric, and it is dragged bodily toward any extreme score. Add one participant earning ten million to a sample of students and the mean income becomes a number no student has. The median would not move at all.
Situation
Use
Why
Symmetric, continuous data
Mean
Uses all information; feeds directly into t, ANOVA, regression.
Skewed data — income, RT, symptom counts
Median
Resistant to the tail; describes the typical case rather than the balance point.
Ordinal data
Median
Requires only order, which is all ordinal data provide.
Nominal data
Mode
The only one defined; 'average diagnosis' is not a thing.
Open-ended categories — '65 and over'
Median
The mean cannot be computed without inventing a value for the top category.
Bimodal data
Report both modes
Any single centre misrepresents a two-population mixture.
The relationship exams ask you to draw
In a perfectly symmetric unimodal distribution, mean = median = mode. In a positively skewed distribution (long right tail) the order is mode < median < mean — the mean is pulled toward the tail. In a negatively skewed distribution it reverses: mean < median < mode. The mean always sits nearest the tail. If you can reconstruct that rule you can answer every skew question on a NET paper, including the ones that give you two of the three and ask for the direction.
The mistake this topic produces
Choosing the median because the data 'look a bit skewed' and then running a t-test on the means anyway. The descriptive you report and the statistic your test analyses should be the same quantity, or your results section contradicts itself.
Do this now · 15 minutes
Compute the mean, median and mode of a variable of yours by hand, then again after removing the single most extreme value. Note how far each one moved. That distance is exactly what 'resistant' means, and you now have the number to quote.
Day 550 minutes · the most important descriptive
Dispersion, and why you divide by n − 1
Two groups can have identical means and be completely different data. Centre without spread is a half-description, and spread is the more consequential half: every inferential test you will ever run is essentially a signal-to-noise ratio, where the noise term is a measure of dispersion. Improving measurement precision is a statistical intervention, not just a methodological nicety.
The range (max − min) uses two values and is hostage to both. The interquartile range (Q3 − Q1) covers the middle fifty per cent and is resistant to outliers, which is why it partners the median. The variance and its square root the standard deviation use every value and are the basis of nearly all parametric statistics.
s² = Σ(X − M)² / (n − 1) s = √s²average squared distance from the mean — then square-root it back into real unitsa deviation: how far one score sits from the meanthe sum of squares, SS — the single most reused quantity in statisticsdegrees of freedom: one is spent estimating M from the same datastandard deviation, in the original units of measurementScores: 4, 7, 7, 9, 13 M = 8
X X − M (X − M)²
4 −4 16
7 −1 1
7 −1 1
9 +1 1
13 +5 25
----
SS = 44
s² = 44 / (5 − 1) = 11.0
s = √11.0 = 3.32
Population version would divide by N = 5: σ² = 8.8 — smaller, and biased low.
Why n − 1, asked properly: your deviations are taken from M, not from μ, and M is the value that makes the sum of squared deviations as small as it can possibly be for this sample. So SS computed around M systematically underestimates SS around the true μ. Dividing by the smaller number n − 1 inflates the estimate by exactly the right amount to remove that bias. The same logic — subtract one degree of freedom per parameter estimated from the data — reappears in every df you will ever compute.
Read the SD, do not just report it
An SD of 3.32 on a 0–10 scale is enormous spread. The same 3.32 on an IQ-type scale is almost no spread at all. Always read the SD relative to the scale and to published norms — and when your SD is dramatically smaller than the literature's, suspect restriction of range in your sample, which will attenuate every correlation you compute (day 52).
The mistake this topic produces
Forgetting to square-root back. Reporting variance where the sentence demands SD — 'scores averaged 24.1 (SD = 42.3)' on a 30-point scale — is an immediate signal to a reader that the results section was not checked. Variance is for computation; SD is for communication.
Do this now · 15 minutes
Compute SS, s² and s by hand for five to eight of your own scores, laying out the four-column table as above. Then get software to do the same and confirm you match. If you do not match, you will find you used n instead of n − 1 — which is precisely the error worth making once, now.
Day 645 minutes · skew, kurtosis, outliers
Shape: skewness, kurtosis, and what to do about outliers
Centre and spread describe a distribution only if you already know its shape. Two more numbers finish the job. Skewness measures asymmetry: zero for symmetric, positive when the long tail runs right, negative when it runs left. Kurtosis measures tail weight relative to a normal curve: positive (leptokurtic) means heavy tails and a sharp peak, negative (platykurtic) means light tails and a flat middle.
Rules of thumb that get you through both exams and reviewers: skewness and kurtosis within ±1 is unremarkable, ±2 is worth mentioning, beyond that you address it explicitly. A more useful check is the ratio of the statistic to its standard error — beyond about ±3 in samples under 300 is a real departure. But in samples above a few hundred, tiny meaningless departures become 'significant', which is why the Kolmogorov-Smirnov and Shapiro-Wilk normality tests are nearly useless at the sample sizes where normality matters least.
Positively skewed
The psychology default
Reaction times, symptom counts, income, latencies to respond. Bounded at zero, unbounded above. Expect it; do not treat it as an error.
Negatively skewed
Ceiling effects
Easy tasks, well-being in healthy samples, satisfaction ratings. Often a sign that your measure lacks headroom rather than that the world is skewed.
Leptokurtic
Heavy tails
Outliers are more common than a normal model predicts. Means become unstable; trimmed means and robust methods earn their keep (day 110).
Platykurtic
Flat, light tails
Sometimes a uniform-ish process; sometimes a coarse measure with too few response options.
An outlier protocol you can defend
One: define the rule before you look — the standard options are beyond 1.5 × IQR from the quartiles, or beyond |z| = 3.29, or the median absolute deviation equivalent. Two: check whether it is an error (impossible value, equipment failure, participant misunderstood) or an extreme but real observation. Errors are corrected or removed with a stated reason. Real extremes stay. Three: if you do remove real data, report the analysis both ways. A finding that exists only after removing three people is not a finding; it is a sentence in your limitations section.
The deeper point: outliers are information. A participant whose reaction times are four SDs slow may be the only person in your sample who was actually doing the task the way you designed it. In clinical samples, the extreme cases are frequently the ones the research is for.
The mistake this topic produces
Deleting outliers until the p-value cooperates, then reporting only the final analysis. This is p-hacking whether or not you intended it as such (day 33), and because the decision is made after seeing results it cannot be undone by describing it honestly afterwards. Preregister the rule, or report both.
Do this now · 15 minutes
Find the skewness and kurtosis of one of your variables, plot it, and apply an outlier rule you write down before looking. Report how many cases the rule flags and whether each is an error or a real extreme. Then decide — in writing, with a reason — what you will do with them.
Day 745 minutes · the universal translator
z-scores: putting everything on one scale
A raw score means nothing on its own. Is 47 good? It depends entirely on the mean and the spread of the distribution it came from. A z-score answers that question by re-expressing the score as a distance from the mean, measured in standard deviations — which makes scores from completely different tests directly comparable.
z = (X − M) / show many standard deviations above or below the mean this score sitsthe raw score you are convertingthe mean of its distribution (μ if you know the population)the standard deviation (σ for a population)the standardised score: signed, unitless, mean 0, SD 1Anxiety score X = 47, sample M = 38, s = 6
z = (47 − 38) / 6 = 9 / 6 = +1.50
This person sits one and a half SDs above the mean.
If the distribution is normal, that is the 93.3rd percentile.
Going backwards, which exams love:
X = M + z·s = 38 + (−2.00 × 6) = 26 is the score at z = −2.
Standardising a whole variable — converting every score to z — leaves the shape of the distribution completely unchanged. This is the most commonly missed point on the topic. Standardising does not normalise. It is a linear transformation: it slides the distribution so the mean is 0 and stretches it so the SD is 1. A skewed distribution of z-scores is exactly as skewed as it was before.
Where z-scores reappear in this course
As standardised regression coefficients (β) in volume 5; as the test statistic in a z-test in volume 4; as the metric for effect sizes such as Cohen's d, which is literally a difference between means in SD units; as the basis of all standard scores — T, IQ, stanine — on day 9; and as the way outliers are defined on day 6. Learning z properly is the highest-leverage forty-five minutes in the volume.
The mistake this topic produces
Treating z-scores as percentiles without the normality assumption. z = 1.5 corresponds to the 93rd percentile only if the distribution is normal. In a skewed distribution the same z can sit at a very different percentile, and reporting it as one is a genuine error, not a rounding issue.
Do this now · 15 minutes
Convert five of your own scores to z by hand. Then take a score from one measure and a score from another with a different scale, standardise both, and state which participant is more extreme on their respective measure. That comparison is impossible with raw scores and trivial with z.
Day 850 minutes · areas under the curve
The normal curve, and what it actually promises
The normal distribution is a mathematical idealisation — symmetric, unimodal, defined entirely by μ and σ, with tails that approach but never touch zero. No real psychological variable is exactly normal. Its importance is not descriptive. It matters because of what happens to sample means when you draw repeatedly from any population, which is the central limit theorem and the business of volume 2.
What the curve gives you immediately is the mapping from a z-score to a proportion of the distribution. The empirical rule — 68 per cent of cases within ±1 SD, 95 per cent within ±1.96, 99.7 per cent within ±3 — is worth committing to memory because it turns any mean and SD into a statement about where people fall.
Drag the observed statistic above to z = 1.96 with a two-tailed α of .05 and watch the blue area exactly fill the terracotta one: that is the definition of the critical value, and the whole of null hypothesis testing is this picture with different distributions. Switch to the t tab and shrink df to 5 — the tails thicken, and the critical value moves out past 2.5. That is the price of estimating σ from a small sample, made visible.
z
Area below (percentile)
Area in both tails
What it is used for
±1.00
84.13% / 15.87%
31.7%
The everyday 'one SD out' benchmark.
±1.645
95% / 5%
10%
One-tailed .05 critical value.
±1.96
97.5% / 2.5%
5%
The two-tailed .05 critical value; the 1.96 in every 95% CI.
±2.58
99.5% / 0.5%
1%
Two-tailed .01; the 99% confidence interval.
±3.29
99.95%
0.1%
A common outlier cutoff, and two-tailed .001.
The assumption, stated precisely
Parametric tests do not require your raw data to be normal. They require the sampling distribution of the statistic to be approximately normal — which the CLT delivers at moderate n even from non-normal data — and, for regression, that the residuals are approximately normal. This single distinction resolves most of the confusion around 'is my data normal enough', and it is the answer that earns full marks.
The mistake this topic produces
Testing your raw data for normality with Shapiro-Wilk and abandoning a t-test when it is significant at n = 400. At that sample size the test detects trivial departures while the CLT has long since made the t-test robust. Look at a Q-Q plot, consider the shape and the n together, and decide — do not outsource the decision to a hypothesis test about the wrong thing.
Do this now · 15 minutes
Using the simulator above and a z-table, answer three questions about your own data: what proportion of your sample falls above the clinical cutoff, what raw score sits at the 90th percentile, and what z your most extreme participant has. Then check the answers against the actual counts in the data, and explain any gap in terms of non-normality.
Day 945 minutes · the psychometric currency
Standard scores: T, stanine, IQ and the rest
z-scores are perfect mathematically and awkward in practice: half of them are negative, and most are small decimals. So psychometrics rescales them. Every standard score you have ever seen on a test report is a z-score with a new mean and SD bolted on, chosen to make the numbers friendly.
Standard score = new mean + z × new SDtake the z, stretch it to the new SD, slide it to the new meanthe standardised score from day 710 for T scores, 15 for IQ, 2 for stanines50 for T, 100 for IQ, 5 for staninesA client scores 2 SDs above the mean: z = +2.00
T score = 50 + (2.00 × 10) = 70
IQ = 100 + (2.00 × 15) = 130
stanine = 5 + (2.00 × 2) = 9
sten = 5.5 + (2.00 × 2) = 9.5
All four numbers describe exactly the same person.
Score type
Mean
SD
Where you meet it
z
0
1
Statistics, effect sizes, standardised betas.
T score
50
10
MMPI, most clinical inventories, many personality measures.
Educational testing; nine bands, deliberately coarse.
Sten
5.5
2
16PF and several personality instruments.
Percentile rank
—
—
Position, not distance. Ordinal: the gap from the 50th to 55th percentile is far smaller than from the 90th to 95th.
Why percentile ranks are not interchangeable with the rest
Percentiles are an ordinal transformation of a possibly interval scale. Because the normal curve is dense in the middle and sparse in the tails, a five-point move near the median reflects a tiny score difference, while the same five points near the 95th reflects a large one. Never average percentile ranks, never compute a difference score from them, and be suspicious of any progress claim reported in percentile gains.
The mistake this topic produces
Comparing standard scores across tests normed on different populations or in different decades as if they were on one scale. A WAIS-IV IQ of 105 and a score of 105 from a test normed in 1978 are not the same quantity — the Flynn effect alone is worth several points a decade (day 86).
Do this now · 15 minutes
Take one participant's raw scores on two or three measures. Convert each to z, then to T scores. Write the profile as a clinician would read it — which score is elevated relative to the others, and by how much in SD units. Then write down the one thing that comparison assumes about the norm groups.
Day 1045 minutes · honest fixes only
Transformations: legitimate repair versus cosmetics
When a distribution violates an assumption badly, one option is to re-express the variable on a different scale. A transformation applies the same monotone function to every score, so it preserves the order of participants while changing the spacing between them — which is precisely how it pulls in a long tail.
Problem
Transformation
Notes
Moderate positive skew
Square root √X
Mild. Requires non-negative values; add a constant first if zeros are present.
Strong positive skew — RT, income
Log (ln or log10)
The workhorse. Undefined at zero, so log(X + 1) is common. Coefficients become multiplicative.
Severe positive skew
Inverse 1/X
Strongest of the three; reverses the order of scores, so reflect afterwards (use −1/X).
Negative skew
Reflect then transform
Compute (max + 1 − X), apply the transformation, and remember the direction has flipped when interpreting.
Proportions near 0 or 1
Arcsine √p, or logit
Stabilises variance at the boundaries. Beta regression is the modern alternative.
Counts
Do not transform
Model them properly with Poisson or negative binomial regression instead (day 63).
The cost you must state
After a log transformation your analysis is no longer about mean reaction time. It is about mean log reaction time, which back-transforms to the geometric mean, not the arithmetic one. Every effect size, every confidence interval and every plotted value inherits that change. Transformations are legitimate; silently reporting the results as if they were on the original scale is not.
The modern alternative is usually better: fit a model that assumes the distribution you actually have. Generalized linear models handle counts and binary outcomes directly (days 62–63); robust methods and bootstrapping (day 110) handle heavy tails without touching the scale; mixed models handle reaction-time distributions honestly. Transformation is a 1960s solution that remains examinable and occasionally still the neatest fix.
The mistake this topic produces
Transforming to chase significance and reporting only the transformed analysis. A defensible transformation is chosen because of the distribution's shape — visible before you test anything — not because the untransformed p was .08.
Do this now · 15 minutes
Take a skewed variable of yours. Plot it raw, log-transformed and square-rooted. Compute skewness for all three. Then write the results sentence you would publish for the version you would choose, making the transformation explicit and stating what the back-transformed mean represents.
Day 1150 minutes · the results section opens here
Reporting descriptives so a reviewer trusts the rest
Your results section opens with descriptives, and experienced readers use them to decide how much of the rest to believe. A clean Table 1 that reconciles with the text buys you goodwill; a table where the ns do not add up costs you all of it, before anyone reaches your hypothesis test.
01
Report n, and every n
Total N, n per group, and how many cases each analysis actually used after exclusions. If those differ, say why here rather than in a footnote.
02
Centre and spread together, always
M (SD) for symmetric data, Mdn [IQR] for skewed. Never a mean without its spread — it is an unfinished sentence.
03
Two decimal places, and stop
Report to a precision your measurement supports. Three decimals on a Likert mean claims a resolution the instrument does not have.
04
Name the missingness
How many missing per variable, and how you handled it. This single sentence pre-empts the most common reviewer question in applied psychology.
05
Give the reliability, in this sample
Cronbach's alpha or omega for each scale, computed on your data — not quoted from the manual (volume 7).
06
Show the distribution somewhere
A histogram, a raincloud, a boxplot by group. A figure of raw data with the summary overlaid is now an expectation, not a flourish.
The APA conventions that get marked
Statistical symbols italic: M, SD, n, p, r, t, F. Greek letters not italic. Numbers below ten spelled out in text unless they are statistics or measurements. No leading zero on quantities that cannot exceed one: p = .03, r = .45 — but 0.87 seconds keeps its zero. Exact p values rather than p < .05, down to .001, below which write p < .001. Tables carry the numbers; the text carries the interpretation; neither repeats the other.
One structural habit worth building now. Write your Table 1 and its accompanying paragraph before you run a single inferential test. It forces you to look at your data, it makes any impossible value obvious while there is still time, and it means the most-read part of your results was written by someone who was not yet hoping for a particular answer.
The mistake this topic produces
A Table 1 whose group ns do not sum to the total N reported in the method, with no explanation. It is the fastest way to tell a reviewer that nobody checked the manuscript — and it makes every subsequent number suspect.
Do this now · 15 minutes
Build a complete Table 1 for your own data: n per group, and for each variable the mean, SD, median, range, number missing, skewness, and scale reliability where relevant. Then write the accompanying paragraph, in APA style, without mentioning any hypothesis. That paragraph is a real deliverable — you will paste it into a thesis.
Reference
The volume 1 formula sheet
Everything in this volume that you may have to compute under exam conditions, with its trap noted. Learn the structure, not the string of symbols — every one of these reappears inside a larger formula later.
Quantity
Formula
Watch out for
Mean
M = ΣX / n
Sensitive to outliers; meaningless for nominal data.
Median position
(n + 1) / 2 th value
With even n, average the two middle scores. Order the data first.
Sum of squares
SS = Σ(X − M)²
The computational form ΣX² − (ΣX)²/n gives the same answer with less rounding.
Sample variance
s² = SS / (n − 1)
n − 1, not n, whenever you are estimating from a sample.
Population variance
σ² = SS / N
Only when you genuinely have the whole population.
Standard deviation
s = √s²
Report this, not the variance, in a results section.
z-score
z = (X − M) / s
Changes location and scale, never shape.
Raw score from z
X = M + z·s
The reverse question, asked constantly in entrance papers.
Standard score
SS = new M + z × new SD
T: 50/10. IQ: 100/15. Stanine: 5/2.
Coefficient of variation
CV = (s / M) × 100
Ratio scales only — it needs a true zero to be interpretable.
Interquartile range
IQR = Q3 − Q1
Partner of the median; basis of the 1.5 × IQR outlier rule.
Reference
Notation you will meet for the next 101 days
Latin letters for sample statistics, Greek for population parameters. Once that convention is automatic, half of statistical notation reads itself.
Concept
Sample
Population
Note
Mean
M or X̄
μ (mu)
APA prefers M in text.
Standard deviation
s or SD
σ (sigma)
Lowercase sigma; the capital Σ means 'sum'.
Variance
s²
σ²
Always the square of the SD.
Correlation
r
ρ (rho)
Spearman's is written r-sub-s or rho depending on the house style.
Proportion
p or p̂
π (pi)
Do not confuse this π with the p-value.
Regression slope
b
β (beta)
Confusingly, β also denotes the standardised slope and the Type II error rate — context decides.
Sample size
n
N
Lowercase for a group, uppercase for the total, in most journals.
Sum
Σ
Σ
Add everything that follows.
Degrees of freedom
df
df
Usually n minus the number of parameters estimated.
Checkpoint
Ten questions before you move on
Answer from memory, then reveal. Below seven, reread the day named in the feedback — volume 2 assumes every idea here without reintroducing it.
{{ quizCounter }}
{{ quizScore }}
{{ quizQ }}
{{ quizFb }}
Next: volume 2, Probability
Eleven days on where the numbers come from — probability rules, Bayes' theorem, the binomial, and then the single most important object in the whole subject: the sampling distribution. Everything inferential is an application of volume 2.