Everything in volumes 4 to 7 was about explanation: which variables matter, how much, and with what uncertainty. Prediction asks something else entirely — how accurately can I guess an outcome for someone I have never seen? The tooling, the validation logic and the reporting standards all change, and conflating the two questions is the most common error in psychological machine-learning papers.
6
days, framing to reporting
2
questions never to conflate
1
rule: never evaluate on training data
Day 9745 minutes · decide which thesis you are writing
Two questions, two sets of rules
A model built to explain and a model built to predict can look identical and be judged by opposite criteria. Explanation cares about unbiased coefficients, interpretable parameters and inference; prediction cares only about accuracy on data the model has never seen. A predictor that works can contain no interpretable parameters at all, and a beautifully specified causal model can predict badly. Deciding which you are doing is the first and most consequential step.
Explanation
Prediction
The question
Does X affect Y, and by how much?
How accurately can I guess Y for a new case?
Judged by
Unbiased estimates, intervals, effect sizes
Out-of-sample error: RMSE, AUC, calibration
Variable selection
Theory and causal reasoning
Whatever improves held-out accuracy
Collinearity
A serious problem for interpretation
Largely irrelevant to accuracy
Overfitting
Shows up as unreliable coefficients
The central enemy; the whole workflow exists to control it
In-sample R²
Reported, with caveats
Meaningless. Use cross-validated metrics
A good result
"Stress predicted depression, b = .41 [.29, .53]"
"AUC = .78 [.71, .84] in held-out data, well calibrated"
The two sentences that reveal the confusion
"Random forest variable importance showed stress is the strongest cause of depression." Importance measures contribution to predictive accuracy, not causal magnitude — correlated predictors share importance arbitrarily, and a proxy can outrank the real cause. "Our model explained 68% of variance (R² in the training data)." That number describes the sample, not the model's usefulness; the honest version is a cross-validated R², which is almost always much lower. Both sentences appear regularly in published psychology, and both are catchable by any reviewer who knows this volume.
prediction is the right frame when
Triage and screening, risk flags, forecasting attrition, recommending an intervention, any deployment decision about an individual.
explanation is the right frame when
Theory testing, intervention mechanisms, group comparisons, anything where you intend to say "because".
both, reported separately
A common and honest structure: a causal model for the mechanism, a predictive model for the applied claim. Two sections, two sets of metrics.
neither, yet
Exploratory description. Say so — it is a legitimate contribution and misdescribing it is not.
Do this now · 15 minutes
Write your research question in one sentence, then label it: explanation, prediction, or both. If it is both, split it into two sentences with two sets of success criteria. Everything in the remaining five days assumes you did this.
Day 9855 minutes · five objects, one workflow
recipe, spec, workflow, fit, predict
tidymodels standardises the modelling steps so you can swap algorithms without rewriting your code. Five objects: a split that holds out test data, a recipe that specifies preprocessing, a model spec that is engine-agnostic, a workflow that bundles the two, and then fit and predict. Learning the skeleton once means every algorithm afterwards is a two-line change.
library(tidymodels)
set.seed(2026)
# 1. Split ONCE, at the very beginning. Do not look at the test set again
# until the final line of the analysis.
split <- initial_split(d, prop = 0.75, strata = relapse)
train <- training(split); test <- testing(split)
# 2. Recipe: preprocessing, learned from training data only
rec <- recipe(relapse ~ ., data = train) |>
step_rm(id, recruitment_site) |>
step_impute_median(all_numeric_predictors()) |>
step_novel(all_nominal_predictors()) |>
step_dummy(all_nominal_predictors()) |>
step_zv(all_predictors()) |>
step_normalize(all_numeric_predictors())
# 3. Model specification: engine-agnostic
spec <- logistic_reg() |> set_engine("glm") |> set_mode("classification")
# 4. Workflow bundles them, so preprocessing travels with the model
wf <- workflow() |> add_recipe(rec) |> add_model(spec)
# 5. Fit and predict
fit1 <- fit(wf, data = train)
predict(fit1, new_data = test, type = "prob") |> head()#> # A tibble: 6 × 2
#> .pred_no .pred_yes
#> <dbl> <dbl>
#> 1 0.812 0.188
#> 2 0.341 0.659
#> 3 0.904 0.096
Why the recipe matters more than the algorithm
Preprocessing must be learned from the training data and applied to the test data, never computed on everything at once. Normalising with the full dataset's mean, imputing with the full dataset's median, or selecting predictors by their correlation with the outcome across all rows are all forms of data leakage: the test set influences the model, held-out performance becomes optimistic, and the number you report is not achievable in deployment. A recipe inside a workflow makes leakage structurally difficult, which is its main value.
Build the five objects for a binary or continuous outcome in your own data. Split first, put every preprocessing step in the recipe, and fit one simple model. Do not evaluate anything yet — that is tomorrow, deliberately.
Day 9955 minutes · the discipline that makes it science
Why in-sample fit lies, and what to do instead
A model evaluated on the data it was fitted to reports how well it memorised, not how well it predicts. The gap is enormous for flexible models — a random forest can reach near-perfect training accuracy on pure noise. Cross-validation estimates out-of-sample performance using only training data, so the held-out test set stays untouched for the single final evaluation.
# Show the problem: fit and evaluate on the same data, on pure noise
set.seed(1)
noise <- as.data.frame(matrix(rnorm(100 * 50), 100, 50)) |>
mutate(y = rnorm(100))
m_noise <- lm(y ~ ., data = noise)
summary(m_noise)$r.squared # 50 random predictors, 100 rows#> [1] 0.5124
#> ← 51% of variance "explained" by predictors that are literally random.
#> Cross-validated R² for the same model is about -0.4: worse than the mean.library(tidymodels)
set.seed(2026)
folds <- vfold_cv(train, v = 10, strata = relapse) # or repeats = 5
res <- fit_resamples(
wf, resamples = folds,
metrics = metric_set(roc_auc, accuracy, brier_class, sens, spec),
control = control_resamples(save_pred = TRUE))
collect_metrics(res)
# Tuning: a grid searched inside the resamples, so the test set stays clean
spec_tune <- logistic_reg(penalty = tune(), mixture = 1) |>
set_engine("glmnet") |> set_mode("classification")
tuned <- tune_grid(workflow() |> add_recipe(rec) |> add_model(spec_tune),
resamples = folds, grid = 30,
metrics = metric_set(roc_auc))
show_best(tuned, metric = "roc_auc", n = 3)#> # A tibble: 5 × 6
#> .metric .estimator mean n std_err
#> accuracy binary 0.741 10 0.0182
#> brier_class binary 0.178 10 0.0091
#> roc_auc binary 0.782 10 0.0214
#> sens binary 0.612 10 0.0341
#> spec binary 0.834 10 0.0208
#>
#> # A tibble: 3 × 7
#> penalty .metric mean n std_err
#> 0.00812 roc_auc 0.791 10 0.0198
#> 0.01412 roc_auc 0.788 10 0.0201
Scheme
When
10-fold CV
The default. Good bias-variance balance for n in the hundreds.
Repeated 10-fold (5 repeats)
Smaller samples, where a single CV estimate is noisy. Cheap and worth it.
Leave-one-out
Very small n. High variance; rarely the best choice.
Grouped CV (group_vfold_cv)
Clustered data — all of one person's rows must sit in the same fold, or you leak.
Time-based (rolling_origin)
Anything temporal. Random folds let the model see the future.
Nested CV
When tuning and evaluating with no separate test set.
The one rule
The test set is opened once, at the end, with the single model you have committed to: last_fit(final_wf, split). Every comparison, every tuning decision, every preprocessing choice happens inside the training folds. If you evaluate on the test set, adjust something, and evaluate again, you have converted it into a validation set and your reported performance is optimistic — and with repeated-measures data, remember that random folds split a person across folds, which leaks that person's idiosyncrasies into the model that predicts them.
Do this now · 25 minutes
Run 10-fold cross-validation on your workflow and compare the cross-validated metric with the in-sample equivalent. If your data are clustered, switch to group_vfold_cv() and note how much the estimate drops — that drop was leakage.
Day 10055 minutes · more predictors than sense
Lasso and ridge: shrinking coefficients on purpose
With 80 candidate predictors and 200 participants, ordinary regression overfits badly and stepwise selection — still common in psychology — produces unstable models, biased coefficients and invalid p-values. Regularisation instead penalises large coefficients, trading a little bias for a large reduction in variance. Lasso can shrink coefficients exactly to zero, giving selection as a by-product.
Many correlated predictors, all plausibly relevant
Lasso (mixture = 1)
Shrinks some to exactly zero
You want a sparse, deployable model
Elastic net (between)
Compromise; handles correlated groups better than lasso
The safe default — tune the mixture
Stepwise selection
Unstable, biased, invalid inference
Never. Regularisation replaced it
Choosing by p-values
Overfits and invalidates inference
Never for prediction; for explanation use theory
What you may and may not say about lasso output
The predictors that survive are the ones that helped predict in this sample at this penalty — swap in a new sample and the set changes, sometimes substantially. So you may say "the model retained nine predictors and achieves cross-validated RMSE 4.18"; you may not say "these nine variables are the important predictors of depression", and you must not report the shrunken coefficients as effect sizes with p-values. Inference after selection needs different machinery (post-selection inference, or an independent sample).
Do this now · 25 minutes
Fit an elastic net on your widest set of predictors, tune both penalty and mixture, and list the surviving terms. Then refit on a different random split and compare the survivor lists — the instability you see is the reason not to interpret them as "the important variables".
Day 10155 minutes · flexible models, careful claims
Random forests: strong baselines, weak explanations
A random forest averages hundreds of decision trees, each grown on a bootstrap sample with a random subset of predictors. It captures interactions and non-linearity without you specifying them, needs little preprocessing, and is a strong baseline on tabular psychological data. The interpretation tools — importance, partial dependence, SHAP — describe the model, not the world, and that distinction is the whole day.
library(tidymodels)
spec_rf <- rand_forest(trees = 1000, mtry = tune(), min_n = tune()) |>
set_engine("ranger", importance = "permutation") |>
set_mode("regression")
set.seed(2026)
tuned_rf <- tune_grid(workflow() |> add_recipe(rec) |> add_model(spec_rf),
resamples = vfold_cv(train, v = 10), grid = 20,
metrics = metric_set(rmse, rsq))
show_best(tuned_rf, metric = "rmse", n = 3)
fit_rf <- finalize_workflow(workflow() |> add_recipe(rec) |> add_model(spec_rf),
select_best(tuned_rf, metric = "rmse")) |> fit(train)
library(vip)
vip(fit_rf, num_features = 15, geom = "point")#> # A tibble: 3 × 8
#> mtry min_n .metric mean n std_err
#> 6 12 rmse 4.021 10 0.0884
#> 9 8 rmse 4.048 10 0.0902
#> ← marginally better than the elastic net (4.18). Report both; a complex
#> model that barely beats a simple one is a finding, not a failure.# What the model does with one predictor, holding others at their observed
# values: partial dependence. Direction and shape, not causal magnitude.
library(DALEXtra)
expl <- explain_tidymodels(fit_rf, data = train |> select(-bdi_post),
y = train$bdi_post)
model_profile(expl, variables = c("rumination", "sleep_hours")) |> plot()
# Per-case explanations, for a clinician looking at one person
predict_parts(expl, new_observation = test[1, ], type = "shap") |> plot()
Four honest limits on variable importance
It is not causal. A downstream proxy of the outcome will outrank the upstream cause. Correlated predictors split their importance arbitrarily — two near-duplicate scales each look half as important as either would alone. Default impurity importance is biased toward continuous and high-cardinality variables; use importance = "permutation". It has no uncertainty attached unless you bootstrap the whole pipeline, so small differences in rank are noise. Say "contributed most to predictive accuracy in this model", never "is the most important determinant of".
Do this now · 25 minutes
Fit a tuned random forest and compare its cross-validated metric against your regularised regression. Produce the permutation importance plot and two partial dependence curves, then write the interpretation sentence twice — once as a predictive claim and once as the causal claim you must not make.
Day 10250 minutes · the last day
Metrics that matter, calibration, and refusing to overclaim
The final day is about the reporting standards that make a predictive paper credible: the right metrics for the decision at stake, calibration alongside discrimination, an honest baseline comparison, and explicit statements about generalisation and fairness. Accuracy alone is almost never sufficient, and on imbalanced outcomes it is actively misleading.
# The single, final evaluation on data never touched before
final <- last_fit(final_wf, split,
metrics = metric_set(roc_auc, brier_class, accuracy,
sens, spec, ppv, npv))
collect_metrics(final)
# Discrimination and calibration, side by side — never one alone
library(probably)
preds <- collect_predictions(final)
roc_curve(preds, truth = relapse, .pred_yes, event_level = "second") |> autoplot()
cal_plot_breaks(preds, truth = relapse, estimate = .pred_yes)#> # A tibble: 7 × 4
#> .metric .estimator .estimate
#> roc_auc binary 0.774
#> brier_class binary 0.181
#> accuracy binary 0.738
#> sens binary 0.604
#> spec binary 0.829
#> ppv binary 0.641
#> npv binary 0.804
#> ← calibration plot shows over-prediction in the top risk decile: the
#> ranking is decent, the probabilities are not yet trustworthy.
Metric
Answers
Accuracy
Proportion correct. Useless when outcomes are imbalanced — 92% accuracy by always predicting "no relapse" at 8% base rate.
Sensitivity / specificity
How many true cases you catch; how many non-cases you leave alone. Pick the threshold from the costs of each error.
PPV / NPV
Given a positive prediction, what is the chance it is right? Depends on base rate, so report it for your population.
ROC AUC
Ranking ability across all thresholds. Insensitive to base rate and to calibration.
PR AUC
Better than ROC for rare outcomes.
Brier score / calibration
Are the predicted probabilities honest? Required for any clinical use.
RMSE / MAE / R²
Continuous outcomes. Always cross-validated or test-set, never in-sample.
Baseline comparison
Beat the majority class, the sample mean, and a simple regression — or the complexity is unjustified.
The four sentences a predictive paper needs
1. "Performance was estimated by 10-fold cross-validation within the training set and confirmed once on a held-out 25% test set." 2. "The model achieved AUC = .77 and a Brier score of .18, compared with .50 and .24 for the base-rate model and .74 and .19 for logistic regression with three preregistered predictors." 3. "Calibration was adequate below the top risk decile, where the model over-predicted." 4. "The sample was recruited from a single university clinic; performance in other settings is unknown, and we make no causal claims about the predictors." Together they pre-empt nearly every reasonable objection.
# Before anyone deploys anything: does it work equally well for everyone?
preds |> group_by(site) |> roc_auc(truth = relapse, .pred_yes,
event_level = "second")
preds |> group_by(gender) |> cal_plot_breaks(truth = relapse,
estimate = .pred_yes)
# And the deployable object, versioned with everything it needs
final_model <- extract_workflow(final)
saveRDS(final_model, "output/models/relapse_model_v1.rds")
# renv::snapshot() — volume 9 day 92 — or the model is not reproducible#> # A tibble: 3 × 4
#> site .metric .estimate
#> clinic_a roc_auc 0.801
#> clinic_b roc_auc 0.784
#> online roc_auc 0.658 ← markedly worse; report it, do not average it away
Day 102, and what comes next
You began 102 days ago not knowing what a vector was. You can now clean data reproducibly, build journal-grade figures, choose and defend the right model for a design, fit multilevel and latent-variable models, report Bayesian estimates, synthesise a literature, plan a study's power by simulation, work reproducibly in Quarto and Git, and build a validated predictive model. The remaining step is not more technique — it is assembling the eight artefacts in the capstone volume into something an admissions panel can read in ten minutes.
Do this now · 25 minutes
Run last_fit() once, report the full metric set alongside a base-rate and a simple-regression baseline, produce the calibration plot, and check performance by subgroup. Then write the four sentences. That is the course finished — go to the capstone.
Reference
Nine steps, in strict order
The order is the method. Reordering it is how optimistic performance estimates happen.
1 · frame
Prediction or explanation? Write the success criterion before you fit anything.
2 · split
initial_split() once. Stratify on the outcome. Then forget the test set exists.
3 · recipe
All preprocessing, learned from training data only. Grouped or time-based splits where the data demand it.
4 · baseline
Majority class or sample mean, plus a simple preregistered regression. Everything is judged against these.
5 · resample
10-fold CV, or grouped/rolling. All comparisons live here.
6 · tune
tune_grid() inside the folds. Never against the test set.
7 · finalise
One model, finalize_workflow(), committed to in writing.
Metrics with baselines, calibration, limits of generalisation, no causal language.
Reference
Install these once
install.packages(c(
"tidymodels", # rsample, recipes, parsnip, workflows, tune, yardstick
"glmnet", # lasso, ridge, elastic net
"ranger", # fast random forests
"xgboost", # gradient boosting
"vip", # variable importance plots
"DALEXtra", # partial dependence, SHAP, model profiles
"probably", # calibration plots and threshold optimisation
"finetune" # racing methods: faster tuning
))
Checkpoint
Six questions to finish the course
{{ quizCounter }}
{{ quizScore }}
{{ quizQ }}
{{ quizFb }}
You have reached day 102
One hundred and two days from installing R to cross-validated prediction. What converts that into a PhD application is the portfolio: eight artefacts, each with a build day, a checklist and a template. The capstone volume is where you assemble them.