Six months after submission a reviewer asks where a number came from, and you cannot find it — that is the normal experience, and it is entirely avoidable. This volume builds the infrastructure: one document that contains text, code and output; tables and manuscripts generated rather than typed; version control; a frozen environment; a pipeline that reruns only what changed; your own functions and package; and a preregistration on OSF. It is also where your application portfolio comes from.
12
days, five portfolio artefacts
0
numbers typed by hand
8
git commands you actually need
Day 8545 minutes · what must be true
What has to hold for someone else to get your numbers
Reproducibility is not a virtue signal; it is the property that someone with your data and code obtains your results. That requires five things to be true, and each has a tool in this volume. The fastest way to find out whether they hold is to try it: move your project folder to a different location, restart R with a blank session, and run everything from scratch.
Requirement
Common failure
Fixed on day
Raw data is untouched and available
Cleaning done by hand in Excel, over the original file
85, 96
Every transformation is in code
"I deleted three outliers" recorded nowhere
85, 93
Results are generated, not typed
A number in the manuscript that no longer matches the model
86, 87, 88
The environment is recoverable
A package updated and the analysis silently changed
92
The history is visible
One file called analysis_final_v3_REAL.R
90, 91
# The project structure that makes all five achievable
# my-thesis/
# ├── my-thesis.Rproj
# ├── data/
# │ ├── raw/ ← READ-ONLY. Never written to by any script.
# │ └── processed/ ← generated; safe to delete and rebuild
# ├── R/
# │ ├── 00_packages.R
# │ ├── 01_import.R
# │ ├── 02_clean.R
# │ ├── 03_score.R
# │ ├── 04_analyse.R
# │ └── functions/ ← your own functions (day 94)
# ├── output/
# │ ├── figures/
# │ └── tables/
# ├── manuscript/
# │ └── paper.qmd
# ├── renv.lock ← day 92
# └── README.md ← day 91
# Enforce it in code, not discipline: this line makes raw data read-only
fs::file_chmod("data/raw", "a-w") # macOS/Linux
# Windows: right-click the folder, Properties, Read-only
The test that tells you the truth
Restart R with a completely clean session (Session → Restart R, and turn off workspace saving in Tools → Global Options), then run every numbered script in order. If anything fails, you have found a hidden dependency on something in your environment — a variable you created interactively, a working directory, a package you forgot to load. Do this weekly. The alternative is finding out in month 34.
never save .RData
Tools → Global Options → uncheck "Restore .RData" and set "Save workspace" to Never. Objects must come from code.
never setwd()
Use the .Rproj file and here::here(). A hard-coded path breaks on every other machine including your own next one.
never edit raw data
Every correction is a line of code in the cleaning script, with a comment saying why.
never copy a number
If it appears in the manuscript, it comes from an inline code chunk. Day 86.
Do this now · 15 minutes
Restructure one existing project to the layout above, make data/raw read-only, turn off workspace saving, and run the clean-session test. Write down what broke — that list is what this volume fixes.
Day 8650 minutes · text and code in one file
Quarto: the document that cannot go out of date
A Quarto document interleaves prose and code, and renders to PDF, Word or HTML with the output computed at render time. The consequence is structural: a number in your text cannot disagree with your analysis, because it is your analysis. Inline code is the feature that matters most — `r round(coef(m)[2], 2)` in a sentence.
---
title: "Stress and mood in daily life"
author: "Your Name"
format:
pdf:
documentclass: article
fontsize: 11pt
docx: default # supervisors want Word; render both
bibliography: references.bib
csl: apa.csl
execute:
echo: false # hide code in the output
warning: false
cache: true # do not refit a 20-minute model on every render
---
## Method
Participants were @sec-sample describes ...
```{r}
#| label: setup
library(tidyverse); library(lme4); library(lmerTest)
source(here::here("R", "02_clean.R"))
m <- lmer(mood ~ stress_cw + (1 + stress_cw | id), data = esm)
b <- round(fixef(m)["stress_cw"], 2)
ci <- round(confint(m, parm = "stress_cw", method = "Wald"), 2)
```
## Results
Within-person stress predicted lower mood, *b* = `r b`,
95% CI [`r ci[1]`, `r ci[2]`].#> Within-person stress predicted lower mood, b = -0.29, 95% CI [-0.33, -0.25].
#> ← that sentence updates itself when the model changes. No typing, no drift.
Chunk option
Effect
#| label: fig-results
Names the chunk; enables cross-references with @fig-results
#| echo: false
Run the code, hide it. Set globally in the YAML for manuscripts
#| fig-width: 3.35
Inches. 3.35in = 85 mm, the single-column journal width
#| fig-cap: "..."
Caption, and makes the figure referenceable
#| cache: true
Reuse results unless the chunk changes. Essential with brms or simr
#| include: false
Run silently — for setup chunks
#| eval: false
Show the code without running it — for teaching materials
Two habits that make Quarto pay
Source your scripts, do not duplicate them. The document's setup chunk should source() the cleaning and scoring scripts, so there is exactly one definition of your dataset. Render to Word for your supervisor. The most common objection to Quarto is "my supervisor wants track changes" — format: docx solves it, and you keep the reproducible source. Add reference-doc: to match their template.
Do this now · 20 minutes
Convert one existing results section to Quarto with every number as inline code, render to both PDF and Word, then change something in your cleaning script and re-render. Watch the text update. That is the whole argument.
Day 8755 minutes · never typed by hand
APA tables, generated from the model
Hand-typed tables are where transcription errors live, and they are the part of a manuscript most often out of date. Three packages cover almost everything psychology needs: gt for full control, flextable for Word, and apaTables or modelsummary for the standard APA layouts.
library(gt); library(tidyverse)
# Table 1: sample descriptives, the one every paper needs
tab1 <- d |>
select(condition, age, gender, bdi_pre, bdi_post) |>
pivot_longer(c(age, bdi_pre, bdi_post)) |>
group_by(condition, name) |>
summarise(M = mean(value, na.rm = TRUE), SD = sd(value, na.rm = TRUE),
.groups = "drop") |>
mutate(cell = sprintf("%.2f (%.2f)", M, SD)) |>
select(-M, -SD) |> pivot_wider(names_from = condition, values_from = cell)
tab1 |> gt() |>
tab_header(title = "Table 1", subtitle = "Descriptive statistics by condition") |>
cols_label(name = "") |>
tab_source_note("Values are M (SD). BDI-II range 0-63.") |>
gtsave("output/tables/table1.docx")# Regression tables, several models side by side
library(modelsummary)
modelsummary(list("Model 1" = m1, "Model 2" = m2, "Model 3" = m3),
statistic = "conf.int", conf_level = 0.95,
stars = FALSE, # APA prefers CIs to stars
coef_rename = c("conditionCBT" = "CBT vs control",
"bdi_pre" = "Baseline BDI-II"),
gof_map = c("nobs", "r.squared", "adj.r.squared"),
notes = "95% confidence intervals in brackets.",
output = "output/tables/table2.docx")
# APA correlation table, straight to Word, in one line
apaTables::apa.cor.table(d[, c("neuro", "agree", "stress", "mood")],
filename = "output/tables/table3.doc",
table.number = 3)#> +--------------------+-------------------+-------------------+
#> | | Model 1 | Model 2 |
#> +--------------------+-------------------+-------------------+
#> | CBT vs control | -3.918 | -3.408 |
#> | | [-5.863, -1.973] | [-4.819, -1.997] |
#> | Baseline BDI-II | | 0.681 |
#> | | | [0.578, 0.784] |
#> | Num.Obs. | 247 | 247 |
#> | R2 | 0.068 | 0.526 |
#> +--------------------+-------------------+-------------------+
Package
Best for
gt
Full control of any table, HTML and Word output, good defaults
flextable
Word-first workflows and complex headers; works inside Quarto docx
modelsummary
Model comparison tables with CIs, to any format
apaTables
Correlation, ANOVA and regression tables in strict APA layout, to Word
papaja::apa_table
Tables inside a papaja manuscript (tomorrow)
sjPlot::tab_model
Mixed models and GLMs, quick and readable (volume 5 day 51)
gtsummary
Clinical-style Table 1 with tests, popular in medical journals
The APA conventions worth knowing
Table number and title above the table, title in italics on its own line; no vertical rules and minimal horizontal ones; general notes then specific then probability notes; means and SDs to two decimals, correlations to two, p-values to three (and "< .001" below that); report exact n per cell when it varies. Modern APA prefers confidence intervals to significance asterisks — and every package above will do both, so choose deliberately.
Do this now · 25 minutes
Generate your Table 1 and your main regression table with code and save them to output/tables/. Delete the hand-typed versions. Then change an exclusion rule and re-render both to confirm nothing needs retyping.
Day 8860 minutes · portfolio artefact one
papaja: title page to references, rendered from your analysis
papaja renders a complete, correctly formatted APA 7 manuscript — title page, running head, abstract, sections, tables, figures, references — from an R Markdown source. Combined with day 86's inline code and day 87's generated tables, you get a document where nothing is typed twice and nothing can drift out of date. This is portfolio artefact one.
install.packages("papaja") # plus a LaTeX distribution:
tinytex::install_tinytex() # tinytex is the painless option
# In RStudio: File → New File → R Markdown → From Template → APA article (papaja)---
title : "Within-person stress and momentary mood"
shorttitle : "STRESS AND MOOD"
author:
- name : "Your Name"
affiliation : "1"
corresponding : yes
email : "you@university.edu"
affiliation:
- id : "1"
institution : "University of Somewhere"
abstract: |
Momentary stress was associated with lower mood within persons ...
keywords : "experience sampling, stress, mood, multilevel"
bibliography : "references.bib"
floatsintext : yes
figurelist : no
documentclass : "apa7"
classoption : "man" # "man" = submission, "doc" = readable draft
output : papaja::apa6_pdf
---
```{r analysis, include = FALSE}
library(papaja); library(lme4); library(lmerTest)
source(here::here("R", "04_analyse.R"))
apa_m <- apa_print(m_main)
```
## Results
Within-person stress predicted mood, `r apa_m$full_result$stress_cw`.
```{r tab1, results = "asis"}
apa_table(tab1, caption = "Descriptive statistics by condition",
note = "Values are M (SD).", escape = FALSE)
```#> Renders to:
#> Within-person stress predicted mood, b = -0.29, 95% CI [-0.33, -0.25],
#> t(88.77) = -14.27, p < .001.
#> ← apa_print() formats the whole result string, correctly, from the model object
apa_print(), the function that saves the most time
Hand it almost any model — lm, aov, t.test, lmer, anova — and it returns APA-formatted strings: estimate, interval, test statistic, df and p, with the right italics, the right number of decimals and "< .001" handled. You stop formatting results, which is both a time saving and an error class eliminated. Combine with apa_table() and a manuscript becomes almost entirely generated.
Reviewer's request
Your response with papaja
"Report exact p-values"
Already exact, everywhere, automatically
"Add the confidence intervals"
One argument to apa_print(); all numbers update
"Exclude participants under 18 and rerun"
Change one filter, re-render: text, tables and figures all update together
"Send a Word version with line numbers"
output: papaja::apa6_docx, classoption: "man"
"Where did this number come from?"
Point to the chunk. It came from the model
Do this now · 30 minutes
Create a papaja manuscript from the RStudio template, move one real results section into it, and render to PDF. Then change an exclusion criterion in the cleaning script and re-render to see every number move together. Artefact one exists.
Day 8945 minutes · a short, high-value day
Zotero, Better BibTeX, and citing R itself
Citation management becomes trivial once set up and stays painful forever if you do not. The configuration is Zotero for the library, Better BibTeX for stable citation keys and an auto-exported .bib file, and a CSL style file for formatting. One afternoon, then never again.
1 · Zotero
Free, with a browser connector that grabs papers and metadata in one click. Check the metadata — publishers supply it badly.
2 · Better BibTeX
A Zotero plugin. Gives stable keys like bergomi2024stress and auto-updates a .bib file whenever your library changes.
3 · Auto-export
Right-click a collection → Export → Better BibTeX → tick "Keep updated". Point it at references.bib in your project.
4 · CSL style
Download apa.csl from the Zotero style repository into your project and name it in the YAML.
---
bibliography: references.bib
csl: apa.csl
---
Momentary designs reduce recall bias [@shiffman2008ecological].
@bolger2013intensive argue that ...
Several authors agree [@shiffman2008ecological; @bolger2013intensive, pp. 45-52].
A suppressed-author citation [-@bolger2013intensive].
# References#> Momentary designs reduce recall bias (Shiffman et al., 2008).
#> Bolger and Laurenceau (2013) argue that ...
#> Several authors agree (Bolger & Laurenceau, 2013, pp. 45-52;
#> Shiffman et al., 2008).# Cite R and every package you used — expected practice, rarely done
citation()
citation("lme4")
# Generate a .bib entry for all loaded packages at once
library(grateful)
cite_packages(output = "paragraph", out.dir = ".", pkgs = "Session")
# Or papaja's built-in version
papaja::r_refs("r-references.bib")
# then in the YAML: bibliography: ["references.bib", "r-references.bib"]#> We used R version 4.4.1 [@rcore2024] and the packages lme4 v1.1-35
#> [@bates2015lme4], brms v2.21.0 [@burkner2017brms], and tidyverse v2.0.0
#> [@wickham2019tidyverse] ...
Why the stable-key part matters
Zotero's default keys change when metadata changes, which silently breaks every citation in a manuscript you wrote six months ago. Better BibTeX pins them. Set the key format once (author + year + first title word), tick "keep updated" on the export, and your .bib is always current without a manual export step — which is the difference between a system that works and one you abandon in month four.
Do this now · 20 minutes
Install Zotero and Better BibTeX, set up a project collection with auto-export to references.bib, and cite three papers plus R and your five main packages in your manuscript. Render and check the reference list formatting.
Day 9050 minutes · version control, minimally
Git, learned as eight commands and one habit
Git is a system for recording the history of a project so you can see what changed, when and why, and return to any earlier state. Its reputation for difficulty comes from the branching model you mostly will not need. For a solo thesis, eight commands and one habit — commit when something works, with a message explaining why — deliver almost all the benefit.
# One-time setup
usethis::use_git_config(user.name = "Your Name", user.email = "you@uni.edu")
# Turn the project into a repository (restarts RStudio, adds the Git pane)
usethis::use_git()
# The eight commands, in the terminal (RStudio's Git pane does the same)
git status # what has changed? Run this constantly
git add R/04_analyse.R # stage a specific file
git add . # stage everything changed
git commit -m "Add sensitivity analysis excluding three influential cases"
git log --oneline # the history, one line per commit
git diff # what exactly changed, line by line
git checkout -- R/04_analyse.R # discard uncommitted changes to a file
git revert <hash> # undo a specific commit, keeping the history#> $ git log --oneline
#> a3f91c2 Add sensitivity analysis excluding three influential cases
#> 8b21e40 Fix reverse-scoring on items 4 and 7
#> 1c94af8 Add figure 2: within-person slopes
#> f02d1b9 Initial commit: project structure and raw data
#> ← this is the document that shows an admissions panel how you think# .gitignore: what must NOT be committed
usethis::use_git_ignore(c(
".Rhistory", ".RData", ".Rproj.user",
"data/raw/participants_identifiable.csv", # NEVER commit identifiable data
"output/figures/*.png", # regenerable; keeps the repo light
"*_cache/", "*_files/"
))
# Committed something sensitive? It stays in history until removed properly.
# Small repo, recent mistake: git reset --soft HEAD~1 (before pushing)
# Already pushed: treat the data as exposed, rotate anything secret, and use
# git filter-repo. This is the one Git mistake that is genuinely expensive.
What makes a good commit message
Not "updates" or "fixed stuff". A commit message answers why: "Reverse-score items 4 and 7; keying was wrong in the original SPSS syntax". Commit once per meaningful change rather than once per day, and never commit something you know is broken without saying so in the message. Six months later this history is how you answer a reviewer's question about a decision you no longer remember making.
Do this now · 20 minutes
Put one project under version control, write a proper .gitignore, and make five commits as you work through today's other tasks. Then run git log --oneline and read your own messages critically — would they mean anything to you in a year?
Day 9150 minutes · portfolio artefact two
A public repository that reads well to a stranger
A GitHub repository is the most easily verified evidence of research competence available to a PhD applicant: it shows real code, a real history, and whether you can write for someone who was not in the room. The technical part is two commands. The work is the README and the structure.
# Two commands: create the remote and push
usethis::use_github() # needs a GitHub PAT
usethis::create_github_token() # opens the browser; then gitcreds::gitcreds_set()
# Thereafter
git push # send local commits to GitHub
git pull # bring down changes (relevant with collaborators)# README.md — the only file most visitors will read
# Within-person stress and momentary mood
Analysis code for [preprint/paper link]. Experience-sampling study of 94
participants over 14 days (2,814 observations).
## Reproducing the analysis
1. Clone this repository and open `stress-mood.Rproj`.
2. `renv::restore()` to install the exact package versions used.
3. `targets::tar_make()` to run the full pipeline (about 4 minutes).
Rendered manuscript: `manuscript/paper.pdf`.
## Structure
| Path | Contents |
|---|---|
| `data/raw/` | De-identified survey export (see codebook.md) |
| `R/` | Numbered analysis scripts |
| `output/` | Generated figures and tables |
| `manuscript/` | Quarto/papaja source |
## Data availability
Item-level data are available on OSF (doi:...). Identifiable variables
(birth date, free-text responses) are withheld per ethics approval #2024-118.
## Citation
Your Name (2026). Within-person stress and momentary mood. https://doi.org/...
A reviewer or panel checks
Make sure
Can I tell what this project is in 20 seconds?
First paragraph of the README says the question, design and n
Could I run it?
Three numbered steps, with renv::restore() named
Is the code readable?
Numbered scripts, comments explaining why, no 200-line unbroken blocks
Is there a real history?
Commits spread over weeks with meaningful messages, not one "initial commit" dump
Is the data handled responsibly?
An explicit data-availability statement; nothing identifiable in the repo
Does anything obviously not work?
Clone it into a fresh folder yourself and run it before you link it on a CV
Portfolio artefact two
One well-presented repository beats five abandoned ones. Pick the project you are proudest of, spend two hours on the README and structure, verify a clean clone runs, and put the link on your CV and in your application. Add a short codebook.md describing every variable — it takes twenty minutes and it is the thing that most signals you have worked with real data.
Do this now · 20 minutes
Publish one repository, write the README to the template above, then clone it into a fresh directory and follow your own instructions exactly. Fix whatever fails. Artefact two exists.
Day 9245 minutes · the silent failure
renv: the package versions your results depend on
Package updates change results. Defaults shift, estimators are corrected, functions are deprecated — and "I reran my thesis analysis and the numbers moved" is a real and common experience. renv gives each project its own library and a lockfile recording exact versions, so the analysis can be reconstructed years later.
library(renv)
init() # project-local library + renv.lock; scans your code for deps
# ...work normally: install.packages() now installs into the project library
snapshot() # record current versions into renv.lock. Commit the lockfile.
status() # is the lockfile in sync with what is installed?
restore() # reinstall exactly what the lockfile says — on any machine
# In the manuscript, the honest minimum
sessionInfo()
sessioninfo::session_info() # nicer, and includes source (CRAN/GitHub)#> * Project '~/thesis' loaded. [renv 1.0.7]
#> The following package(s) will be installed:
#> lme4 [1.1-35.5]
#> brms [2.21.0]
#> ...
#> * Lockfile written to '~/thesis/renv.lock'.
#>
#> R version 4.4.1 (2024-06-14)
#> Platform: aarch64-apple-darwin20
#> attached base packages: stats graphics grDevices utils datasets methods base
#> other attached packages:
#> lme4_1.1-35.5 Matrix_1.7-0 lmerTest_3.1-3 tidyverse_2.0.0
Level
Gives you
sessionInfo() in the paper
A record of what you used. The minimum acceptable practice
renv lockfile, committed
Someone else can install exactly your versions. The realistic target
Docker or rocker image
The full system frozen, including R and system libraries. For methods papers and long-lived pipelines
Binder / Codespaces
A one-click runnable environment in the browser. Impressive on an application, half a day's work
Two practical notes
Commit renv.lock, not the library. The lockfile is a small text file; renv/library/ is hundreds of megabytes and is gitignored automatically. Snapshot deliberately. Run snapshot() when the analysis reaches a state you might need to return to — after a submission, before a major revision — not continuously. And when a reviewer asks you to add an analysis two years later, restore() first, so you are working with the environment the results came from.
Do this now · 20 minutes
Run renv::init() and snapshot() on your main project, commit the lockfile, then delete the project library and run restore() to prove it works. Add session_info() output to your manuscript's supplement.
Day 9355 minutes · portfolio artefact six
targets: rerun only what changed, and prove the chain is intact
Numbered scripts have a flaw: nothing stops you running script 4 after editing script 2, and nothing tells you which outputs are now stale. targets makes the dependency graph explicit — it knows that your model depends on your cleaned data which depends on your raw file, so changing the cleaning invalidates exactly the downstream objects and nothing else.
# _targets.R — the whole pipeline, declared not run
library(targets); library(tarchetypes)
tar_option_set(packages = c("tidyverse", "lme4", "lmerTest", "ggplot2"))
tar_source("R/functions") # your own functions (day 94)
list(
tar_file(raw_file, "data/raw/esm_export.csv"), # tracked by content hash
tar_target(raw, readr::read_csv(raw_file)),
tar_target(clean, clean_esm(raw)),
tar_target(scored, score_scales(clean)),
tar_target(m_main, fit_main_model(scored)),
tar_target(m_lag, fit_lagged_model(scored)),
tar_target(fig_main, make_slope_figure(scored, m_main)),
tar_target(tab_main, make_model_table(m_main, m_lag)),
tar_quarto(paper, "manuscript/paper.qmd")
)library(targets)
tar_visnetwork() # the dependency graph, colour-coded by what is outdated
tar_make() # run only the outdated targets
tar_read(m_main) # pull any intermediate object into your session
tar_load(scored) # load it under its own name, for interactive work
tar_outdated() # what would rerun, and why
# Parallel execution for slow branches (brms models, simulations)
tar_make_future(workers = 4)#> + raw dispatched
#> + clean dispatched
#> - scored skipped ← unchanged
#> - m_main skipped
#> + fig_main dispatched ← only because the figure function changed
#> + paper dispatched
#> ✔ ended pipeline [3.412 seconds]
Portfolio artefact six: the reusable template
Once your pipeline works, strip the content and keep the skeleton: _targets.R, the folder structure, your function files, the Quarto manuscript template, renv.lock, the README. Put it on GitHub as research-template. Your second study then starts on day one with working infrastructure instead of month one — which is the real reason experienced researchers publish faster, and it is a portfolio artefact in its own right.
Symptom
targets' answer
"Did I rerun the analysis after fixing the cleaning?"
tar_outdated() tells you, and tar_make() fixes it
"Rerunning everything takes 40 minutes"
Only changed targets rerun; the brms model is not refitted for a figure tweak
"Which script produced this figure?"
The graph shows it explicitly
"My collaborator's numbers differ from mine"
Same pipeline, same lockfile, same results
"Is my manuscript current?"
tar_quarto() makes the document a target like any other
Do this now · 25 minutes
Convert your numbered scripts into a _targets.R pipeline with at least six targets, run tar_visnetwork(), then change one function and watch what reruns. Strip a copy into a template repository.
Day 9455 minutes · when copy-paste becomes a function
Functions: the rule of three, arguments, and defaults
The moment to write a function is the third time you copy a block of code. Copy-paste duplicates bugs, and fixing one instance silently leaves the others wrong — which is how two figures in the same paper end up using different exclusion rules. A function has one definition, one place to fix, and a name that documents intent.
# R/functions/scoring.R
#' Score a scale from item columns
#'
#' @param data A data frame containing the items.
#' @param items Unquoted item column names, tidyselect style.
#' @param reverse Names of items to reverse-score.
#' @param max_scale Highest possible response option (for reversing).
#' @param min_valid Minimum number of answered items required; below this,
#' the score is NA rather than a mean of very few items.
#' @return The data frame with one added column, `score`.
score_scale <- function(data, items, reverse = NULL, max_scale = 5,
min_valid = 3) {
stopifnot(is.data.frame(data), min_valid >= 1)
it <- dplyr::select(data, {{ items }})
if (!is.null(reverse)) {
it[reverse] <- (max_scale + 1) - it[reverse]
}
n_valid <- rowSums(!is.na(it))
score <- rowMeans(it, na.rm = TRUE)
score[n_valid < min_valid] <- NA_real_
dplyr::mutate(data, score = score)
}
# Used three times, defined once
d |> score_scale(N1:N5, reverse = "N4") |> dplyr::pull(score) |> head()#> [1] 2.60 3.40 2.20 NA 4.00 3.20
#> ← the NA is a participant who answered only two of five items: the
#> min_valid rule made that explicit instead of silently averaging two items
The common case should work with no arguments beyond the data
Fail early and loudly
stopifnot() or rlang::abort() beats a silent wrong answer
No global variables inside
Everything the function needs comes in as an argument, or it is not reproducible
Document the why
Roxygen comments above; they become help pages tomorrow
Test it on a case you know
One tiny example with a hand-computed answer catches most mistakes
# Functions turn repetition into iteration
library(purrr)
scales <- list(neuro = c("N1","N2","N3","N4","N5"),
agree = c("A1","A2","A3","A4","A5"))
scored <- imap(scales, \(items, nm) {
score_scale(d, all_of(items)) |> dplyr::select(id, !!nm := score)
}) |> reduce(dplyr::left_join, by = "id")
# And they can be tested — which is how you know a change did not break them
library(testthat)
test_that("score_scale honours min_valid", {
df <- data.frame(a = c(1, 1), b = c(5, NA), c = c(3, NA))
out <- score_scale(df, a:c, min_valid = 3)
expect_equal(out$score[1], 3)
expect_true(is.na(out$score[2]))
})#> Test passed 🎉
The habit that compounds
Keep a personal R/functions/ folder in every project and move anything you have written twice into it. Within a year you will have thirty functions that encode your own conventions — your theme, your scoring rules, your table format — and each new project starts faster than the last. When that folder gets large enough to be worth sharing, it becomes tomorrow's package.
Do this now · 25 minutes
Find a block of code you have copied at least twice and turn it into a documented function with defaults and a guard clause. Write one testthat test for it with a hand-computed expected value, and use it in your pipeline.
Day 9560 minutes · portfolio artefact four
A package with three working functions
Almost no psychology PhD applicant has written a package, and doing so signals engineering literacy that transfers directly to research-assistant and postdoc work. It does not need to be novel or on CRAN. Three documented, tested functions that solve a real problem in your own workflow — scoring your lab's questionnaires, applying your theme, cleaning your platform's export — are enough.
library(usethis); library(devtools)
create_package("~/dev/labtools") # opens a new project
use_git()
use_mit_license("Your Name")
use_r("score_scale") # creates R/score_scale.R
# paste your function in, put the cursor inside it, then:
# Code → Insert Roxygen Skeleton (or Ctrl/Cmd+Alt+Shift+R)
use_package("dplyr") # declares a dependency in DESCRIPTION
use_pipe() # if you want %>% available
use_testthat(); use_test("score_scale")
document() # roxygen comments → help pages and NAMESPACE
load_all() # load the package as if installed, for interactive testing
test() # run the test suite
check() # the full CRAN-style check: this is the real bar#> ── R CMD check results ─────────────────── labtools 0.0.0.9000 ────
#> Duration: 24.1s
#>
#> 0 errors ✔ | 0 warnings ✔ | 0 notes ✔
#> ← that line is the achievement. It means documentation, dependencies,
#> examples and tests are all consistent.# The DESCRIPTION file — your package's identity
Package: labtools
Title: Scoring and Plotting Helpers for the Somewhere Lab
Version: 0.1.0
Authors@R:
person("Your", "Name", email = "you@uni.edu", role = c("aut", "cre"),
comment = c(ORCID = "0000-0000-0000-0000"))
Description: Utilities used across projects in the Somewhere Lab: scale
scoring with explicit missing-data rules, a consistent ggplot2 theme,
and importers for Qualtrics and m-Path exports.
License: MIT + file LICENSE
Encoding: UTF-8
Imports: dplyr, ggplot2, rlang
Suggests: testthat (>= 3.0.0)
# Then make it installable by anyone, in one line from their console:
# remotes::install_github("yourname/labtools")
use_readme_rmd(); use_github_action("check-standard") # CI badge
pkgdown::build_site() # a documentation website
Portfolio artefact four
Three functions, documented, tested, checking clean, installable from GitHub, with a README showing one worked example. That is a complete artefact, buildable in an afternoon on top of day 94's work, and it is the single most distinctive item most psychology applicants could add to a CV. Do not wait until you have something impressive — labtools with a scoring function and a theme is genuinely useful to you and genuinely evidence to a panel.
Do this now · 30 minutes
Create a package, move three functions from day 94 into it with roxygen documentation and tests, and get check() to pass with zero errors, warnings and notes. Push it to GitHub with a README example. Artefact four exists.
Day 9650 minutes · artefacts three and eight
Preregistration, sharing, and the ethics of both
A preregistration is a time-stamped statement of your hypotheses and analysis plan, made before you see the data. It does not prevent exploration; it distinguishes exploration from confirmation, which is the distinction that makes a p-value meaningful at all. OSF hosts it free, and the same workflow handles data and materials sharing.
# The osfr package: manage OSF from R
library(osfr)
osf_auth(Sys.getenv("OSF_PAT")) # token from osf.io/settings/tokens
proj <- osf_create_project("Stress and mood in daily life")
comp <- osf_create_component(proj, "Analysis code")
osf_upload(comp, c("R/", "_targets.R", "renv.lock"), recurse = TRUE)
# Preregistrations are created on the website (they are time-stamped and
# frozen); the AsPredicted and PRP-QUANT templates are the usual choices.
Preregistration section
What goes in it
Hypotheses
Directional, numbered, each mapped to one analysis. "H2: the within-person stress-mood association is stronger at higher neuroticism"
Design and sample
Recruitment, inclusion criteria, stopping rule, and the target n with its justification
Power analysis
Volume 8's simulation script and power curve, with the effect size and its source
Measures
Every instrument, and how scales will be scored (reverse items, minimum valid responses)
Exclusions
Participant- and observation-level rules, stated in advance and in code where possible
Analysis plan
The model formula, the software, the inference method, the correction for multiplicity
Contingencies
What you will do if the model fails to converge or assumptions are violated
Exploratory work
Say that you will do it, and that it will be labelled as such
Two things preregistration is not
It is not a cage: deviations are allowed and expected, and the requirement is only that you report them and their reasons. A paper that says "we planned X, the model would not converge, so we did Y; results with both are in the supplement" is stronger than one where nobody can tell. And it is not a quality mark — a preregistered underpowered study with a poor design is still a poor study. It buys credibility for confirmatory claims, nothing more and nothing less.
# Sharing data responsibly: de-identify BEFORE upload, in code
d_share <- d |>
select(-name, -email, -birth_date, -free_text_response) |>
mutate(id = as.integer(factor(id)), # sequential, non-linkable IDs
age_band = cut(age, c(17, 25, 35, 50, 100)),
.keep = "unused") |>
select(-age)
readr::write_csv(d_share, "data/shared/esm_deidentified.csv")
# And a codebook, which is what makes shared data usable
library(codebook)
codebook_table(d_share) |> gt::gt() |> gt::gtsave("data/shared/codebook.html")
Artefacts three and eight
Three: a preregistration with a simulation-based power analysis attached — the clearest single signal that you understand design, and something you can discuss for ten minutes in an interview. Eight: a workshop handout. Take any two days of this course — "reproducible manuscripts in Quarto", "multilevel models for diary data" — write them up as a two-hour Quarto tutorial with worked exercises, and deposit it. Teaching materials demonstrate that you can TA, which is how a large share of PhD positions are funded.
Do this now · 20 minutes
Write a full preregistration for your next study using the AsPredicted or PRP-QUANT template, attach day 79's power simulation, and register it on OSF. Then de-identify one dataset in code, write its codebook, and deposit both. Artefacts three and eight exist.
Reference
Five of the eight, in twelve days
Nothing in this volume required a new statistical idea. That is the point — the portfolio is infrastructure, not cleverness.
Artefact 1 · reproducible manuscript
Quarto + papaja, every number from the model. Day 88.
Artefact 2 · public repository
Readable structure, honest README, real commit history. Day 91.
Artefact 3 · preregistration
Hypotheses, analysis plan, simulation-based power curve. Day 96.
Artefact 4 · your own package
Three documented, tested functions; check() clean. Day 95.
Artefact 6 · reusable template
targets pipeline plus renv, stripped of content. Day 93.
Artefact 8 · workshop handout
Two hours of this course, written up as a Quarto tutorial. Day 96.
Reference
Before any manuscript leaves your machine
Check
How
Clean session runs the whole analysis
Restart R, tar_make() or run scripts 00–04 in order
No number in the text was typed
Search the source for hard-coded digits in results sentences
Every figure and table is generated
They live in output/ and are gitignored, because they rebuild
Package versions recorded
renv.lock committed; session_info() in the supplement
Raw data untouched
data/raw is read-only and no script writes to it
Deviations from preregistration reported
A short subsection, not a footnote
Data availability statement present
What is shared, where, and what is withheld and why
Repository public and runnable
Clone into a fresh folder and follow your own README
The last volume, and a change of question: prediction rather than explanation. Six days on tidymodels, resampling, regularisation, forests, and how to report machine learning in psychology without overclaiming.