Seven days to the point where R stops being the obstacle.
This week has one job: make the mechanics automatic so that from day 8 you are thinking about your research question instead of about quotation marks. Nothing statistical happens here. That is deliberate — almost everyone who bounces off R bounces in week one, on file paths and data types, not on statistics.
7
days of 45–60 min
36
functions, all you need
10
errors, decoded
1
working pipeline at the end
How a day works
Read, run, break, tick
Each day is a short read, three to six code blocks you type out yourself, a table of the mistakes that specific topic produces, and one concrete task. Forty-five minutes is enough if you resist the urge to read ahead.
01
Type the code
Into a script, not the console. Typing is where the syntax lands in your fingers; copying skips that entirely.
02
Break it on purpose
Remove a comma. Misspell a column. Read what R says. You are building a dictionary of error messages, which is most of fluency.
03 · Compare your output
Every block prints what it really produced. If yours differs, that difference is the lesson — usually a data type, a package version, or a missing value.
04 · Tick the day
Only after the code ran on your machine. Reading a day does not count, and the progress bar is only useful if it tells the truth.
Day 145 minutes · install nothing statistical yet
The four panes, and the two things people confuse forever
R is the engine. RStudio is the dashboard you drive it from. You install both, you open RStudio, and you never open R directly again. Download R from cran.r-project.org first, then RStudio Desktop from posit.co/downloads — in that order, because RStudio looks for an engine that already exists.
RStudio shows four panes. Bottom-left is the console: R runs whatever you type there immediately and forgets it. Top-left is your script: a plain text file of instructions you keep. Top-right is the environment: the objects currently in memory. Bottom-right is files, plots and help. Press Cmd/Ctrl + Enter to send the current line of your script to the console. That single keystroke is the whole workflow.
The one habit that prevents a hundred problems
Work inside an RStudio Project. File → New Project → New Directory, call it r-course. R then treats that folder as the centre of the universe, so read_csv("data/raw.csv") works on any machine, including a reviewer's. If you ever find yourself typing setwd("C:/Users/..."), you have created a script that only works on your computer on this particular Tuesday.
# Anything after a hash is a comment. R ignores it; your future self does not.
2 + 2
sqrt(144)
(3 + 5) / 2
# Functions take arguments inside brackets
round(3.14159, digits = 2)
# R is case-sensitive and unforgiving about it
Round(3.14159, 2)#> [1] 4
#> [1] 12
#> [1] 4
#> [1] 3.14
#> Error in Round(3.14159, 2) : could not find function "Round"
That last error is worth provoking on purpose. Could not find function almost always means one of three things: a typo, wrong capitalisation, or a package you have not loaded. It never means R is broken.
# Installing DOWNLOADS the package to your machine. Do this once, ever.
install.packages("tidyverse")
# Loading makes it available in THIS session. Do this at the top of every script.
library(tidyverse)
# Where am I, and what version am I running? Put this in every project's notes.
getwd()
R.version.string#> ── Attaching core tidyverse packages ─────────────── tidyverse 2.0.0 ──
#> ✔ dplyr 1.1.4 ✔ readr 2.1.5
#> ✔ forcats 1.0.0 ✔ stringr 1.5.1
#> ✔ ggplot2 3.5.1 ✔ tibble 3.2.1
#> ✔ lubridate 1.9.3 ✔ tidyr 1.3.1
#> ── Conflicts ──────────────────────────────── tidyverse_conflicts() ──
#> ✖ dplyr::filter() masks stats::filter()
#> ✖ dplyr::lag() masks stats::lag()
#> [1] "/Users/you/r-course"
#> [1] "R version 4.4.1 (2024-06-14)"
The conflict message is not a warning about your code — it is R telling you that filter() now means dplyr's version rather than the time-series one in stats. That is what you want. When two packages genuinely clash, name the package explicitly: dplyr::filter().
The mistakes this day produces
What you see
What it means
What to do
could not find function "select"
The package is installed but not loaded in this session.
Put library(tidyverse) at the top of the script and rerun from line 1.
there is no package called 'dplyr'
Never installed, or installed under a different R version.
install.packages("dplyr") once, then library().
A lonely + in the console
R is waiting — you left a bracket or quote open.
Press Esc, then fix the unclosed pair.
Everything worked yesterday, nothing today
You relied on objects in memory rather than on the script.
Restart R (Cmd/Ctrl + Shift + F10) and run the script top to bottom. It must survive that.
Do this now · 15 minutes
Create the project r-course with folders data/, scripts/ and output/. In scripts/01-first-steps.R, write five comment lines describing your actual research question, then library(tidyverse) and sessionInfo(). Restart R and run the whole file. If it runs clean from a cold start, day 1 is done.
Day 245 minutes · the day that prevents the most silent errors
Objects, vectors, and the four types your data can be
Everything in R is an object with a name, and almost every object is built from vectors — ordered sets of values that are all the same type. That last clause is the whole day. When one character sneaks into a numeric column, R silently converts the entire column to text, your means stop working, and nothing anywhere says error.
# <- is assignment. Read it as "gets".
age <- 34
name <- "Tanuj"
# c() combines values into a vector. Same type throughout.
scores <- c(12, 15, 9, 21, 17)
groups <- c("control", "control", "cbt", "cbt", "cbt")
mean(scores)
length(scores)
scores * 2 # vectorised: no loop needed
scores[2] # second element — R counts from 1, not 0
scores[c(1, 5)] # first and fifth
scores[scores > 14] # logical subsetting: the basis of filter()#> [1] 14.8
#> [1] 5
#> [1] 24 30 18 42 34
#> [1] 15
#> [1] 12 17
#> [1] 15 21 17
Note the last line. scores > 14 produces FALSE TRUE FALSE TRUE TRUE, and putting that inside brackets keeps the TRUEs. Every filtering operation you will write for the rest of your career is this idea with better syntax.
mixed <- c(1, 2, "3") # one character value in the vector
class(mixed)
mean(mixed)
# The fix: find out why a character got in, then convert deliberately
as.numeric(mixed)
# The classic disaster: "missing" typed as a word in a spreadsheet
bdi <- c(18, 22, "missing", 31)
mean(as.numeric(bdi), na.rm = TRUE)#> [1] "character"
#> Error in mean.default(mixed) : argument is not numeric or logical: returning NA
#> [1] 1 2 3
#> Warning: NAs introduced by coercion
#> [1] 23.66667
That warning — NAs introduced by coercion — is a gift. It tells you exactly how many cells in your real dataset contain something you did not expect. Never suppress it; count it.
cond <- factor(c("cbt", "control", "cbt", "waitlist"))
levels(cond) # alphabetical by default — this becomes your reference group
# Set the reference level yourself. In a model, the first level is the baseline.
cond <- factor(cond, levels = c("control", "cbt", "waitlist"))
levels(cond)
# The single most destructive line in beginner R:
fake <- factor(c("10", "20", "30"))
as.numeric(fake) # gives level POSITIONS, not values
as.numeric(as.character(fake)) # correct#> [1] "cbt" "control" "waitlist"
#> [1] "control" "cbt" "waitlist"
#> [1] 1 2 3
#> [1] 10 20 30bdi <- c(18, 22, NA, 31)
mean(bdi) # NA, and that is correct behaviour
mean(bdi, na.rm = TRUE)
is.na(bdi)
sum(is.na(bdi)) # how many missing — report this in every paper
bdi == NA # never do this; NA is not a value to compare against#> [1] NA
#> [1] 23.66667
#> [1] FALSE FALSE TRUE FALSE
#> [1] 1
#> [1] NA NA NA NA
Do this now · 15 minutes
Build three vectors describing six imaginary participants of your own study: an ID (character), a group (factor with your real condition names, reference level set deliberately), and an outcome (numeric with one NA). Print the class of each, the mean of the outcome with and without na.rm, and the count of missings.
Day 345 minutes · the shape all of your data will take
Data frames: one row is one observation
A data frame is a set of equal-length vectors side by side: columns are variables, rows are observations. A tibble is the tidyverse's data frame — same thing, better printing, fewer surprises. Psychology data arrives as a data frame and leaves as a data frame; the middle of your career is spent reshaping it.
library(tidyverse)
d <- tibble(
id = c("p01", "p02", "p03", "p04"),
group = factor(c("cbt", "cbt", "control", "control")),
pre = c(28, 31, 27, 33),
post = c(19, 22, 26, 31)
)
d
dim(d) # rows, columns
names(d)
glimpse(d)#> # A tibble: 4 × 4
#> id group pre post
#> <chr> <fct> <dbl> <dbl>
#> 1 p01 cbt 28 19
#> 2 p02 cbt 31 22
#> 3 p03 control 27 26
#> 4 p04 control 33 31
#> [1] 4 4
#> [1] "id" "group" "pre" "post"
#> Rows: 4
#> Columns: 4
#> $ id <chr> "p01", "p02", "p03", "p04"
#> $ group <fct> cbt, cbt, control, control
#> $ pre <dbl> 28, 31, 27, 33
#> $ post <dbl> 19, 22, 26, 31
Read the type abbreviations under the column names every single time you import data: <dbl> numeric, <chr> text, <fct> factor, <date>, <lgl> logical. Ninety per cent of import problems are visible right there, in the first five seconds.
d$pre # a column, as a vector — the workhorse
d[["pre"]] # same thing, works when the name is in a variable
d[2, ] # second row, all columns
d[, c("id", "post")] # all rows, two columns
d[d$group == "cbt", ] # base-R filtering — you will replace this with filter()
# Built-in datasets are how you practise without owning data
data(mtcars)
head(mtcars, 3)
psych::describe(mtcars[, c("mpg", "hp")])#> [1] 28 31 27 33
#> [1] 28 31 27 33
#> # A tibble: 1 × 4
#> id group pre post
#> 1 p02 cbt 31 22
#> mpg hp
#> Mazda RX4 21.0 110
#> Mazda RX4 Wag 21.0 110
#> Datsun 710 22.8 93
#> vars n mean sd median trimmed mad min max range skew kurtosis
#> mpg 1 32 20.09 6.03 19.20 19.70 5.41 10.4 33.9 23.5 0.61 -0.37
#> hp 2 32 146.69 68.56 123.00 141.19 77.10 52.0 335.0 283.0 0.73 -0.14
Three datasets to practise on all week
psych::bfi — 2,800 people, 25 personality items plus age, gender, education. A real psychometrics dataset with real missingness. lavaan::HolzingerSwineford1939 — the canonical CFA dataset, nine cognitive tests in two schools. datasets::sleep — a tiny two-condition within-subject dataset, perfect for paired tests. All three are already on your machine once the packages are installed; none needs downloading.
Do this now · 15 minutes
Load psych::bfi. Report its dimensions, the class of every column via glimpse(), how many values are missing in item A1, and the mean age. Then write one sentence in a comment describing what a single row of this dataset represents. If you cannot write that sentence, you do not yet know the dataset.
Day 445 minutes · bring your own file today
Getting your data in without losing half of it
Importing is where most beginners quietly corrupt a dataset: a decimal comma read as text, an SPSS label turned into a number, an Excel sheet whose header sits on row three. Read the type line after every import and compare the row count to what you expect.
library(tidyverse); library(here)
# CSV — the format you should store your own data in
d <- read_csv(here("data", "study1_raw.csv"))
# Force the types you expect instead of hoping
d <- read_csv(here("data", "study1_raw.csv"),
col_types = cols(id = col_character(),
condition = col_factor(),
bdi_total = col_double()),
na = c("", "NA", "missing", "-99"))
# Excel — say which sheet, and where the header actually is
library(readxl)
x <- read_excel(here("data", "clinic.xlsx"), sheet = "wave1", skip = 2)
# SPSS — keeps variable and value labels
library(haven)
s <- read_sav(here("data", "survey.sav"))
s$condition <- as_factor(s$condition) # labels become factor levels
attr(s$bdi_total, "label") # the question wording, preserved#> Rows: 248 Columns: 14
#> ── Column specification ────────────────────────────────────────────
#> chr (3): id, condition, notes
#> dbl (10): age, bdi_total, gad_total, ...
#> date (1): session_date
#> [1] "BDI-II sum score at intake"
The na = argument earns its place immediately: survey platforms export -99, "missing", "n/a" and blank cells in the same file. Declaring them at import means you never discover a mean of −41 three weeks later.
nrow(d) # does this match your recruitment log?
sum(duplicated(d$id)) # duplicated IDs = a merge or export bug
colSums(is.na(d)) |> sort(decreasing = TRUE) |> head()
summary(d$bdi_total) # is the range possible for this scale?#> [1] 248
#> [1] 0
#> notes gad_total bdi_total age id condition
#> 31 7 4 0 0 0
#> Min. 1st Qu. Median Mean 3rd Qu. Max. NA's
#> 0.00 11.00 19.00 19.84 28.00 61.00 4
BDI-II runs 0–63, so a maximum of 61 is plausible and a maximum of 630 would mean a data-entry error or a decimal problem. Knowing your instruments' possible ranges is a statistical skill, not a clinical one.
The mistakes this day produces
What you see
What it means
What to do
cannot open file 'data.csv': No such file or directory
R is looking somewhere else, or the extension is hidden.
Open the project, use here("data", "file.csv"), and check list.files("data").
A numeric column imported as <chr>
Something non-numeric is in it — a comma decimal, a footnote, a "missing".
Find it: d |> filter(is.na(as.numeric(col))), then fix at source or via na =.
Row count is one less than expected
Your header row was treated as data, or vice versa.
Import a real file of your own — a past dataset, a course dataset, anything with more than fifty rows. Run the four-line sanity check. Write the results as comments: expected n, actual n, duplicated IDs, and the three columns with the most missing values. Save the script as scripts/02-import.R. Never edit the raw file again.
Day 550 minutes · the four verbs you will use daily for years
filter, select, mutate, arrange — and the pipe
dplyr gives each data operation a verb, and the pipe |> chains them so the code reads in the order you think: take the data, then keep these rows, then make this variable, then sort. Read |> aloud as "and then".
library(tidyverse)
bfi <- psych::bfi |> as_tibble()
bfi |>
filter(age >= 18, gender == 2) |> # commas mean AND
select(age, education, A1:A5) |> # a range of columns
mutate(agree_mean = rowMeans(across(A1:A5), na.rm = TRUE)) |>
arrange(desc(agree_mean)) |>
head(4)#> # A tibble: 4 × 8
#> age education A1 A2 A3 A4 A5 agree_mean
#> <int> <int> <int> <int> <int> <int> <int> <dbl>
#> 1 21 3 1 6 6 6 6 5
#> 2 30 2 1 6 6 6 6 5
#> 3 19 NA 2 6 6 6 6 5.2
#> 4 26 3 1 6 6 6 6 5d |> filter(bdi_total > 20) # greater than
d |> filter(condition == "cbt") # EQUALS is two equals signs
d |> filter(condition != "waitlist") # not equal
d |> filter(condition %in% c("cbt", "act")) # one of several
d |> filter(bdi_total > 20 & age < 30) # both
d |> filter(bdi_total > 20 | gad_total > 15) # either
d |> filter(!is.na(bdi_total)) # drop missings on one variable
d |> filter(between(age, 18, 65)) # inclusive range#> # each returns a tibble with the matching rows only
The single-equals mistake — filter(condition = "cbt") — produces a confusing error about arguments rather than a wrong answer, which is mercy. Remember: one equals sign assigns, two compare.
Notice that the four missing BDI scores appear as their own category in the count. Any recoding you do must account for NA deliberately — case_when without .default silently produces missings, and silent missings become a mysterious sample-size drop in your results table.
Do this now · 20 minutes
On your own imported data, write one piped chain that filters to your analysis sample, selects only the variables you actually need, creates at least one new variable with mutate, and one clinical or theoretical banding with case_when. Then count() the banding and check the total equals your filtered n.
Day 650 minutes · your first real results table
Groups, summaries, and the missing-value decision
Descriptives by group is Table 1 of every paper you will write. group_by declares the grouping, summarise collapses each group to one row, and the .by argument does both inline when you do not want the grouping to persist.
Report n_miss per group, always. A reviewer who sees means without missingness counts assumes the worst, and is often right to.
bfi |>
mutate(gender = factor(gender, levels = c(1, 2), labels = c("male", "female"))) |>
group_by(gender) |>
summarise(across(c(A1, A2, A3), \(x) mean(x, na.rm = TRUE)),
n = n(),
.groups = "drop")
# Proportions within groups
d |> count(condition, severity) |>
mutate(pct = round(100 * n / sum(n), 1), .by = condition)#> # A tibble: 3 × 5
#> gender A1 A2 A3 n
#> <fct> <dbl> <dbl> <dbl> <int>
#> 1 male 2.71 4.61 4.55 919
#> 2 female 2.29 4.87 4.85 1881
#> 3 NA NaN NaN NaN 0
#> # A tibble: 12 × 4
#> condition severity n pct
#> 1 cbt mild 19 22.6
#> 2 cbt minimal 16 19.0
#> 3 cbt moderate 31 36.9
na.rm = TRUE is a decision, not a convenience
Writing na.rm = TRUE means "compute this on whoever is left". That is often right for a descriptive and often wrong for an analysis, because different variables lose different people and your effective sample changes from row to row of your own results table. Volume 2 gives missingness a full day. For now: every time you type na.rm = TRUE, also print how many cases it dropped.
Do this now · 20 minutes
Produce a real Table 1 for your own data: n, missing count, mean, SD and median for two variables, split by your grouping variable. Round it. Then write the sentence you would put under it in a paper. Save as scripts/03-descriptives.R.
Day 760 minutes · the whole week, in one script
One script: raw data in, figure and table out
Today you build the shape every analysis you ever write will have. Nothing here is new; the point is that it runs from a cold start, end to end, without your hands. That property — and not any particular technique — is what makes research reproducible.
Restart R with Cmd/Ctrl + Shift + F10, which wipes every object from memory, then run the whole script with Cmd/Ctrl + Shift + S. If a table and a 300-dpi figure appear in output/ without a single error, you have a pipeline. From here on, every volume just replaces the middle section.
Do this now · 30 minutes
Write the same five-section script — setup, import, clean, summarise, export — for your own data, with your own scale and your own grouping variable. Print how many cases your cleaning removed. Restart R and run it cold. This script becomes the skeleton you reuse for the next 95 days.
Reference
Error decoder
R's error messages are terse but not arbitrary. These ten cover the overwhelming majority of week-one stoppages. Bookmark this section; you will return to it for months.
Message
Real cause
Fix
object 'bdi' not found
Typo, or you never ran the line that created it.
Check spelling, then rerun the script from the top.
could not find function "filter"
Package not loaded this session.
library(tidyverse) at the top of the file.
unexpected ')' in ...
Unbalanced brackets or a stray comma.
Let RStudio indent the block; the misalignment shows where.
argument is not numeric or logical
The column is character or factor, not numeric.
class(d$col), then convert deliberately.
NAs introduced by coercion
Non-numeric text inside a numeric column.
Find those cells before converting; do not suppress the warning.
undefined columns selected
Base-R bracket indexing with a name that does not exist.
names(d) and check capitalisation and spaces.
arguments imply differing number of rows
Vectors of unequal length in a data frame.
length() each one; usually a filtered vector snuck in.
subscript out of bounds
Asking for element 6 of a 5-element object.
Check length() or dim() first.
$ operator is invalid for atomic vectors
You used $ on a vector, not a data frame.
Use [ ], or check that the object is what you think.
there is no package called 'x'
Not installed, or installed under an older R.
install.packages("x"); after an R upgrade, reinstall.
When the message is not here: paste it into a search engine with the function name and without your variable names. Someone had the same problem in 2013 and the answer is on Stack Overflow. That is a legitimate research skill, not cheating.
Reference
The twenty functions that carry the first month
{{ c.fn }}
{{ c.what }}
Checkpoint
Eight questions before you move on
Answer without running anything. If you score below six, reread the day named in the feedback rather than pushing on — volume 2 assumes all of this.
{{ quizCounter }}
{{ quizScore }}
{{ quizQ }}
{{ quizFb }}
Next: volume 2
Wrangling is nine days on the work that consumes eighty per cent of real analysis time: getting messy, wide, multi-file survey data into the one shape every statistical model in R demands. It is also where your own data replaces the examples permanently.