# code to generated and clean data from data_transformation_1:
library(truffle)
library(dplyr)
library(stringr)
set.seed(42)
dat_demographics_messy <- data.frame(id = 1:500) %>%
truffle::truffle_demographics() %>%
truffle::dirt_demographics()
dat_gender_tidier <- dat_demographics_messy %>%
# convert to lower case and remove non-letters
mutate(gender = str_to_lower(gender),
gender = str_remove_all(gender, "[^A-Za-z]")) %>%
# remove everything except numbers and decimal points from age.
# note that ages written as words (e.g., "thirty") become empty strings, and therefore NA below.
# these could be recovered (see the Strings and factors chapter), but we accept losing them for now.
mutate(age = str_remove_all(age, "[^0-9\\.]"),
age = na_if(age, ""),
age = as.numeric(age)) %>%
# tidy up cases.
# note that spaces were already removed above, so the misspellings here contain no spaces.
mutate(gender = case_when(gender %in% c("female", "femail", "femle", "femlae") ~ "female",
gender %in% c("male", "malee", "mal") ~ "male",
gender == "nonbinary" ~ "non-binary",
gender == "" ~ NA_character_, # if an empty character string, change to NA_character_
TRUE ~ gender)) # if none of the above apply, keep original value8 Data transformation III
This chapter covers “aggregating” functions. They aggregate results across rows to create a new data frame that does not contain any of the rows of the original. For example, to convert the ‘age’ column that contains participants’ ages to ‘age_mean’, which only has one row containing the average age.
8.1 Summarizing across rows
It is very common that we need to create summaries across rows. For example, to create the mean and standard deviation of a column like age. This can be done with summarize(). Remember: mutate() creates new columns or modifies the contents of existing columns, but does not change the number of rows. Whereas summarize() reduces a data frame down to one row.
library(dplyr)
library(tidyr)
library(janitor)
# mean
dat_gender_tidier %>%
summarize(mean_age = mean(age, na.rm = TRUE))| mean_age |
|---|
| 31.72093 |
dat_gender_tidier %>%
summarize(mean_age = mean(age))| mean_age |
|---|
| NA |
# SD
dat_gender_tidier %>%
summarize(sd_age = sd(age, na.rm = TRUE))| sd_age |
|---|
| 8.157242 |
# mean and SD with rounding, illustrating how multiple summarizes can be done in one function call
dat_gender_tidier %>%
summarize(mean_age = mean(age, na.rm = TRUE),
sd_age = sd(age, na.rm = TRUE)) %>%
mutate(mean_age = round_half_up(mean_age, digits = 2),
sd_age = round_half_up(sd_age, digits = 2))| mean_age | sd_age |
|---|---|
| 31.72 | 8.16 |
8.1.1 group_by()
Often, we don’t want to reduce a data frame down to a single row / summarize the whole dataset, but instead we want to create a summary for each (sub)group. For example
# illustrate use of group_by() and summarize()
dat_gender_tidier %>%
summarize(mean_age = mean(age, na.rm = TRUE),
sd_age = sd(age, na.rm = TRUE))| mean_age | sd_age |
|---|---|
| 31.72093 | 8.157242 |
dat_gender_tidier %>%
group_by(gender) %>%
summarize(mean_age = round_half_up(mean(age, na.rm = TRUE), digits = 2),
sd_age = round_half_up(sd(age, na.rm = TRUE), digits = 2))| gender | mean_age | sd_age |
|---|---|---|
| female | 32.30 | 8.40 |
| male | 30.89 | 7.69 |
| non-binary | 29.56 | 7.94 |
| NA | 30.39 | 7.23 |
dat <- dat_gender_tidier %>%
group_by(gender)
dat %>%
summarize(mean_age = round_half_up(mean(age, na.rm = TRUE), digits = 2),
sd_age = round_half_up(sd(age, na.rm = TRUE), digits = 2))| gender | mean_age | sd_age |
|---|---|---|
| female | 32.30 | 8.40 |
| male | 30.89 | 7.69 |
| non-binary | 29.56 | 7.94 |
| NA | 30.39 | 7.23 |
8.1.2 ungroup()
Grouping is an invisible property of a data frame or tibble. You can’t find a value inside it that tells you it’s grouped, but the grouping persists until you ungroup(). This can cause unexpected behavior, so watch out for it.
For example, imagine that earlier in your script you grouped the data by gender, and saved the result:
dat_grouped <- dat_gender_tidier %>%
group_by(gender)Later in your script, you want to find the three youngest participants in the whole sample:
dat_grouped %>%
slice_min(age, n = 3, with_ties = FALSE)| id | age | gender |
|---|---|---|
| 33 | 18 | female |
| 52 | 18 | female |
| 59 | 18 | female |
| 66 | 18 | male |
| 174 | 18 | male |
| 256 | 18 | male |
| 305 | 19 | non-binary |
| 278 | 21 | non-binary |
| 32 | 23 | non-binary |
| 428 | 18 | NA |
| 22 | 19 | NA |
| 412 | 19 | NA |
Instead of three rows, you get three rows per gender group, because the data frame is still grouped. slice_min(), like many other functions, operates within each group when the data is grouped. To get the intended result, ungroup() first:
dat_grouped %>%
ungroup() %>%
slice_min(age, n = 3, with_ties = FALSE)| id | age | gender |
|---|---|---|
| 33 | 18 | female |
| 52 | 18 | female |
| 59 | 18 | female |
When a grouped tibble is printed in the console, the first line of the output tells you about its groups, e.g., # Groups: gender [4]. However, this is easy to miss, and it isn’t shown at all when you print a table with kable(). You can also check with group_vars(), which returns the names of the grouping columns (or nothing if there are none):
dat_grouped %>%
group_vars()[1] "gender"
dat_grouped %>%
ungroup() %>%
group_vars()character(0)
Note also that summarize() removes only the last level of grouping. If you group_by(gender, condition) and then summarize(), the result is still grouped by gender. {dplyr} prints a message telling you this. You can avoid it with summarize(..., .groups = "drop"), which removes all grouping.
A good habit is to ungroup() at the end of any pipe that uses group_by().
8.1.2.1 The .by argument
Newer versions of {dplyr} offer an alternative that avoids this problem entirely: the .by argument. Instead of a separate group_by() step, you specify the grouping inside the function itself. The grouping only applies to that one function call, and the result is never grouped:
dat_gender_tidier %>%
summarize(mean_age = mean(age, na.rm = TRUE),
sd_age = sd(age, na.rm = TRUE),
.by = gender)| gender | mean_age | sd_age |
|---|---|---|
| female | 32.29508 | 8.401019 |
| male | 30.88710 | 7.694586 |
| NA | 30.39286 | 7.233355 |
| non-binary | 29.56250 | 7.941190 |
.by works with summarize(), mutate(), filter(), and the slice_() functions (where it is called by, without the dot). One small difference is that group_by() sorts the output by the groups, whereas .by keeps the groups in the order they first appear in the data.
You will see both styles in other people’s code. This book mostly uses group_by() because it makes the grouping step easy to see in a pipe, but .by is a good choice when you only need grouping for a single step.
8.1.3 n()
n() calculates the number of rows, i.e., the N. It can be useful in summarize.
# summarize n
dat_gender_tidier %>%
summarize(n_age = n())| n_age |
|---|
| 500 |
# summarize n per gender group
dat_gender_tidier %>%
group_by(gender) %>%
summarize(n_age = n())| gender | n_age |
|---|---|
| female | 323 |
| male | 131 |
| non-binary | 16 |
| NA | 30 |
# dat_gender_tidier %>%
# summarize(n_age = n())
#
# dat_gender_tidier %>%
# n()
#
# dat_gender_tidier %>%
# mutate(n_age = n())8.1.4 Realise that count() is just a wrapper function for summarize()
count() is just the combination of group_by() and summarize() and n(). They produce the same results as above.
# summarize n
dat_gender_tidier %>%
count()| n |
|---|
| 500 |
# summarize n per gender group
dat_gender_tidier %>%
count(gender)| gender | n |
|---|---|
| female | 323 |
| male | 131 |
| non-binary | 16 |
| NA | 30 |
8.2 distinct()
Distinct is a variation on count() that doesn’t return the counts, just the distinct values.
# summarize n per gender group
dat_gender_tidier %>%
count(gender)| gender | n |
|---|---|
| female | 323 |
| male | 131 |
| non-binary | 16 |
| NA | 30 |
# summarize n per gender group
dat_gender_tidier %>%
distinct(gender)| gender |
|---|
| female |
| male |
| NA |
| non-binary |
8.2.1 Handling duplicates with multiple-counts() and distinct()
It can be used to remove duplicates, but it’s important to remember that you don’t know which row is retained vs dropped.
# create duplicates that need cleaning
dat_gender_tidier_with_duplicates <- dat_gender_tidier %>%
dirt_duplicates(prop = 0.10)
# count rows
dat_gender_tidier_with_duplicates %>%
count()| n |
|---|
| 550 |
# number of rows for each participant id (first 10 shown)
dat_gender_tidier_with_duplicates %>%
count(id, name = "n_rows") %>%
head(10)| id | n_rows |
|---|---|
| 1 | 1 |
| 2 | 1 |
| 3 | 1 |
| 4 | 2 |
| 5 | 1 |
| 6 | 1 |
| 7 | 1 |
| 8 | 1 |
| 9 | 1 |
| 10 | 1 |
# count unique participant ids and the number of rows they have
dat_gender_tidier_with_duplicates %>%
count(id, name = "n_rows") %>%
count(n_rows, name = "n_ids_with_n_rows")| n_rows | n_ids_with_n_rows |
|---|---|
| 1 | 452 |
| 2 | 46 |
| 3 | 2 |
# count unique participant ids and the number of rows they have
dat_gender_tidier_duplicates_removed <- dat_gender_tidier_with_duplicates %>%
distinct(id, .keep_all = TRUE) # keep all other columns
dat_gender_tidier_duplicates_removed %>%
count(id, name = "n_rows") %>%
count(n_rows, name = "n_ids_with_n_rows")| n_rows | n_ids_with_n_rows |
|---|---|
| 1 | 500 |
# alternative, fewer assumptions
dat_gender_tidier_duplicates_removed2 <- dat_gender_tidier_with_duplicates %>%
distinct(id, age, gender)
dat_gender_tidier_duplicates_removed2 %>%
count(id, name = "n_rows") %>%
count(n_rows, name = "n_ids_with_n_rows")| n_rows | n_ids_with_n_rows |
|---|---|
| 1 | 500 |
8.2.2 More complex summarizations
Like mutate, the operation you do to summarize can also be more complex, such as finding the mean result of a logical test to calculate a proportion. For example, the proportion of participants who are less than 25 years old:
dat_gender_tidier %>%
summarize(proportion_less_than_25 = mean(age < 25, na.rm = TRUE)) %>%
mutate(percent_less_than_25 = round_half_up(proportion_less_than_25 * 100, 1))| proportion_less_than_25 | percent_less_than_25 |
|---|---|
| 0.2367865 | 23.7 |
dat_gender_tidier %>%
mutate(less_than_25 = if_else(age < 25, 1, 0)) %>%
summarize(percent_below_25 = mean(less_than_25, na.rm = TRUE)*100)| percent_below_25 |
|---|
| 23.67865 |
dat_gender_tidier %>%
mutate(less_than_25 = age < 25) %>%
head()| id | age | gender | less_than_25 |
|---|---|---|---|
| 1 | 39 | female | FALSE |
| 2 | 35 | female | FALSE |
| 3 | 37 | female | FALSE |
| 4 | 44 | female | FALSE |
| 5 | 28 | female | FALSE |
| 6 | NA | male | NA |
8.2.3 Practice using summarize()
Calculate the N, min, max, mean, and SD of age in dat_gender_tidier.
There are two ways to deal with the missing ages: remove the rows with missing ages first with drop_na(age), or use na.rm = TRUE in each function. Both are shown below, combined into one table with bind_rows() so they can be compared.
bind_rows(
dat_gender_tidier %>%
drop_na(age) %>%
summarize(N = n(),
min = min(age),
max = max(age),
mean = mean(age),
sd = sd(age)) %>%
round_half_up(2),
dat_gender_tidier %>%
summarize(N = n(),
min = min(age, na.rm = TRUE),
max = max(age, na.rm = TRUE),
mean = mean(age, na.rm = TRUE),
sd = sd(age, na.rm = TRUE)) %>%
round_half_up(2)
)| N | min | max | mean | sd |
|---|---|---|---|---|
| 473 | 18 | 45 | 31.72 | 8.16 |
| 500 | 18 | 45 | 31.72 | 8.16 |
The min, max, mean, and SD are identical, but the N is not. n() counts rows, not non-missing values, and it has no na.rm argument. In the second version, the participants with missing ages are still in the data frame, so they are counted in the N even though they don’t contribute to any of the other statistics. If you report this N alongside this mean, you would be overstating the sample size the mean is based on.
To count the non-missing values in a column without dropping rows, use sum(!is.na(age)). Remember from the previous chapter that TRUE counts as 1 when summed.
8.3 Grouped mutate()
group_by() doesn’t only change how summarize() works. It also changes how mutate() works: calculations are done separately within each group.
The key difference between the two is the same as without grouping: summarize() reduces each group down to one row, whereas mutate() keeps every row. If you use an aggregating function such as mean() inside a grouped mutate(), the group’s mean is calculated and then repeated in every row of that group:
dat_gender_tidier %>%
group_by(gender) %>%
mutate(mean_age_gender = mean(age, na.rm = TRUE),
# how much older or younger is each participant than the mean of their gender group?
age_relative_to_gender = age - mean_age_gender) %>%
ungroup() %>%
head(10)| id | age | gender | mean_age_gender | age_relative_to_gender |
|---|---|---|---|---|
| 1 | 39 | female | 32.29508 | 6.7049180 |
| 2 | 35 | female | 32.29508 | 2.7049180 |
| 3 | 37 | female | 32.29508 | 4.7049180 |
| 4 | 44 | female | 32.29508 | 11.7049180 |
| 5 | 28 | female | 32.29508 | -4.2950820 |
| 6 | NA | male | 30.88710 | NA |
| 7 | 27 | female | 32.29508 | -5.2950820 |
| 8 | 22 | male | 30.88710 | -8.8870968 |
| 9 | 30 | male | 30.88710 | -0.8870968 |
| 10 | 38 | male | 30.88710 | 7.1129032 |
Compare this to the summarize() version in the group_by() section above, which returned one row per gender.
8.3.1 Calculations within each participant
Grouped mutate() is especially useful for trial-level data, where each participant has many rows. In the previous chapter, we saw that lag(), lead(), cumsum(), and fill() don’t know anything about participants, so their results can leak from one participant into the next. group_by(id) solves this.
library(tibble) # for tribble()
dat_trials <- tribble(
~id, ~trial, ~rt, ~correct,
1, 1, 650, 1,
1, 2, 590, 1,
1, 3, 720, 0,
1, 4, 810, 1,
2, 1, 880, 1,
2, 2, 910, 0,
2, 3, 1020, 1,
2, 4, 870, 0
)
# without grouping: participant 2's first trial is given participant 1's last trial as its "previous" trial
dat_trials %>%
arrange(id, trial) %>%
mutate(correct_previous = lag(correct),
n_errors_so_far = cumsum(correct == 0))| id | trial | rt | correct | correct_previous | n_errors_so_far |
|---|---|---|---|---|---|
| 1 | 1 | 650 | 1 | NA | 0 |
| 1 | 2 | 590 | 1 | 1 | 0 |
| 1 | 3 | 720 | 0 | 1 | 1 |
| 1 | 4 | 810 | 1 | 0 | 1 |
| 2 | 1 | 880 | 1 | 1 | 1 |
| 2 | 2 | 910 | 0 | 1 | 2 |
| 2 | 3 | 1020 | 1 | 0 | 2 |
| 2 | 4 | 870 | 0 | 1 | 3 |
# with grouping: lag() and cumsum() start again for each participant
dat_trials %>%
arrange(id, trial) %>%
group_by(id) %>%
mutate(correct_previous = lag(correct),
n_errors_so_far = cumsum(correct == 0)) %>%
ungroup()| id | trial | rt | correct | correct_previous | n_errors_so_far |
|---|---|---|---|---|---|
| 1 | 1 | 650 | 1 | NA | 0 |
| 1 | 2 | 590 | 1 | 1 | 0 |
| 1 | 3 | 720 | 0 | 1 | 1 |
| 1 | 4 | 810 | 1 | 0 | 1 |
| 2 | 1 | 880 | 1 | NA | 0 |
| 2 | 2 | 910 | 0 | 1 | 1 |
| 2 | 3 | 1020 | 1 | 0 | 1 |
| 2 | 4 | 870 | 0 | 1 | 2 |
In the grouped version, each participant’s first trial correctly has NA for ‘correct_previous’, and each participant’s error count starts from 0.
Grouping also allows you to calculate values relative to each participant’s own data. For example, each participant’s proportion of correct responses, and each trial’s RT relative to that participant’s mean RT. This is often called “within-participant centering”, and the z-score version is often called “within-participant standardization”:
dat_trials %>%
group_by(id) %>%
mutate(proportion_correct = mean(correct),
rt_mean = mean(rt),
rt_centered = rt - rt_mean,
rt_z = (rt - rt_mean) / sd(rt)) %>%
ungroup()| id | trial | rt | correct | proportion_correct | rt_mean | rt_centered | rt_z |
|---|---|---|---|---|---|---|---|
| 1 | 1 | 650 | 1 | 0.75 | 692.5 | -42.5 | -0.4490300 |
| 1 | 2 | 590 | 1 | 0.75 | 692.5 | -102.5 | -1.0829546 |
| 1 | 3 | 720 | 0 | 0.75 | 692.5 | 27.5 | 0.2905488 |
| 1 | 4 | 810 | 1 | 0.75 | 692.5 | 117.5 | 1.2414358 |
| 2 | 1 | 880 | 1 | 0.50 | 920.0 | -40.0 | -0.5814019 |
| 2 | 2 | 910 | 0 | 0.50 | 920.0 | -10.0 | -0.1453505 |
| 2 | 3 | 1020 | 1 | 0.50 | 920.0 | 100.0 | 1.4535047 |
| 2 | 4 | 870 | 0 | 0.50 | 920.0 | -50.0 | -0.7267524 |
Participant 2 is slower overall than participant 1, but centering shows which of each participant’s trials were fast or slow for them.
8.3.2 Grouped fill()
fill() also respects grouping. In the example below, participant 2’s first block label was never recorded. Without grouping, fill() gives it participant 1’s last block label, “test”. With grouping, it correctly remains NA:
dat_blocks <- tribble(
~id, ~block, ~trial, ~rt,
1, "practice", 1, 812,
1, NA, 2, 1043,
1, "test", 1, 612,
1, NA, 2, 788,
2, NA, 1, 1150,
2, NA, 2, 1320,
2, "test", 1, 760,
2, NA, 2, 655
)
# without grouping
dat_blocks %>%
fill(block)| id | block | trial | rt |
|---|---|---|---|
| 1 | practice | 1 | 812 |
| 1 | practice | 2 | 1043 |
| 1 | test | 1 | 612 |
| 1 | test | 2 | 788 |
| 2 | test | 1 | 1150 |
| 2 | test | 2 | 1320 |
| 2 | test | 1 | 760 |
| 2 | test | 2 | 655 |
# with grouping
dat_blocks %>%
group_by(id) %>%
fill(block) %>%
ungroup()| id | block | trial | rt |
|---|---|---|---|
| 1 | practice | 1 | 812 |
| 1 | practice | 2 | 1043 |
| 1 | test | 1 | 612 |
| 1 | test | 2 | 788 |
| 2 | NA | 1 | 1150 |
| 2 | NA | 2 | 1320 |
| 2 | test | 1 | 760 |
| 2 | test | 2 | 655 |
Note that you can’t group by a column that itself still needs to be filled. If ‘id’ is also only written on the first row of each participant’s data, as in the previous chapter, fill ‘id’ first without grouping, and then group_by(id) to fill the other columns.
8.3.2.1 Exercise
Using ‘dat_trials’, write code that creates the following columns, calculated within each participant:
- ‘trial_check’ : the row number of each trial within each participant, using
row_number(). - ‘rt_change’ : the difference between the RT on each trial and the RT on the participant’s previous trial.
- ‘rt_slower_than_own_mean’ : TRUE if the trial’s RT is slower than the participant’s mean RT, and FALSE otherwise.
dat_trials %>%
arrange(id, trial) %>%
group_by(id) %>%
mutate(trial_check = row_number(),
rt_change = rt - lag(rt),
rt_slower_than_own_mean = rt > mean(rt)) %>%
ungroup()| id | trial | rt | correct | trial_check | rt_change | rt_slower_than_own_mean |
|---|---|---|---|---|---|---|
| 1 | 1 | 650 | 1 | 1 | NA | FALSE |
| 1 | 2 | 590 | 1 | 2 | -60 | FALSE |
| 1 | 3 | 720 | 0 | 3 | 130 | TRUE |
| 1 | 4 | 810 | 1 | 4 | 90 | TRUE |
| 2 | 1 | 880 | 1 | 1 | NA | FALSE |
| 2 | 2 | 910 | 0 | 2 | 30 | FALSE |
| 2 | 3 | 1020 | 1 | 3 | 110 | TRUE |
| 2 | 4 | 870 | 0 | 4 | -150 | FALSE |
8.3.2.2 Exercise
What is the difference between the outputs of these two chunks? How many rows does each return? Predict first, then run them to check.
dat_trials %>%
group_by(id) %>%
summarize(mean_rt = mean(rt))
dat_trials %>%
group_by(id) %>%
mutate(mean_rt = mean(rt))dat_trials %>%
group_by(id) %>%
summarize(mean_rt = mean(rt))| id | mean_rt |
|---|---|
| 1 | 692.5 |
| 2 | 920.0 |
dat_trials %>%
group_by(id) %>%
mutate(mean_rt = mean(rt))| id | trial | rt | correct | mean_rt |
|---|---|---|---|---|
| 1 | 1 | 650 | 1 | 692.5 |
| 1 | 2 | 590 | 1 | 692.5 |
| 1 | 3 | 720 | 0 | 692.5 |
| 1 | 4 | 810 | 1 | 692.5 |
| 2 | 1 | 880 | 1 | 920.0 |
| 2 | 2 | 910 | 0 | 920.0 |
| 2 | 3 | 1020 | 1 | 920.0 |
| 2 | 4 | 870 | 0 | 920.0 |
Both calculate the same mean RT for each participant. summarize() returns one row per participant (2 rows) containing only the grouping column and the summary. mutate() returns all of the original rows and columns (8 rows), with each participant’s mean repeated on each of their rows. Note also that the mutate() version is still grouped by id.
8.4 Grouped filter() and slice_()
filter() and the slice_() functions also work within groups. This lets you keep or drop rows based on calculations done within each group, e.g., based on a participant’s data as a whole rather than one trial at a time.
To illustrate this, we’ll use a larger simulated data set from a reaction time task. Each row is one trial. There are 20 participants who each completed up to 60 trials.
dat_rt %>%
head(10)| id | trial | correct | rt |
|---|---|---|---|
| 1 | 1 | 0 | 620 |
| 1 | 2 | 0 | 778 |
| 1 | 3 | 1 | 880 |
| 1 | 4 | 1 | 790 |
| 1 | 5 | 1 | 570 |
| 1 | 6 | 1 | 947 |
| 1 | 7 | 1 | 479 |
| 1 | 8 | 1 | 796 |
| 1 | 9 | 1 | 587 |
| 1 | 10 | 1 | 456 |
Some examples of grouped filters:
# keep only participants who completed all 60 trials
dat_rt %>%
group_by(id) %>%
filter(n() == 60) %>%
ungroup() %>%
distinct(id)| id |
|---|
| 1 |
| 2 |
| 3 |
| 4 |
| 5 |
| 6 |
| 7 |
| 8 |
| 9 |
| 10 |
| 11 |
| 13 |
| 14 |
| 15 |
| 16 |
| 17 |
| 18 |
| 19 |
| 20 |
# keep each participant's single fastest trial
dat_rt %>%
group_by(id) %>%
slice_min(rt, n = 1, with_ties = FALSE) %>%
ungroup()| id | trial | correct | rt |
|---|---|---|---|
| 1 | 47 | 1 | 450 |
| 2 | 38 | 1 | 437 |
| 3 | 35 | 1 | 432 |
| 4 | 41 | 1 | 415 |
| 5 | 8 | 1 | 511 |
| 6 | 6 | 0 | 393 |
| 7 | 15 | 1 | 108 |
| 8 | 24 | 1 | 411 |
| 9 | 41 | 1 | 496 |
| 10 | 20 | 1 | 435 |
| 11 | 25 | 1 | 421 |
| 12 | 15 | 1 | 453 |
| 13 | 27 | 1 | 416 |
| 14 | 10 | 1 | 466 |
| 15 | 57 | 1 | 112 |
| 16 | 50 | 1 | 412 |
| 17 | 40 | 1 | 473 |
| 18 | 52 | 1 | 444 |
| 19 | 54 | 1 | 439 |
| 20 | 33 | 1 | 434 |
# drop each participant's first 5 trials, e.g., to remove 'warm up' trials
dat_rt %>%
group_by(id) %>%
filter(row_number() > 5) %>%
ungroup() %>%
head(10)| id | trial | correct | rt |
|---|---|---|---|
| 1 | 6 | 1 | 947 |
| 1 | 7 | 1 | 479 |
| 1 | 8 | 1 | 796 |
| 1 | 9 | 1 | 587 |
| 1 | 10 | 1 | 456 |
| 1 | 11 | 1 | 460 |
| 1 | 12 | 1 | 690 |
| 1 | 13 | 0 | 1129 |
| 1 | 14 | 1 | 754 |
| 1 | 15 | 1 | 849 |
Note that the last example assumes that the rows are already in trial order within each participant. If you’re not sure, arrange(id, trial) first, or use filter(trial > 5), which doesn’t depend on row order at all.
8.4.1 Participant-level exclusions
A very common use of grouped filters is applying exclusion criteria that are defined at the participant level but applied to trial-level data. For example, a preregistration might state: “Participants who responded faster than 200 ms on more than 10% of trials will be excluded, as this suggests they were not paying attention”.
Remember from the previous chapter that the mean of a logical test is a proportion. So mean(rt < 200) is the proportion of a participant’s trials that were too fast:
# inspect the proportion of too-fast trials for each participant
dat_rt %>%
group_by(id) %>%
summarize(proportion_too_fast = mean(rt < 200)) %>%
arrange(desc(proportion_too_fast)) %>%
head()| id | proportion_too_fast |
|---|---|
| 7 | 0.25 |
| 15 | 0.20 |
| 1 | 0.00 |
| 2 | 0.00 |
| 3 | 0.00 |
| 4 | 0.00 |
# apply the exclusion
dat_rt_exclusions_applied <- dat_rt %>%
group_by(id) %>%
filter(mean(rt < 200) <= 0.10) %>%
ungroup()
# check how many participants remain
dat_rt_exclusions_applied %>%
distinct(id) %>%
count()| n |
|---|
| 18 |
This removes all the trials of participants 7 and 15, not just their too-fast trials. Compare this to an ungrouped filter(rt >= 200), which would remove the too-fast trials of every participant but keep all participants. These are different exclusion rules, and preregistrations often contain both: one applied to trials and one to participants. It’s important to implement each one at the correct level.
8.4.1.1 Exercise
Using ‘dat_rt’, write code that excludes participants whose proportion of correct responses is lower than 0.85. How many participants remain?
dat_rt %>%
group_by(id) %>%
filter(mean(correct) >= 0.85) %>%
ungroup() %>%
distinct(id) %>%
count()| n |
|---|
| 17 |
8.4.1.2 Exercise
Using ‘dat_rt’, write code that keeps each participant’s three slowest trials.
dat_rt %>%
group_by(id) %>%
slice_max(rt, n = 3, with_ties = FALSE) %>%
ungroup() %>%
# show only the first three participants' rows
head(9)| id | trial | correct | rt |
|---|---|---|---|
| 1 | 13 | 0 | 1129 |
| 1 | 6 | 1 | 947 |
| 1 | 35 | 1 | 907 |
| 2 | 52 | 1 | 989 |
| 2 | 5 | 1 | 947 |
| 2 | 57 | 1 | 902 |
| 3 | 22 | 1 | 984 |
| 3 | 51 | 1 | 943 |
| 3 | 2 | 1 | 924 |
8.5 Putting it together: post-error slowing
The previous chapter introduced “post-error slowing”: people tend to respond more slowly on the trial immediately after making an error. We now have all the tools needed to test this in ‘dat_rt’. This requires processing at three levels: trials, participants, and the sample.
dat_post_error_participants <- dat_rt %>%
# trial level: was the previous trial an error?
# this must be done *before* any trials are removed, otherwise lag() would refer to the wrong trial
arrange(id, trial) %>%
group_by(id) %>%
mutate(after_error = if_else(lag(correct) == 0, "post-error", "post-correct")) %>%
# participant level: exclude participants with more than 10% too-fast trials
filter(mean(rt < 200) <= 0.10) %>%
ungroup() %>%
# trial level: exclude implausible RTs, errors, and each participant's first trial (which has no previous trial)
filter(rt >= 200 & rt <= 3000) %>%
filter(correct == 1) %>%
drop_na(after_error) %>%
# participant level: mean RT for each participant on each type of trial
group_by(id, after_error) %>%
summarize(mean_rt = mean(rt),
n_trials = n(),
.groups = "drop")
dat_post_error_participants %>%
head(10)| id | after_error | mean_rt | n_trials |
|---|---|---|---|
| 1 | post-correct | 671.5405 | 37 |
| 1 | post-error | 714.4000 | 10 |
| 2 | post-correct | 653.0980 | 51 |
| 2 | post-error | 647.6667 | 3 |
| 3 | post-correct | 654.6875 | 48 |
| 3 | post-error | 694.0000 | 5 |
| 4 | post-correct | 668.0200 | 50 |
| 4 | post-error | 724.6000 | 5 |
| 5 | post-correct | 714.1961 | 51 |
| 5 | post-error | 709.2500 | 4 |
# sample level: mean and SD of the participants' mean RTs on each type of trial
dat_post_error_participants %>%
group_by(after_error) %>%
summarize(n_participants = n(),
mean_of_mean_rts = round_half_up(mean(mean_rt), 1),
sd_of_mean_rts = round_half_up(sd(mean_rt), 1))| after_error | n_participants | mean_of_mean_rts | sd_of_mean_rts |
|---|---|---|---|
| post-correct | 18 | 672.3 | 27.3 |
| post-error | 18 | 740.3 | 68.3 |
On average, participants responded more slowly on trials after an error than after a correct response.
A few things to note about this pipe:
- The order of the steps matters.
lag()is applied before any trials are excluded. If error trials were removed first,lag()would compare each trial to the previous remaining trial, and no trial could ever follow an error. - The participant-level exclusion is done with a grouped
filter(), while the trial-level exclusions are done with ungroupedfilter()s. - The ‘n_trials’ column shows that some participants’ mean post-error RTs are based on very few trials, because errors were rare. Means based on few trials are much noisier, and it is common to report this or to set a minimum number of trials.
- The sample-level results are a “mean of means”: each participant’s mean counts equally, regardless of how many trials it is based on. This is usually what you want, but it is not the same as the mean of all trials pooled together.
- To calculate each participant’s post-error slowing as a single difference score (post-error minus post-correct), the data would need to be reshaped so that the two means are in separate columns. This is covered in the Reshaping and pivots chapter.
8.5.0.1 Exercise
Modify the pipe above so that, in addition to excluding participants with more than 10% too-fast trials, it also excludes participants who did not complete all 60 trials. How many participants are included in the final results?
dat_rt %>%
arrange(id, trial) %>%
group_by(id) %>%
mutate(after_error = if_else(lag(correct) == 0, "post-error", "post-correct")) %>%
filter(mean(rt < 200) <= 0.10,
n() == 60) %>%
ungroup() %>%
filter(rt >= 200 & rt <= 3000) %>%
filter(correct == 1) %>%
drop_na(after_error) %>%
group_by(id, after_error) %>%
summarize(mean_rt = mean(rt),
n_trials = n(),
.groups = "drop") %>%
group_by(after_error) %>%
summarize(n_participants = n(),
mean_of_mean_rts = round_half_up(mean(mean_rt), 1),
sd_of_mean_rts = round_half_up(sd(mean_rt), 1))| after_error | n_participants | mean_of_mean_rts | sd_of_mean_rts |
|---|---|---|---|
| post-correct | 17 | 674.1 | 27.1 |
| post-error | 17 | 736.5 | 68.4 |
17 participants: participants 7 and 15 are excluded for too-fast responding, and participant 12 for not completing all trials. Note that the participant-level exclusions are applied before any trials are removed, so n() counts all the trials each participant completed.
8.6 Frequncies of sets of columns
Note that it is also possible to use count to obtain the frequencies of sets of unique values across columns, e.g., unique combinations of item and response.
# data_demographics_trimmed %>%
# count(item)
#
# data_demographics_trimmed %>%
# count(response)
#
# data_demographics_trimmed %>%
# count(item, response)It can be useful to arrange the output by the frequencies.
# data_demographics_trimmed %>%
# count(item, response) %>%
# arrange(desc(n)) # arrange in descending order8.6.1 summarize(across())
8.7 Exercises
8.7.1 mutate() vs. summarize()
What is the difference between mutate() and summarize()? If you use the wrong one, will you get the same answer? E.g., mutate(mean_age = mean(age, na.rm = TRUE)) vs. summarize(mean_age = mean(age, na.rm = TRUE)).
Both calculate the same mean age, but they return it differently. summarize() reduces the data frame to one row (or one row per group) containing only the summary (and any grouping columns). mutate() keeps every row and every column, and adds the mean as a new column, repeated on every row (or on every row of each group).
So the value is the same, but the shape of the result is very different. If you use mutate() when you meant summarize(), any later step that assumes one row per participant or group, e.g., calculating an N with n() or a mean of means, will be wrong.
8.7.2 group_by()
What does group_by() change about how summarize(), mutate(), filter(), and the slice_() functions work?
They all do their calculations separately within each group:
summarize()returns one row per group rather than one row in total.mutate()calculates aggregating functions (e.g.,mean()) and functions that use other rows (e.g.,lag(),cumsum()) within each group.filter()keeps or drops rows based on calculations within each group, e.g.,filter(n() == 60)keeps groups with exactly 60 rows.slice_()functions return rows from each group, e.g.,slice_head(n = 3)returns the first 3 rows of each group.
8.7.3 ungroup() and .by
Why is it important to ungroup()? What is an alternative that avoids the problem?
Grouping persists until it is removed, and it’s easy to miss because it isn’t visible in the data itself. Later functions in the pipe, or in later code using the saved data frame, will silently operate within groups, e.g., slice_min(age, n = 3) returns three rows per group rather than three rows in total.
Using the .by argument (e.g., summarize(..., .by = gender)) groups the data for that one function call only, so the result is never left grouped.
8.7.4 Counting
What is the difference between n() and sum(!is.na(age)) inside summarize()? Which should you report alongside mean(age, na.rm = TRUE)?
n() counts rows, including those where ‘age’ is NA. sum(!is.na(age)) counts only the non-missing values of ‘age’. The mean is calculated only from the non-missing values, so sum(!is.na(age)) is the N it is based on.
count(gender) is a wrapper for which other functions?
group_by(gender) %>% summarize(n = n()).
How can you calculate a proportion inside summarize(), e.g., the proportion of participants younger than 25?
Take the mean of a logical test, e.g., mean(age < 25, na.rm = TRUE). TRUE counts as 1 and FALSE as 0, so the mean is the proportion of TRUEs.
8.7.5 Duplicates
What is the difference between distinct() and distinct(id, .keep_all = TRUE) when removing duplicate rows? Which is safer?
distinct() only removes rows that are identical in every column. distinct(id, .keep_all = TRUE) keeps the first row for each id, regardless of whether that participant’s rows contain different values in other columns. You don’t control which row is kept, and conflicting values are silently discarded. distinct() is safer because any remaining duplicate ids reveal conflicts that you need to look at and resolve deliberately.
8.7.6 Exclusions
What is the difference between a trial-level and a participant-level exclusion rule? How is each implemented?
A trial-level rule removes individual trials, e.g., filter(rt >= 200) removes all trials faster than 200 ms, from any participant. A participant-level rule removes all of a participant’s trials based on a calculation across their trials, e.g., group_by(id) %>% filter(mean(rt < 200) <= 0.10) removes participants with more than 10% too-fast trials. Participant-level rules require grouping.
Why does the order of processing steps matter when using lag() or participant-level exclusion rules?
lag() refers to the previous row, so if trials are removed first, it refers to the previous remaining trial rather than the actual previous trial. Similarly, a participant-level rule like “more than 10% too-fast trials” should usually be calculated on all of the participant’s trials. If the too-fast trials are removed first, no participant would have any too-fast trials left, and no one would be excluded. In general, calculate values that depend on the full trial sequence, and apply participant-level rules, before removing any trials.
8.7.7 Interactive exercises
Complete the interactive exercises for:
8.7.8 Practice summarizing a demographics table
In your local version of this .qmd file:
- Use ‘dat_gender_tidier’ to create a table with one row per gender.
- Include the number of participants, the percentage of participants, and the mean and SD of age, rounded to 1 decimal place.
- Make sure the N reported for the mean age only includes participants with a known age.
- Arrange the table from the most to the least common gender.
Note: no solution provided here as many different ones are possible.
8.7.9 Practice processing trial-level data
In your local version of this .qmd file:
- Use ‘dat_rt’ to create a data frame called ‘dat_rt_participants’, with one row per participant.
- Calculate each participant’s number of trials, proportion of correct responses, proportion of too-fast trials (< 200 ms), and mean RT on correct trials only. Hint: all of these can be calculated in one
summarize()if you first usemutate()andif_else()to create a column containing the RT on correct trials andNAon error trials. - Add a column ‘exclude’ that is “exclude” if the participant has more than 10% too-fast trials or did not complete all 60 trials, and “include” otherwise.
- Then, using ‘dat_rt_participants’, calculate the number and percentage of participants who are excluded, and the mean and SD of the included participants’ mean RTs.
Note: no solution provided here to reduce the temptation to peek or copy-paste.