8  Data transformation III

This chapter covers “aggregating” functions. They aggregate results across rows to create a new data frame that does not contain any of the rows of the original. For example, to convert the ‘age’ column that contains participants’ ages to ‘age_mean’, which only has one row containing the average age.

8.1 Summarizing across rows

It is very common that we need to create summaries across rows. For example, to create the mean and standard deviation of a column like age. This can be done with summarize(). Remember: mutate() creates new columns or modifies the contents of existing columns, but does not change the number of rows. Whereas summarize() reduces a data frame down to one row.

# code to generated and clean data from data_transformation_1:
library(truffle)
library(dplyr)
library(stringr)

set.seed(42)

dat_demographics_messy <- data.frame(id = 1:500) %>%
  truffle::truffle_demographics() %>%
  truffle::dirt_demographics()

dat_gender_tidier <- dat_demographics_messy %>%
  # convert to lower case and remove non-letters
  mutate(gender = str_to_lower(gender),
         gender = str_remove_all(gender, "[^A-Za-z]")) %>%
  # remove everything except numbers and decimal points from age.
  # note that ages written as words (e.g., "thirty") become empty strings, and therefore NA below. 
  # these could be recovered (see the Strings and factors chapter), but we accept losing them for now.
  mutate(age = str_remove_all(age, "[^0-9\\.]"),
         age = na_if(age, ""),
         age = as.numeric(age)) %>%
  # tidy up cases. 
  # note that spaces were already removed above, so the misspellings here contain no spaces.
  mutate(gender = case_when(gender %in% c("female", "femail", "femle", "femlae") ~ "female",
                            gender %in% c("male", "malee", "mal") ~ "male",
                            gender == "nonbinary" ~ "non-binary",
                            gender == "" ~ NA_character_, # if an empty character string, change to NA_character_
                            TRUE ~ gender)) # if none of the above apply, keep original value
library(dplyr)
library(tidyr)
library(janitor)

# mean
dat_gender_tidier %>%
  summarize(mean_age = mean(age, na.rm = TRUE))
mean_age
31.72093
dat_gender_tidier %>%
  summarize(mean_age = mean(age))
mean_age
NA
# SD
dat_gender_tidier %>%
  summarize(sd_age = sd(age, na.rm = TRUE))
sd_age
8.157242
# mean and SD with rounding, illustrating how multiple summarizes can be done in one function call
dat_gender_tidier %>%
  summarize(mean_age = mean(age, na.rm = TRUE),
            sd_age = sd(age, na.rm = TRUE)) %>%
  mutate(mean_age = round_half_up(mean_age, digits = 2),
         sd_age = round_half_up(sd_age, digits = 2))
mean_age sd_age
31.72 8.16

8.1.1 group_by()

Often, we don’t want to reduce a data frame down to a single row / summarize the whole dataset, but instead we want to create a summary for each (sub)group. For example

# illustrate use of group_by() and summarize()
dat_gender_tidier %>%
  summarize(mean_age = mean(age, na.rm = TRUE),
            sd_age = sd(age, na.rm = TRUE))
mean_age sd_age
31.72093 8.157242
dat_gender_tidier %>%
  group_by(gender) %>%
  summarize(mean_age = round_half_up(mean(age, na.rm = TRUE), digits = 2),
            sd_age = round_half_up(sd(age, na.rm = TRUE), digits = 2))
gender mean_age sd_age
female 32.30 8.40
male 30.89 7.69
non-binary 29.56 7.94
NA 30.39 7.23
dat <- dat_gender_tidier %>%
  group_by(gender)

dat  %>%
  summarize(mean_age = round_half_up(mean(age, na.rm = TRUE), digits = 2),
            sd_age = round_half_up(sd(age, na.rm = TRUE), digits = 2))
gender mean_age sd_age
female 32.30 8.40
male 30.89 7.69
non-binary 29.56 7.94
NA 30.39 7.23

8.1.2 ungroup()

Grouping is an invisible property of a data frame or tibble. You can’t find a value inside it that tells you it’s grouped, but the grouping persists until you ungroup(). This can cause unexpected behavior, so watch out for it.

For example, imagine that earlier in your script you grouped the data by gender, and saved the result:

dat_grouped <- dat_gender_tidier %>%
  group_by(gender)

Later in your script, you want to find the three youngest participants in the whole sample:

dat_grouped %>%
  slice_min(age, n = 3, with_ties = FALSE)
id age gender
33 18 female
52 18 female
59 18 female
66 18 male
174 18 male
256 18 male
305 19 non-binary
278 21 non-binary
32 23 non-binary
428 18 NA
22 19 NA
412 19 NA

Instead of three rows, you get three rows per gender group, because the data frame is still grouped. slice_min(), like many other functions, operates within each group when the data is grouped. To get the intended result, ungroup() first:

dat_grouped %>%
  ungroup() %>%
  slice_min(age, n = 3, with_ties = FALSE)
id age gender
33 18 female
52 18 female
59 18 female

When a grouped tibble is printed in the console, the first line of the output tells you about its groups, e.g., # Groups: gender [4]. However, this is easy to miss, and it isn’t shown at all when you print a table with kable(). You can also check with group_vars(), which returns the names of the grouping columns (or nothing if there are none):

dat_grouped %>%
  group_vars()
[1] "gender"
dat_grouped %>%
  ungroup() %>%
  group_vars()
character(0)

Note also that summarize() removes only the last level of grouping. If you group_by(gender, condition) and then summarize(), the result is still grouped by gender. {dplyr} prints a message telling you this. You can avoid it with summarize(..., .groups = "drop"), which removes all grouping.

A good habit is to ungroup() at the end of any pipe that uses group_by().

8.1.2.1 The .by argument

Newer versions of {dplyr} offer an alternative that avoids this problem entirely: the .by argument. Instead of a separate group_by() step, you specify the grouping inside the function itself. The grouping only applies to that one function call, and the result is never grouped:

dat_gender_tidier %>%
  summarize(mean_age = mean(age, na.rm = TRUE),
            sd_age = sd(age, na.rm = TRUE),
            .by = gender)
gender mean_age sd_age
female 32.29508 8.401019
male 30.88710 7.694586
NA 30.39286 7.233355
non-binary 29.56250 7.941190

.by works with summarize(), mutate(), filter(), and the slice_() functions (where it is called by, without the dot). One small difference is that group_by() sorts the output by the groups, whereas .by keeps the groups in the order they first appear in the data.

You will see both styles in other people’s code. This book mostly uses group_by() because it makes the grouping step easy to see in a pipe, but .by is a good choice when you only need grouping for a single step.

8.1.3 n()

n() calculates the number of rows, i.e., the N. It can be useful in summarize.

# summarize n
dat_gender_tidier %>%
  summarize(n_age = n())
n_age
500
# summarize n per gender group
dat_gender_tidier %>%
  group_by(gender) %>%
  summarize(n_age = n())
gender n_age
female 323
male 131
non-binary 16
NA 30
# dat_gender_tidier %>%
#   summarize(n_age = n())
# 
# dat_gender_tidier %>%
#   n()
# 
# dat_gender_tidier %>%
#   mutate(n_age = n())

8.1.4 Realise that count() is just a wrapper function for summarize()

count() is just the combination of group_by() and summarize() and n(). They produce the same results as above.

# summarize n
dat_gender_tidier %>%
  count()
n
500
# summarize n per gender group
dat_gender_tidier %>%
  count(gender)
gender n
female 323
male 131
non-binary 16
NA 30

8.2 distinct()

Distinct is a variation on count() that doesn’t return the counts, just the distinct values.

# summarize n per gender group
dat_gender_tidier %>%
  count(gender)
gender n
female 323
male 131
non-binary 16
NA 30
# summarize n per gender group
dat_gender_tidier %>%
  distinct(gender)
gender
female
male
NA
non-binary

8.2.1 Handling duplicates with multiple-counts() and distinct()

It can be used to remove duplicates, but it’s important to remember that you don’t know which row is retained vs dropped.

# create duplicates that need cleaning
dat_gender_tidier_with_duplicates <- dat_gender_tidier %>%
  dirt_duplicates(prop = 0.10)
  
# count rows
dat_gender_tidier_with_duplicates %>%
  count()
n
550
# number of rows for each participant id (first 10 shown)
dat_gender_tidier_with_duplicates %>%
  count(id, name = "n_rows") %>%
  head(10)
id n_rows
1 1
2 1
3 1
4 2
5 1
6 1
7 1
8 1
9 1
10 1
# count unique participant ids and the number of rows they have
dat_gender_tidier_with_duplicates %>%
  count(id, name = "n_rows") %>%
  count(n_rows, name = "n_ids_with_n_rows")
n_rows n_ids_with_n_rows
1 452
2 46
3 2
# count unique participant ids and the number of rows they have
dat_gender_tidier_duplicates_removed <- dat_gender_tidier_with_duplicates %>%
  distinct(id, .keep_all = TRUE)  # keep all other columns

dat_gender_tidier_duplicates_removed %>% 
  count(id, name = "n_rows") %>%
  count(n_rows, name = "n_ids_with_n_rows")
n_rows n_ids_with_n_rows
1 500
# alternative, fewer assumptions
dat_gender_tidier_duplicates_removed2 <- dat_gender_tidier_with_duplicates %>%
  distinct(id, age, gender)

dat_gender_tidier_duplicates_removed2 %>% 
  count(id, name = "n_rows") %>%
  count(n_rows, name = "n_ids_with_n_rows")
n_rows n_ids_with_n_rows
1 500

8.2.2 More complex summarizations

Like mutate, the operation you do to summarize can also be more complex, such as finding the mean result of a logical test to calculate a proportion. For example, the proportion of participants who are less than 25 years old:

dat_gender_tidier %>%
  summarize(proportion_less_than_25 = mean(age < 25, na.rm = TRUE)) %>%
  mutate(percent_less_than_25 = round_half_up(proportion_less_than_25 * 100, 1))
proportion_less_than_25 percent_less_than_25
0.2367865 23.7
dat_gender_tidier %>%
  mutate(less_than_25 = if_else(age < 25, 1, 0)) %>%
  summarize(percent_below_25 = mean(less_than_25, na.rm = TRUE)*100)
percent_below_25
23.67865
dat_gender_tidier %>%
  mutate(less_than_25 = age < 25) %>%
  head()
id age gender less_than_25
1 39 female FALSE
2 35 female FALSE
3 37 female FALSE
4 44 female FALSE
5 28 female FALSE
6 NA male NA

8.2.3 Practice using summarize()

Calculate the N, min, max, mean, and SD of age in dat_gender_tidier.

There are two ways to deal with the missing ages: remove the rows with missing ages first with drop_na(age), or use na.rm = TRUE in each function. Both are shown below, combined into one table with bind_rows() so they can be compared.

bind_rows(
dat_gender_tidier %>%
  drop_na(age) %>%
  summarize(N = n(),
            min = min(age),
            max = max(age),
            mean = mean(age),
            sd = sd(age)) %>%
  round_half_up(2),

dat_gender_tidier %>%
  summarize(N = n(),
            min = min(age, na.rm = TRUE),
            max = max(age, na.rm = TRUE),
            mean = mean(age, na.rm = TRUE),
            sd = sd(age, na.rm = TRUE)) %>%
  round_half_up(2)
)
N min max mean sd
473 18 45 31.72 8.16
500 18 45 31.72 8.16

The min, max, mean, and SD are identical, but the N is not. n() counts rows, not non-missing values, and it has no na.rm argument. In the second version, the participants with missing ages are still in the data frame, so they are counted in the N even though they don’t contribute to any of the other statistics. If you report this N alongside this mean, you would be overstating the sample size the mean is based on.

To count the non-missing values in a column without dropping rows, use sum(!is.na(age)). Remember from the previous chapter that TRUE counts as 1 when summed.

8.3 Grouped mutate()

group_by() doesn’t only change how summarize() works. It also changes how mutate() works: calculations are done separately within each group.

The key difference between the two is the same as without grouping: summarize() reduces each group down to one row, whereas mutate() keeps every row. If you use an aggregating function such as mean() inside a grouped mutate(), the group’s mean is calculated and then repeated in every row of that group:

dat_gender_tidier %>%
  group_by(gender) %>%
  mutate(mean_age_gender = mean(age, na.rm = TRUE),
         # how much older or younger is each participant than the mean of their gender group?
         age_relative_to_gender = age - mean_age_gender) %>%
  ungroup() %>%
  head(10)
id age gender mean_age_gender age_relative_to_gender
1 39 female 32.29508 6.7049180
2 35 female 32.29508 2.7049180
3 37 female 32.29508 4.7049180
4 44 female 32.29508 11.7049180
5 28 female 32.29508 -4.2950820
6 NA male 30.88710 NA
7 27 female 32.29508 -5.2950820
8 22 male 30.88710 -8.8870968
9 30 male 30.88710 -0.8870968
10 38 male 30.88710 7.1129032

Compare this to the summarize() version in the group_by() section above, which returned one row per gender.

8.3.1 Calculations within each participant

Grouped mutate() is especially useful for trial-level data, where each participant has many rows. In the previous chapter, we saw that lag(), lead(), cumsum(), and fill() don’t know anything about participants, so their results can leak from one participant into the next. group_by(id) solves this.

library(tibble) # for tribble()

dat_trials <- tribble(
  ~id, ~trial, ~rt,  ~correct,
  1,   1,      650,  1,
  1,   2,      590,  1,
  1,   3,      720,  0,
  1,   4,      810,  1,
  2,   1,      880,  1,
  2,   2,      910,  0,
  2,   3,      1020, 1,
  2,   4,      870,  0
)

# without grouping: participant 2's first trial is given participant 1's last trial as its "previous" trial
dat_trials %>%
  arrange(id, trial) %>%
  mutate(correct_previous = lag(correct),
         n_errors_so_far = cumsum(correct == 0))
id trial rt correct correct_previous n_errors_so_far
1 1 650 1 NA 0
1 2 590 1 1 0
1 3 720 0 1 1
1 4 810 1 0 1
2 1 880 1 1 1
2 2 910 0 1 2
2 3 1020 1 0 2
2 4 870 0 1 3
# with grouping: lag() and cumsum() start again for each participant
dat_trials %>%
  arrange(id, trial) %>%
  group_by(id) %>%
  mutate(correct_previous = lag(correct),
         n_errors_so_far = cumsum(correct == 0)) %>%
  ungroup()
id trial rt correct correct_previous n_errors_so_far
1 1 650 1 NA 0
1 2 590 1 1 0
1 3 720 0 1 1
1 4 810 1 0 1
2 1 880 1 NA 0
2 2 910 0 1 1
2 3 1020 1 0 1
2 4 870 0 1 2

In the grouped version, each participant’s first trial correctly has NA for ‘correct_previous’, and each participant’s error count starts from 0.

Grouping also allows you to calculate values relative to each participant’s own data. For example, each participant’s proportion of correct responses, and each trial’s RT relative to that participant’s mean RT. This is often called “within-participant centering”, and the z-score version is often called “within-participant standardization”:

dat_trials %>%
  group_by(id) %>%
  mutate(proportion_correct = mean(correct),
         rt_mean = mean(rt),
         rt_centered = rt - rt_mean,
         rt_z = (rt - rt_mean) / sd(rt)) %>%
  ungroup()
id trial rt correct proportion_correct rt_mean rt_centered rt_z
1 1 650 1 0.75 692.5 -42.5 -0.4490300
1 2 590 1 0.75 692.5 -102.5 -1.0829546
1 3 720 0 0.75 692.5 27.5 0.2905488
1 4 810 1 0.75 692.5 117.5 1.2414358
2 1 880 1 0.50 920.0 -40.0 -0.5814019
2 2 910 0 0.50 920.0 -10.0 -0.1453505
2 3 1020 1 0.50 920.0 100.0 1.4535047
2 4 870 0 0.50 920.0 -50.0 -0.7267524

Participant 2 is slower overall than participant 1, but centering shows which of each participant’s trials were fast or slow for them.

8.3.2 Grouped fill()

fill() also respects grouping. In the example below, participant 2’s first block label was never recorded. Without grouping, fill() gives it participant 1’s last block label, “test”. With grouping, it correctly remains NA:

dat_blocks <- tribble(
  ~id, ~block,     ~trial, ~rt,
  1,   "practice", 1,      812,
  1,   NA,         2,      1043,
  1,   "test",     1,      612,
  1,   NA,         2,      788,
  2,   NA,         1,      1150,
  2,   NA,         2,      1320,
  2,   "test",     1,      760,
  2,   NA,         2,      655
)

# without grouping
dat_blocks %>%
  fill(block)
id block trial rt
1 practice 1 812
1 practice 2 1043
1 test 1 612
1 test 2 788
2 test 1 1150
2 test 2 1320
2 test 1 760
2 test 2 655
# with grouping
dat_blocks %>%
  group_by(id) %>%
  fill(block) %>%
  ungroup()
id block trial rt
1 practice 1 812
1 practice 2 1043
1 test 1 612
1 test 2 788
2 NA 1 1150
2 NA 2 1320
2 test 1 760
2 test 2 655

Note that you can’t group by a column that itself still needs to be filled. If ‘id’ is also only written on the first row of each participant’s data, as in the previous chapter, fill ‘id’ first without grouping, and then group_by(id) to fill the other columns.

8.3.2.1 Exercise

Using ‘dat_trials’, write code that creates the following columns, calculated within each participant:

  • ‘trial_check’ : the row number of each trial within each participant, using row_number().
  • ‘rt_change’ : the difference between the RT on each trial and the RT on the participant’s previous trial.
  • ‘rt_slower_than_own_mean’ : TRUE if the trial’s RT is slower than the participant’s mean RT, and FALSE otherwise.
dat_trials %>%
  arrange(id, trial) %>%
  group_by(id) %>%
  mutate(trial_check = row_number(),
         rt_change = rt - lag(rt),
         rt_slower_than_own_mean = rt > mean(rt)) %>%
  ungroup()
id trial rt correct trial_check rt_change rt_slower_than_own_mean
1 1 650 1 1 NA FALSE
1 2 590 1 2 -60 FALSE
1 3 720 0 3 130 TRUE
1 4 810 1 4 90 TRUE
2 1 880 1 1 NA FALSE
2 2 910 0 2 30 FALSE
2 3 1020 1 3 110 TRUE
2 4 870 0 4 -150 FALSE

8.3.2.2 Exercise

What is the difference between the outputs of these two chunks? How many rows does each return? Predict first, then run them to check.

dat_trials %>%
  group_by(id) %>%
  summarize(mean_rt = mean(rt))

dat_trials %>%
  group_by(id) %>%
  mutate(mean_rt = mean(rt))
dat_trials %>%
  group_by(id) %>%
  summarize(mean_rt = mean(rt))
id mean_rt
1 692.5
2 920.0
dat_trials %>%
  group_by(id) %>%
  mutate(mean_rt = mean(rt))
id trial rt correct mean_rt
1 1 650 1 692.5
1 2 590 1 692.5
1 3 720 0 692.5
1 4 810 1 692.5
2 1 880 1 920.0
2 2 910 0 920.0
2 3 1020 1 920.0
2 4 870 0 920.0

Both calculate the same mean RT for each participant. summarize() returns one row per participant (2 rows) containing only the grouping column and the summary. mutate() returns all of the original rows and columns (8 rows), with each participant’s mean repeated on each of their rows. Note also that the mutate() version is still grouped by id.

8.4 Grouped filter() and slice_()

filter() and the slice_() functions also work within groups. This lets you keep or drop rows based on calculations done within each group, e.g., based on a participant’s data as a whole rather than one trial at a time.

To illustrate this, we’ll use a larger simulated data set from a reaction time task. Each row is one trial. There are 20 participants who each completed up to 60 trials.

dat_rt %>%
  head(10)
id trial correct rt
1 1 0 620
1 2 0 778
1 3 1 880
1 4 1 790
1 5 1 570
1 6 1 947
1 7 1 479
1 8 1 796
1 9 1 587
1 10 1 456

Some examples of grouped filters:

# keep only participants who completed all 60 trials
dat_rt %>%
  group_by(id) %>%
  filter(n() == 60) %>%
  ungroup() %>%
  distinct(id)
id
1
2
3
4
5
6
7
8
9
10
11
13
14
15
16
17
18
19
20
# keep each participant's single fastest trial
dat_rt %>%
  group_by(id) %>%
  slice_min(rt, n = 1, with_ties = FALSE) %>%
  ungroup()
id trial correct rt
1 47 1 450
2 38 1 437
3 35 1 432
4 41 1 415
5 8 1 511
6 6 0 393
7 15 1 108
8 24 1 411
9 41 1 496
10 20 1 435
11 25 1 421
12 15 1 453
13 27 1 416
14 10 1 466
15 57 1 112
16 50 1 412
17 40 1 473
18 52 1 444
19 54 1 439
20 33 1 434
# drop each participant's first 5 trials, e.g., to remove 'warm up' trials
dat_rt %>%
  group_by(id) %>%
  filter(row_number() > 5) %>%
  ungroup() %>%
  head(10)
id trial correct rt
1 6 1 947
1 7 1 479
1 8 1 796
1 9 1 587
1 10 1 456
1 11 1 460
1 12 1 690
1 13 0 1129
1 14 1 754
1 15 1 849

Note that the last example assumes that the rows are already in trial order within each participant. If you’re not sure, arrange(id, trial) first, or use filter(trial > 5), which doesn’t depend on row order at all.

8.4.1 Participant-level exclusions

A very common use of grouped filters is applying exclusion criteria that are defined at the participant level but applied to trial-level data. For example, a preregistration might state: “Participants who responded faster than 200 ms on more than 10% of trials will be excluded, as this suggests they were not paying attention”.

Remember from the previous chapter that the mean of a logical test is a proportion. So mean(rt < 200) is the proportion of a participant’s trials that were too fast:

# inspect the proportion of too-fast trials for each participant
dat_rt %>%
  group_by(id) %>%
  summarize(proportion_too_fast = mean(rt < 200)) %>%
  arrange(desc(proportion_too_fast)) %>%
  head()
id proportion_too_fast
7 0.25
15 0.20
1 0.00
2 0.00
3 0.00
4 0.00
# apply the exclusion
dat_rt_exclusions_applied <- dat_rt %>%
  group_by(id) %>%
  filter(mean(rt < 200) <= 0.10) %>%
  ungroup()

# check how many participants remain
dat_rt_exclusions_applied %>%
  distinct(id) %>%
  count()
n
18

This removes all the trials of participants 7 and 15, not just their too-fast trials. Compare this to an ungrouped filter(rt >= 200), which would remove the too-fast trials of every participant but keep all participants. These are different exclusion rules, and preregistrations often contain both: one applied to trials and one to participants. It’s important to implement each one at the correct level.

8.4.1.1 Exercise

Using ‘dat_rt’, write code that excludes participants whose proportion of correct responses is lower than 0.85. How many participants remain?

dat_rt %>%
  group_by(id) %>%
  filter(mean(correct) >= 0.85) %>%
  ungroup() %>%
  distinct(id) %>%
  count()
n
17

8.4.1.2 Exercise

Using ‘dat_rt’, write code that keeps each participant’s three slowest trials.

dat_rt %>%
  group_by(id) %>%
  slice_max(rt, n = 3, with_ties = FALSE) %>%
  ungroup() %>%
  # show only the first three participants' rows
  head(9)
id trial correct rt
1 13 0 1129
1 6 1 947
1 35 1 907
2 52 1 989
2 5 1 947
2 57 1 902
3 22 1 984
3 51 1 943
3 2 1 924

8.5 Putting it together: post-error slowing

The previous chapter introduced “post-error slowing”: people tend to respond more slowly on the trial immediately after making an error. We now have all the tools needed to test this in ‘dat_rt’. This requires processing at three levels: trials, participants, and the sample.

dat_post_error_participants <- dat_rt %>%
  # trial level: was the previous trial an error? 
  # this must be done *before* any trials are removed, otherwise lag() would refer to the wrong trial
  arrange(id, trial) %>%
  group_by(id) %>%
  mutate(after_error = if_else(lag(correct) == 0, "post-error", "post-correct")) %>%
  # participant level: exclude participants with more than 10% too-fast trials
  filter(mean(rt < 200) <= 0.10) %>%
  ungroup() %>%
  # trial level: exclude implausible RTs, errors, and each participant's first trial (which has no previous trial)
  filter(rt >= 200 & rt <= 3000) %>%
  filter(correct == 1) %>%
  drop_na(after_error) %>%
  # participant level: mean RT for each participant on each type of trial
  group_by(id, after_error) %>%
  summarize(mean_rt = mean(rt),
            n_trials = n(),
            .groups = "drop")

dat_post_error_participants %>%
  head(10)
id after_error mean_rt n_trials
1 post-correct 671.5405 37
1 post-error 714.4000 10
2 post-correct 653.0980 51
2 post-error 647.6667 3
3 post-correct 654.6875 48
3 post-error 694.0000 5
4 post-correct 668.0200 50
4 post-error 724.6000 5
5 post-correct 714.1961 51
5 post-error 709.2500 4
# sample level: mean and SD of the participants' mean RTs on each type of trial
dat_post_error_participants %>%
  group_by(after_error) %>%
  summarize(n_participants = n(),
            mean_of_mean_rts = round_half_up(mean(mean_rt), 1),
            sd_of_mean_rts = round_half_up(sd(mean_rt), 1))
after_error n_participants mean_of_mean_rts sd_of_mean_rts
post-correct 18 672.3 27.3
post-error 18 740.3 68.3

On average, participants responded more slowly on trials after an error than after a correct response.

A few things to note about this pipe:

  • The order of the steps matters. lag() is applied before any trials are excluded. If error trials were removed first, lag() would compare each trial to the previous remaining trial, and no trial could ever follow an error.
  • The participant-level exclusion is done with a grouped filter(), while the trial-level exclusions are done with ungrouped filter()s.
  • The ‘n_trials’ column shows that some participants’ mean post-error RTs are based on very few trials, because errors were rare. Means based on few trials are much noisier, and it is common to report this or to set a minimum number of trials.
  • The sample-level results are a “mean of means”: each participant’s mean counts equally, regardless of how many trials it is based on. This is usually what you want, but it is not the same as the mean of all trials pooled together.
  • To calculate each participant’s post-error slowing as a single difference score (post-error minus post-correct), the data would need to be reshaped so that the two means are in separate columns. This is covered in the Reshaping and pivots chapter.

8.5.0.1 Exercise

Modify the pipe above so that, in addition to excluding participants with more than 10% too-fast trials, it also excludes participants who did not complete all 60 trials. How many participants are included in the final results?

dat_rt %>%
  arrange(id, trial) %>%
  group_by(id) %>%
  mutate(after_error = if_else(lag(correct) == 0, "post-error", "post-correct")) %>%
  filter(mean(rt < 200) <= 0.10,
         n() == 60) %>%
  ungroup() %>%
  filter(rt >= 200 & rt <= 3000) %>%
  filter(correct == 1) %>%
  drop_na(after_error) %>%
  group_by(id, after_error) %>%
  summarize(mean_rt = mean(rt),
            n_trials = n(),
            .groups = "drop") %>%
  group_by(after_error) %>%
  summarize(n_participants = n(),
            mean_of_mean_rts = round_half_up(mean(mean_rt), 1),
            sd_of_mean_rts = round_half_up(sd(mean_rt), 1))
after_error n_participants mean_of_mean_rts sd_of_mean_rts
post-correct 17 674.1 27.1
post-error 17 736.5 68.4

17 participants: participants 7 and 15 are excluded for too-fast responding, and participant 12 for not completing all trials. Note that the participant-level exclusions are applied before any trials are removed, so n() counts all the trials each participant completed.

8.6 Frequncies of sets of columns

Note that it is also possible to use count to obtain the frequencies of sets of unique values across columns, e.g., unique combinations of item and response.

# data_demographics_trimmed %>%
#   count(item)
# 
# data_demographics_trimmed %>%
#   count(response)
# 
# data_demographics_trimmed %>%
#   count(item, response)

It can be useful to arrange the output by the frequencies.

# data_demographics_trimmed %>%
#   count(item, response) %>%
#   arrange(desc(n)) # arrange in descending order

8.6.1 summarize(across())

8.7 Exercises

8.7.1 mutate() vs. summarize()

What is the difference between mutate() and summarize()? If you use the wrong one, will you get the same answer? E.g., mutate(mean_age = mean(age, na.rm = TRUE)) vs. summarize(mean_age = mean(age, na.rm = TRUE)).

Both calculate the same mean age, but they return it differently. summarize() reduces the data frame to one row (or one row per group) containing only the summary (and any grouping columns). mutate() keeps every row and every column, and adds the mean as a new column, repeated on every row (or on every row of each group).

So the value is the same, but the shape of the result is very different. If you use mutate() when you meant summarize(), any later step that assumes one row per participant or group, e.g., calculating an N with n() or a mean of means, will be wrong.

8.7.2 group_by()

What does group_by() change about how summarize(), mutate(), filter(), and the slice_() functions work?

They all do their calculations separately within each group:

  • summarize() returns one row per group rather than one row in total.
  • mutate() calculates aggregating functions (e.g., mean()) and functions that use other rows (e.g., lag(), cumsum()) within each group.
  • filter() keeps or drops rows based on calculations within each group, e.g., filter(n() == 60) keeps groups with exactly 60 rows.
  • slice_() functions return rows from each group, e.g., slice_head(n = 3) returns the first 3 rows of each group.

8.7.3 ungroup() and .by

Why is it important to ungroup()? What is an alternative that avoids the problem?

Grouping persists until it is removed, and it’s easy to miss because it isn’t visible in the data itself. Later functions in the pipe, or in later code using the saved data frame, will silently operate within groups, e.g., slice_min(age, n = 3) returns three rows per group rather than three rows in total.

Using the .by argument (e.g., summarize(..., .by = gender)) groups the data for that one function call only, so the result is never left grouped.

8.7.4 Counting

What is the difference between n() and sum(!is.na(age)) inside summarize()? Which should you report alongside mean(age, na.rm = TRUE)?

n() counts rows, including those where ‘age’ is NA. sum(!is.na(age)) counts only the non-missing values of ‘age’. The mean is calculated only from the non-missing values, so sum(!is.na(age)) is the N it is based on.

count(gender) is a wrapper for which other functions?

group_by(gender) %>% summarize(n = n()).

How can you calculate a proportion inside summarize(), e.g., the proportion of participants younger than 25?

Take the mean of a logical test, e.g., mean(age < 25, na.rm = TRUE). TRUE counts as 1 and FALSE as 0, so the mean is the proportion of TRUEs.

8.7.5 Duplicates

What is the difference between distinct() and distinct(id, .keep_all = TRUE) when removing duplicate rows? Which is safer?

distinct() only removes rows that are identical in every column. distinct(id, .keep_all = TRUE) keeps the first row for each id, regardless of whether that participant’s rows contain different values in other columns. You don’t control which row is kept, and conflicting values are silently discarded. distinct() is safer because any remaining duplicate ids reveal conflicts that you need to look at and resolve deliberately.

8.7.6 Exclusions

What is the difference between a trial-level and a participant-level exclusion rule? How is each implemented?

A trial-level rule removes individual trials, e.g., filter(rt >= 200) removes all trials faster than 200 ms, from any participant. A participant-level rule removes all of a participant’s trials based on a calculation across their trials, e.g., group_by(id) %>% filter(mean(rt < 200) <= 0.10) removes participants with more than 10% too-fast trials. Participant-level rules require grouping.

Why does the order of processing steps matter when using lag() or participant-level exclusion rules?

lag() refers to the previous row, so if trials are removed first, it refers to the previous remaining trial rather than the actual previous trial. Similarly, a participant-level rule like “more than 10% too-fast trials” should usually be calculated on all of the participant’s trials. If the too-fast trials are removed first, no participant would have any too-fast trials left, and no one would be excluded. In general, calculate values that depend on the full trial sequence, and apply participant-level rules, before removing any trials.

8.7.7 Interactive exercises

Complete the interactive exercises for:

8.7.8 Practice summarizing a demographics table

In your local version of this .qmd file:

  • Use ‘dat_gender_tidier’ to create a table with one row per gender.
  • Include the number of participants, the percentage of participants, and the mean and SD of age, rounded to 1 decimal place.
  • Make sure the N reported for the mean age only includes participants with a known age.
  • Arrange the table from the most to the least common gender.

Note: no solution provided here as many different ones are possible.

8.7.9 Practice processing trial-level data

In your local version of this .qmd file:

  • Use ‘dat_rt’ to create a data frame called ‘dat_rt_participants’, with one row per participant.
  • Calculate each participant’s number of trials, proportion of correct responses, proportion of too-fast trials (< 200 ms), and mean RT on correct trials only. Hint: all of these can be calculated in one summarize() if you first use mutate() and if_else() to create a column containing the RT on correct trials and NA on error trials.
  • Add a column ‘exclude’ that is “exclude” if the participant has more than 10% too-fast trials or did not complete all 60 trials, and “include” otherwise.
  • Then, using ‘dat_rt_participants’, calculate the number and percentage of participants who are excluded, and the mean and SD of the included participants’ mean RTs.

Note: no solution provided here to reduce the temptation to peek or copy-paste.