library(dplyr)
library(tidyr)
library(tibble)
library(stringr)
library(forcats)
library(knitr)
library(kableExtra)11 Strings and factors
This chapter covers two data types that are especially common in psychology data: character strings (text), using the {stringr} package, and factors (categorical variables), using the {forcats} package. Both packages are part of the {tidyverse}.
The previous chapters have already used a few {stringr} functions, such as str_to_lower() and str_remove_all(), along with some cryptic-looking patterns like "[^A-Za-z]". This chapter explains how they work.
11.1 Strings with {stringr}
A string is any text value, written between quotation marks, e.g., "female", "23", or "I enjoyed the study!". Note that "23" is a string, not a number, because of the quotation marks.
All {stringr} functions start with str_, which makes them easy to find with auto-complete: type str_ in RStudio and wait for the list of functions to appear. They all take the string(s) as their first argument, so they work well inside mutate().
We’ll use some made-up free-text responses to illustrate:
dat_text <- tribble(
~id, ~gender, ~age, ~handedness,
1, "Female", "23", "right",
2, " male", "31 years", "Right-handed",
3, "FEMALE ", "about 40", "left",
4, "non binary", "19", "RIGHT",
5, "Woman", "twenty-eight", "ambidextrous",
6, "m", "27", "R"
)
dat_text %>%
kable() %>%
kable_classic(full_width = FALSE)| id | gender | age | handedness |
|---|---|---|---|
| 1 | Female | 23 | right |
| 2 | male | 31 years | Right-handed |
| 3 | FEMALE | about 40 | left |
| 4 | non binary | 19 | RIGHT |
| 5 | Woman | twenty-eight | ambidextrous |
| 6 | m | 27 | R |
11.1.1 Case
str_to_lower(), str_to_upper(), str_to_title(), and str_to_sentence() change the case of letters. Converting to lower case is usually the first step in cleaning free-text responses, so that “Female”, “FEMALE”, and “female” are treated as the same value:
dat_text %>%
mutate(gender_lower = str_to_lower(gender),
gender_title = str_to_title(gender)) %>%
select(id, gender, gender_lower, gender_title)| id | gender | gender_lower | gender_title |
|---|---|---|---|
| 1 | Female | female | Female |
| 2 | male | male | Male |
| 3 | FEMALE | female | Female |
| 4 | non binary | non binary | Non Binary |
| 5 | Woman | woman | Woman |
| 6 | m | m | M |
11.1.2 Whitespace
Free-text responses often contain extra spaces, which are hard to see but make values different. str_trim() removes spaces at the start and end, and str_squish() also reduces repeated spaces in the middle to a single space:
dat_text %>%
mutate(gender_trimmed = str_trim(gender),
gender_squished = str_squish(gender)) %>%
# spaces are hard to see when printed, so count the number of characters with str_length() to make them visible
mutate(length_original = str_length(gender),
length_trimmed = str_length(gender_trimmed),
length_squished = str_length(gender_squished)) %>%
select(id, gender, length_original, length_trimmed, length_squished)| id | gender | length_original | length_trimmed | length_squished |
|---|---|---|---|---|
| 1 | Female | 6 | 6 | 6 |
| 2 | male | 5 | 4 | 4 |
| 3 | FEMALE | 7 | 6 | 6 |
| 4 | non binary | 11 | 11 | 10 |
| 5 | Woman | 5 | 5 | 5 |
| 6 | m | 1 | 1 | 1 |
For example, ” male” has 5 characters because of the leading space, and “non binary” has 11 because of the double space. After str_squish(), they have 4 and 10.
11.1.3 Combining strings
str_c() combines strings, e.g., to create labels for tables and plots. Base R’s paste0() does the same thing, and you’ll see both in other people’s code.
dat_text %>%
mutate(label = str_c("Participant ", id, ": ", str_to_lower(str_squish(gender)))) %>%
select(id, label)| id | label |
|---|---|
| 1 | Participant 1: female |
| 2 | Participant 2: male |
| 3 | Participant 3: female |
| 4 | Participant 4: non binary |
| 5 | Participant 5: woman |
| 6 | Participant 6: m |
11.1.4 Padding
Participant ids that are stored as text sort alphabetically, not numerically, so “p10” comes before “p2”:
sort(c("p1", "p2", "p10"))[1] "p1" "p10" "p2"
str_pad() adds characters to make strings a fixed width, e.g., adding leading zeros to ids so that they sort correctly:
str_pad(c("1", "2", "10"), width = 3, side = "left", pad = "0")[1] "001" "002" "010"
sort(str_c("p", str_pad(c("1", "2", "10"), width = 3, side = "left", pad = "0")))[1] "p001" "p002" "p010"
11.1.5 Detecting patterns
str_detect() returns TRUE or FALSE for whether each string contains a pattern. This makes it useful inside filter(), if_else(), and case_when(). str_starts() and str_ends() check only the start or end of the string.
dat_text %>%
mutate(handedness_lower = str_to_lower(handedness),
right_handed = str_detect(handedness_lower, "right"),
starts_with_r = str_starts(handedness_lower, "r")) %>%
select(id, handedness, right_handed, starts_with_r)| id | handedness | right_handed | starts_with_r |
|---|---|---|---|
| 1 | right | TRUE | TRUE |
| 2 | Right-handed | TRUE | TRUE |
| 3 | left | FALSE | FALSE |
| 4 | RIGHT | TRUE | TRUE |
| 5 | ambidextrous | FALSE | FALSE |
| 6 | R | FALSE | TRUE |
Note that pattern matching is case sensitive, which is why the responses were converted to lower case first.
Be careful with str_detect(): it matches the pattern anywhere in the string. For example, str_detect(gender, "male") is also TRUE for “female”! Patterns need to be specific enough to match only what you intend, which is what the next section on regular expressions is about.
11.1.6 Removing and replacing
str_remove() removes the first match of a pattern, and str_remove_all() removes every match. str_replace() and str_replace_all() replace matches with something else.
dat_text %>%
mutate(age_no_years = str_remove(age, " years"),
handedness_clean = str_replace(str_to_lower(handedness), "-handed", "")) %>%
select(id, age, age_no_years, handedness, handedness_clean)| id | age | age_no_years | handedness | handedness_clean |
|---|---|---|---|---|
| 1 | 23 | 23 | right | right |
| 2 | 31 years | 31 | Right-handed | right |
| 3 | about 40 | about 40 | left | left |
| 4 | 19 | 19 | RIGHT | right |
| 5 | twenty-eight | twenty-eight | ambidextrous | ambidextrous |
| 6 | 27 | 27 | R | r |
11.1.7 Extracting
str_extract() returns the part of each string that matches a pattern, and NA if there is no match. For example, extracting the number from age responses (the pattern "[0-9]+" means “one or more digits”, explained below):
dat_text %>%
mutate(age_number = str_extract(age, "[0-9]+"),
age_number = as.numeric(age_number)) %>%
select(id, age, age_number)| id | age | age_number |
|---|---|---|
| 1 | 23 | 23 |
| 2 | 31 years | 31 |
| 3 | about 40 | 40 |
| 4 | 19 | 19 |
| 5 | twenty-eight | NA |
| 6 | 27 | 27 |
Note that this extracts 40 from “about 40”, which might or might not be what you want, and can’t handle “twenty-eight”. Always check the results of extractions.
str_sub() extracts characters by position, e.g., str_sub(x, start = 1, end = 3) returns the first three characters.
11.1.8 Splitting
str_split() splits strings into pieces at a separator. In data frames, it’s usually easier to use {tidyr}’s separate_wider_delim() or separate(), which split one column into several columns. These are covered in the Reshaping and pivots chapter.
11.1.8.1 Exercise
Using ‘dat_text’, write code to clean the ‘gender’ column so that it contains only “female”, “male”, or “non-binary”. Use {stringr} functions to deal with case and whitespace first, then case_when() for the rest.
dat_text %>%
mutate(gender = str_to_lower(gender),
gender = str_squish(gender),
gender = case_when(gender %in% c("female", "woman", "f") ~ "female",
gender %in% c("male", "man", "m") ~ "male",
gender %in% c("non binary", "nonbinary", "non-binary") ~ "non-binary",
TRUE ~ gender)) %>%
select(id, gender)| id | gender |
|---|---|
| 1 | female |
| 2 | male |
| 3 | female |
| 4 | non-binary |
| 5 | female |
| 6 | male |
Note that the order of the cleaning steps matters: “non binary” (two spaces) only matches “non binary” after str_squish().
11.2 Regular expressions
The patterns used by {stringr} functions are “regular expressions”, or “regex”. A regular expression is a small language for describing patterns in text, which is then called by other languages including R.
regex is notoriously difficult to write. I suggest you make use of AI to help you write (and then validate!) complex regex searches! Nonetheless, I provide an introduction here.
Most characters simply match themselves, e.g., the pattern "male" matches the letters m, a, l, e in that order. Some characters have special meanings.
str_view() is very useful for learning regex: it shows exactly which parts of each string match a pattern, highlighted with <>:
str_view(c("female", "male", "Male"), "male")[1] │ fe<male>
[2] │ <male>
11.2.1 Character classes
Square brackets match any one of the characters inside them:
[abc]: a, b, or c.[a-z]: any lower case letter.[A-Z]is any upper case letter, and[A-Za-z]is any letter.[0-9]: any digit.[^...]: a^inside the brackets means “not”, e.g.,[^0-9]is any character that is not a digit.
This explains the patterns used in previous chapters:
str_remove_all(gender, "[^A-Za-z]")removes every character that is not a letter, e.g., spaces, hyphens, and numbers.str_remove_all(age, "[^0-9\\.]")removes every character that is not a digit or a decimal point.
str_remove_all(c("non-binary", "fe male", "male!"), "[^A-Za-z]")[1] "nonbinary" "female" "male"
str_remove_all(c("23 years", "about 40", "31.5"), "[^0-9\\.]")[1] "23" "40" "31.5"
There are also shortcuts for common classes, e.g., \\d for any digit and \\s for any whitespace.
11.2.2 Quantifiers
Quantifiers say how many times the previous character or class should be matched:
+: one or more times, e.g.,[0-9]+matches “4”, “40”, and “400”.*: zero or more times.?: zero or one times, i.e., optional, e.g.,"non-?binary"matches both “nonbinary” and “non-binary”.{n}: exactly n times, e.g.,[0-9]{4}matches four digits in a row.
str_view(c("nonbinary", "non-binary", "non binary"), "non-?binary")[1] │ <nonbinary>
[2] │ <non-binary>
11.2.3 Anchors
Anchors match positions rather than characters:
^: the start of the string (when outside square brackets).$: the end of the string.
Anchors are the solution to the “female contains male” problem above. "^male$" only matches strings that are exactly “male”, from start to end:
str_detect(c("male", "female", "male "), "male")[1] TRUE TRUE TRUE
str_detect(c("male", "female", "male "), "^male$")[1] TRUE FALSE FALSE
11.2.4 Special characters
Characters that have special meanings, such as . (which matches any character), ?, +, (, and $, need to be “escaped” with two backslashes if you want to match them literally. For example, "\\." matches a full stop, and "\\?" matches a question mark.
str_view(c("3.5", "305"), "3.5")[1] │ <3.5>
[2] │ <305>
str_view(c("3.5", "305"), "3\\.5")[1] │ <3.5>
Regex can become very complicated very quickly. Keep your patterns as simple as possible, test them on small examples with str_view() before using them on your data, and add a comment explaining what each pattern is meant to match.
11.2.4.1 Exercise
Write a regular expression that matches strings that start with “p” followed by exactly three digits, and nothing else, e.g., “p001” and “p123” but not “p12”, “p1234”, or “xp001”. Test it on the vector below with str_detect().
ids <- c("p001", "p123", "p12", "p1234", "xp001", "P001")str_detect(ids, "^p[0-9]{3}$")[1] TRUE TRUE FALSE FALSE FALSE FALSE
^ anchors the match to the start, p matches a literal “p”, [0-9]{3} matches exactly three digits, and $ anchors to the end. “P001” doesn’t match because regex is case sensitive.
11.2.5 A more complex example: extracting p-values from text
The real power of regular expressions is that a single pattern can match many variations of the same thing. For example, imagine you are conducting a meta-analysis or checking the results reported in a set of articles, and you have extracted the results sentences into a data frame. You want to extract the p-value from each one. Authors report p-values in many slightly different ways:
- With or without spaces, e.g., “p = .025” vs. “p=.025”.
- With or without a leading zero, e.g., “.025” vs. “0.025”.
- With upper or lower case, e.g., “p” vs. “P”.
- With different comparison operators, e.g., “=”, “<”, “>”, “<=”.
- Different numbers of decimal places, e.g., “.05” vs. “.050”.
dat_results_text <- tibble(
result = c("t(48) = 2.31, p = .025",
"F(1, 120) = 8.45, p<0.01",
"r = .12, p = 0.21",
"chi2(2) = 15.20, P < .001",
"z = 1.20, p=.230",
"t(30) = 1.70, p > .05",
"b = 0.31, SE = 0.10, p <= .002",
"Step 2 was significant, p = .041")
)Without regex, you would need to anticipate every combination of these variations and write a separate condition for each, e.g., in case_when() or with many str_detect() calls:
# don't do this!
case_when(str_detect(result, "p = .") ~ ...,
str_detect(result, "p=.") ~ ...,
str_detect(result, "p = 0.") ~ ...,
str_detect(result, "p=0.") ~ ...,
str_detect(result, "P = .") ~ ...,
str_detect(result, "p < .") ~ ...,
str_detect(result, "p<.") ~ ...,
# ... and so on, for every combination of case, spacing, leading zero, and operator
)Even with just the five variations listed above, there are dozens of combinations, and it’s easy to miss one. A regular expression can describe all of them in a single pattern.
Long patterns are hard to read, so {stringr}’s regex() function has a comments = TRUE option, which allows you to spread the pattern over several lines and add a comment explaining each part (spaces in the pattern itself are then ignored, so spaces you want to match have to be written as \\s). ignore_case = TRUE makes the pattern match both upper and lower case:
pattern_p <- regex("
\\b p # the letter p (or P) as a separate word, so not the p in 'Step'
\\s* # zero or more spaces
([<>=]{1,2}) # the comparison: =, <, >, <=, or >= (group 1)
\\s* # zero or more spaces
( # the value (group 2), made up of:
0? # an optional leading zero
\\. # a decimal point
([0-9]+) # one or more digits, i.e., the decimal places (group 3)
)
", ignore_case = TRUE, comments = TRUE)
str_view(dat_results_text$result, pattern_p)[1] │ t(48) = 2.31, <p = .025>
[2] │ F(1, 120) = 8.45, <p<0.01>
[3] │ r = .12, <p = 0.21>
[4] │ chi2(2) = 15.20, <P < .001>
[5] │ z = 1.20, <p=.230>
[6] │ t(30) = 1.70, <p > .05>
[7] │ b = 0.31, SE = 0.10, <p <= .002>
[8] │ Step 2 was significant, <p = .041>
Two new features are used here:
\\bmatches a “word boundary”, i.e., the start or end of a word. Without it, the pattern could match the “p” at the end of a word like “Step” if it happened to be followed by “=” and a number.- Parentheses create “groups”: parts of the match that can be extracted separately.
str_extract()’sgroupargument returns just that part of the match. Groups can be nested inside each other, and are numbered in the order of their opening brackets. Here, group 3 (the decimal places) is inside group 2 (the whole value).
dat_results_text %>%
mutate(p_text = str_extract(result, pattern_p),
p_operator = str_extract(result, pattern_p, group = 1),
p_value = as.numeric(str_extract(result, pattern_p, group = 2)))| result | p_text | p_operator | p_value |
|---|---|---|---|
| t(48) = 2.31, p = .025 | p = .025 | = | 0.025 |
| F(1, 120) = 8.45, p<0.01 | p<0.01 | < | 0.010 |
| r = .12, p = 0.21 | p = 0.21 | = | 0.210 |
| chi2(2) = 15.20, P < .001 | P < .001 | < | 0.001 |
| z = 1.20, p=.230 | p=.230 | = | 0.230 |
| t(30) = 1.70, p > .05 | p > .05 | > | 0.050 |
| b = 0.31, SE = 0.10, p <= .002 | p <= .002 | <= | 0.002 |
| Step 2 was significant, p = .041 | p = .041 | = | 0.041 |
One pattern has correctly extracted the operator and value from every variation.
11.2.6 Extracting the precision of reported p-values
The decimal places in group 3 also tell us the precision with which each p-value was reported. This matters because a reported p-value is a rounded version of the true one, and the number of decimals determines which true values it could have been rounded from. For example, “p = .05” could have been rounded from any value from .045 up to (but not including) .055, which includes both significant and non-significant results. “p = .050” narrows this down to .0495 to .0505.
This range is useful when checking whether reported results are internally consistent, e.g., whether a reported p-value could have come from the reported test statistic and degrees of freedom. The range is only meaningful for exact p-values, i.e., when the operator is “=”:
dat_results_text %>%
mutate(p_operator = str_extract(result, pattern_p, group = 1),
p_value = as.numeric(str_extract(result, pattern_p, group = 2)),
# the number of decimal places reported
p_decimals = str_length(str_extract(result, pattern_p, group = 3)),
# the range of values that would round to the reported value
p_min = if_else(p_operator == "=", p_value - 0.5 / 10^p_decimals, NA_real_),
p_max = if_else(p_operator == "=", p_value + 0.5 / 10^p_decimals, NA_real_)) %>%
select(-result)| p_operator | p_value | p_decimals | p_min | p_max |
|---|---|---|---|---|
| = | 0.025 | 3 | 0.0245 | 0.0255 |
| < | 0.010 | 2 | NA | NA |
| = | 0.210 | 2 | 0.2050 | 0.2150 |
| < | 0.001 | 3 | NA | NA |
| = | 0.230 | 3 | 0.2295 | 0.2305 |
| > | 0.050 | 2 | NA | NA |
| <= | 0.002 | 3 | NA | NA |
| = | 0.041 | 3 | 0.0405 | 0.0415 |
Note that the precision has to be extracted from the text. Once a value has been converted to a number, “.05” and “.050” are identical, and the information about how precisely it was reported is lost.
Note what the pattern does not handle, though: symbols such as “≤” (which can be added inside the square brackets, although non-English characters can behave differently depending on your computer’s language settings), p-values reported in scientific notation (e.g., “p = 1.2e-5”), p-values of exactly 1, or plural forms such as “ps < .001”. This is typical of regex: always check the results, e.g., by counting how many rows returned NA, and inspecting the text of those that did.
11.2.7 Worked example: recovering ages written as words
In the Data transformation III chapter, some participants’ ages were written as words, e.g., “thirty five”, and these became NA when age was converted to numeric. With {stringr}, we can recover them.
library(truffle)
set.seed(42)
dat_demographics_messy <- data.frame(id = 1:500) %>%
truffle_demographics() %>%
dirt_demographics()
# which age responses contain letters?
dat_demographics_messy %>%
filter(str_detect(age, "[A-Za-z]")) %>%
count(age)| age | n |
|---|---|
| female | 1 |
| forty | 3 |
| forty five | 1 |
| nineteen | 1 |
| thirty | 2 |
| thirty eight | 2 |
| thirty five | 3 |
| thirty four | 1 |
| thirty one | 1 |
| thirty seven | 1 |
| thirty two | 1 |
| twenty eight | 1 |
| twenty five | 1 |
| twenty four | 2 |
| twenty nine | 2 |
| twenty one | 2 |
| twenty three | 1 |
| twenty two | 1 |
Most of these are a “tens” word (e.g., “thirty”) optionally followed by a “units” word (e.g., “five”). {stringr}’s word() extracts words by position: word(x, 1) returns the first word and word(x, 2) the second (or NA if there is no second word). We can convert each word to a number with case_when(), and add them up:
dat_age_tidier <- dat_demographics_messy %>%
mutate(age_raw = str_squish(str_to_lower(age)),
# ages written as numbers
age_from_numbers = as.numeric(str_extract(age_raw, "^[0-9]+$")),
# ages written as words
tens_word = word(age_raw, 1),
units_word = word(age_raw, 2),
tens = case_when(tens_word == "twenty" ~ 20,
tens_word == "thirty" ~ 30,
tens_word == "forty" ~ 40,
TRUE ~ NA_real_),
units = case_when(is.na(units_word) ~ 0,
units_word == "one" ~ 1,
units_word == "two" ~ 2,
units_word == "three" ~ 3,
units_word == "four" ~ 4,
units_word == "five" ~ 5,
units_word == "six" ~ 6,
units_word == "seven" ~ 7,
units_word == "eight" ~ 8,
units_word == "nine" ~ 9,
TRUE ~ NA_real_),
# special cases that don't follow the tens + units pattern
age_from_words = case_when(age_raw == "nineteen" ~ 19,
TRUE ~ tens + units),
# use the number if there is one, otherwise the value from the words
age = if_else(!is.na(age_from_numbers), age_from_numbers, age_from_words))
# check the conversions of the word responses
dat_age_tidier %>%
filter(str_detect(age_raw, "[a-z]")) %>%
count(age_raw, tens, units, age)| age_raw | tens | units | age | n |
|---|---|---|---|---|
| female | NA | 0 | NA | 1 |
| forty | 40 | 0 | 40 | 3 |
| forty five | 40 | 5 | 45 | 1 |
| nineteen | NA | 0 | 19 | 1 |
| thirty | 30 | 0 | 30 | 2 |
| thirty eight | 30 | 8 | 38 | 2 |
| thirty five | 30 | 5 | 35 | 3 |
| thirty four | 30 | 4 | 34 | 1 |
| thirty one | 30 | 1 | 31 | 1 |
| thirty seven | 30 | 7 | 37 | 1 |
| thirty two | 30 | 2 | 32 | 1 |
| twenty eight | 20 | 8 | 28 | 1 |
| twenty five | 20 | 5 | 25 | 1 |
| twenty four | 20 | 4 | 24 | 2 |
| twenty nine | 20 | 9 | 29 | 2 |
| twenty one | 20 | 1 | 21 | 2 |
| twenty three | 20 | 3 | 23 | 1 |
| twenty two | 20 | 2 | 22 | 1 |
All of the ages written as words have been recovered, except “female”, which was presumably entered in the wrong field and remains NA. Note that the code only handles the words that are actually present in the data. If new data were added, you would need to check for new cases, e.g., “fifty”, which would currently become NA. This is why the last line checks the result.
# how many ages are missing before vs. after?
dat_age_tidier %>%
summarize(n_missing_numbers_only = sum(is.na(age_from_numbers)),
n_missing_after_recovery = sum(is.na(age)))| n_missing_numbers_only | n_missing_after_recovery |
|---|---|
| 27 | 1 |
11.2.7.1 Exercise
Some participants in ‘dat_demographics_messy’ typed numbers into the gender question. Use str_detect() to find how many ‘gender’ responses contain digits. What do you think happened?
dat_demographics_messy %>%
filter(str_detect(gender, "[0-9]")) %>%
count(gender)| gender | n |
|---|---|
| 18 | 2 |
| 19 | 1 |
| 22 | 1 |
| 23 | 1 |
| 25 | 3 |
| 26 | 2 |
| 27 | 3 |
| 30 | 3 |
| 31 | 1 |
| 36 | 1 |
| 38 | 2 |
| 40 | 2 |
| 41 | 1 |
| 42 | 1 |
| 43 | 3 |
| 45 | 3 |
dat_demographics_messy %>%
filter(str_detect(gender, "[0-9]")) %>%
nrow()[1] 30
These look like ages. These participants probably typed their age into the gender question. In the previous chapters, these values were removed by str_remove_all(gender, "[^A-Za-z]") and ended up as NA. Whether you could recover their gender depends on whether their age column contains it (only one ‘age’ response, “female”, does).
11.3 Factors with {forcats}
11.3.1 What factors are and why they matter
A factor is R’s type for categorical variables. It looks like a character variable, but it also stores the set of possible values, called “levels”, in a specific order.
The order matters because many functions use it. By default, character variables are sorted alphabetically. For example, condition names sort as “control”, “intervention”, “waitlist”. That might be fine, but it’s often not the order you want, e.g., for timepoints (“baseline”, “followup”, “post”) or Likert responses (“agree”, “disagree”, “neutral”).
Factor levels control:
- The order of rows in tables made with
count()andsummarize(). - The order of categories in plots, e.g., the bars of a bar chart. See the Visualization chapter.
- The reference category in regression models, i.e., which group the others are compared to. See The linear model chapter.
dat_timepoints <- tribble(
~id, ~timepoint, ~score,
1, "baseline", 10,
1, "post", 14,
1, "followup", 13,
2, "baseline", 12,
2, "post", 15,
2, "followup", 15
)
# character: alphabetical order
dat_timepoints %>%
group_by(timepoint) %>%
summarize(mean_score = mean(score))| timepoint | mean_score |
|---|---|
| baseline | 11.0 |
| followup | 14.0 |
| post | 14.5 |
# factor: the order of the levels
dat_timepoints %>%
mutate(timepoint = factor(timepoint, levels = c("baseline", "post", "followup"))) %>%
group_by(timepoint) %>%
summarize(mean_score = mean(score))| timepoint | mean_score |
|---|---|
| baseline | 11.0 |
| post | 14.5 |
| followup | 14.0 |
You can see the levels of a factor with levels().
11.3.2 Creating factors safely: factor() vs. fct()
Base R’s factor() has a dangerous behavior: values that aren’t in the levels you specify silently become NA.
gender <- c("female", "male", "non-binary", "female")
factor(gender, levels = c("female", "male"))[1] female male <NA> female
Levels: female male
“non-binary” has silently become NA, perhaps because of a typo in the levels or because a category was forgotten. {forcats}’ fct() instead throws an error if any values are not in the levels:
fct(gender, levels = c("female", "male"))Error in `fct()`:
! All values of `x` must appear in `levels` or `na`
ℹ Missing level: "non-binary"
fct(gender, levels = c("female", "male", "non-binary"))[1] female male non-binary female
Levels: female male non-binary
As with the typed NAs in the Data types chapter, this error protects you from silent mistakes. Prefer fct() over factor().
11.3.3 Changing the order of levels
fct_relevel()moves one or more levels to the front, e.g., to make “control” the reference category.fct_infreq()orders levels from most to least frequent, which is useful for tables and bar charts.fct_rev()reverses the order of the levels.
dat_conditions <- tibble(condition = c("waitlist", "intervention", "control", "intervention", "control", "intervention"))
dat_conditions %>%
mutate(condition = fct(condition)) %>%
pull(condition) %>%
levels()[1] "waitlist" "intervention" "control"
dat_conditions %>%
mutate(condition = fct_relevel(condition, "control")) %>%
pull(condition) %>%
levels()[1] "control" "intervention" "waitlist"
dat_conditions %>%
mutate(condition = fct_infreq(condition)) %>%
count(condition)| condition | n |
|---|---|
| intervention | 3 |
| control | 2 |
| waitlist | 1 |
Note that fct() without specified levels uses the order in which the values first appear in the data, whereas factor() sorts them alphabetically. pull() extracts a single column from a data frame as a vector.
11.3.4 Changing the values of levels
fct_recode()renames levels, written asnew_name = "old name"(the same direction asrename()).fct_collapse()combines several levels into one.fct_lump_n()andfct_lump_min()combine rare levels into an “Other” category, keeping the n most frequent levels, or the levels with at leastminobservations.
dat_handedness <- tibble(handedness = c("right", "right", "left", "right", "ambidextrous", "right", "left", "unsure"))
dat_handedness %>%
mutate(handedness = fct_recode(handedness, "Right-handed" = "right", "Left-handed" = "left")) %>%
count(handedness)| handedness | n |
|---|---|
| ambidextrous | 1 |
| Left-handed | 2 |
| Right-handed | 4 |
| unsure | 1 |
dat_handedness %>%
mutate(handedness = fct_collapse(handedness, "not right" = c("left", "ambidextrous", "unsure"))) %>%
count(handedness)| handedness | n |
|---|---|
| not right | 4 |
| right | 4 |
dat_handedness %>%
mutate(handedness = fct_lump_min(handedness, min = 2)) %>%
count(handedness)| handedness | n |
|---|---|
| left | 2 |
| right | 4 |
| Other | 2 |
Be careful with lumping categories together, especially for demographic variables like gender and ethnicity. It can be useful for analyses with small cell sizes, but it can also erase participants’ identities. Report what was combined and why.
11.3.5 Missing values as a level
NA is not a level of a factor, so it is dropped from some outputs, e.g., the legends of plots. fct_na_value_to_level() turns NA into an explicit level:
dat_gender <- tibble(gender = fct(c("female", "male", NA, "female"), levels = c("female", "male")))
levels(dat_gender$gender)[1] "female" "male"
dat_gender %>%
mutate(gender = fct_na_value_to_level(gender, level = "Not reported")) %>%
count(gender)| gender | n |
|---|---|
| female | 2 |
| male | 1 |
| Not reported | 1 |
11.3.6 Unused levels
Levels remain even when there are no longer any observations with that value, e.g., after filter(). Some functions, such as count() with .drop = FALSE and many plotting functions, will show these empty levels. fct_drop() removes unused levels:
dat_conditions_filtered <- dat_conditions %>%
mutate(condition = fct(condition)) %>%
filter(condition != "waitlist")
levels(dat_conditions_filtered$condition)[1] "waitlist" "intervention" "control"
dat_conditions_filtered %>%
count(condition, .drop = FALSE)| condition | n |
|---|---|
| waitlist | 0 |
| intervention | 3 |
| control | 2 |
dat_conditions_filtered %>%
mutate(condition = fct_drop(condition)) %>%
pull(condition) %>%
levels()[1] "intervention" "control"
Whether you want empty levels depends on the context. Showing that no participants were in a category can be informative in a table, but for a filtered analysis they usually should be dropped.
11.3.7 Ordered factors
Ordered factors are factors whose levels have a meaningful order, e.g., Likert responses. They allow comparisons like < and > between levels:
likert <- factor(c("disagree", "agree", "neutral", "agree"),
levels = c("disagree", "neutral", "agree"),
ordered = TRUE)
likert[1] disagree agree neutral agree
Levels: disagree < neutral < agree
likert > "neutral"[1] FALSE TRUE FALSE TRUE
Note that ordered factors are treated differently by some functions, e.g., lm() uses a different coding scheme for them, so only use them when you specifically need ordering comparisons. For most purposes, a regular factor with levels in the right order is enough.
11.3.8 Converting factors to numbers
A very common and dangerous mistake is converting a factor to a number with as.numeric(). Internally, factors are stored as integer codes (1 for the first level, 2 for the second, etc.), and as.numeric() returns these codes rather than the values you see:
response <- factor(c("10", "20", "30", "10"))
response[1] 10 20 30 10
Levels: 10 20 30
as.numeric(response)[1] 1 2 3 1
To get the actual values, convert to character first:
as.numeric(as.character(response))[1] 10 20 30 10
This mostly happens when numbers have been stored as a factor, e.g., when a data file was read with older code that converted all text to factors, or when a stray non-numeric value made a numeric column character and it was then converted to a factor. In general, keep numbers as numbers, and only use factors for categorical variables.
11.3.8.1 Exercise
Using ‘dat_timepoints’:
- Convert ‘timepoint’ to a factor with the levels in the order “baseline”, “post”, “followup”, using
fct(). - Rename the levels to “Baseline”, “Post-intervention”, and “Follow-up”.
- Calculate the mean score at each timepoint.
dat_timepoints %>%
mutate(timepoint = fct(timepoint, levels = c("baseline", "post", "followup")),
timepoint = fct_recode(timepoint,
"Baseline" = "baseline",
"Post-intervention" = "post",
"Follow-up" = "followup")) %>%
group_by(timepoint) %>%
summarize(mean_score = mean(score))| timepoint | mean_score |
|---|---|
| Baseline | 11.0 |
| Post-intervention | 14.5 |
| Follow-up | 14.0 |
11.4 Exercises
11.4.1 {stringr}
What is the difference between str_trim() and str_squish()?
str_trim() removes whitespace from the start and end of strings. str_squish() does the same, and also replaces repeated whitespace inside the string with a single space.
Why is filter(str_detect(gender, "male")) a bad way to find male participants? What would be better?
str_detect() matches the pattern anywhere in the string, so it also returns “female”. Better options are an exact comparison, filter(gender == "male"), after cleaning the column, or an anchored pattern, str_detect(gender, "^male$").
11.4.2 Regular expressions
What does each of these patterns match? "[^A-Za-z]", "[0-9]+", "^p", "\\.".
"[^A-Za-z]": any single character that is not a letter."[0-9]+": one or more digits in a row."^p": a “p” at the start of the string."\\.": a literal full stop. Without the backslashes,.matches any character.
11.4.3 Factors
What is the difference between a character variable and a factor? Name two things that the order of factor levels affects.
A factor stores a defined set of possible values (levels) in a specific order, whereas a character variable is just text, which is sorted alphabetically. The order of levels affects the order of rows in tables (e.g., from count() and summarize()), the order of categories in plots, and the reference category in regression models.
Why should you prefer fct() over factor()?
factor() silently converts any values that are not in the specified levels to NA. fct() throws an error instead, so that you notice typos in the levels or categories you forgot.
What does as.numeric() return when applied to a factor? How do you get the values instead?
It returns the internal integer codes of the levels (1, 2, 3, …), not the values that are displayed. Use as.numeric(as.character(x)) to get the values.
11.4.4 Practice
In your local version of this .qmd file:
- Use ‘dat_demographics_messy’ to create ‘dat_demographics_clean’.
- Clean the ‘gender’ column using {stringr} functions and
case_when(), so that it contains only “female”, “male”, “non-binary”, orNA. - Clean the ‘age’ column, including recovering ages written as words, as in the worked example above.
- Convert ‘gender’ to a factor using
fct(), with the levels ordered from most to least frequent and missing values as an explicit “Not reported” level. - Create a table of the number of participants and mean age for each gender.
Note: no solution provided here to reduce the temptation to peek or copy-paste.