flowchart LR
A[("data/raw/")] --> B["code/processing.qmd"]
B --> C[("data/processed/")]
C --> D["code/analysis.qmd"]
D --> E[("data/outputs/")]
9 Structuring projects ✎ Rough draft
So far, we have mostly worked with single .qmd files. Real research projects contain many files: raw data from several sources, code to process and analyze it, plots and tables, study materials, a preregistration, and a manuscript. How these files are organized determines whether the project can be understood, checked, and reproduced by anyone else, including yourself in six months’ time.
This chapter covers:
- The general principles behind a well-structured project, which apply whatever tools you use.
- Data standards, and psych-DS specifically, which turn these principles into a shared set of rules.
- How to use the {psychdsish} R package to create projects that follow these rules from the start, and to check that they still do.
9.1 Why structure matters
Here is a (lightly fictionalized) project folder of the kind I receive from students, collaborators, and, if I’m honest, my past self:
thesis stuff/
├── analysis FINAL.R
├── analysis FINAL v2 (use this one).R
├── data.xlsx
├── data_cleaned.xlsx
├── data_cleaned_NEW.csv
├── Figure1.png
├── fig1_revised.png
├── lab meeting notes.docx
├── model.rds
├── output.html
├── questionnaire.pdf
└── old/
├── analysis.R
└── data.xlsx
Try to answer the following questions about it:
- Which file contains the data as it was originally collected?
- Was
data.xlsxedited by hand after it was downloaded? Which of the twodata.xlsxfiles is the original? - Which script created
data_cleaned_NEW.csv, and which script reads it? - Does
analysis FINAL v2 (use this one).RproduceFigure1.pngorfig1_revised.png? Which one is in the manuscript? - If I delete
model.rds, can I get it back?
You can’t answer any of these questions from the folder alone, and often the person who created it can’t either. None of these problems are caused by bad code. They are caused by the absence of a structure that makes the answers obvious.
This matters for more than tidiness. A project that can’t be understood can’t be checked, so errors go unnoticed (see the case studies in the chapter on Loading data). A project that can’t be re-run from the raw data can’t be reproduced, which is the minimum standard for computational work. And a project that only its creator can navigate becomes unusable when that person leaves, forgets, or loses their laptop.
9.2 General principles
The following principles are not specific to R, psych-DS, or any particular tool. They are distilled from standards like psych-DS (discussed below) and from guides such as Good enough practices in scientific computing (Wilson et al., 2017).
9.2.1 1. One project, one folder
Everything needed to understand and reproduce a project lives in a single folder, and nothing outside that folder is needed. That folder can be moved anywhere on your computer, zipped and emailed, or uploaded to GitHub or OSF, and it will still work.
This is only possible if code refers to files using relative paths (e.g., ../data/raw/data.csv) rather than absolute paths (e.g., C:/Users/ian/Documents/thesis/data.csv) or setwd(). See the section on relative vs. absolute paths to refresh your knowledge.
In RStudio, opening a project folder via its .Rproj file also sets the console’s working directory to the project folder, and keeps each project’s environment and history separate.
9.2.2 2. Separate files by their role
Different kinds of files have different jobs, and should live in different places. At a minimum:
| Role | What it contains | Can it be regenerated? |
|---|---|---|
| Raw data | Data as it was collected or received | No. Must be preserved |
| Code | Scripts that process and analyze the data | No. This is your work |
| Processed data | Cleaned data created by code | Yes, by re-running the code |
| Outputs | Plots, tables, fitted models created by code | Yes, by re-running the code |
| Materials | Measures, experiment files, procedures | No |
| Documents | Preregistration, manuscript, slides | No |
The final column is the most important one. Files that cannot be regenerated must be protected. Files that can be regenerated are disposable: if in doubt, delete them and re-run the code. Keeping these two kinds of files apart makes it obvious which is which.
9.2.3 3. Raw data is read-only
The earliest form of the data you have access to must be preserved exactly as it was received. Never edit it by hand, never overwrite it with code, and never save over it after opening it in Excel. Every change to the data is instead made by code that reads the raw data and writes a new file somewhere else.
The only exception is the removal of private or identifying information before sharing, which is covered in the chapter on Sharing and privacy. Even then, the original is kept (securely, and not shared) and the removal is done with code.
9.2.4 4. Data flows in one direction, and only code moves it
Data moves from raw, through processing code, to processed data, and then through analysis code to outputs. Nothing flows backward: analysis code never writes to data/raw/, and processing code never reads from data/outputs/.
Because every arrow is code, every transformation is documented, and anyone can see exactly how the numbers in a paper were derived from the data.
9.2.5 5. Everything can be rebuilt from scratch, in order
A good test of a project is: if I delete everything in data/processed/ and data/outputs/, can I get it all back by running the code? If the answer is “yes, but only if you run these three scripts in the right order and remember to skip chunk 7”, the project is not reproducible yet.
The code should therefore run from start to finish without manual intervention, and the order in which scripts run should be written down, ideally in a form the computer can follow (see Rendering the whole project below).
9.2.6 6. Name files for computers and for humans
File names should be:
- Machine readable: no spaces, no special characters (
&,?,!,(, accents), and consistent case. Spaces in particular break many tools. - Human readable: the name says what is in the file.
- Sortable: when files sort alphabetically, they appear in a useful order. Dates follow ISO 8601 (
2026-09-23, not23.9.26), and numbers are zero-padded (01,02, …10, not1,2, …10, which sorts as1,10,2).
| Bad | Better |
|---|---|
analysis FINAL v2 (use this one).R |
analysis.qmd (and use version control for versions) |
data_cleaned_NEW.csv |
study-1_stage-processed_data.csv |
Figure1.png, fig1_revised.png |
plot_01_self_reports.png |
23.9.26 pilot.csv |
2026-09-23_pilot_data.csv |
Note that “final”, “v2”, and “NEW” in file names are attempts to do version control by hand. They reliably fail, because there is always a “final_v3”. Version control with git and GitHub, covered in the chapter on Sharing and privacy, solves this properly: there is only ever one analysis.qmd, and git remembers every previous version of it.
See Jenny Bryan’s How to name files for more on this.
9.2.7 7. Document the project
A project needs at least:
- A README: what the project is, how it is organized, and how to reproduce its results. It is the first file anyone opens, and on GitHub it is displayed on the repository’s front page.
- A codebook (data dictionary) for each data file: what each variable is, its units, and what its values mean. A column called
q7_rwith values from 1 to 5 is meaningless without one. - A licence: without one, others are not legally allowed to reuse your work, even if it is public.
9.2.8 8. Follow a convention rather than inventing one
You could follow all of the above principles and still organize your project differently from everyone else. If everyone invents their own structure, everyone else has to learn it before they can find anything. The final principle is therefore to use a structure that others already know, i.e., a standard.
9.3 Data standards
A standard is a specification of what must be done, what must not be done, and what may be done to allow flexibility. For project structure, this means rules about which folders exist, which files go where, how files are named, and what documentation is required.
Standards have two big advantages over good intentions:
- Shared expectations: anyone who knows the standard can navigate any project that follows it, without needing to read the README first. Raw data is always in the same place.
- Automatic checking: because the rules are written down precisely, a computer can check whether a project follows them. This is much more reliable than trying to remember all the rules yourself, and it catches problems as they arise rather than when a reviewer asks for your data.
Several standards and conventions exist for different fields, for example BIDS for neuroimaging data, the TIER Protocol for social science projects, and Cookiecutter Data Science for Python data science projects. They differ in the details but share the principles above.
9.3.1 psych-DS
psych-DS is a data standard for psychological data, led by Melissa Kline Struhl. It was inspired by BIDS but aims to be much lighter-weight, so that it can be used for the huge variety of data collected in psychology. I have contributed a very small amount to the debate around it.
psych-DS is built on four principles, which you will recognize from the previous section:
- The earliest form of data must be preserved.
- Original data should never be modified.
- Different versions of the data should be kept separate.
- All transformations should be documented.
In practice, a psych-DS dataset has:
- A
data/folder containing the data files, which are .csv files. - Data file names made of
key-valuepairs separated by underscores and ending in_data.csv, e.g.,study-1_task-stroop_data.csv. This makes the contents of each file clear from its name, and lets software find, for example, all the Stroop task data across studies. - A
dataset_description.jsonfile in the project root, containing machine-readable metadata about the dataset, including a description of every variable.
psych-DS provides an online validator that checks whether a dataset complies with the standard.
9.3.2 Why psych-DS-ish?
I really like the concept of psych-DS, but I am not yet convinced by some of its choices, at least as a starting point for most researchers and for students on courses like this one:
- The
.jsonmetadata file is required. .json files are a pain to write by hand, and very few psychology workflows currently use them. I didn’t want creating one to be the price of entry. - psych-DS is deliberately light on what it requires, so that it can apply to as many datasets as possible. For my own projects and for teaching, I’m happy to be more prescriptive, e.g., about where code, outputs, and study materials go, not just data.
- psych-DS focuses on checking whether a project is compliant, not on helping people set up a compliant project in the first place. Tidying up a project after the fact is much harder than starting with a template.
For the moment, I therefore recommend partial compliance with psych-DS: its high-level principles and file naming conventions make for a very well structured project, but other parts are (currently) high-effort-low-reward. That is what {psychdsish} implements.
9.4 psych-DS-ish

{psychdsish} (‘psych-DS-ish’) is an R package I wrote to make the principles above easy to follow. It does three main things:
create_project_skeleton()creates a new project with the standard folder structure, plus templates for the README, processing and analysis code, a licence, a citation file, and more.validator()checks whether a project follows the standard and, if not, tells you how to fix it.write_dataset_description()optionally creates the psych-DSdataset_description.jsonfrom your codebooks, for when you want full psych-DS metadata.
It is ‘compliant-ish’ with psych-DS: it follows its principles and naming conventions, but makes the .json file optional.
9.4.1 Installation
# install.packages("remotes")
remotes::install_github("ianhussey/psychdsish")If you use RStudio, restart it after installing so that the New Project wizard and the Addins menu pick up the package.
9.4.2 Creating a new project
9.4.2.1 In RStudio
Go to File > New Project > New Directory > psych-DS-ish Project. Give the project a directory name and choose where to create it. The other options can usually be left at their defaults:
- Create _quarto.yml: on by default. This lets you render the whole project in order (see below).
- Number of studies and Multi-study layout: for projects with more than one study (see Multi-study projects).
Click Create Project. RStudio creates the project, opens it, and opens README.md, code/processing.qmd, and code/analysis.qmd ready for you to start work.
Here’s a demo of the project creator in action:
9.4.2.2 From the console
In Positron, VS Code, or any other editor, or if you prefer code to menus, run this from the R console and then open the folder:
psychdsish::create_project_skeleton(project_root = "~/git/my_project")create_project_skeleton() never overwrites existing files unless you set overwrite = TRUE, so it is safe to run on a folder that already exists (see Restructuring an existing project).
9.4.3 What you get
Let’s create a project in a temporary folder and look at what’s in it. fs::dir_tree() prints a folder’s contents as a tree, and all = TRUE includes hidden files, whose names start with a .:
library(psychdsish)
project <- file.path(tempdir(), "my_project")
create_project_skeleton(project_root = project)
fs::dir_tree(project, all = TRUE)/var/folders/45/d07jd4jn4756zs6q38rgl4xw0000gp/T//RtmpLiPBkh/my_project
├── .gitattributes
├── .gitignore
├── CITATION.cff
├── LICENSE
├── README.md
├── _quarto.yml
├── code
│ ├── analysis.qmd
│ └── processing.qmd
├── data
│ ├── outputs
│ │ ├── fitted_models
│ │ │ └── .gitkeep
│ │ ├── plots
│ │ │ └── .gitkeep
│ │ └── results
│ │ └── .gitkeep
│ ├── processed
│ │ └── .gitkeep
│ └── raw
│ └── .gitkeep
├── methods
│ └── .gitkeep
├── my_project.Rproj
├── preregistration
│ └── .gitkeep
├── reports
│ └── .gitkeep
└── tools
├── project_creator.qmd
├── project_validator.qmd
└── style_all_files.qmd
(The first line is the location of the temporary folder on the computer that rendered this book. Yours will be wherever you created the project.)
The empty folders each contain a .gitkeep file. These are empty placeholders that exist only so that git, which ignores empty folders, keeps the folder structure when you put the project on GitHub.
Each folder has a clear role:
| Folder | What goes in it | Examples |
|---|---|---|
code/ |
Processing and analysis code, and the .html reports they render | processing.qmd, analysis.qmd, analysis.html |
data/raw/ |
Data as it was collected, never modified, and its codebooks | Qualtrics or lab.js exports, e.g., study-1_task-selfreports_stage-raw_data.csv |
data/processed/ |
Cleaned data written by processing.qmd, and their codebooks |
study-1_stage-processed_data.csv, study-1_stage-processed_codebook.csv |
data/outputs/plots/ |
Plots written by analysis.qmd |
plot_01_self_reports.png |
data/outputs/fitted_models/ |
Fitted model objects, which can be slow to re-fit | fit_model_1.rds (e.g., from {lme4}, {brms}, {lavaan}) |
data/outputs/results/ |
Tables and other results | descriptives.csv, cor_matrix.csv |
methods/ |
Study materials | Item wordings, Qualtrics .qsf files, PsychoPy or lab.js files, procedure documents |
preregistration/ |
Preregistration documents | preregistration.pdf |
reports/ |
Anything written about the project | Thesis, manuscript, preprint, slides, posters |
tools/ |
Utilities for managing the project, not part of the analysis | project_validator.qmd, style_all_files.qmd |
And the files in the project root:
| File | What it’s for |
|---|---|
README.md |
What the project is, how it’s organized, and how to reproduce it. Contains placeholder text to replace with your own. |
LICENSE |
CC BY 4.0: others may reuse your work as long as they credit you. |
CITATION.cff |
Machine-readable citation information. On GitHub, it adds a Cite this repository button that gives APA and BibTeX citations. |
_quarto.yml |
Lists which .qmd files to render, and in which order, when rendering the whole project. |
my_project.Rproj |
The RStudio project file. Double-click it to open the project. It is set not to save or restore your workspace, so that your code, not leftover objects, determines your results. |
.gitignore |
Tells git which files not to track, e.g., R history files, caches, .DS_Store files, and large outputs. |
.gitattributes |
Stops GitHub from labeling your repository as an HTML project because of the rendered reports. |
Note that .gitignore excludes data/outputs/plots/ and data/outputs/fitted_models/ by default, because they can be regenerated from the code and model objects can be very large. If you want your plots to appear on GitHub, delete those lines from .gitignore.
9.4.3.1 Where does this file go?
When you’re unsure, ask two questions: Did code create it? and Could I recreate it if it were deleted?
- Data you received or collected, that code did not create:
data/raw/. - Data that code created from other data:
data/processed/. - Plots, tables, and model objects that code created:
data/outputs/. - Code that does the processing or analysis:
code/. - Things participants saw or did:
methods/. - Things you wrote about the project, before it (
preregistration/) or after it (reports/).
9.4.4 Working in the project
The workflow follows the one-directional data flow described above:
- Put the raw data in
data/raw/, and don’t touch it again. If you can still choose its name (i.e., it hasn’t been shared or committed to git yet), follow the psych-DS convention, e.g.,study-1_task-selfreports_stage-raw_data.csv. - Write
code/processing.qmdto read the raw data, clean it, and write the result todata/processed/. - Write
code/analysis.qmdto read the processed data and write plots, fitted models, and tables todata/outputs/.
Each .qmd file runs with its own folder as the working directory. Because both code files are in code/, paths from them always start by going ‘up’ one level with ../:
library(readr)
# in code/processing.qmd
data_raw <- read_csv("../data/raw/study-1_task-selfreports_stage-raw_data.csv")
# ... processing ...
write_csv(data_processed, "../data/processed/study-1_stage-processed_data.csv")library(readr)
library(ggplot2)
# in code/analysis.qmd
data_processed <- read_csv("../data/processed/study-1_stage-processed_data.csv")
# ... analysis ...
ggsave("../data/outputs/plots/plot_01_self_reports.png", p_self_reports)It is fine to split processing or analysis across several files if they get long, e.g., processing_selfreports.qmd and processing_behavioral.qmd. Keep them in code/, and add them to _quarto.yml (see next section).
9.4.5 Rendering the whole project
Clicking Render in a .qmd file renders only that file. While you’re developing, this is what you want. But before you share results, you need to know that they are produced by running all the code, from the raw data, in the right order. Otherwise, your analysis might be using processed data from an older version of your processing code.
_quarto.yml makes this a single step. It lists the files to render, in order:
project:
type: default
render:
- code/processing.qmd
- code/analysis.qmd
# run each file with its own folder as the working directory
execute-dir: fileTo render the whole project, use any of these:
- RStudio: click Build > Render Project in the Build pane (top right).
- R console:
quarto::quarto_render(), with the working directory set to the project root (automatic when you have opened the.Rprojfile). - Terminal:
quarto renderfrom the project root.
Rendering stops at the first error, so your analysis never runs on stale or half-processed data. If you add a new .qmd file, add it to the render: list in the position it should run, or it won’t be rendered with the rest of the project.
9.4.6 Codebooks
Every data file should have a codebook that describes each of its variables. {psychdsish} expects codebooks to be named after the data file they describe, with _codebook in place of _data, and stored next to it, e.g.:
data/processed/
├── study-1_stage-processed_data.csv
└── study-1_stage-processed_codebook.csv
You don’t need to write codebooks from scratch. The processing.qmd template contains a chunk that creates the codebook from your processed data. Change the data_processed object and file names in that chunk to match your own. Each time you render the file, it fills in what can be worked out automatically: each variable’s type, number of missing values (n_missing), and range or unique values. It adds three columns for you to complete by hand, marked “TO BE COMPLETED MANUALLY”:
description: what the variable is, e.g., the item wording, or how a score was calculated.units: e.g., years or milliseconds. Write “none” if it doesn’t apply.coding: what the values mean, e.g., “1 = strongly disagree to 7 = strongly agree”, which items are reverse-scored, or codes for missing values such as -99.
For example, a completed codebook might look like this:
| variable | type | n_missing | values | description | units | coding |
|---|---|---|---|---|---|---|
| id | character | 0 | 214 unique values, e.g., p001; p002; p003 | Participant ID | none | none |
| age | numeric | 3 | 18 to 64 | Self-reported age | years | none |
| condition | character | 0 | control; intervention | Experimental condition, randomly assigned | none | none |
| bdi_sum | numeric | 5 | 0 to 49 | Sum score of the 21 BDI-II items | none | Higher = more depressive symptoms. NA if any item missing |
Open the .csv (e.g., in Excel), replace every placeholder, and save it, still as .csv. When you re-render processing.qmd, your entries are kept, new variables are added, and variables no longer in the data are removed.
Raw data should have a codebook too, although it is often supplied with the data (e.g., exported from Qualtrics) rather than generated by you.
AI assistants can help draft descriptions, but only from information they can actually see. For example, you can ask one to read processing.qmd and describe how each variable was created. It can’t know what your items said or what your codes mean, and it will guess convincingly if asked. Check every entry against your study materials in methods/, and don’t keep any description you can’t verify.
9.4.6.1 Optional: psych-DS metadata
If you want your project to be fully psych-DS compliant, write_dataset_description() creates dataset_description.json in the project root from your codebooks, so each variable is still only described once, in the codebook:
psychdsish::write_dataset_description(
project_root = "..",
name = "My study",
description = "What the dataset contains"
)The processing.qmd template includes this chunk, set not to run by default. You only need the codebook .csv files or the .json, not both, and validator() accepts either. Use the psych-DS validator to check compliance with psych-DS itself.
9.4.7 Validating a project
Projects drift. Someone saves a plot into code/, downloads a data file with a space in its name, or adds a setwd() “just to test something”. validator() checks the project against the rules of the standard and tells you what to fix.
9.4.7.1 Running the validator
- RStudio: with the project open, click Addins > Validate psych-DS-ish project in the toolbar. The results are printed in the console and shown as a colour-coded table in the Viewer pane. You can also assign it a keyboard shortcut via Tools > Modify Keyboard Shortcuts, searching for “psych-DS-ish”.
- R console: from the project root, run
psychdsish::validator("."). - Report: render
tools/project_validator.qmdfor an .html report.
Each check is reported as one of:
PASS: the rule is followed.FAIL: the rule is broken. The guidance says what to change, and which files are affected.WARN: something that isn’t wrong but falls short of full psych-DS compliance, e.g., a raw data file name that doesn’t follow thekey-valueconvention. Raw data files often can’t be renamed, so this isn’t a failure.SKIP: the check couldn’t be run, e.g., there are no rendered .html files to check yet.
9.4.7.2 A freshly created project
Let’s validate the project we just created:
results <- validator(project)
summary(results)[c("n_pass", "n_fail", "n_warn", "n_skip")]$n_pass
[1] 32
$n_fail
[1] 1
$n_warn
[1] 0
$n_skip
[1] 6
Even a brand-new project fails one check:
results[results$Status == "FAIL", ]| Test | Status | Details / Guidance |
|---|---|---|
| README has been customised (no template placeholders) | FAIL | Replace the template text in README.md: ‘# Project Title’; ‘Add aims, data sources, and reproduction steps.’; ‘Authors (Year). Title. URL.’ |
This is deliberate. The README template contains placeholder text, and a project isn’t finished until you’ve replaced it with a description of your project.
9.4.7.3 A messy project
Now let’s break the project in some common ways: save a data file in code/, add raw data with a space in its name, add processed data without a codebook, put an .R script in the project root, and use setwd() and an absolute path in the processing code.
# a data file saved in the code folder
write.csv(mtcars, file.path(project, "code", "mtcars.csv"))
# raw data with a space in its name
write.csv(mtcars, file.path(project, "data", "raw", "final data v2.csv"))
# processed data without a codebook
write.csv(mtcars,
file.path(project, "data", "processed", "study-1_stage-processed_data.csv"),
row.names = FALSE)
# an .R script outside code/
writeLines("x <- 1", file.path(project, "analysis.R"))
# setwd() and an absolute path in the processing code
writeLines(c("```{r}",
"setwd('/Users/ian/Desktop/thesis')",
"dat <- read.csv('C:/Users/ian/data.csv')",
"```"),
file.path(project, "code", "processing.qmd"))Then validate it again, this time only showing the checks that didn’t pass:
results_messy <- validator(project)
results_messy[results_messy$Status %in% c("FAIL", "WARN"), ]| Test | Status | Details / Guidance |
|---|---|---|
| All .R files reside in code, tools | FAIL | Move/remove: analysis.R |
| All .csv files reside in data | FAIL | Move/remove: code/mtcars.csv |
| Every data file has a codebook | FAIL | Document: data/processed/study-1_stage-processed_data.csv. Add ’ |
| No absolute file paths in code | FAIL | Replace absolute paths with relative ones (e.g., ../data/raw/) in: code/processing.qmd:2; code/processing.qmd:3 |
| No data files stored under code | FAIL | Move: code/mtcars.csv |
| No setwd() calls in code | FAIL | Remove setwd() from: code/processing.qmd:2. Each .qmd runs from its own folder, so use relative paths (e.g., ../data/raw/). |
| No spaces in filenames | FAIL | Rename: final data v2.csv |
| README has been customised (no template placeholders) | FAIL | Replace the template text in README.md: ‘# Project Title’; ‘Add aims, data sources, and reproduction steps.’; ‘Authors (Year). Title. URL.’ |
| Every raw data file has a codebook | WARN | Undocumented: data/raw/final data v2.csv. Raw data should have a codebook too, e.g., ’ |
| Raw data file names follow the psych-DS convention | WARN | These do not follow the psych-DS convention (e.g., ‘study-1_stage-raw_data.csv’): data/raw/final data v2.csv. Rename them only if the raw data have not been shared or committed yet; otherwise leave them as they are and note the names in the codebook. |
Each problem is identified, along with the files (and for code, the line numbers) involved and what to do about it. Note that some rules are checked from more than one angle, e.g., code/mtcars.csv fails both because .csv files belong in data/ and because data files shouldn’t be stored under code/.
9.4.7.4 What the validator checks
The full list of rules is given in the psychdsish README. In summary, it checks that:
- The standard folders exist, along with a README, a licence, and a
.gitignore. - Files are in the folder their type belongs in: code in
code/ortools/, data indata/, plots indata/outputs/plots/, and so on. Data files are never stored undercode/. - No file names contain spaces, and data file names follow the psych-DS
key-valueconvention (as a warning). - The code contains no
setwd()calls or absolute paths. - Every data file has a codebook (or is described in
dataset_description.json), and no codebook still contains “TO BE COMPLETED MANUALLY”. - Rendered .html files are newer than their .qmd, i.e., the reports you share reflect the current code.
- The README has been customised.
- Raw data hasn’t been changed since it was first committed to git.
The last check is only possible if the project uses git, which records every version of every file. If the project isn’t a git repository, the check is reported as SKIP. We’ll cover git and GitHub in the chapter on Sharing and privacy; from then on, this check enforces the rule that raw data is read-only.
9.4.7.5 When to validate
Validate whenever you reach a milestone: before a lab meeting, before sending the project to a collaborator or supervisor, before submitting a thesis or paper, and before making a repository public. It takes a second, and it is much less embarrassing than a reviewer finding the problem.
If you use GitHub Actions, validator(".", strict = TRUE) throws an error if any check fails, so that the validator can run automatically every time you push changes.
9.4.8 Restructuring an existing project
The easiest time to structure a project is at the start. But if you already have a project like thesis stuff/ above, you can retrofit the structure:
- Make a backup copy of the whole project folder first.
- Run
create_project_skeleton()on the existing folder. Existing files are never overwritten (unless you setoverwrite = TRUE), so this only adds the missing folders and template files. - Move each file into its place, using Where does this file go? Work out which data files are truly raw. Delete duplicate or obsolete versions once you’re sure they aren’t needed, or leave them in the backup.
- Update the paths in your code to the new locations, and remove any
setwd()calls. - Render the whole project to check that everything still runs from the raw data.
- Run
validator(), fix what it reports, and repeat until it passes.
9.4.9 Multi-study projects
Many papers report several studies. Set studies (or Number of studies in the RStudio wizard) and choose a layout:
# one folder per study, each with the full structure (default)
psychdsish::create_project_skeleton("~/git/my_project", studies = 2, layout = "by_study")
# one set of folders, with a subfolder per study inside each
psychdsish::create_project_skeleton("~/git/my_project", studies = 2, layout = "by_type")layout = "by_study" gives each study its own copy of the single-study structure, with shared reports/ and tools/ folders:
my_project/
├── study_1/
│ ├── code/
│ ├── data/
│ ├── methods/
│ └── preregistration/
├── study_2/
│ └── ... (same as study_1/)
├── reports/
└── tools/
layout = "by_type" keeps one code/, data/, etc., with a subfolder per study inside each:
my_project/
├── code/
│ ├── study_1/
│ └── study_2/
├── data/
│ ├── raw/
│ │ ├── study_1/
│ │ └── study_2/
│ └── ...
├── methods/
│ ├── study_1/
│ └── study_2/
└── ...
Which to choose? by_study is usually simpler: each study is self-contained, and the paths in its code are the same as in a single-study project (e.g., ../data/raw/). by_type suits studies that share a lot, e.g., the same materials and processing steps, or analyses that combine studies, but paths gain a level (e.g., ../../data/raw/study_1/).
In both layouts, the README, licence, _quarto.yml, and other root files are shared, and _quarto.yml renders each study’s processing and analysis files in turn. validator() detects the layout automatically. To add a study later, increase studies in tools/project_creator.qmd and render it, then add the new files to _quarto.yml.
9.4.10 Other tools
The tools/ folder also contains:
project_creator.qmd: re-runscreate_project_skeleton()with the settings the project was created with, e.g., to add a study, or to restore a template file you deleted.style_all_files.qmd: applies the tidyverse code style to every .qmd, .Rmd, and .R file in the project. See the section on Code styling to refresh your knowledge.
9.5 Exercises
Check your learning with the following questions.
9.5.1 Where does each file go?
For each file from the thesis stuff/ folder above, which folder should it go in within a psych-DS-ish project? Assume that old/data.xlsx is the data as it was originally downloaded, and that data.xlsx was later edited by hand.
analysis FINAL v2 (use this one).Rold/data.xlsxdata_cleaned_NEW.csvfig1_revised.pngmodel.rdsquestionnaire.pdflab meeting notes.docx
analysis FINAL v2 (use this one).R:code/, ideally converted to a .qmd file and renamed (e.g.,analysis.qmd). The other versions can be deleted once you’ve checked they aren’t needed.old/data.xlsx:data/raw/. It is the earliest form of the data, so it’s the one to keep.data.xlsxwas edited by hand, so it can’t be trusted: rewrite those edits as code inprocessing.qmd, working from the original.data_cleaned_NEW.csv:data/processed/, but only if code creates it. If you don’t know which code created it, delete it and regenerate it from the raw data withprocessing.qmd.fig1_revised.png:data/outputs/plots/, and it should be written there byanalysis.qmd, e.g., usingggsave().model.rds:data/outputs/fitted_models/.questionnaire.pdf:methods/.lab meeting notes.docx:reports/, if it is about the project. Arguably it doesn’t belong in the project at all.
9.5.2 Rename these files
Rename the following files to follow the principles in this chapter, and the psych-DS convention for data files.
Pilot Data 23.9.26.csv(raw self-report data from a pilot study)Study 2 - cleaned!.csv(processed data from study 2)plot 1.png,plot 2.png, …plot 12.png
There are many good answers. For example:
study-pilot_task-selfreports_stage-raw_data.csv, with the date the data was collected noted in the README or codebook. Remember that raw data should only be renamed if it hasn’t been shared or committed to git yet.study-2_stage-processed_data.csvplot_01.png,plot_02.png, …plot_12.png, so that they sort in the correct order. Even better, add what each plot shows, e.g.,plot_01_self_reports.png.
9.5.3 Why is raw data read-only?
A colleague notices that a few participants entered their age as “999” and fixes this directly in the raw data file in Excel, arguing that it’s quicker than writing code. What problems does this cause?
- The change is undocumented: nobody, including your colleague in six months, can see that the data was changed, how, or why.
- It can’t be undone: the original values are lost, so if the decision was wrong (e.g., “999” was actually a missing data code that should have been handled differently), it can’t be revisited.
- It can’t be reproduced or checked: anyone else starting from the original data will get different results.
- Opening and re-saving data in Excel can silently change other values too (see the chapter on Loading data).
Instead, write code in processing.qmd to recode the values, e.g., mutate(age = if_else(age == 999, NA, age)), with a comment explaining why.
9.5.4 Create your own project
- Create a new psych-DS-ish project called
likert_project, using either the RStudio wizard orcreate_project_skeleton(). - Copy this book’s
data/raw/data_likert.csvinto the new project’sdata/raw/folder, and rename it to follow the psych-DS convention. - In
code/processing.qmd, load the raw data, make at least one change to it (e.g., rename a column), assign it todata_processed, and write it todata/processed/with a psych-DS file name. Update the codebook chunk to use the matching codebook name. - Render the whole project, then complete the codebook.
- Replace the placeholder text in the README.
- Run the validator, and fix anything it reports until every check passes or is skipped.
The raw data file could be named study-1_task-likert_stage-raw_data.csv and the processed data study-1_stage-processed_data.csv. Its codebook is then study-1_stage-processed_codebook.csv.
The validator will warn that the raw data has no codebook. Writing one by hand for a small file is good practice.
The processing code in code/processing.qmd might look like this:
library(readr)
library(dplyr)
data_raw <- read_csv("../data/raw/study-1_task-likert_stage-raw_data.csv")
data_processed <- data_raw |>
rename_with(tolower)
write_csv(data_processed, "../data/processed/study-1_stage-processed_data.csv")And in the codebook chunk:
codebook_path <- "../data/processed/study-1_stage-processed_codebook.csv"Render the project with Build > Render Project, fill in the description, units, and coding columns in the codebook, edit the README, and run Addins > Validate psych-DS-ish project.
9.5.5 Fix a messy project
Run the following code to create a messy project in your home folder, then use the validator to find the problems and fix all of them.
library(psychdsish)
messy <- "~/messy_project"
create_project_skeleton(messy)
write.csv(mtcars, file.path(messy, "code", "mtcars.csv"))
write.csv(iris, file.path(messy, "data", "raw", "iris data.csv"), row.names = FALSE)
writeLines("x <- 1", file.path(messy, "analysis.R"))
writeLines(c("```{r}", "setwd('~/Desktop')", "```"), file.path(messy, "code", "processing.qmd"))
validator(messy)When you’re done, delete the folder.
- Move
code/mtcars.csvtodata/raw/, and rename it (e.g.,study-1_task-cars_stage-raw_data.csv). - Rename
iris data.csvto remove the space (e.g.,study-1_task-flowers_stage-raw_data.csv). - Move
analysis.Rintocode/. - Remove the
setwd()call fromcode/processing.qmd. - Replace the README’s placeholder text.
- Optionally, add codebooks for the raw data files to clear the warnings.