9  Structuring projects ✎ Rough draft

So far, we have mostly worked with single .qmd files. Real research projects contain many files: raw data from several sources, code to process and analyze it, plots and tables, study materials, a preregistration, and a manuscript. How these files are organized determines whether the project can be understood, checked, and reproduced by anyone else, including yourself in six months’ time.

This chapter covers:

  1. The general principles behind a well-structured project, which apply whatever tools you use.
  2. Data standards, and psych-DS specifically, which turn these principles into a shared set of rules.
  3. How to use the {psychdsish} R package to create projects that follow these rules from the start, and to check that they still do.

9.1 Why structure matters

Here is a (lightly fictionalized) project folder of the kind I receive from students, collaborators, and, if I’m honest, my past self:

thesis stuff/
├── analysis FINAL.R
├── analysis FINAL v2 (use this one).R
├── data.xlsx
├── data_cleaned.xlsx
├── data_cleaned_NEW.csv
├── Figure1.png
├── fig1_revised.png
├── lab meeting notes.docx
├── model.rds
├── output.html
├── questionnaire.pdf
└── old/
    ├── analysis.R
    └── data.xlsx

Try to answer the following questions about it:

  • Which file contains the data as it was originally collected?
  • Was data.xlsx edited by hand after it was downloaded? Which of the two data.xlsx files is the original?
  • Which script created data_cleaned_NEW.csv, and which script reads it?
  • Does analysis FINAL v2 (use this one).R produce Figure1.png or fig1_revised.png? Which one is in the manuscript?
  • If I delete model.rds, can I get it back?

You can’t answer any of these questions from the folder alone, and often the person who created it can’t either. None of these problems are caused by bad code. They are caused by the absence of a structure that makes the answers obvious.

This matters for more than tidiness. A project that can’t be understood can’t be checked, so errors go unnoticed (see the case studies in the chapter on Loading data). A project that can’t be re-run from the raw data can’t be reproduced, which is the minimum standard for computational work. And a project that only its creator can navigate becomes unusable when that person leaves, forgets, or loses their laptop.

9.2 General principles

The following principles are not specific to R, psych-DS, or any particular tool. They are distilled from standards like psych-DS (discussed below) and from guides such as Good enough practices in scientific computing (Wilson et al., 2017).

9.2.1 1. One project, one folder

Everything needed to understand and reproduce a project lives in a single folder, and nothing outside that folder is needed. That folder can be moved anywhere on your computer, zipped and emailed, or uploaded to GitHub or OSF, and it will still work.

This is only possible if code refers to files using relative paths (e.g., ../data/raw/data.csv) rather than absolute paths (e.g., C:/Users/ian/Documents/thesis/data.csv) or setwd(). See the section on relative vs. absolute paths to refresh your knowledge.

In RStudio, opening a project folder via its .Rproj file also sets the console’s working directory to the project folder, and keeps each project’s environment and history separate.

9.2.2 2. Separate files by their role

Different kinds of files have different jobs, and should live in different places. At a minimum:

Role What it contains Can it be regenerated?
Raw data Data as it was collected or received No. Must be preserved
Code Scripts that process and analyze the data No. This is your work
Processed data Cleaned data created by code Yes, by re-running the code
Outputs Plots, tables, fitted models created by code Yes, by re-running the code
Materials Measures, experiment files, procedures No
Documents Preregistration, manuscript, slides No

The final column is the most important one. Files that cannot be regenerated must be protected. Files that can be regenerated are disposable: if in doubt, delete them and re-run the code. Keeping these two kinds of files apart makes it obvious which is which.

9.2.3 3. Raw data is read-only

The earliest form of the data you have access to must be preserved exactly as it was received. Never edit it by hand, never overwrite it with code, and never save over it after opening it in Excel. Every change to the data is instead made by code that reads the raw data and writes a new file somewhere else.

The only exception is the removal of private or identifying information before sharing, which is covered in the chapter on Sharing and privacy. Even then, the original is kept (securely, and not shared) and the removal is done with code.

9.2.4 4. Data flows in one direction, and only code moves it

Data moves from raw, through processing code, to processed data, and then through analysis code to outputs. Nothing flows backward: analysis code never writes to data/raw/, and processing code never reads from data/outputs/.

flowchart LR
  A[("data/raw/")] --> B["code/processing.qmd"]
  B --> C[("data/processed/")]
  C --> D["code/analysis.qmd"]
  D --> E[("data/outputs/")]

Because every arrow is code, every transformation is documented, and anyone can see exactly how the numbers in a paper were derived from the data.

9.2.5 5. Everything can be rebuilt from scratch, in order

A good test of a project is: if I delete everything in data/processed/ and data/outputs/, can I get it all back by running the code? If the answer is “yes, but only if you run these three scripts in the right order and remember to skip chunk 7”, the project is not reproducible yet.

The code should therefore run from start to finish without manual intervention, and the order in which scripts run should be written down, ideally in a form the computer can follow (see Rendering the whole project below).

9.2.6 6. Name files for computers and for humans

File names should be:

  • Machine readable: no spaces, no special characters (&, ?, !, (, accents), and consistent case. Spaces in particular break many tools.
  • Human readable: the name says what is in the file.
  • Sortable: when files sort alphabetically, they appear in a useful order. Dates follow ISO 8601 (2026-09-23, not 23.9.26), and numbers are zero-padded (01, 02, … 10, not 1, 2, … 10, which sorts as 1, 10, 2).
Bad Better
analysis FINAL v2 (use this one).R analysis.qmd (and use version control for versions)
data_cleaned_NEW.csv study-1_stage-processed_data.csv
Figure1.png, fig1_revised.png plot_01_self_reports.png
23.9.26 pilot.csv 2026-09-23_pilot_data.csv

Note that “final”, “v2”, and “NEW” in file names are attempts to do version control by hand. They reliably fail, because there is always a “final_v3”. Version control with git and GitHub, covered in the chapter on Sharing and privacy, solves this properly: there is only ever one analysis.qmd, and git remembers every previous version of it.

See Jenny Bryan’s How to name files for more on this.

9.2.7 7. Document the project

A project needs at least:

  • A README: what the project is, how it is organized, and how to reproduce its results. It is the first file anyone opens, and on GitHub it is displayed on the repository’s front page.
  • A codebook (data dictionary) for each data file: what each variable is, its units, and what its values mean. A column called q7_r with values from 1 to 5 is meaningless without one.
  • A licence: without one, others are not legally allowed to reuse your work, even if it is public.

9.2.8 8. Follow a convention rather than inventing one

You could follow all of the above principles and still organize your project differently from everyone else. If everyone invents their own structure, everyone else has to learn it before they can find anything. The final principle is therefore to use a structure that others already know, i.e., a standard.

9.3 Data standards

A standard is a specification of what must be done, what must not be done, and what may be done to allow flexibility. For project structure, this means rules about which folders exist, which files go where, how files are named, and what documentation is required.

Standards have two big advantages over good intentions:

  1. Shared expectations: anyone who knows the standard can navigate any project that follows it, without needing to read the README first. Raw data is always in the same place.
  2. Automatic checking: because the rules are written down precisely, a computer can check whether a project follows them. This is much more reliable than trying to remember all the rules yourself, and it catches problems as they arise rather than when a reviewer asks for your data.

Several standards and conventions exist for different fields, for example BIDS for neuroimaging data, the TIER Protocol for social science projects, and Cookiecutter Data Science for Python data science projects. They differ in the details but share the principles above.

9.3.1 psych-DS

psych-DS is a data standard for psychological data, led by Melissa Kline Struhl. It was inspired by BIDS but aims to be much lighter-weight, so that it can be used for the huge variety of data collected in psychology. I have contributed a very small amount to the debate around it.

psych-DS is built on four principles, which you will recognize from the previous section:

  • The earliest form of data must be preserved.
  • Original data should never be modified.
  • Different versions of the data should be kept separate.
  • All transformations should be documented.

In practice, a psych-DS dataset has:

  • A data/ folder containing the data files, which are .csv files.
  • Data file names made of key-value pairs separated by underscores and ending in _data.csv, e.g., study-1_task-stroop_data.csv. This makes the contents of each file clear from its name, and lets software find, for example, all the Stroop task data across studies.
  • A dataset_description.json file in the project root, containing machine-readable metadata about the dataset, including a description of every variable.

psych-DS provides an online validator that checks whether a dataset complies with the standard.

9.3.2 Why psych-DS-ish?

I really like the concept of psych-DS, but I am not yet convinced by some of its choices, at least as a starting point for most researchers and for students on courses like this one:

  1. The .json metadata file is required. .json files are a pain to write by hand, and very few psychology workflows currently use them. I didn’t want creating one to be the price of entry.
  2. psych-DS is deliberately light on what it requires, so that it can apply to as many datasets as possible. For my own projects and for teaching, I’m happy to be more prescriptive, e.g., about where code, outputs, and study materials go, not just data.
  3. psych-DS focuses on checking whether a project is compliant, not on helping people set up a compliant project in the first place. Tidying up a project after the fact is much harder than starting with a template.

For the moment, I therefore recommend partial compliance with psych-DS: its high-level principles and file naming conventions make for a very well structured project, but other parts are (currently) high-effort-low-reward. That is what {psychdsish} implements.

9.4 psych-DS-ish

{psychdsish} (‘psych-DS-ish’) is an R package I wrote to make the principles above easy to follow. It does three main things:

  1. create_project_skeleton() creates a new project with the standard folder structure, plus templates for the README, processing and analysis code, a licence, a citation file, and more.
  2. validator() checks whether a project follows the standard and, if not, tells you how to fix it.
  3. write_dataset_description() optionally creates the psych-DS dataset_description.json from your codebooks, for when you want full psych-DS metadata.

It is ‘compliant-ish’ with psych-DS: it follows its principles and naming conventions, but makes the .json file optional.

9.4.1 Installation

# install.packages("remotes")
remotes::install_github("ianhussey/psychdsish")

If you use RStudio, restart it after installing so that the New Project wizard and the Addins menu pick up the package.

9.4.2 Creating a new project

9.4.2.1 In RStudio

Go to File > New Project > New Directory > psych-DS-ish Project. Give the project a directory name and choose where to create it. The other options can usually be left at their defaults:

  • Create _quarto.yml: on by default. This lets you render the whole project in order (see below).
  • Number of studies and Multi-study layout: for projects with more than one study (see Multi-study projects).

Click Create Project. RStudio creates the project, opens it, and opens README.md, code/processing.qmd, and code/analysis.qmd ready for you to start work.

Here’s a demo of the project creator in action:

9.4.2.2 From the console

In Positron, VS Code, or any other editor, or if you prefer code to menus, run this from the R console and then open the folder:

psychdsish::create_project_skeleton(project_root = "~/git/my_project")

create_project_skeleton() never overwrites existing files unless you set overwrite = TRUE, so it is safe to run on a folder that already exists (see Restructuring an existing project).

9.4.3 What you get

Let’s create a project in a temporary folder and look at what’s in it. fs::dir_tree() prints a folder’s contents as a tree, and all = TRUE includes hidden files, whose names start with a .:

library(psychdsish)

project <- file.path(tempdir(), "my_project")
create_project_skeleton(project_root = project)

fs::dir_tree(project, all = TRUE)
/var/folders/45/d07jd4jn4756zs6q38rgl4xw0000gp/T//RtmpLiPBkh/my_project
├── .gitattributes
├── .gitignore
├── CITATION.cff
├── LICENSE
├── README.md
├── _quarto.yml
├── code
│   ├── analysis.qmd
│   └── processing.qmd
├── data
│   ├── outputs
│   │   ├── fitted_models
│   │   │   └── .gitkeep
│   │   ├── plots
│   │   │   └── .gitkeep
│   │   └── results
│   │       └── .gitkeep
│   ├── processed
│   │   └── .gitkeep
│   └── raw
│       └── .gitkeep
├── methods
│   └── .gitkeep
├── my_project.Rproj
├── preregistration
│   └── .gitkeep
├── reports
│   └── .gitkeep
└── tools
    ├── project_creator.qmd
    ├── project_validator.qmd
    └── style_all_files.qmd

(The first line is the location of the temporary folder on the computer that rendered this book. Yours will be wherever you created the project.)

The empty folders each contain a .gitkeep file. These are empty placeholders that exist only so that git, which ignores empty folders, keeps the folder structure when you put the project on GitHub.

Each folder has a clear role:

Folder What goes in it Examples
code/ Processing and analysis code, and the .html reports they render processing.qmd, analysis.qmd, analysis.html
data/raw/ Data as it was collected, never modified, and its codebooks Qualtrics or lab.js exports, e.g., study-1_task-selfreports_stage-raw_data.csv
data/processed/ Cleaned data written by processing.qmd, and their codebooks study-1_stage-processed_data.csv, study-1_stage-processed_codebook.csv
data/outputs/plots/ Plots written by analysis.qmd plot_01_self_reports.png
data/outputs/fitted_models/ Fitted model objects, which can be slow to re-fit fit_model_1.rds (e.g., from {lme4}, {brms}, {lavaan})
data/outputs/results/ Tables and other results descriptives.csv, cor_matrix.csv
methods/ Study materials Item wordings, Qualtrics .qsf files, PsychoPy or lab.js files, procedure documents
preregistration/ Preregistration documents preregistration.pdf
reports/ Anything written about the project Thesis, manuscript, preprint, slides, posters
tools/ Utilities for managing the project, not part of the analysis project_validator.qmd, style_all_files.qmd

And the files in the project root:

File What it’s for
README.md What the project is, how it’s organized, and how to reproduce it. Contains placeholder text to replace with your own.
LICENSE CC BY 4.0: others may reuse your work as long as they credit you.
CITATION.cff Machine-readable citation information. On GitHub, it adds a Cite this repository button that gives APA and BibTeX citations.
_quarto.yml Lists which .qmd files to render, and in which order, when rendering the whole project.
my_project.Rproj The RStudio project file. Double-click it to open the project. It is set not to save or restore your workspace, so that your code, not leftover objects, determines your results.
.gitignore Tells git which files not to track, e.g., R history files, caches, .DS_Store files, and large outputs.
.gitattributes Stops GitHub from labeling your repository as an HTML project because of the rendered reports.

Note that .gitignore excludes data/outputs/plots/ and data/outputs/fitted_models/ by default, because they can be regenerated from the code and model objects can be very large. If you want your plots to appear on GitHub, delete those lines from .gitignore.

9.4.3.1 Where does this file go?

When you’re unsure, ask two questions: Did code create it? and Could I recreate it if it were deleted?

  • Data you received or collected, that code did not create: data/raw/.
  • Data that code created from other data: data/processed/.
  • Plots, tables, and model objects that code created: data/outputs/.
  • Code that does the processing or analysis: code/.
  • Things participants saw or did: methods/.
  • Things you wrote about the project, before it (preregistration/) or after it (reports/).

9.4.4 Working in the project

The workflow follows the one-directional data flow described above:

  1. Put the raw data in data/raw/, and don’t touch it again. If you can still choose its name (i.e., it hasn’t been shared or committed to git yet), follow the psych-DS convention, e.g., study-1_task-selfreports_stage-raw_data.csv.
  2. Write code/processing.qmd to read the raw data, clean it, and write the result to data/processed/.
  3. Write code/analysis.qmd to read the processed data and write plots, fitted models, and tables to data/outputs/.

Each .qmd file runs with its own folder as the working directory. Because both code files are in code/, paths from them always start by going ‘up’ one level with ../:

library(readr)

# in code/processing.qmd
data_raw <- read_csv("../data/raw/study-1_task-selfreports_stage-raw_data.csv")

# ... processing ...

write_csv(data_processed, "../data/processed/study-1_stage-processed_data.csv")
library(readr)
library(ggplot2)

# in code/analysis.qmd
data_processed <- read_csv("../data/processed/study-1_stage-processed_data.csv")

# ... analysis ...

ggsave("../data/outputs/plots/plot_01_self_reports.png", p_self_reports)

It is fine to split processing or analysis across several files if they get long, e.g., processing_selfreports.qmd and processing_behavioral.qmd. Keep them in code/, and add them to _quarto.yml (see next section).

9.4.5 Rendering the whole project

Clicking Render in a .qmd file renders only that file. While you’re developing, this is what you want. But before you share results, you need to know that they are produced by running all the code, from the raw data, in the right order. Otherwise, your analysis might be using processed data from an older version of your processing code.

_quarto.yml makes this a single step. It lists the files to render, in order:

project:
  type: default
  render:
    - code/processing.qmd
    - code/analysis.qmd
  # run each file with its own folder as the working directory
  execute-dir: file

To render the whole project, use any of these:

  • RStudio: click Build > Render Project in the Build pane (top right).
  • R console: quarto::quarto_render(), with the working directory set to the project root (automatic when you have opened the .Rproj file).
  • Terminal: quarto render from the project root.

Rendering stops at the first error, so your analysis never runs on stale or half-processed data. If you add a new .qmd file, add it to the render: list in the position it should run, or it won’t be rendered with the rest of the project.

9.4.6 Codebooks

Every data file should have a codebook that describes each of its variables. {psychdsish} expects codebooks to be named after the data file they describe, with _codebook in place of _data, and stored next to it, e.g.:

data/processed/
├── study-1_stage-processed_data.csv
└── study-1_stage-processed_codebook.csv

You don’t need to write codebooks from scratch. The processing.qmd template contains a chunk that creates the codebook from your processed data. Change the data_processed object and file names in that chunk to match your own. Each time you render the file, it fills in what can be worked out automatically: each variable’s type, number of missing values (n_missing), and range or unique values. It adds three columns for you to complete by hand, marked “TO BE COMPLETED MANUALLY”:

  • description: what the variable is, e.g., the item wording, or how a score was calculated.
  • units: e.g., years or milliseconds. Write “none” if it doesn’t apply.
  • coding: what the values mean, e.g., “1 = strongly disagree to 7 = strongly agree”, which items are reverse-scored, or codes for missing values such as -99.

For example, a completed codebook might look like this:

variable type n_missing values description units coding
id character 0 214 unique values, e.g., p001; p002; p003 Participant ID none none
age numeric 3 18 to 64 Self-reported age years none
condition character 0 control; intervention Experimental condition, randomly assigned none none
bdi_sum numeric 5 0 to 49 Sum score of the 21 BDI-II items none Higher = more depressive symptoms. NA if any item missing

Open the .csv (e.g., in Excel), replace every placeholder, and save it, still as .csv. When you re-render processing.qmd, your entries are kept, new variables are added, and variables no longer in the data are removed.

Raw data should have a codebook too, although it is often supplied with the data (e.g., exported from Qualtrics) rather than generated by you.

WarningUsing AI to write codebooks

AI assistants can help draft descriptions, but only from information they can actually see. For example, you can ask one to read processing.qmd and describe how each variable was created. It can’t know what your items said or what your codes mean, and it will guess convincingly if asked. Check every entry against your study materials in methods/, and don’t keep any description you can’t verify.

9.4.6.1 Optional: psych-DS metadata

If you want your project to be fully psych-DS compliant, write_dataset_description() creates dataset_description.json in the project root from your codebooks, so each variable is still only described once, in the codebook:

psychdsish::write_dataset_description(
  project_root = "..",
  name = "My study",
  description = "What the dataset contains"
)

The processing.qmd template includes this chunk, set not to run by default. You only need the codebook .csv files or the .json, not both, and validator() accepts either. Use the psych-DS validator to check compliance with psych-DS itself.

9.4.7 Validating a project

Projects drift. Someone saves a plot into code/, downloads a data file with a space in its name, or adds a setwd() “just to test something”. validator() checks the project against the rules of the standard and tells you what to fix.

9.4.7.1 Running the validator

  • RStudio: with the project open, click Addins > Validate psych-DS-ish project in the toolbar. The results are printed in the console and shown as a colour-coded table in the Viewer pane. You can also assign it a keyboard shortcut via Tools > Modify Keyboard Shortcuts, searching for “psych-DS-ish”.
  • R console: from the project root, run psychdsish::validator(".").
  • Report: render tools/project_validator.qmd for an .html report.

Each check is reported as one of:

  • PASS: the rule is followed.
  • FAIL: the rule is broken. The guidance says what to change, and which files are affected.
  • WARN: something that isn’t wrong but falls short of full psych-DS compliance, e.g., a raw data file name that doesn’t follow the key-value convention. Raw data files often can’t be renamed, so this isn’t a failure.
  • SKIP: the check couldn’t be run, e.g., there are no rendered .html files to check yet.

9.4.7.2 A freshly created project

Let’s validate the project we just created:

results <- validator(project)

summary(results)[c("n_pass", "n_fail", "n_warn", "n_skip")]
$n_pass
[1] 32

$n_fail
[1] 1

$n_warn
[1] 0

$n_skip
[1] 6

Even a brand-new project fails one check:

results[results$Status == "FAIL", ]
Test Status Details / Guidance
README has been customised (no template placeholders) FAIL Replace the template text in README.md: ‘# Project Title’; ‘Add aims, data sources, and reproduction steps.’; ‘Authors (Year). Title. URL.’

This is deliberate. The README template contains placeholder text, and a project isn’t finished until you’ve replaced it with a description of your project.

9.4.7.3 A messy project

Now let’s break the project in some common ways: save a data file in code/, add raw data with a space in its name, add processed data without a codebook, put an .R script in the project root, and use setwd() and an absolute path in the processing code.

# a data file saved in the code folder
write.csv(mtcars, file.path(project, "code", "mtcars.csv"))

# raw data with a space in its name
write.csv(mtcars, file.path(project, "data", "raw", "final data v2.csv"))

# processed data without a codebook
write.csv(mtcars,
          file.path(project, "data", "processed", "study-1_stage-processed_data.csv"),
          row.names = FALSE)

# an .R script outside code/
writeLines("x <- 1", file.path(project, "analysis.R"))

# setwd() and an absolute path in the processing code
writeLines(c("```{r}",
             "setwd('/Users/ian/Desktop/thesis')",
             "dat <- read.csv('C:/Users/ian/data.csv')",
             "```"),
           file.path(project, "code", "processing.qmd"))

Then validate it again, this time only showing the checks that didn’t pass:

results_messy <- validator(project)

results_messy[results_messy$Status %in% c("FAIL", "WARN"), ]
Test Status Details / Guidance
All .R files reside in code, tools FAIL Move/remove: analysis.R
All .csv files reside in data FAIL Move/remove: code/mtcars.csv
Every data file has a codebook FAIL Document: data/processed/study-1_stage-processed_data.csv. Add ’_codebook.csv’ (or .xlsx) next to the data file, or describe it in dataset_description.json. The processing.qmd template contains a chunk that creates a codebook.
No absolute file paths in code FAIL Replace absolute paths with relative ones (e.g., ../data/raw/) in: code/processing.qmd:2; code/processing.qmd:3
No data files stored under code FAIL Move: code/mtcars.csv
No setwd() calls in code FAIL Remove setwd() from: code/processing.qmd:2. Each .qmd runs from its own folder, so use relative paths (e.g., ../data/raw/).
No spaces in filenames FAIL Rename: final data v2.csv
README has been customised (no template placeholders) FAIL Replace the template text in README.md: ‘# Project Title’; ‘Add aims, data sources, and reproduction steps.’; ‘Authors (Year). Title. URL.’
Every raw data file has a codebook WARN Undocumented: data/raw/final data v2.csv. Raw data should have a codebook too, e.g., ’_codebook.csv’ or the codebook supplied with the data.
Raw data file names follow the psych-DS convention WARN These do not follow the psych-DS convention (e.g., ‘study-1_stage-raw_data.csv’): data/raw/final data v2.csv. Rename them only if the raw data have not been shared or committed yet; otherwise leave them as they are and note the names in the codebook.

Each problem is identified, along with the files (and for code, the line numbers) involved and what to do about it. Note that some rules are checked from more than one angle, e.g., code/mtcars.csv fails both because .csv files belong in data/ and because data files shouldn’t be stored under code/.

9.4.7.4 What the validator checks

The full list of rules is given in the psychdsish README. In summary, it checks that:

  • The standard folders exist, along with a README, a licence, and a .gitignore.
  • Files are in the folder their type belongs in: code in code/ or tools/, data in data/, plots in data/outputs/plots/, and so on. Data files are never stored under code/.
  • No file names contain spaces, and data file names follow the psych-DS key-value convention (as a warning).
  • The code contains no setwd() calls or absolute paths.
  • Every data file has a codebook (or is described in dataset_description.json), and no codebook still contains “TO BE COMPLETED MANUALLY”.
  • Rendered .html files are newer than their .qmd, i.e., the reports you share reflect the current code.
  • The README has been customised.
  • Raw data hasn’t been changed since it was first committed to git.

The last check is only possible if the project uses git, which records every version of every file. If the project isn’t a git repository, the check is reported as SKIP. We’ll cover git and GitHub in the chapter on Sharing and privacy; from then on, this check enforces the rule that raw data is read-only.

9.4.7.5 When to validate

Validate whenever you reach a milestone: before a lab meeting, before sending the project to a collaborator or supervisor, before submitting a thesis or paper, and before making a repository public. It takes a second, and it is much less embarrassing than a reviewer finding the problem.

If you use GitHub Actions, validator(".", strict = TRUE) throws an error if any check fails, so that the validator can run automatically every time you push changes.

9.4.8 Restructuring an existing project

The easiest time to structure a project is at the start. But if you already have a project like thesis stuff/ above, you can retrofit the structure:

  1. Make a backup copy of the whole project folder first.
  2. Run create_project_skeleton() on the existing folder. Existing files are never overwritten (unless you set overwrite = TRUE), so this only adds the missing folders and template files.
  3. Move each file into its place, using Where does this file go? Work out which data files are truly raw. Delete duplicate or obsolete versions once you’re sure they aren’t needed, or leave them in the backup.
  4. Update the paths in your code to the new locations, and remove any setwd() calls.
  5. Render the whole project to check that everything still runs from the raw data.
  6. Run validator(), fix what it reports, and repeat until it passes.

9.4.9 Multi-study projects

Many papers report several studies. Set studies (or Number of studies in the RStudio wizard) and choose a layout:

# one folder per study, each with the full structure (default)
psychdsish::create_project_skeleton("~/git/my_project", studies = 2, layout = "by_study")

# one set of folders, with a subfolder per study inside each
psychdsish::create_project_skeleton("~/git/my_project", studies = 2, layout = "by_type")

layout = "by_study" gives each study its own copy of the single-study structure, with shared reports/ and tools/ folders:

my_project/
├── study_1/
│   ├── code/
│   ├── data/
│   ├── methods/
│   └── preregistration/
├── study_2/
│   └── ... (same as study_1/)
├── reports/
└── tools/

layout = "by_type" keeps one code/, data/, etc., with a subfolder per study inside each:

my_project/
├── code/
│   ├── study_1/
│   └── study_2/
├── data/
│   ├── raw/
│   │   ├── study_1/
│   │   └── study_2/
│   └── ...
├── methods/
│   ├── study_1/
│   └── study_2/
└── ...

Which to choose? by_study is usually simpler: each study is self-contained, and the paths in its code are the same as in a single-study project (e.g., ../data/raw/). by_type suits studies that share a lot, e.g., the same materials and processing steps, or analyses that combine studies, but paths gain a level (e.g., ../../data/raw/study_1/).

In both layouts, the README, licence, _quarto.yml, and other root files are shared, and _quarto.yml renders each study’s processing and analysis files in turn. validator() detects the layout automatically. To add a study later, increase studies in tools/project_creator.qmd and render it, then add the new files to _quarto.yml.

9.4.10 Other tools

The tools/ folder also contains:

  • project_creator.qmd: re-runs create_project_skeleton() with the settings the project was created with, e.g., to add a study, or to restore a template file you deleted.
  • style_all_files.qmd: applies the tidyverse code style to every .qmd, .Rmd, and .R file in the project. See the section on Code styling to refresh your knowledge.

9.5 Exercises

Check your learning with the following questions.

9.5.1 Where does each file go?

For each file from the thesis stuff/ folder above, which folder should it go in within a psych-DS-ish project? Assume that old/data.xlsx is the data as it was originally downloaded, and that data.xlsx was later edited by hand.

  • analysis FINAL v2 (use this one).R
  • old/data.xlsx
  • data_cleaned_NEW.csv
  • fig1_revised.png
  • model.rds
  • questionnaire.pdf
  • lab meeting notes.docx
  • analysis FINAL v2 (use this one).R: code/, ideally converted to a .qmd file and renamed (e.g., analysis.qmd). The other versions can be deleted once you’ve checked they aren’t needed.
  • old/data.xlsx: data/raw/. It is the earliest form of the data, so it’s the one to keep. data.xlsx was edited by hand, so it can’t be trusted: rewrite those edits as code in processing.qmd, working from the original.
  • data_cleaned_NEW.csv: data/processed/, but only if code creates it. If you don’t know which code created it, delete it and regenerate it from the raw data with processing.qmd.
  • fig1_revised.png: data/outputs/plots/, and it should be written there by analysis.qmd, e.g., using ggsave().
  • model.rds: data/outputs/fitted_models/.
  • questionnaire.pdf: methods/.
  • lab meeting notes.docx: reports/, if it is about the project. Arguably it doesn’t belong in the project at all.

9.5.2 Rename these files

Rename the following files to follow the principles in this chapter, and the psych-DS convention for data files.

  • Pilot Data 23.9.26.csv (raw self-report data from a pilot study)
  • Study 2 - cleaned!.csv (processed data from study 2)
  • plot 1.png, plot 2.png, … plot 12.png

There are many good answers. For example:

  • study-pilot_task-selfreports_stage-raw_data.csv, with the date the data was collected noted in the README or codebook. Remember that raw data should only be renamed if it hasn’t been shared or committed to git yet.
  • study-2_stage-processed_data.csv
  • plot_01.png, plot_02.png, … plot_12.png, so that they sort in the correct order. Even better, add what each plot shows, e.g., plot_01_self_reports.png.

9.5.3 Why is raw data read-only?

A colleague notices that a few participants entered their age as “999” and fixes this directly in the raw data file in Excel, arguing that it’s quicker than writing code. What problems does this cause?

  • The change is undocumented: nobody, including your colleague in six months, can see that the data was changed, how, or why.
  • It can’t be undone: the original values are lost, so if the decision was wrong (e.g., “999” was actually a missing data code that should have been handled differently), it can’t be revisited.
  • It can’t be reproduced or checked: anyone else starting from the original data will get different results.
  • Opening and re-saving data in Excel can silently change other values too (see the chapter on Loading data).

Instead, write code in processing.qmd to recode the values, e.g., mutate(age = if_else(age == 999, NA, age)), with a comment explaining why.

9.5.4 Create your own project

  1. Create a new psych-DS-ish project called likert_project, using either the RStudio wizard or create_project_skeleton().
  2. Copy this book’s data/raw/data_likert.csv into the new project’s data/raw/ folder, and rename it to follow the psych-DS convention.
  3. In code/processing.qmd, load the raw data, make at least one change to it (e.g., rename a column), assign it to data_processed, and write it to data/processed/ with a psych-DS file name. Update the codebook chunk to use the matching codebook name.
  4. Render the whole project, then complete the codebook.
  5. Replace the placeholder text in the README.
  6. Run the validator, and fix anything it reports until every check passes or is skipped.

The raw data file could be named study-1_task-likert_stage-raw_data.csv and the processed data study-1_stage-processed_data.csv. Its codebook is then study-1_stage-processed_codebook.csv.

The validator will warn that the raw data has no codebook. Writing one by hand for a small file is good practice.

The processing code in code/processing.qmd might look like this:

library(readr)
library(dplyr)

data_raw <- read_csv("../data/raw/study-1_task-likert_stage-raw_data.csv")

data_processed <- data_raw |>
  rename_with(tolower)

write_csv(data_processed, "../data/processed/study-1_stage-processed_data.csv")

And in the codebook chunk:

codebook_path <- "../data/processed/study-1_stage-processed_codebook.csv"

Render the project with Build > Render Project, fill in the description, units, and coding columns in the codebook, edit the README, and run Addins > Validate psych-DS-ish project.

9.5.5 Fix a messy project

Run the following code to create a messy project in your home folder, then use the validator to find the problems and fix all of them.

library(psychdsish)

messy <- "~/messy_project"
create_project_skeleton(messy)

write.csv(mtcars, file.path(messy, "code", "mtcars.csv"))
write.csv(iris, file.path(messy, "data", "raw", "iris data.csv"), row.names = FALSE)
writeLines("x <- 1", file.path(messy, "analysis.R"))
writeLines(c("```{r}", "setwd('~/Desktop')", "```"), file.path(messy, "code", "processing.qmd"))

validator(messy)

When you’re done, delete the folder.

  • Move code/mtcars.csv to data/raw/, and rename it (e.g., study-1_task-cars_stage-raw_data.csv).
  • Rename iris data.csv to remove the space (e.g., study-1_task-flowers_stage-raw_data.csv).
  • Move analysis.R into code/.
  • Remove the setwd() call from code/processing.qmd.
  • Replace the README’s placeholder text.
  • Optionally, add codebooks for the raw data files to clear the warnings.