6  Writing Code with AI

This chapter closes the coding section because writing code with AI represents the highest-risk, highest-reward use case. By this point, you should be able to judge when AI suggestions are useful scaffolds and when they are dangerous shortcuts.

If you write code for a living you would probably use an agentic coding tool (Claude Code, Codex, Cursor and the like) or one of the packages that integrate an LLM into RStudio, rather than a chat window. However, unless coding is the main part of your job, most people are likely to use AI through a generic platform so we’ll stick with regular Copilot.

It’s also worth being clear about where this chapter sits. Most professional programmers now use AI to write at least some of their code, and code is probably the thing these models are best at: there is an enormous amount of it in the training data and, unlike an essay, code either runs or it doesn’t. So this chapter is not arguing that you should never use AI to write code. The point is that professional programmers using AI today learned to code without it. They can read what it produces, spot when it’s wrong, and describe what they want precisely enough to get it, and none of that is possible without having done the work first. The skill you’re building now is what makes AI useful later. Skip it and you’re not a programmer with a faster tool, you’re someone who can’t check the tool’s work.

Additionally, if you’ve worked through this entire book then hopefully what you’ve learned is that AI is very useful but you should also have a healthy mistrust of anything it produces which is why this chapter is the last in the coding section.

6.1 The problem

The dataset used in this exercise is described in the paper Data from an International Multi-Centre Study of Statistics and Mathematics Anxieties and Related Variables in University Students (the SMARVUS Dataset) from the Journal of Open Psychology Data.

The dataset is massive so as our first task, we want to write code that narrows it down and creates three different objects.

  • First, we want to create an object named demo that has the demographic information, age and gender. Age data has transformed into categories (e.g., 18-21) for anonymisation purposes. The option ‘Implausible’ describes values that were 99 or higher or 17 or lower.

  • Second, we want an object named STARS which will represent each participant’s score on the Statistics Anxiety Rating Scale (STARS; Cruise et al., 1985). Each item describes a situation involving statistics such as “Doing an examination in a statistics course” (test and class anxiety), “Interpreting the meaning of a table in a journal article” (interpretation anxiety), or “Going to ask my statistics teacher for individual help with material I am having difficulty understanding” (fear of asking for help).

  • Third, we want an object named, IUS which will represent each participant’s score on the Intolerance of Uncertainty Scale – Short Form (IUS-SF; Carleton et al., 2007). The scale contains 2 subscales, Prospective Anxiety and Inhibitory Anxiety, each with 6 items. The Prospective Anxiety subscale includes statements such as, “The smallest doubt can stop me from acting”. The Inhibitory Anxiety subscale includes statements such as, “It frustrates me not having all the information I need”.

For the two scales, we need to pivot it to long-form and then calculate the mean score for each participant.

We would need this information if we wanted to try and answer questions like:

  • What is the demographic make-up of the sample in terms of age and gender?
  • How is statistics anxiety and intolerance of uncertainty related?
  • Do men and women differ in their level of statistics anxiety?

But we’ll start with the wrangling and see how far we get.

NoteActivity 1

First, download the dataset from the Open Science Framework. Create a new Quarto or Rmd document and ensure the dataset is in the same folder, as this document, and then run the below code in a new code chunk. This will load in the data and then it will create the three objects using the code I have written and the approach I have taken so that we can compare with how Copilot gets on.

This is such a large dataset that we will just select the columns we need, but bear in mind I have already made the task easier for the AI by doing so.

library(tidyverse)

# load in the data and just select the columns we will need

dat <- read_csv("smarvus_complete_050524.csv") |>
  select(unique_id, age, gender, Q7.1_1:Q7.1_24, Q14.1_1:Q14.1_12)

# pull out demographic data
demo <- dat %>%
  select(unique_id, age, gender) |>
  drop_na()

# create statistics anxiety scale

STARS <- dat %>%
  select(unique_id, Q7.1_1:Q7.1_24) %>% # select the STARS columns
  filter(Q7.1_24 == 1) %>% # remove those who failed the attention check
  select(-Q7.1_24) %>% # remove the attention check column
  pivot_longer(cols = Q7.1_1:Q7.1_23, names_to = "item", values_to = "score") %>% # transform to long-form
  group_by(unique_id) %>% # group by participant
  summarise(stars_score = mean(score)) %>% # calculate mean STARS score for each ppt
  ungroup() %>% # ungroup so it doesn't mess things up later
  drop_na() #  get rid of NAs so they don't cause us havoc

# Intolerance of Uncertainty.

IUS <- dat %>%
  select(unique_id, Q14.1_1:Q14.1_12) %>% # select the IUS columns
  pivot_longer(cols = Q14.1_1:Q14.1_12, names_to = "item", values_to = "score") %>% # transform to long-form
  group_by(unique_id) %>% # group by participant
  summarise(ius_score = mean(score)) %>% # calculate mean IUS-SF score for each ppt
  ungroup() %>% # ungroup so it doesn't mess things up later
  drop_na() #  get rid of NAs so they don't cause us havoc


# join it all together

dat_joined <- left_join(demo, IUS) |>
  left_join(STARS)

6.2 Prompting

Now we want to do some set-up to prompt the AI to help us write code that works and does what we want.

NoteActivity 2

In Copilot, enter the following prompt:

You are an expert R programmer. Help me write tidyverse-first code. Before writing code, ask any clarifying questions needed about data, objectives, constraints, and outputs.When you do write code:

Use tidyverse idioms (|>, dplyr, ggplot2, tibble), no deprecated functions. Comment clearly and explain non-obvious choices (e.g., NA handling, binning, model specs). Be safe and reproducible: no setwd(), no absolute paths, no silent package installs; show any set.seed() if randomness is used. Handle missing data explicitly and avoid destructive transformations.

When you input this starting prompt you’ll get a restatement of your instructions back at you. When I ran it in September 2026 it said “Understood” and then listed what it would do (tidyverse-first, native pipe, ask clarifying questions about data structure, objective, missing data and outputs, no setwd(), and so on), and then asked what I wanted to do. Earlier versions opened with an enthusiastic offer to help and a promise of a “tailored and effective” script. If you get that, it’s sycophancy and you can ignore it. Mine didn’t, because my custom instructions (chapter 2) tell it not to, which is a small demonstration that they work. It also announced that it would use British spelling, which I have never asked for in this chat; it had picked that up from my other conversations. That is the memory feature from chapter 2 at work, and it’s worth noticing, because it will be shaping your answers too.

6.3 Describing your dataset

As I may have mentioned once or twice in this book, there is no substitute for knowing your data. For the dataset provided here, your first step as an analyst would be to read the paper in full so that you understand how and what data was collected, the scales of measurement, and information about missing data etc.

In addition to this domain knowledge, you can also use information R provides about the dataset to help you.

summary() is useful because it provides a list of all variables with some descriptive statistics so that the AI has a sense of the type and range of data:

summary(dat)
  unique_id             age               gender              Q7.1_1     
 Length:12570       Length:12570       Length:12570       Min.   :1.000  
 Class :character   Class :character   Class :character   1st Qu.:2.000  
 Mode  :character   Mode  :character   Mode  :character   Median :3.000  
                                                          Mean   :3.229  
                                                          3rd Qu.:4.000  
                                                          Max.   :5.000  
                                                          NA's   :457    
     Q7.1_2          Q7.1_3          Q7.1_4          Q7.1_5     
 Min.   :1.000   Min.   :1.000   Min.   :1.000   Min.   :1.000  
 1st Qu.:2.000   1st Qu.:2.000   1st Qu.:2.000   1st Qu.:2.000  
 Median :3.000   Median :3.000   Median :3.000   Median :3.000  
 Mean   :2.696   Mean   :2.817   Mean   :2.858   Mean   :2.687  
 3rd Qu.:4.000   3rd Qu.:4.000   3rd Qu.:4.000   3rd Qu.:4.000  
 Max.   :5.000   Max.   :5.000   Max.   :5.000   Max.   :5.000  
 NA's   :458     NA's   :453     NA's   :451     NA's   :457    
     Q7.1_6          Q7.1_7          Q7.1_8          Q7.1_9        Q7.1_10     
 Min.   :1.000   Min.   :1.000   Min.   :1.000   Min.   :1.00   Min.   :1.000  
 1st Qu.:1.000   1st Qu.:2.000   1st Qu.:3.000   1st Qu.:1.00   1st Qu.:3.000  
 Median :2.000   Median :3.000   Median :4.000   Median :2.00   Median :4.000  
 Mean   :2.424   Mean   :3.167   Mean   :3.841   Mean   :2.03   Mean   :3.647  
 3rd Qu.:3.000   3rd Qu.:4.000   3rd Qu.:5.000   3rd Qu.:3.00   3rd Qu.:5.000  
 Max.   :5.000   Max.   :5.000   Max.   :5.000   Max.   :5.00   Max.   :5.000  
 NA's   :467     NA's   :461     NA's   :454     NA's   :457    NA's   :459    
    Q7.1_11         Q7.1_12         Q7.1_13         Q7.1_14     
 Min.   :1.000   Min.   :1.000   Min.   :1.000   Min.   :1.000  
 1st Qu.:2.000   1st Qu.:1.000   1st Qu.:2.000   1st Qu.:2.000  
 Median :3.000   Median :2.000   Median :3.000   Median :3.000  
 Mean   :2.698   Mean   :2.583   Mean   :3.374   Mean   :2.654  
 3rd Qu.:4.000   3rd Qu.:4.000   3rd Qu.:4.000   3rd Qu.:4.000  
 Max.   :5.000   Max.   :5.000   Max.   :5.000   Max.   :5.000  
 NA's   :461     NA's   :454     NA's   :462     NA's   :459    
    Q7.1_15         Q7.1_16         Q7.1_17         Q7.1_18     
 Min.   :1.000   Min.   :1.000   Min.   :1.000   Min.   :1.000  
 1st Qu.:3.000   1st Qu.:2.000   1st Qu.:1.000   1st Qu.:1.000  
 Median :4.000   Median :3.000   Median :2.000   Median :2.000  
 Mean   :3.608   Mean   :2.711   Mean   :2.362   Mean   :2.461  
 3rd Qu.:5.000   3rd Qu.:4.000   3rd Qu.:3.000   3rd Qu.:3.000  
 Max.   :5.000   Max.   :5.000   Max.   :5.000   Max.   :5.000  
 NA's   :452     NA's   :463     NA's   :457     NA's   :460    
    Q7.1_19         Q7.1_20         Q7.1_21         Q7.1_22     
 Min.   :1.000   Min.   :1.000   Min.   :1.000   Min.   :1.000  
 1st Qu.:2.000   1st Qu.:2.000   1st Qu.:1.000   1st Qu.:2.000  
 Median :2.000   Median :3.000   Median :2.000   Median :3.000  
 Mean   :2.612   Mean   :2.734   Mean   :2.576   Mean   :3.048  
 3rd Qu.:4.000   3rd Qu.:4.000   3rd Qu.:4.000   3rd Qu.:4.000  
 Max.   :5.000   Max.   :5.000   Max.   :5.000   Max.   :5.000  
 NA's   :453     NA's   :458     NA's   :458     NA's   :463    
    Q7.1_23         Q7.1_24         Q14.1_1         Q14.1_2     
 Min.   :1.000   Min.   :1.000   Min.   :1.000   Min.   :1.000  
 1st Qu.:1.000   1st Qu.:1.000   1st Qu.:2.000   1st Qu.:3.000  
 Median :2.000   Median :1.000   Median :3.000   Median :4.000  
 Mean   :2.314   Mean   :1.161   Mean   :2.927   Mean   :3.465  
 3rd Qu.:3.000   3rd Qu.:1.000   3rd Qu.:4.000   3rd Qu.:4.000  
 Max.   :5.000   Max.   :5.000   Max.   :5.000   Max.   :5.000  
 NA's   :460     NA's   :1306    NA's   :1497    NA's   :1500   
    Q14.1_3         Q14.1_4         Q14.1_5         Q14.1_6     
 Min.   :1.000   Min.   :1.000   Min.   :1.000   Min.   :1.000  
 1st Qu.:2.000   1st Qu.:2.000   1st Qu.:2.000   1st Qu.:1.000  
 Median :3.000   Median :3.000   Median :3.000   Median :2.000  
 Mean   :2.681   Mean   :2.974   Mean   :2.702   Mean   :2.527  
 3rd Qu.:4.000   3rd Qu.:4.000   3rd Qu.:4.000   3rd Qu.:3.000  
 Max.   :5.000   Max.   :5.000   Max.   :5.000   Max.   :5.000  
 NA's   :1500    NA's   :1497    NA's   :1499    NA's   :1498   
    Q14.1_7         Q14.1_8         Q14.1_9         Q14.1_10    
 Min.   :1.000   Min.   :1.000   Min.   :1.000   Min.   :1.000  
 1st Qu.:2.000   1st Qu.:2.000   1st Qu.:2.000   1st Qu.:2.000  
 Median :3.000   Median :3.000   Median :3.000   Median :3.000  
 Mean   :3.008   Mean   :3.168   Mean   :2.667   Mean   :2.676  
 3rd Qu.:4.000   3rd Qu.:4.000   3rd Qu.:4.000   3rd Qu.:4.000  
 Max.   :5.000   Max.   :5.000   Max.   :5.000   Max.   :5.000  
 NA's   :1498    NA's   :1500    NA's   :1499    NA's   :1496   
    Q14.1_11        Q14.1_12   
 Min.   :1.000   Min.   :1.00  
 1st Qu.:2.000   1st Qu.:2.00  
 Median :3.000   Median :3.00  
 Mean   :3.228   Mean   :2.64  
 3rd Qu.:4.000   3rd Qu.:4.00  
 Max.   :5.000   Max.   :5.00  
 NA's   :1496    NA's   :1500  

str() is also useful because it lists the variables, their data type, and the initial values for each variable. However, that means that you are giving it at least some of the raw data so you have to be very careful if you have sensitive / confidential data and you must ensure that any use of AI is in line with your data management plan. Using Copilot Enterprise means the data won’t be stored and used to train the AI further so it’s potentially the best option (which is not to say it’s safe or problem free, please be careful and critical!).

str(dat)
tibble [12,570 × 39] (S3: tbl_df/tbl/data.frame)
 $ unique_id: chr [1:12570] "01057178" "0300b5f2" "03f6503b" "0601d699" ...
 $ age      : chr [1:12570] "18-21" "18-21" "22-25" NA ...
 $ gender   : chr [1:12570] "Male/Man" "Female/Woman" "Female/Woman" NA ...
 $ Q7.1_1   : num [1:12570] 3 3 4 4 5 1 1 3 4 4 ...
 $ Q7.1_2   : num [1:12570] 5 4 4 4 4 1 2 1 4 3 ...
 $ Q7.1_3   : num [1:12570] 3 1 3 4 5 2 2 1 5 2 ...
 $ Q7.1_4   : num [1:12570] 4 5 4 5 5 2 3 3 3 3 ...
 $ Q7.1_5   : num [1:12570] 4 5 4 4 4 1 1 2 4 2 ...
 $ Q7.1_6   : num [1:12570] 4 4 3 2 4 1 1 2 1 3 ...
 $ Q7.1_7   : num [1:12570] 5 4 4 4 5 3 1 5 5 4 ...
 $ Q7.1_8   : num [1:12570] 5 2 5 5 5 3 4 4 3 4 ...
 $ Q7.1_9   : num [1:12570] 2 1 1 1 1 1 1 1 1 4 ...
 $ Q7.1_10  : num [1:12570] 4 3 4 4 5 2 1 2 2 4 ...
 $ Q7.1_11  : num [1:12570] 4 4 4 4 4 1 1 1 5 4 ...
 $ Q7.1_12  : num [1:12570] 2 1 1 2 4 1 4 1 4 3 ...
 $ Q7.1_13  : num [1:12570] 4 2 5 3 5 2 1 2 4 4 ...
 $ Q7.1_14  : num [1:12570] 4 5 4 2 5 2 1 4 2 5 ...
 $ Q7.1_15  : num [1:12570] 3 3 3 4 5 2 2 1 3 2 ...
 $ Q7.1_16  : num [1:12570] 2 4 2 4 5 1 2 2 5 3 ...
 $ Q7.1_17  : num [1:12570] 2 2 2 2 4 1 1 1 3 5 ...
 $ Q7.1_18  : num [1:12570] 4 1 5 2 5 1 1 3 4 2 ...
 $ Q7.1_19  : num [1:12570] 4 3 2 2 5 1 1 1 4 1 ...
 $ Q7.1_20  : num [1:12570] 5 4 4 4 4 1 1 3 4 2 ...
 $ Q7.1_21  : num [1:12570] 2 3 3 4 4 1 1 4 1 3 ...
 $ Q7.1_22  : num [1:12570] 5 4 4 5 4 1 1 2 4 3 ...
 $ Q7.1_23  : num [1:12570] 4 4 2 1 5 1 1 2 4 1 ...
 $ Q7.1_24  : num [1:12570] 1 1 1 1 1 1 1 1 1 1 ...
 $ Q14.1_1  : num [1:12570] 3 2 3 2 5 2 1 1 4 2 ...
 $ Q14.1_2  : num [1:12570] 3 2 4 5 5 3 4 1 2 3 ...
 $ Q14.1_3  : num [1:12570] 2 5 3 2 5 2 2 2 5 2 ...
 $ Q14.1_4  : num [1:12570] 1 2 3 5 4 2 5 1 4 2 ...
 $ Q14.1_5  : num [1:12570] 3 2 4 4 5 1 2 2 4 2 ...
 $ Q14.1_6  : num [1:12570] 2 4 3 2 3 1 2 5 3 2 ...
 $ Q14.1_7  : num [1:12570] 4 4 2 4 5 2 3 2 1 4 ...
 $ Q14.1_8  : num [1:12570] 3 4 4 5 4 3 4 5 5 2 ...
 $ Q14.1_9  : num [1:12570] 1 3 2 4 5 2 1 3 1 2 ...
 $ Q14.1_10 : num [1:12570] 2 2 2 5 4 1 1 2 4 3 ...
 $ Q14.1_11 : num [1:12570] 4 3 2 5 5 2 5 4 4 3 ...
 $ Q14.1_12 : num [1:12570] 2 3 2 2 5 1 1 1 4 2 ...

Finally, ls() provides a list of all the variables in a given object. It doesn’t provide any info on the variable type or sample, but that does mean it’s the most secure and depending on the task, this might be all the info you really need to give the AI. I would suggest starting with ls() and only scaling up if necessary (and your data isn’t sensitive):

ls(dat)
 [1] "age"       "gender"    "Q14.1_1"   "Q14.1_10"  "Q14.1_11"  "Q14.1_12" 
 [7] "Q14.1_2"   "Q14.1_3"   "Q14.1_4"   "Q14.1_5"   "Q14.1_6"   "Q14.1_7"  
[13] "Q14.1_8"   "Q14.1_9"   "Q7.1_1"    "Q7.1_10"   "Q7.1_11"   "Q7.1_12"  
[19] "Q7.1_13"   "Q7.1_14"   "Q7.1_15"   "Q7.1_16"   "Q7.1_17"   "Q7.1_18"  
[25] "Q7.1_19"   "Q7.1_2"    "Q7.1_20"   "Q7.1_21"   "Q7.1_22"   "Q7.1_23"  
[31] "Q7.1_24"   "Q7.1_3"    "Q7.1_4"    "Q7.1_5"    "Q7.1_6"    "Q7.1_7"   
[37] "Q7.1_8"    "Q7.1_9"    "unique_id"
NoteActivity 3

Now we’re going to tell the AI about our dataset and what we’d like. For educational value, we’ll give it a purposefully brief description.

I have a dataset from a paper on statistics anxiety. I want to pull out the demographic information, age and gender and I want to calculate each participant’s Statistic Anxiety score and their Intolerance of Uncertainty

here are the variables in my data set

[1] “age” “gender” “Q14.1_1” “Q14.1_10” “Q14.1_11” “Q14.1_12” “Q14.1_2”
[8] “Q14.1_3” “Q14.1_4” “Q14.1_5” “Q14.1_6” “Q14.1_7” “Q14.1_8” “Q14.1_9”
[15] “Q7.1_1” “Q7.1_10” “Q7.1_11” “Q7.1_12” “Q7.1_13” “Q7.1_14” “Q7.1_15”
[22] “Q7.1_16” “Q7.1_17” “Q7.1_18” “Q7.1_19” “Q7.1_2” “Q7.1_20” “Q7.1_21”
[29] “Q7.1_22” “Q7.1_23” “Q7.1_24” “Q7.1_3” “Q7.1_4” “Q7.1_5” “Q7.1_6”
[36] “Q7.1_7” “Q7.1_8” “Q7.1_9” “unique_id”

The questions Copilot asks in response to this prompt hammer home my point that there is no substitute for knowing your data. In September 2026 it worked out from the variable names that there were a 12-item scale and a 24-item scale, then said it did not know which scale was which, whether any items were reverse scored, whether scores should be summed or averaged, or what the response scale was, and asked for the paper’s DOI or the scoring instructions. Earlier versions asked even more (how to handle missing data, whether age was numeric, how gender was coded); this one asked less, and didn’t ask about missing data at all despite the set-up prompt telling it to.

It then did something worth dwelling on. Having said it needed that information before writing code, it wrote code anyway: a “pattern” for how the scoring would work “if, for example” the 12-item scale was statistics anxiety. That guess was wrong (the 12-item scale is intolerance of uncertainty), and the example code mixed the native pipe |> with the magrittr placeholder . inside select(), which doesn’t run. It was clearly labelled as an example rather than an answer, but this is exactly the sort of thing that ends up copied into a script by someone in a hurry. An AI that says “I need more information” and then produces code regardless is not asking a question, it’s covering itself.

NoteActivity 4

Again for educational value, we’re going to be lazy and give it minimal information. The questions it asks you will be slightly different to the ones above but they will be similar enough that this response should still work. Reply to Copilot with the following:

Questions 7 are the Statistics Anxiety items and Q14 are the Intolerance of Uncertainty items. I only want complete cases. I don’t know anything else.

Copilot then produces the code based on the minimal information you’ve given it. In my September 2026 run it prefaced the code with a sensible warning: it would be “cautious about claiming these are the published scale scores” because both scales might contain reverse-scored items, and it recommended checking the paper’s methods before finalising anything. Then it gave me the code anyway. We’ll get to whether it works in the next step but regardless of whether it does, I find this hugely problematic. Psychology and many other fields have spent the last decade dealing with a replication and reproducibility crisis stemming in part because of questionable research practices. As a researcher, you should be making informed decisions as to how you analyse your data and even when I have admitted I don’t know almost anything about my data, it has provided the code.

“Vibe coding” like this is going to increase p-hacking and atheoretical, exploratory-as-confirmatory nonsense. What happens when the example code the AI spits out without being asked turns out to be a significant regression model that you would never have predicted or run yourself? Are you going to delete it? Or convince yourself that you were going to run it anyway and there’s a perfectly logical explanation?

Before I have a full blown existential crisis, let’s get back on track.

NoteActivity 5

Here’s the code Copilot produced in September 2026. Run it and see if it works. If you’ve prompted it yourself you can also use the version it produced for you, which will be slightly different to mine. Two things to watch: Copilot called the data data and I’ve changed that to dat so it runs, and check that whatever it produces has a new object name rather than overwriting dat, because we need to compare them.

In the first version of this book, the code Copilot produced at this point was 80 lines long, and it didn’t run. This version is a quarter of the length, and it does. Look at it before you run it and see if you can spot what’s wrong.

library(tidyverse)

# Item names
stats_anxiety_items <- paste0("Q7.1_", 1:24)
iu_items <- paste0("Q14.1_", 1:12)

# Keep only participants with complete data on demographics and both scales
scored_data <- dat |>
  filter(
    if_all(
      c(age, gender, all_of(stats_anxiety_items), all_of(iu_items)),
      ~ !is.na(.x)
    )
  ) |>
  mutate(
    statistics_anxiety = rowSums(
      pick(all_of(stats_anxiety_items))
    ),
    intolerance_uncertainty = rowSums(
      pick(all_of(iu_items))
    )
  ) |>
  select(
    unique_id,
    age,
    gender,
    statistics_anxiety,
    intolerance_uncertainty
  )

summary(scored_data)
  unique_id             age               gender          statistics_anxiety
 Length:9117        Length:9117        Length:9117        Min.   : 24.00    
 Class :character   Class :character   Class :character   1st Qu.: 52.00    
 Mode  :character   Mode  :character   Mode  :character   Median : 67.00    
                                                          Mean   : 66.53    
                                                          3rd Qu.: 80.00    
                                                          Max.   :120.00    
 intolerance_uncertainty
 Min.   :12.00          
 1st Qu.:27.00          
 Median :34.00          
 Mean   :34.29          
 3rd Qu.:42.00          
 Max.   :60.00          

The code runs, which is progress. The problems are all in what it calculated.

  • It calculated sum scores, not mean scores. It did offer a mean version as an alternative (“if you would prefer mean scores rather than totals”) but the sum version came first and, if you were in a hurry, is the one you’d copy. I spotted this because I know the plausible range of a mean on a 1–5 scale, and 24 to 120 isn’t it. Here’s the mean version:
scored_means <- dat |>
  filter(
    if_all(
      c(age, gender, all_of(stats_anxiety_items), all_of(iu_items)),
      ~ !is.na(.x)
    )
  ) |>
  mutate(
    statistics_anxiety = rowMeans(
      pick(all_of(stats_anxiety_items))
    ),
    intolerance_uncertainty = rowMeans(
      pick(all_of(iu_items))
    )
  ) |>
  select(
    unique_id,
    age,
    gender,
    statistics_anxiety,
    intolerance_uncertainty
  )

summary(scored_means)
  unique_id             age               gender          statistics_anxiety
 Length:9117        Length:9117        Length:9117        Min.   :1.000     
 Class :character   Class :character   Class :character   1st Qu.:2.167     
 Mode  :character   Mode  :character   Mode  :character   Median :2.792     
                                                          Mean   :2.772     
                                                          3rd Qu.:3.333     
                                                          Max.   :5.000     
 intolerance_uncertainty
 Min.   :1.000          
 1st Qu.:2.250          
 Median :2.833          
 Mean   :2.857          
 3rd Qu.:3.500          
 Max.   :5.000          
  • The means are also wrong, and this is the dangerous one, because they look entirely plausible. I only know they’re wrong because I have the correct scores from my own code. Let’s compare the STARS scores for every participant who appears in both:
check <- scored_means |>
  inner_join(STARS, by = "unique_id") |>
  mutate(diff = statistics_anxiety - stars_score)

# how many participants have a different score, and how big is the difference?
check |>
  summarise(
    n = n(),
    n_different = sum(abs(diff) > 0.001),
    max_diff = max(abs(diff))
  )
n n_different max_diff
8659 8628 0.1666667
  • Almost every participant has a different STARS score, by up to 0.17 on a 1–5 scale. That’s because I failed to tell it that item Q7.1_24 is an attention check, not a scale item, so it averaged over 24 items instead of 23 and kept the participants who failed the check (458 of them). That’s my fault for not giving it enough information, and it’s the whole point: the AI cannot know what you don’t tell it, and it won’t tell you it doesn’t know. The intolerance of uncertainty scores, where there was no attention check to miss, match mine exactly.

  • The other difference is in how it handled missing data. It kept only participants with complete data on age, gender, and both scales together, so someone with a missing gender loses their STARS score too. My code scored each scale separately and dropped participants from each one only where their own data was missing. Neither is wrong, but it’s a choice, and it may not be a choice you intended to make. Notice that it had told me at the start it would “handle missing data explicitly” and then never asked me how.

6.4 Break it down

One of the reasons this is going so wrong is because I have asked it to do too much at once. A better approach is to break down the code into small chunks. Once you can verify each chunk is working, you can then combine and refactor it (if you’re starting to think that it might just be easier to write the code yourself and then review it, yes, yes it would).

NoteActivity 6

We want to start fresh so start a new Copilot chat then enter the prompt from Activity 2 again.

Now we’re going to try and do the same thing again but we’re going to break down the steps and ask for smaller chunks each time. Follow-up on the set-up prompt with the following:

Here is my dataset. First, I want to create an object named demo_copilot that has the participant id, age, and gender. Age is categorical. Remove any participant who is missing any information.

ls(dat) [1] “age” “gender” “Q14.1_1” “Q14.1_10” “Q14.1_11” “Q14.1_12” “Q14.1_2”
[8] “Q14.1_3” “Q14.1_4” “Q14.1_5” “Q14.1_6” “Q14.1_7” “Q14.1_8” “Q14.1_9”
[15] “Q7.1_1” “Q7.1_10” “Q7.1_11” “Q7.1_12” “Q7.1_13” “Q7.1_14” “Q7.1_15”
[22] “Q7.1_16” “Q7.1_17” “Q7.1_18” “Q7.1_19” “Q7.1_2” “Q7.1_20” “Q7.1_21”
[29] “Q7.1_22” “Q7.1_23” “Q7.1_24” “Q7.1_3” “Q7.1_4” “Q7.1_5” “Q7.1_6”
[36] “Q7.1_7” “Q7.1_8” “Q7.1_9” “unique_id”

demo_copilot <- dat |>
  select(
    unique_id,
    age,
    gender
  ) |>
  drop_na()

In September 2026 it said “I have enough information to write this without additional questions”, produced the above, and explained each line. In the first version of this book, the same request produced code that converted age and gender to factors and renamed unique_id to participant_id, neither of which I’d asked for. Now check if this matches our original code:

all.equal(demo, demo_copilot)
[1] TRUE

It does. It is, give or take the pipe, exactly the code I wrote myself. Given a small, fully specified task, it did the task.

NoteActivity 7

Now try and reproduce the objects IUS and STARS using AI to write the code to create IUS_copilot and STARS_copilot version.

If you manage to get this to work, reflect on how much information you had to give it and how you acquired this knowledge.

If you can’t get them to match, reflect on what knowledge you might be missing to get it to do the task

Here’s how it went for me. I asked for STARS_copilot with each participant’s mean score on items Q7.1_1 to Q7.1_24, pivoted to long form, with participants with missing data removed. This is what it produced:

STARS_copilot <- dat |>
  select(unique_id, starts_with("Q7.1_")) |>
  
  # Remove participants with any missing STARS item
  drop_na() |>
  
  # Convert to long format
  pivot_longer(
    cols = starts_with("Q7.1_"),
    names_to = "item",
    values_to = "score"
  ) |>
  
  # Calculate participant mean
  group_by(unique_id) |>
  summarise(
    STARS_mean = mean(score),
    .groups = "drop"
  )

all.equal(STARS, STARS_copilot)
[1] "Names: 1 string mismatch"                                                              
[2] "Attributes: < Component \"row.names\": Numeric: lengths (10420, 11123) differ >"       
[3] "Component \"unique_id\": Lengths (10420, 11123) differ (string compare on first 10420)"
[4] "Component \"unique_id\": 10403 string mismatches"                                      
[5] "Component 2: Numeric: lengths (10420, 11123) differ"                                   

It doesn’t match, and it doesn’t match for the same reason as before: I made the same mistake and didn’t mention the attention check, so it has 700 more participants than it should and the wrong mean for almost all of them. The code is fine. The information was wrong. So I told it:

I forgot to tell you that Q7.1_24 is an attention check, not a STARS item. Only participants who answered 1 to Q7.1_24 should be kept, and Q7.1_24 should not be included in the mean. Update the code.

STARS_copilot <- dat |>
  # Keep only participants who passed the attention check
  filter(Q7.1_24 == 1) |>
  
  # Keep ID and STARS items only (excluding the attention check)
  select(
    unique_id,
    starts_with("Q7.1_"),
    -Q7.1_24
  ) |>
  
  # Remove participants with any missing STARS item
  drop_na() |>
  
  # Convert items to long format
  pivot_longer(
    cols = starts_with("Q7.1_"),
    names_to = "item",
    values_to = "score"
  ) |>
  
  # Calculate mean STARS score for each participant
  group_by(unique_id) |>
  summarise(
    STARS_mean = mean(score),
    .groups = "drop"
  )

all.equal(STARS, STARS_copilot)
[1] "Names: 1 string mismatch"

Still not equal, but read the message: the only difference now is that it called the column STARS_mean and I called it stars_score. The numbers are identical. If you rename the column, all.equal() returns TRUE:

all.equal(STARS, rename(STARS_copilot, stars_score = STARS_mean))
[1] TRUE

IUS_copilot was the same code with the Q14 items and no attention check to forget, and it matched first time (apart from the column name).

So it can be done. But look at what I had to know to get there: which items belong to which scale, that one of them is an attention check, what the pass value is, that the score is a mean not a sum, and what the plausible range of that mean is so that I could tell when it was wrong. None of that came from Copilot. All of it came from reading the paper, which is the thing the AI was supposed to save me from doing.

6.5 Be critical

In terms of your psychological and cognitive development, the warnings of this chapter are similar to those in previous chapters: if you don’t write the code yourself, you won’t gain those mastery experiences that support the development of self-efficacy and competence. Additionally, working through the code yourself requires you to understand your data and is essentially a form of self-explanation which will further impact your competence and ability to produce anything autonomously.

But my biggest warning for this chapter is not about your psychological development, but rather the integrity of your work. If you worked through Activity 7 properly you’ll understand that the amount of information you have to provide in order to get AI to write correct code is so extensive, you might as well just have written it yourself and used AI to fix any resulting errors.

The evidence on whether AI actually makes programmers faster is more mixed than the marketing. In a randomised trial with experienced open-source developers working on their own codebases, having AI tools made them 19% slower, even though they believed afterwards that they had been 20% faster (Becker et al., 2025, a preprint). Other studies with novices find speed gains without obvious harm to learning (Kazemitabaar et al., 2023). The unsatisfying summary is that it depends on who you are and what you’re doing, and that people are bad at judging which case they’re in. What the experienced developers were doing with the time was exactly what you’ve just done: reading, checking and fixing what the AI produced.

This integrity issue isn’t just about students cheating on their homework. The scientific literature is already full of AI garbage, and there is nothing to stop analyses that were vibe-coded and never checked from ending up in the evidence base that treatments get prescribed on. Code written by AI is significantly increasing security risks.

The world has changed and for those of us that value integrity and expertise it feels like we are fighting a losing battle. My greatest challenge to you isn’t about coding, instead, it is to resist every psychological urge that tells all of us to take the easy route. Respect yourself and your abilities. Spend hours figuring out what the bloody hell is wrong with your data. Learn how to do things properly. If nothing else, you’ll be one of an increasingly small pool who can and that is going to become very valuable, very quickly.

NoteKey takeaways
  • Professional programmers use AI to write code, and they can because they learned to code first. You still need the knowledge and the skills to check its work.
  • There is no substitute for knowing your data. At minimum, you need to know what your variables are, what the range of plausible values are, how you want to handle missing data, and you need to know what you intended to do and why.
  • If you ask AI to write code from scratch, break it down into small chunks.
  • The more specific you can be, the better the AI will do.
  • Always verify what it produces. Sometimes you can do this with code, sometimes you need your domain knowledge of the data.

This is a depressing way to end this book so I gave Copilot the prompt:

Write me a fun piece of R code.

And it took me to Vegas (ok I actually quite like this):

# Simple Slot Machine in R
set.seed(Sys.time()) # Seed for randomness based on current time

# Function to spin the slot machine
spin_slot_machine <- function() {
  fruits <- c("🍒", "🍋", "🍊", "🍉", "🍇", "🍓")
  spin <- sample(fruits, size = 3, replace = TRUE)
  cat("Spinning... You got:", spin, "\n")
  
  if (length(unique(spin)) == 1) {
    cat("Congratulations! You won! 🎉\n")
  } else {
    cat("Try again! 🍀\n")
  }
}

# Spin the slot machine
spin_slot_machine()
Spinning... You got: 🍋 🍋 🍓 
Try again! 🍀