5  Code review

In this chapter you’ll learn how to use AI to perform a code review and to add comments to your code. As you’ve already hopefully learned by working through this book, you have to be critical about anything the AI produces or suggests because it has no expert knowledge, but it can be a useful tool for checking and improving your code.

DeBruine et al’s Code Check Guide details what a comprehensive code check refers to:

However, some of these steps cannot (and should not) be performed by an AI. Unless you have specific ethical approval and have included this in your data management plan, you should never upload your research data to an AI tool. This means that assessing reproducibility is difficult. The AI also doesn’t know what you intended to do, and why, and has no subject knowledge so it can’t advise on anything theoretical without you giving it that information explicitly.

Therefore, what we’ll focus on in this chapter is two components of code review: comments and refactoring your code.

5.1 Code comments

Code comments are lines or sections of text added within the code itself that are ignored by the computer when the program runs. They’re there for human readers, not machines. In R, you add comments to code by adding # to the start of the string:

# this is a comment

# compute the mean of three numbers
mean(c(1,2,3))

Comments are useful for several reasons:

  • Clarification: They explain what certain parts of the code do, making it easier for others (and yourself) to understand the logic and flow of the code.
  • Documentation: They provide information on how the code works or why certain decisions were made, which is helpful for future reference.
  • Debugging: Temporarily commenting out parts of code can help isolate sections that may be causing errors, without deleting the code.
  • Collaboration: In team projects, comments can be used to communicate with other developers about the status or purpose of the code.

Overall, comments are a crucial part of writing clean, maintainable, and collaborative code. They help make the code more accessible and understandable to anyone who might work on it in the future.

5.2 Adding comments with AI

First we’ll use use the palmerpenguins dataset again.

library(tidyverse)
library(palmerpenguins)
data("penguins")

You can use AI tools to help add comments to your code. Previous experience has taught us that we need to be a bit specific with what we want the AI to do so we’ll give it a clear prompt.

CautionCaution!!!

For this chapter, be very, very careful about which model you are using. In the model selector at the top left of the chat, choose Think deeper (or GPT 5.6 Think deeper) rather than Quick response or Auto. When I first wrote this chapter, the older model then available (GPT-4) introduced errors into the code and wrote comments that were wrong, which then led it to reason its way into further errors. The newer models are better at this task but the slower, reasoning mode is the one to use for anything involving your code.

NoteActivity 1

In Copilot, input the below code with the prompt:

Please add concise, helpful comments to this code. Explain what each main step does, not every line. If you are unsure about what a section of code is doing or why it exists, ask me a clarifying question. Do not change the code itself in any way.

penguins_clean <- penguins %>%
  filter(complete.cases(.)) %>%
  mutate(across(where(is.factor), as.character)) %>%
  mutate(species_sex = interaction(species, sex, sep = "_"))

Here’s its suggested comments (September 2026, Think deeper):

# Start with the penguins dataset and create a cleaned version
penguins_clean <- penguins %>%
  
  # Keep only rows with no missing values in any variable
  filter(complete.cases(.)) %>%
  
  # Convert all factor variables to character variables
  mutate(across(where(is.factor), as.character)) %>%
  
  # Create a combined variable representing each species and sex combination
  mutate(species_sex = interaction(species, sex, sep = "_"))

It put the comments on their own lines above each step rather than at the end of the line, added blank lines between the steps, and told me that “the pipeline structure, functions, arguments, object names and execution order have not been changed”, which is true. Adding blank lines inside a pipeline is a formatting change rather than a code change, but notice that it did it without being asked, and if you’d asked it not to touch anything it would still count.

However, the AI can only tell you what the code is doing, it can’t say why you chose to do that. The more complex your analysis, the more crucial it becomes to explain the rationale and so ensuring your prompt encourages the AI to ask for clarification is vital and in order to be able to answer these questions you need to know your data and what you’re trying to achieve.

5.2.1 “If you are unsure, ask me”

The prompt told it to ask if it was unsure, and it didn’t ask anything. When I ran the same prompt a year earlier it followed the commented code with three “questions to confirm intent”: did I want to drop rows that were missing on any column, or only on specific variables like species, sex and the bill and flipper measures; should species_sex be a character vector rather than a factor for whatever I was doing next; and did any of the factor columns need to stay ordered, or was converting all of them to character intentional. Those are exactly the questions a good code reviewer would ask, and every one of them is about intent rather than syntax.

This is the most important lesson in the chapter, so it’s worth slowing down on. When a human reviewer is unsure, it’s because they’ve noticed a gap between what the code does and what they would need to know to say why. An AI has no access to why. It doesn’t know what you intended, it can’t know, and it doesn’t experience that as uncertainty. It will produce a fluent, confident comment for a line of code whose purpose it has no way of knowing, because producing fluent, confident text is the only thing it does. “If you are unsure, ask me” therefore asks it to detect something it can’t detect. The questions it asked a year ago were a happy accident of that model, not a sign that it had noticed the gap, and the newer model, which is better at almost everything else, simply didn’t ask. The same goes for every “tell me if you don’t know” instruction you will be tempted to write: an AI’s willingness to answer is not evidence that it knows.

If you need it to stop and ask, you have to make asking the task. Here’s the same code with a different prompt:

Before you write any comments on this code, list the questions you would need me to answer in order to explain why each step exists, not just what it does. Do not write any comments until I have answered.

This time (September 2026, Think deeper) it produced twelve questions and no comments, and finished with “I will not add comments to the code until you have answered these questions.” Some of the questions were the same ones the older model had asked: “Why are participants with any missing value removed? Does the planned analysis require complete data across every variable, or only across a particular subset of variables?” and “Are there any variables that should remain factors because they will be used in statistical models, plots, or ordered comparisons?”. Some were better than I would have thought to ask: “Why is interaction() being used instead of a character-combining function such as paste() or str_c()? Is it important that species_sex is created as a factor with combination levels?” is a good question, because I have just converted every factor to character and then immediately created a new factor, and I’m not sure I could defend that. It also asked what the cleaned dataset was for, whether complete-case filtering was appropriate given the pattern of missingness, and why I’d chosen an underscore as the separator.

Notice what changed. The first prompt made commenting the task and asking an optional extra, conditional on a feeling the AI doesn’t have. The second made asking the task. Same model, same code, same afternoon.

NoteActivity 1b

Run the second prompt with the code from Activity 1. Answer its questions (if you can’t, that tells you something), then tell it to write the comments. Compare them with the comments it produced first time. The difference is the information you gave it, not anything the AI knew.

5.3 Review existing comments

In addition to asking AI to comment your code, you can also ask it to review comments you’ve made yourself. To see how this works with a more complex example, and as an act of masochism, I gave the AI some code I wrote for a publication. The full paper is here if you’re interested - the quant analyses ended up being punted to the online appendix because of word count.

NoteActivity 2

Load in the dataset yourself with this code:

# read in data but skip rows 2 and 3
col_names <- names(read_csv("https://osf.io/download/tf3xs/", n_max = 0))
dat_raw <- read_csv("https://osf.io/download/tf3xs/", col_names = col_names, skip = 3)

The first section of my code involves quite a complicated and long bit of wrangling, all done in a single pipeline. The purpose of the code is to clean up data collected on the survey platform Qualtrics and recode some of the demographic variables. This is actually a shortened version because the original hit the character limit for Copilot. I did put some effort into writing comments before publication but there are almost certainly improvements to be made.

NoteActivity 3

Provide the code with the following prompt followed by the below code:

Please review the comments in my code and improve them where needed. Make comments clear, concise, and useful for someone reading the code for the first time. Keep the meaning of existing comments, but reword or simplify them for better readability.

Add comments only where they genuinely help understanding (e.g., explaining intent or logic, not obvious code). Do not change any of the code itself. After editing, explain your reasoning for each change: briefly describe why the original comment needed improvement (e.g., too long, unclear, redundant, missing context, etc.).

dat <- dat_raw%>%
  filter(Progress > 94, # remove incomplete responses
         DistributionChannel != "preview") %>% # Remove Emily's preview data
  select(ResponseId, "duration" = 5, Q5:Q21) %>%
  # replace NAs with "none" for disability info
  mutate(disability_nos = replace_na(disability_nos, "None"),
         physical_chronic = replace_na(physical_chronic, "None"),
         mental_health = replace_na(mental_health, "None"),
) %>% # recode gender data
  mutate(gender_cleaned = case_when(Q6 %in% c("Female", "female", "Woman",
                                              "woman", 
                                              "Cisgender woman",
                                              "female (she/her)", 
                                              "F", "f", "Womxn", 
                                              "Woman (tranas)") ~ "Woman",
                                    Q6 %in% c("Man", "man", "M", "m", 
                                              "Male (he/him)", "Male",
                                              "male", "Trans man.") ~
                                      "Man",
                                    Q6 %in% c("Agender", "Genderfluid",
                                    "GNC", "NB", "non-binary", 
                                    "   Non-binary", "Non-Binary",
                                    "Non-binary femme", "non-binary male",
                                    "non binary", "Non binary",
                                    "Nonbinary", "Queer", "Transmasculine",
                                    "Non-binary") ~ "Non-binary",
                            TRUE ~ "Not stated")) %>%
  # select necessary columns and tidy up the names
        select(ResponseId,
             "age" = Q5,
             "gender" = Q6,
             "mature" = Q7,
             "level_study" = Q8,
             "country" = Q9,
             "subject" = Q10,
             "english_first" = Q11,
             "neurotype_open" = Q13, 
             "disability_open" = Q14,
             "why_open" = Q18,
             "how_open" = Q23,
             "advantages" = Q20,
             "disadvantages" = Q21,
             everything()) 
CautionHave I mentioned you need to be careful?
  1. In an earlier version of this book, providing this prompt resulted in it changing the code without telling me so that not only did it not do what I intended, it also didn’t work so you need to be very, very careful.

  2. Using an older model (GPT-4, in 2024) for this task resulted in incorrect comments, which then affected subsequent reasoning and introduced errors into the code. Newer models are better, but “better” is not “reliable”, which is why you check.

If you use AI without thinking and accidentally commit research fraud, don’t blame me, I did try and warn you :)

To check that it hasn’t changed any code you can run all.equal() to compare two datasets. If it returns true, it means that the result of your initial code and the new code are identical. This is a really important check.

Here’s the code with the new comments (September 2026, Think deeper). It opened with “I have revised only the comments, leaving the code unchanged”, which is the claim we’re about to check.

# this is the code copied from Copilot with the edited comments

dat_copilot <- dat_raw%>%
  filter(Progress > 94, # Keep completed responses
         DistributionChannel != "preview") %>% # Exclude preview responses
  select(ResponseId, "duration" = 5, Q5:Q21) %>%
  # Record missing disability information as "None"
  mutate(disability_nos = replace_na(disability_nos, "None"),
         physical_chronic = replace_na(physical_chronic, "None"),
         mental_health = replace_na(mental_health, "None"),
) %>%
  # Group gender responses into consistent categories
  mutate(gender_cleaned = case_when(Q6 %in% c("Female", "female", "Woman",
                                              "woman", 
                                              "Cisgender woman",
                                              "female (she/her)", 
                                              "F", "f", "Womxn", 
                                              "Woman (tranas)") ~ "Woman",
                                    Q6 %in% c("Man", "man", "M", "m", 
                                              "Male (he/him)", "Male",
                                              "male", "Trans man.") ~
                                      "Man",
                                    Q6 %in% c("Agender", "Genderfluid",
                                    "GNC", "NB", "non-binary", 
                                    "    Non-binary", "Non-Binary",
                                    "Non-binary femme", "non-binary male",
                                    "non binary", "Non binary",
                                    "Nonbinary", "Queer", "Transmasculine",
                                    "Non-binary") ~ "Non-binary",
                            TRUE ~ "Not stated")) %>%
  # Select and rename variables for analysis
        select(ResponseId,
             "age" = Q5,
             "gender" = Q6,
             "mature" = Q7,
             "level_study" = Q8,
             "country" = Q9,
             "subject" = Q10,
             "english_first" = Q11,
             "neurotype_open" = Q13, 
             "disability_open" = Q14,
             "why_open" = Q18,
             "how_open" = Q23,
             "advantages" = Q20,
             "disadvantages" = Q21,
             everything())


# then we can test if the two objects are identical to ensure it hasn't changed anything

all.equal(dat, dat_copilot)
[1] TRUE

all.equal() returns true which means the two datasets are identical. Its explanation of each change was sensible: it removed my name from “Remove Emily’s preview data” because that’s “person-specific context that is not necessary for understanding the code”; it changed “replace NAs with”none”” to use the capitalisation the code actually uses; it moved the “recode gender data” comment from the end of the previous line to its own line above the mutate() it describes, because where I’d put it wasn’t clear which step it belonged to. All fair, and the last one is a mistake I make constantly.

One of its changes is arguably worse than the original, and it’s a good example of why you read the comments rather than just checking the code. It changed “remove incomplete responses” to “Keep completed responses”, reasoning that the comment should describe “what the filter retains rather than what it removes”. But the filter is Progress > 94, which keeps anyone who got past 94% of the survey. Those responses aren’t “completed”, and my original comment, whilst not perfect, at least didn’t claim they were. The comment now says something the code doesn’t do. The AI improved the style and degraded the accuracy, and it did so confidently and with a justification.

5.4 Refactoring code

Refactoring means improving the structure of your code without changing what it does. You are not adding new features or fixing bugs, you are just making the code cleaner, easier to read, and easier to work with. Refactoring is like tidying your desk: everything still works the same, but it is more organised and makes future work easier.

Refactoring:

  • Makes your code easier to read so that you (and others) can quickly understand what it does later on.
  • Simplifies your code by breaking big, messy pieces into smaller, easier-to-manage parts.
  • Keeps your code easy to update as clean, consistent code makes it easier to fix bugs or add new features later.
  • Helps you avoid and remove duplicated or unnecessary code.
  • Sometimes make it faster. Tidier code can help your program run more efficiently.
  • Can help spot hidden problems. While cleaning up, you might notice bugs or logic errors that you had missed before.
CautionCaution!!!

Are you using Think deeper as instructed? The next section can go very badly with the older or faster models.

NoteActivity 4

We’ll stick with the example of code from my publication because it’s complicated enough that there are likely improvements or alternative solutions. Enter the following prompt followed by the code (use the original, not the Copilot commented version):

I would like you to refactor the following code. Keep the behaviour and output exactly the same, but make the code cleaner, easier to read, and more efficient. Please:

Add helpful comments where needed. Use clear and consistent variable and function names. Remove any unnecessary repetition. Break long sections into smaller, logical parts if possible. Follow good coding style for R and use a tidyverse approach where possible. After refactoring, briefly explain what changes you made and why.

Here’s the refactored code it gave me (September 2026, Think deeper). One change from me: Copilot called the result dat, which would overwrite the original and make the check below meaningless (all.equal(dat, dat) is always true). I’ve renamed it dat_copilot. Watch for this.

# Define gender responses once to keep the data-processing pipeline readable
woman_responses <- c(
  "Female", "female", "Woman", "woman", "Cisgender woman",
  "female (she/her)", "F", "f", "Womxn", "Woman (tranas)"
)

man_responses <- c(
  "Man", "man", "M", "m", "Male (he/him)", "Male",
  "male", "Trans man."
)

non_binary_responses <- c(
  "Agender", "Genderfluid", "GNC", "NB", "non-binary",
  "    Non-binary", "Non-Binary", "Non-binary femme",
  "non-binary male", "non binary", "Non binary",
  "Nonbinary", "Queer", "Transmasculine", "Non-binary"
)

dat_copilot <- dat_raw %>%
  # Retain sufficiently complete responses and exclude preview data
  filter(
    Progress > 94,
    DistributionChannel != "preview"
  ) %>%
  
  # Retain the response metadata and questionnaire variables
  select(
    ResponseId,
    duration = 5,
    Q5:Q21
  ) %>%
  
  # Replace missing disability information with "None"
  mutate(
    across(
      c(disability_nos, physical_chronic, mental_health),
      ~ replace_na(.x, "None")
    )
  ) %>%
  
  # Recode free-text gender responses into broader categories
  mutate(
    gender_cleaned = case_when(
      Q6 %in% woman_responses ~ "Woman",
      Q6 %in% man_responses ~ "Man",
      Q6 %in% non_binary_responses ~ "Non-binary",
      TRUE ~ "Not stated"
    )
  ) %>%
  
  # Rename and reorder the variables required for analysis
  select(
    ResponseId,
    age = Q5,
    gender = Q6,
    mature = Q7,
    level_study = Q8,
    country = Q9,
    subject = Q10,
    english_first = Q11,
    neurotype_open = Q13,
    disability_open = Q14,
    why_open = Q18,
    how_open = Q23,
    advantages = Q20,
    disadvantages = Q21,
    everything()
  )

And here’s the explanation:

Changes made

  • Moved the three sets of gender responses into clearly named character vectors. This makes the case_when() statement substantially easier to read and update.
  • Replaced three nearly identical replace_na() calls with across(), removing unnecessary repetition.
  • Split filter(), select(), and mutate() arguments across lines consistently.
  • Used comments to describe the purpose of each logical stage rather than commenting on individual lines.
  • Removed the trailing comma from the first mutate() call.
  • Retained all original response labels exactly, including “Woman (tranas)”, leading spaces in ” Non-binary”, and punctuation in “Trans man.”, because correcting these would technically change the code’s behaviour.

That last point is worth pausing on. “Woman (tranas)” is a typo in a participant’s response, not in my code, and the leading whitespace is what someone actually typed, so the code has to match it exactly. It spotted that these looked like errors and correctly left them alone because I’d asked it to preserve behaviour. An earlier model would have “fixed” them and silently reassigned those participants to “Not stated”.

It then added a section headed Important issue to check: the first select() keeps Q5:Q21, but the final select() refers to Q23, so “unless Q23 happens to be positioned between Q5 and Q21 in dat_raw, it will have been removed before the final select(), causing an error”. It offered a corrected version but said it had “not incorporated that correction into the main refactor because you asked for the behaviour to remain exactly the same.”

This is a good catch and a wrong diagnosis at the same time. Q23 does sit between Q5 and Q21 in the raw data: Qualtrics numbers questions in the order they were created rather than the order they appear in the survey, and Q23 comes just before Q20 in the export, so the code works and always has. But the AI couldn’t know that, it flagged the risk clearly, it hedged correctly, and it didn’t change my code on the strength of a guess. That is what you want from a reviewer. What you don’t want is to click the “Add a version that fixes Q23” button underneath without checking whether there’s anything to fix.

Again, it’s very important to check that this refactored code has the same result as the original. There is no shortcut for knowing and checking your data.

all.equal(dat, dat_copilot)
[1] TRUE

5.5 Be critical

Writing your own code comments works like taking your own lecture notes: explaining what your code does, in your own words, is retrieval practice and elaboration rolled into one, and it’s what turns a pipeline you copied from last week’s script into something you understand. The specific risk with AI-written comments and “cleaner” code is the fluency illusion from chapter 2: the code reads smoothly, the comments are articulate, so it feels understood. Reading a good explanation of your code is not the same as being able to produce one, and the difference only shows up when something breaks.

In practice: write your own comments first, then use AI to critique or refine them. When refactoring, make sure you can explain what each variable represents, why each step exists, and how the data should behave before you let AI suggest changes, because otherwise you can’t check its work, and as this chapter has shown, you have to check its work.

TipKey takeaways

I hadn’t used AI to perform these types of tasks before writing this book so here’s my takeaways:

  • DID I MENTION? CHECK EVERYTHING.
  • If you give an AI code, you simply cannot trust that it won’t change your code, even if that’s not the task you ask it to do. If you use AI to add or review comments, you must check the output. Tools like all.equal() can help perform these checks.
  • You also can’t trust that the comments will be accurate. Anything an AI writes must be checked before you use it. If you don’t know if it’s right, don’t use it.
  • Because you have to check what it does so carefully, don’t give it a big dump of code. Smaller chunks will end up taking less time.