8  Iteration

D&D alignment chart showing different ways to iterate in R

Intended Learning Outcomes

Setup

  1. Open your reprores project
  2. Create a new quarto file called 08-iteration.qmd
  3. Update the YAML header
  4. Replace the setup chunk with the one below:
```{r}
#| label: setup
#| include: false
library(tidyverse)  # loads purrr for iteration
library(broom)      # converts test output to tidy tables

set.seed(8675309) # makes sure random numbers are reproducible
```

Download the Apply functions with purrr cheat sheet.

8.1 Simple iteration functions

In the next two lectures, we are going to learn more about iteration (doing the same commands over and over) and custom functions through a data simulation exercise, which will also prepare us more traditional statistical topics. We first learned about the two basic iteration functions, rep() and seq() in the Working with Data chapter.

8.1.1 rep()

The function rep() lets you repeat the first argument a number of times.

Use rep() to create a vector of alternating "A" and "B" values of length 24.

rep(c("A", "B"), 12)
 [1] "A" "B" "A" "B" "A" "B" "A" "B" "A" "B" "A" "B" "A" "B" "A" "B" "A" "B" "A"
[20] "B" "A" "B" "A" "B"

If you don’t specify what the second argument is, it defaults to times, repeating the vector in the first argument that many times. Make the same vector as above, setting the second argument explicitly.

rep(c("A", "B"), times = 12)
 [1] "A" "B" "A" "B" "A" "B" "A" "B" "A" "B" "A" "B" "A" "B" "A" "B" "A" "B" "A"
[20] "B" "A" "B" "A" "B"

If the second argument is a vector that is the same length as the first argument, each element in the first vector is repeated than many times. Use rep() to create a vector of 11 "A" values followed by 3 "B" values.

rep(c("A", "B"), c(11, 3))
 [1] "A" "A" "A" "A" "A" "A" "A" "A" "A" "A" "A" "B" "B" "B"

You can repeat each element of the vector a sepcified number of times using the each argument, Use rep() to create a vector of 12 "A" values followed by 12 "B" values.

rep(c("A", "B"), each = 12)
 [1] "A" "A" "A" "A" "A" "A" "A" "A" "A" "A" "A" "A" "B" "B" "B" "B" "B" "B" "B"
[20] "B" "B" "B" "B" "B"

What do you think will happen if you set both times to 3 and each to 2?

rep(c("A", "B"), times = 3, each = 2)
 [1] "A" "A" "B" "B" "A" "A" "B" "B" "A" "A" "B" "B"

8.1.2 seq()

The function seq() is useful for generating a sequence of numbers with some pattern.

Use seq() to create a vector of the integers 0 to 10.

seq(0, 10)
 [1]  0  1  2  3  4  5  6  7  8  9 10

You can set the by argument to count by numbers other than 1 (the default). Use seq() to create a vector of the numbers 0 to 100 by 10s.

seq(0, 100, by = 10)
 [1]   0  10  20  30  40  50  60  70  80  90 100

The argument length.out is useful if you know how many steps you want to divide something into. Use seq() to create a vector that starts with 0, ends with 100, and has 12 equally spaced steps (hint: how many numbers would be in a vector with 2 steps?).

seq(0, 100, length.out = 13)
 [1]   0.000000   8.333333  16.666667  25.000000  33.333333  41.666667
 [7]  50.000000  58.333333  66.666667  75.000000  83.333333  91.666667
[13] 100.000000

8.1.3 replicate()

You can use the replicate() function to run a function n times.

For example, you can get 3 sets of 5 numbers from a random normal distribution by setting n to 3 and expr to rnorm(5).

replicate(n = 3, expr = rnorm(5))
           [,1]       [,2]       [,3]
[1,]  1.0240772  1.0030439  1.5912356
[2,] -0.4327152  0.2342534  1.0063919
[3,]  0.4384483 -0.1217803 -1.7631158
[4,] -1.4106443 -0.3635024 -1.0988308
[5,]  0.9016948  0.5873456  0.7136377

By default, replicate() simplifies your result into a matrix that is easy to convert into a table if your function returns vectors that are the same length. If you’d rather have a list of vectors, set simplify = FALSE.

replicate(n = 3, expr = rnorm(5), simplify = FALSE)
[[1]]
[1] -0.5988848  0.8261559  0.7951251  1.1043681 -2.9012889

[[2]]
[1]  0.6554194  1.0975918  0.9308105 -0.3094711 -0.1582473

[[3]]
[1] -0.1497308  1.0791335 -1.1998364 -0.9865283 -0.9067866

8.2 Complex iteration functions

purrr::map() and lapply() return a list of the same length as a vector or list, each element of which is the result of applying a function to the corresponding element. They function much the same, but purrr functions have some optimisations for working with the tidyverse. We’ll be working mostly with purrr functions in this course, but apply functions are very common in code that you might see in examples on the web.

Imagine you want to calculate the power for a two-sample t-test with a mean difference of 0.2 and SD of 1, for all the sample sizes 100 to 1000 (by 100s). You could run the power.t.test() function 20 times and extract the values for “power” from the resulting list and put it in a table.

p100 <- power.t.test(n = 100, delta = 0.2, sd = 1, type="two.sample")
# 18 more lines
p1000 <- power.t.test(n = 500, delta = 0.2, sd = 1, type="two.sample")

tibble(
  n = c(100, "...", 1000),
  power = c(p100$power, "...", p1000$power)
)
n power
100 0.290266404572216
… …
1000 0.884788352886661

However, the apply() and map() functions allow you to perform a function on each item in a vector or list. First make an object n that is the vector of the sample sizes you want to test, then use lapply() or map() to run the function power.t.test() on each item. You can set other arguments to power.t.test() after the function argument.

n <- seq(100, 1000, 100)
pcalc <- lapply(n, power.t.test, 
                delta = 0.2, sd = 1, type="two.sample")
# or
pcalc <- purrr::map(n, power.t.test, 
                delta = 0.2, sd = 1, type="two.sample")

These functions return a list where each item is the result of power.t.test(), which returns a list of results that includes the named item “power”. This is a special list that has a summary format if you just print it directly:

pcalc[[1]]

     Two-sample t test power calculation 

              n = 100
          delta = 0.2
             sd = 1
      sig.level = 0.05
          power = 0.2902664
    alternative = two.sided

NOTE: n is number in *each* group

But you can see the individual items using the str() function.

pcalc[[1]] |> str()
List of 8
 $ n          : num 100
 $ delta      : num 0.2
 $ sd         : num 1
 $ sig.level  : num 0.05
 $ power      : num 0.29
 $ alternative: chr "two.sided"
 $ note       : chr "n is number in *each* group"
 $ method     : chr "Two-sample t test power calculation"
 - attr(*, "class")= chr "power.htest"

sapply() is a version of lapply() that returns a vector or array instead of a list, where appropriate. The corresponding purrr functions are map_dbl(), map_chr(), map_int() and map_lgl(), which return vectors with the corresponding data type.

You can extract a value from a list with the function [[. You usually see this written as pcalc[[1]], but if you put it inside backticks, you can use it in apply and map functions.

sapply(pcalc, `[[`, "power")
 [1] 0.2902664 0.5140434 0.6863712 0.8064964 0.8847884 0.9333687 0.9623901
 [8] 0.9792066 0.9887083 0.9939638

We use map_dbl() here because the value for “power” is a double.

purrr::map_dbl(pcalc, `[[`, "power")
 [1] 0.2902664 0.5140434 0.6863712 0.8064964 0.8847884 0.9333687 0.9623901
 [8] 0.9792066 0.9887083 0.9939638

We can use the map() functions inside a mutate() function to run the power.t.test() function on the value of n from each row of a table, then extract the value for “power”, and delete the column with the power calculations.

mypower <- tibble(
  n = seq(100, 1000, 100)) |>
  mutate(pcalc = purrr::map(n, power.t.test, 
                            delta = 0.2, 
                            sd = 1, 
                            type="two.sample"),
         power = purrr::map_dbl(pcalc, `[[`, "power")) |>
  select(-pcalc)
A plot with 'n' (0-1000) on the x-axis and 'power' (0 to 1.0) on the y-axis. There is a horizontal red reference line at power = 0.8 and points plotted every 100 along the x-axis with a blue curve through them. Power starts around 0.3 for n = 100 and increases to about 0.8 for n = 400 and 1.0 for n = 1000.
Figure 8.1: Power for a two-sample t-test with d = 0.2

8.3 Exercises

  1. Create the following sequences using rep(), seq() and LETTERS.
  • “A” “B” “C” “D” “A” “B” “C” “D”
  • “A” “A” “B” “B” “C” “C” “D” “D”
  • “A” “A” “B” “B” “C” “C” “A” “A” “B” “B” “C” “C”
  • 4 9 14 19 24 29 34 39 44 49 54 59
  • 0.00000 33.33333 66.66667 100.00000
  • 0 20 40 60 80 100 0 20 40 60 80 100
  1. How many different ways can you think of to create the following table?
id condition manipulation
101 ctl A
201 ctl B
301 ctl C
401 ctl D
501 exp A
601 exp B
701 exp C
801 exp D
  1. Explain why this meme is funny.

See the text below for an accessible version

By arthurwelle on r/rstatsmemes
  • Lawful good r library(purrr) walk(rep("R!", 5), print)
  • Neutral good r for (i in 1:5) print("R!")
  • Chaotic good r for (i in 1:5) print("R!")
  • Lawful neutral r print("R!") print("R!") print("R!") print("R!") print("R!")
  • True neutral r for (i in 1:5) { print("R!") }
  • Chaotic neutral r mapply(\(x) print(x), rep("R!", 5)) |> invisible()
  • Lawful evil r i <<- 1 while (i < 6) { print("R!") i + 1 ->> i }
  • Neutral evil r f <- function(...) { do.call(print, list("R!")) } lapply(LETTERS[1:5], f) |> invisible()
  • Chaotic evil r apply(data.frame(rep("R!", 5)), 1, \(x){cat(x); cat("\n")}) |> invisible()

Glossary

term definition
data-type The kind of data represented by an object.
double A data type representing a real decimal number
function A named section of code that can be reused.
iteration Repeating a process or function
matrix A container data type consisting of numbers arranged into a fixed number of rows and columns

Further Resources