7  Haywood and Stiller (2024): Natural Environment and Mental Health

In this data analysis journey, we will wrangle a dataset, and conduct and interpret a multiple regression. These chapters are designed to help you become more independent in your data analysis pipelines and rely less on very structured instructions.

7.1 Task preparation

7.1.1 Introduction to the data set

For this task, we are using open data from a study looking at the mental health and well-being of allotment owners and non allotment owners. Here is a short explanation of the project:

Research has shown that spending time in nature has an important impact on wellbeing (Russell et al., 2013). Sessions such as gardening have been shown to encourage rehabilitation from mental and physical health conditions (Martin et al., 2020; Rappe, Koivunen & Korpela, 2008). Even just exposure to the natural environment has a significant impact on health and has been shown to increase positive emotions. This data file contains demographic information along with several established wellbeing scale scores of 515 both allotment and non-allotment owners to explore their time in nature and the impact of this on mental health.

You can read full details of the study here.

In today’s analysis, we are looking at whether allotment ownership (yes/no), connectedness to nature (measured by Relatedness to Nature Scale Short Form), and time spent in nature (measured in hours) are associated with better mental health (measured by Warwick-Edinburgh short wellbeing scale). We will add loneliness (measured by de Jong Short Form) as a control measure into the model.

To begin with, answer the following questions:

  • What are the predictor variables here?

  • What are their levels or are they continuous?

  • What is the outcome variable?

No variables are experimentally manipulated in this design. We have four predictors: allotment ownership (between-subjects: yes/no), time spent in nature (continuous), connectedness to nature (continuous), and loneliness (continuous). The outcome variable is participants’ mental health scores.

We will analyse these data with a multiple regression model. If you want to explore further, you can keep some of the other variables.

7.1.2 Organising your files and project for the task

Before we can get started, you need to organise your files and project for the task, so your working directory is in order.

  1. Create a new folder for the data analysis journey called Haywood_analysis. Within this folder, create two new folders called data and figures.

  2. Create an R Project for Haywood_analysis as an existing directory for your chapter folder. This should now be your working directory.

  3. Create a new R Markdown document and give it a sensible title describing the chapter, such as Haywood and Stiller (2024): Natural Environment and Mental Health. Delete everything below line 10 so you have a blank file to work with and save the file in your Haywood_analysis folder.

  4. We are working with one datafile: (Haywood_2024.csv). Right click the link and select “save link as”, or clicking the links will save the files to your Downloads. Save or copy the file to your data/ folder within Haywood_analysis.

You are now ready to start working on the task!

7.2 Overview

7.2.1 Load required packages and read in the data file

Before we explore what wrangling we need to do, load the relevant packages and read the data file.

# Load the required packages below
library(tidyverse)
library(performance)
library(broom)

# Load in the data file 
unprocessed_data <- ?

You should have the following in a code chunk:

# Load the tidyverse package below
library(tidyverse)
library(performance)
library(broom)

# This should be the raw_data.csv file 
unprocessed_data <- read_csv("data/Haywood_2024.csv")

7.2.2 Explore unprocessed_data

Rows: 543
Columns: 42
$ D_Age        <dbl> 34, 47, 56, 32, 39, 48, 44, 62, 59, 47, 51, 65, 46, 51, 5…
$ D_Gender     <chr> "2", "1", "2", "2", "2", "2", "2", "1", "1", "2", "2", "2…
$ D_Nation     <chr> "British", "English", "British", "British", "British", "B…
$ D_hours      <dbl> 5, 12, 6, 20, 10, 6, 21, 18, 25, 10, 20, 6, 10, 12, 20, 6…
$ D_activities <chr> "Walks, gardening", "Allotment", "Gardening, allotments, …
$ D_allot      <dbl> 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, …
$ D_group      <dbl> 2, 1, 2, 1, 2, 2, 2, 1, 2, 1, 2, 1, 2, 2, 2, 2, 1, 2, 1, …
$ MH_1         <dbl> 4, 3, 4, 4, 4, 4, 4, 4, 5, 3, 3, 2, 3, 4, 4, 4, 4, 4, 4, …
$ MH_2         <dbl> 4, 4, 4, 4, 4, 5, 4, 4, 5, 4, 3, 4, 2, 4, 4, 4, 3, 4, 4, …
$ MH_3         <dbl> 4, 4, 4, 4, 3, 3, 4, 4, 5, 3, 4, 3, 2, 3, 4, 4, 4, 3, 4, …
$ MH_4         <dbl> 3, 4, 4, 4, 4, 4, 3, 4, 4, 4, 2, 4, 3, 4, 3, 4, 4, 5, 4, …
$ MH_5         <dbl> 4, 4, 4, 4, 3, 4, 4, 3, 5, 4, 2, 4, 3, 4, 3, 4, 4, 5, 4, …
$ MH_6         <dbl> 4, 4, 4, 4, 4, 3, 4, 3, 5, 3, 1, 4, 3, 4, 4, 3, 4, 4, 3, …
$ MH_7         <dbl> 4, 5, 4, 5, 4, 5, 4, 4, 5, 4, 3, 3, 3, 4, 4, 5, 5, 5, 3, …
$ SE_1         <dbl> 2, 4, 4, 3, 4, 3, 2, 3, 5, 3, 1, 2, 1, 3, 1, 3, 4, 4, 3, …
$ PH_1         <dbl> 2, 2, 2, 3, 2, 4, 2, 4, 2, 5, 5, 2, 3, 4, 4, 2, 2, 4, 4, …
$ PH_2         <dbl> 2, 1, 1, 3, 2, 4, 2, 4, 1, 5, 5, 2, 2, 4, 4, 2, 2, 4, 3, …
$ Nat_1        <dbl> 5, 4, 4, 5, 5, 4, 5, 3, 4, 3, 5, 4, 5, 5, 5, 5, 3, 5, 4, …
$ Nat_2        <dbl> 5, 4, 5, 5, 5, 4, 5, 4, 5, 5, 5, 4, 4, 5, 4, 4, 4, 5, 4, …
$ Nat_3        <dbl> 5, 5, 5, 5, 5, 4, 5, 4, 5, 5, 5, 4, 5, 4, 4, 5, 4, 5, 4, …
$ Nat_4        <dbl> 5, 5, 0, 5, 5, 4, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 3, 5, 5, …
$ Nat_5        <dbl> 5, 4, 5, 5, 5, 3, 5, 4, 5, 5, 5, 5, 5, 5, 4, 5, 4, 5, 4, …
$ Nat_6        <dbl> 5, 5, 5, 5, 5, 4, 5, 4, 5, 5, 5, 4, 4, 5, 4, 5, 3, 5, 5, …
$ SocID_1      <dbl> 4, 5, 6, 6, 5, 2, 5, 5, 6, 4, 0, 7, 6, 4, 5, 3, 6, 4, 6, …
$ SocID_2      <dbl> 4, 6, 7, 6, 5, 2, 5, 5, 6, 4, 0, 6, 7, 4, 6, 0, 6, 4, 6, …
$ SocID_3      <dbl> 5, 7, 6, 7, 5, 2, 5, 6, 6, 6, 0, 6, 5, 4, 6, 0, 6, 4, 6, …
$ SocID_4      <dbl> 4, 5, 6, 7, 5, 1, 5, 6, 6, 3, 0, 7, 6, 4, 6, 0, 6, 4, 6, …
$ Lon_1        <dbl> 2, 3, 3, 3, 3, 3, 3, 3, 3, 3, 1, 3, 3, 3, 3, 3, 3, 3, 3, …
$ Lon_2        <dbl> 3, 1, 1, 1, 3, 2, 3, 3, 1, 3, 3, 1, 3, 1, 1, 3, 1, 2, 2, …
$ Lon_3        <dbl> 1, 2, 2, 1, 3, 3, 3, 3, 2, 3, 3, 1, 1, 1, 2, 3, 1, 3, 2, …
$ Lon_4        <dbl> 1, 1, 3, 3, 1, 3, 3, 2, 1, 3, 1, 2, 3, 1, 2, 3, 2, 3, 3, …
$ Lon_5        <dbl> 2, 1, 1, 1, 2, 1, 2, 3, 1, 1, 3, 1, 3, 1, 3, 2, 1, 2, 2, …
$ Lon_6        <dbl> 3, 3, 3, 3, 3, 3, 1, 3, 3, 3, 1, 3, 3, 3, 1, 3, 3, 3, 3, …
$ SocSup_1     <dbl> 4, 5, 5, 4, 4, 4, 3, 2, 7, 2, 1, 6, 3, 7, 5, 4, 5, 4, 5, …
$ SocSup_2     <dbl> 4, 5, 5, 5, 3, 4, 3, 3, 6, 2, 1, 7, 2, 7, 5, 4, 6, 6, 4, …
$ SocSup_3     <dbl> 4, 6, 4, 5, 7, 4, 3, 3, 6, 2, 1, 6, 2, 7, 5, 5, 5, 4, 4, …
$ SocSup_4     <dbl> 4, 6, 6, 6, 5, 4, 3, 4, 6, 2, 2, 7, 2, 7, 5, 3, 4, 5, 4, …
$ SEff_1       <dbl> 3, 3, 3, 4, 2, 3, 3, 3, 4, 2, 3, 3, 2, 4, 3, 3, 4, 4, 3, …
$ SEff_2       <dbl> 2, 4, 3, 4, 3, 3, 2, 3, 4, 3, 2, 4, 2, 4, 4, 4, 4, 4, 3, …
$ SEff_3       <dbl> 2, 4, 4, 4, 3, 4, 3, 3, 4, 2, 3, 3, 3, 4, 4, 3, 4, 3, 3, …
$ SEff_4       <dbl> 3, 4, 3, 4, 3, 4, 3, 3, 4, 4, 3, 3, 2, 4, 4, 3, 4, 4, 3, …
$ SEff_5       <dbl> 2, 4, 2, 3, 3, 3, 3, 2, 2, 1, 2, 4, 2, 3, 3, 3, 3, 2, 3, …

For the purposes of today’s exercise, the columns that matter to us are:

Variable Type Description
D_Age double Participant age.
D_Gender character Participant gender, 1 for male, 2 for female
D_hours double Number of hours spent in nature per week
D_allot double Do you currently own/work on an allotment site? 1 for yes, 2 for no
D_activities character What activities in nature do you participate in? Open text response
MH_1 to MH_7 double The short wellbeing scale. Participant answers each statement from 1=none of the time to 5=all of the time. The overall score is calculated as a sum score, with higher scores indicating higher wellbeing
Nat_1 to Nat_6 double The short nature connectedness scale. Participant answers each statement from 1=disagree strongly to 5=agree strongly. The overall score is calculated as a sum score, with higher scores indicating stronger nature connectedness
Lon_1 to Lon_6 double De Jong Short Form loneliness questionnaire. Participant answers each statement from 1=yes, 2=more or less, 3=no. The overall score is calculated as a sum score, with higher scores indicating higher loneliness. Items 2,3 and 5 are reverse coded
TipTry this

Next, spend some time exploring the data. For example, open the data object as a tab to scroll around, explore with glimpse(), and try quickly plotting some of the variables with ggplot2.

Can you see any responses in there that are going to cause problems?

7.3 First data wrangling steps

Rows: 543
Columns: 42
$ D_Age        <dbl> 34, 47, 56, 32, 39, 48, 44, 62, 59, 47, 51, 65, 46, 51, 5…
$ D_Gender     <chr> "2", "1", "2", "2", "2", "2", "2", "1", "1", "2", "2", "2…
$ D_Nation     <chr> "British", "English", "British", "British", "British", "B…
$ D_hours      <dbl> 5, 12, 6, 20, 10, 6, 21, 18, 25, 10, 20, 6, 10, 12, 20, 6…
$ D_activities <chr> "Walks, gardening", "Allotment", "Gardening, allotments, …
$ D_allot      <dbl> 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, …
$ D_group      <dbl> 2, 1, 2, 1, 2, 2, 2, 1, 2, 1, 2, 1, 2, 2, 2, 2, 1, 2, 1, …
$ MH_1         <dbl> 4, 3, 4, 4, 4, 4, 4, 4, 5, 3, 3, 2, 3, 4, 4, 4, 4, 4, 4, …
$ MH_2         <dbl> 4, 4, 4, 4, 4, 5, 4, 4, 5, 4, 3, 4, 2, 4, 4, 4, 3, 4, 4, …
$ MH_3         <dbl> 4, 4, 4, 4, 3, 3, 4, 4, 5, 3, 4, 3, 2, 3, 4, 4, 4, 3, 4, …
$ MH_4         <dbl> 3, 4, 4, 4, 4, 4, 3, 4, 4, 4, 2, 4, 3, 4, 3, 4, 4, 5, 4, …
$ MH_5         <dbl> 4, 4, 4, 4, 3, 4, 4, 3, 5, 4, 2, 4, 3, 4, 3, 4, 4, 5, 4, …
$ MH_6         <dbl> 4, 4, 4, 4, 4, 3, 4, 3, 5, 3, 1, 4, 3, 4, 4, 3, 4, 4, 3, …
$ MH_7         <dbl> 4, 5, 4, 5, 4, 5, 4, 4, 5, 4, 3, 3, 3, 4, 4, 5, 5, 5, 3, …
$ SE_1         <dbl> 2, 4, 4, 3, 4, 3, 2, 3, 5, 3, 1, 2, 1, 3, 1, 3, 4, 4, 3, …
$ PH_1         <dbl> 2, 2, 2, 3, 2, 4, 2, 4, 2, 5, 5, 2, 3, 4, 4, 2, 2, 4, 4, …
$ PH_2         <dbl> 2, 1, 1, 3, 2, 4, 2, 4, 1, 5, 5, 2, 2, 4, 4, 2, 2, 4, 3, …
$ Nat_1        <dbl> 5, 4, 4, 5, 5, 4, 5, 3, 4, 3, 5, 4, 5, 5, 5, 5, 3, 5, 4, …
$ Nat_2        <dbl> 5, 4, 5, 5, 5, 4, 5, 4, 5, 5, 5, 4, 4, 5, 4, 4, 4, 5, 4, …
$ Nat_3        <dbl> 5, 5, 5, 5, 5, 4, 5, 4, 5, 5, 5, 4, 5, 4, 4, 5, 4, 5, 4, …
$ Nat_4        <dbl> 5, 5, 0, 5, 5, 4, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 3, 5, 5, …
$ Nat_5        <dbl> 5, 4, 5, 5, 5, 3, 5, 4, 5, 5, 5, 5, 5, 5, 4, 5, 4, 5, 4, …
$ Nat_6        <dbl> 5, 5, 5, 5, 5, 4, 5, 4, 5, 5, 5, 4, 4, 5, 4, 5, 3, 5, 5, …
$ SocID_1      <dbl> 4, 5, 6, 6, 5, 2, 5, 5, 6, 4, 0, 7, 6, 4, 5, 3, 6, 4, 6, …
$ SocID_2      <dbl> 4, 6, 7, 6, 5, 2, 5, 5, 6, 4, 0, 6, 7, 4, 6, 0, 6, 4, 6, …
$ SocID_3      <dbl> 5, 7, 6, 7, 5, 2, 5, 6, 6, 6, 0, 6, 5, 4, 6, 0, 6, 4, 6, …
$ SocID_4      <dbl> 4, 5, 6, 7, 5, 1, 5, 6, 6, 3, 0, 7, 6, 4, 6, 0, 6, 4, 6, …
$ Lon_1        <dbl> 2, 3, 3, 3, 3, 3, 3, 3, 3, 3, 1, 3, 3, 3, 3, 3, 3, 3, 3, …
$ Lon_2        <dbl> 3, 1, 1, 1, 3, 2, 3, 3, 1, 3, 3, 1, 3, 1, 1, 3, 1, 2, 2, …
$ Lon_3        <dbl> 1, 2, 2, 1, 3, 3, 3, 3, 2, 3, 3, 1, 1, 1, 2, 3, 1, 3, 2, …
$ Lon_4        <dbl> 1, 1, 3, 3, 1, 3, 3, 2, 1, 3, 1, 2, 3, 1, 2, 3, 2, 3, 3, …
$ Lon_5        <dbl> 2, 1, 1, 1, 2, 1, 2, 3, 1, 1, 3, 1, 3, 1, 3, 2, 1, 2, 2, …
$ Lon_6        <dbl> 3, 3, 3, 3, 3, 3, 1, 3, 3, 3, 1, 3, 3, 3, 1, 3, 3, 3, 3, …
$ SocSup_1     <dbl> 4, 5, 5, 4, 4, 4, 3, 2, 7, 2, 1, 6, 3, 7, 5, 4, 5, 4, 5, …
$ SocSup_2     <dbl> 4, 5, 5, 5, 3, 4, 3, 3, 6, 2, 1, 7, 2, 7, 5, 4, 6, 6, 4, …
$ SocSup_3     <dbl> 4, 6, 4, 5, 7, 4, 3, 3, 6, 2, 1, 6, 2, 7, 5, 5, 5, 4, 4, …
$ SocSup_4     <dbl> 4, 6, 6, 6, 5, 4, 3, 4, 6, 2, 2, 7, 2, 7, 5, 3, 4, 5, 4, …
$ SEff_1       <dbl> 3, 3, 3, 4, 2, 3, 3, 3, 4, 2, 3, 3, 2, 4, 3, 3, 4, 4, 3, …
$ SEff_2       <dbl> 2, 4, 3, 4, 3, 3, 2, 3, 4, 3, 2, 4, 2, 4, 4, 4, 4, 4, 3, …
$ SEff_3       <dbl> 2, 4, 4, 4, 3, 4, 3, 3, 4, 2, 3, 3, 3, 4, 4, 3, 4, 3, 3, …
$ SEff_4       <dbl> 3, 4, 3, 4, 3, 4, 3, 3, 4, 4, 3, 3, 2, 4, 4, 3, 4, 4, 3, …
$ SEff_5       <dbl> 2, 4, 2, 3, 3, 3, 3, 2, 2, 1, 2, 4, 2, 3, 3, 3, 3, 2, 3, …
Rows: 515
Columns: 25
$ D_Age        <dbl> 34, 47, 56, 32, 39, 48, 44, 62, 59, 47, 51, 65, 46, 51, 5…
$ D_Gender     <chr> "2", "1", "2", "2", "2", "2", "2", "1", "1", "2", "2", "2…
$ D_hours      <dbl> 5, 12, 6, 20, 10, 6, 21, 18, 25, 10, 20, 6, 10, 12, 20, 6…
$ D_activities <chr> "Walks, gardening", "Allotment", "Gardening, allotments, …
$ D_allot      <dbl> 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, …
$ MH_1         <dbl> 4, 3, 4, 4, 4, 4, 4, 4, 5, 3, 3, 2, 3, 4, 4, 4, 4, 4, 4, …
$ MH_2         <dbl> 4, 4, 4, 4, 4, 5, 4, 4, 5, 4, 3, 4, 2, 4, 4, 4, 3, 4, 4, …
$ MH_3         <dbl> 4, 4, 4, 4, 3, 3, 4, 4, 5, 3, 4, 3, 2, 3, 4, 4, 4, 3, 4, …
$ MH_4         <dbl> 3, 4, 4, 4, 4, 4, 3, 4, 4, 4, 2, 4, 3, 4, 3, 4, 4, 5, 4, …
$ MH_5         <dbl> 4, 4, 4, 4, 3, 4, 4, 3, 5, 4, 2, 4, 3, 4, 3, 4, 4, 5, 4, …
$ MH_6         <dbl> 4, 4, 4, 4, 4, 3, 4, 3, 5, 3, 1, 4, 3, 4, 4, 3, 4, 4, 3, …
$ MH_7         <dbl> 4, 5, 4, 5, 4, 5, 4, 4, 5, 4, 3, 3, 3, 4, 4, 5, 5, 5, 3, …
$ Nat_1        <dbl> 5, 4, 4, 5, 5, 4, 5, 3, 4, 3, 5, 4, 5, 5, 5, 5, 3, 5, 4, …
$ Nat_2        <dbl> 5, 4, 5, 5, 5, 4, 5, 4, 5, 5, 5, 4, 4, 5, 4, 4, 4, 5, 4, …
$ Nat_3        <dbl> 5, 5, 5, 5, 5, 4, 5, 4, 5, 5, 5, 4, 5, 4, 4, 5, 4, 5, 4, …
$ Nat_4        <dbl> 5, 5, 0, 5, 5, 4, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 3, 5, 5, …
$ Nat_5        <dbl> 5, 4, 5, 5, 5, 3, 5, 4, 5, 5, 5, 5, 5, 5, 4, 5, 4, 5, 4, …
$ Nat_6        <dbl> 5, 5, 5, 5, 5, 4, 5, 4, 5, 5, 5, 4, 4, 5, 4, 5, 3, 5, 5, …
$ Lon_1        <dbl> 2, 3, 3, 3, 3, 3, 3, 3, 3, 3, 1, 3, 3, 3, 3, 3, 3, 3, 3, …
$ Lon_2        <dbl> 3, 1, 1, 1, 3, 2, 3, 3, 1, 3, 3, 1, 3, 1, 1, 3, 1, 2, 2, …
$ Lon_3        <dbl> 1, 2, 2, 1, 3, 3, 3, 3, 2, 3, 3, 1, 1, 1, 2, 3, 1, 3, 2, …
$ Lon_4        <dbl> 1, 1, 3, 3, 1, 3, 3, 2, 1, 3, 1, 2, 3, 1, 2, 3, 2, 3, 3, …
$ Lon_5        <dbl> 2, 1, 1, 1, 2, 1, 2, 3, 1, 1, 3, 1, 3, 1, 3, 2, 1, 2, 2, …
$ Lon_6        <dbl> 3, 3, 3, 3, 3, 3, 1, 3, 3, 3, 1, 3, 3, 3, 1, 3, 3, 3, 3, …
$ nas          <dbl> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, …
TipTry this

Before we give you a task list, try and switch between the raw data and the wrangled data. Make a list of all the differences you can see between the two data objects.

  1. What has changed from the raw data?

  2. What number of observations do we have? What responses have been taken out and why?

When you get as far as you can, check the task list which explains all the steps we applied, but not how to do them. Then, you can check the solution for our code.

7.3.1 Task list

These are all the steps we applied to create the wrangled data object:

  1. Select the columns we need.

  2. Remove NAs. Here we removed only rows which have all NAs (some participants had one NA, so we can still use their data). In order to do this, we are adding a new column that shows the number of NAs for each row (participant), and then filter the dataset based on this column.

It’s not always easy to decide which strategy to use for missing data, particularly if you are collecting primary data and have a smaller sample size than anticipated. This is a discussion you should have with your dissertation supervisor, and usually it’s good to decide in advance (you can detail this in your pre-registration).

7.3.2 Solution

This is the code we used to tidy up the dataset before calculating demographic information. As long as you get the same end result, the exact code is not important - there are always multiple ways of doing the same thing.

data_processed_1 <- unprocessed_data %>% 
  select(D_Age, 
         D_Gender, 
         D_hours, 
         D_activities, 
         D_allot, 
         MH_1:MH_7, 
         Nat_1:Nat_6, 
         Lon_1:Lon_6) %>% 
  mutate(nas = rowSums(is.na(.))) %>% 
  filter(nas < 24)

7.4 Potential side quest

If you would like to explore how to process messy text data, follow the side quest below.

As you may see on the unprocessed_data file, the D_activities column lists all the different nature activities our participants do. It is an open text response column where the participant responses are indeed quite messy. As your side quest (and a cautionary tale in open text responses), see if you can get to a list of different activities and how many participants have said this

You want the data to look something like we show below:

activity n
gardening 419
walking 360
allotment 194
dog walking 65
visiting green spaces 42
cycling 41
running 22
bird watching 17
photography 17
keeping livestock 15
hiking 11
fishing 8
conservation work 6
horse riding 6
working 6
beach activities 5
sitting 5
eating 4
foraging 4
shooting 4
bee-keeping 3
camping 3
feeding wildlife 3
golf 3
meditation 3
outdoor holidays 3
rock climbing 3
swimming 3
bug activity 2
forest school 2
kayaking 2
living in a rural area 2
studying 2
taking children out 2
wildlife surveying 2
archaeology 1
bootcamp 1
bowls competitions 1
cutting firewood 1
forrestry 1
geological mapping 1
identifying wild plants 1
instagram 1
live web cam viewing (shetland) 1
outdoor play 1
outdoor sports 1
reading about wildlife 1
scouting 1
shopping 1
skiing 1
snow shoeing 1
spectator football 1
teaching 1
tennis 1
volunteering 1
watch wildlife 1

Here is one suggested way to do this which is not perfect but gives you an idea of how you might approach something like this - dealing with text data can be quite messy and time-consuming!

If we used these responses for actual hypothesis-testing analysis, we would need to pay more attention to getting things exactly right.

Sometimes, if your open text data are really messy, fixing things manually on excel can be quicker and easier than trying to write the code for every eventuality. The danger in this approach is that it’s not reproducible and it’s more prone to human error.

#First select the column you need and make all the letters lower case - R treats upper and lower case letters as different characters
activities <- data_processed_1 %>% 
  select(D_activities) %>% 
  mutate(D_activities=tolower(D_activities))

#Most people have separated their different activities with a comma but some have used different characters. This code tries to scope the most common separators and make them into commas.
activities$D_activities <- str_replace_all(activities$D_activities, 
                                           c("&" =",", "/"=",", " and"=",", "\\."=",", "/+"=",", ";"=",", "-"=","))

#This didn't work perfectly as some people have not used any separators apart from spaces. Filtering a list of people with no commas so we can manually change some of the separators
activities <- mutate(activities, 
                     row_n = row_number())

separators_missing <- activities %>% 
  filter(!str_detect(D_activities, ","))

#Do a couple of manual changes for separators 
activities[23,1] <- "dog walking,gardening,cycling,shooting"
activities[38,1] <- "walking,gardening,allotment,park"
activities[44,1] <- "gardening,walking,contemplation,birdwatching"
activities[45,1] <- "gardening, dog walking"
activities[48,1] <- "gardening,walking,teaching"
activities[60,1] <- "walks,gardening,photography"
activities[62,1] <- "walking dogs,gardening"
activities[64,1] <- "walks,gardening,allotment"
activities[93,1] <- "walking,gardening,swimming"
activities[96,1] <- "gardening,cycling"
activities[103,1]<- "walking,allotment,hiking,scouting"
activities[140,1]<- "gardening,walking,allotment"
activities[150,1]<- "gardening,walking,allotment"
activities[162,1]<- "walking,gardening"
activities[176,1]<- "walking,gardening"
activities[185,1]<- "walking,gardening"
activities[195,1]<- "walking,gardening"
activities[197,1]<- "dog walking,gardening"
activities[199,1]<- "walking,gardening,cycling"
activities[245,1]<- "walking,running,gardening"
activities[246,1]<- "walking,gardening,cycling"
activities[253,1]<- "gardening,walking,sightseeing when on holiday"
activities[260,1]<- "alllotment,walking,gardening"
activities[273,1]<- "walks,gardening,allotment,cycling"
activities[303,1]<- "walks,gardening"
activities[311,1]<- "walking,gardening,allotment"
activities[315,1]<- "walks,gardening"
activities[317,1]<- "walks,gardening,allotment"
activities[321,1]<- "walking the dog,gardening"
activities[334,1]<- "walks,allotment"
activities[358,1]<- "walking,gardening,sitting,working"
activities[401,1]<- "walks,gardening"
activities[418,1]<- "walking,gardening,cycling"
activities[441,1]<- "walking,gardening"
activities[446,1]<- "walking,gardening"
activities[458,1]<- "walking,gardening"
activities[463,1]<- "walking,gardening"
activities[510,1]<- "walking,gardening"

# Then we can separate multiple activities into their own columns, so that all activities get counted
activities <- activities %>% 
  separate(D_activities, 
           into=c("a1", "a2", "a3", "a4", "a5"), sep=",") %>% 
  pivot_longer(a1:a5, 
               names_to="a_n", 
               values_to = "activity") %>% 
  drop_na()

#Use str_squish to remove white space from the character strings. Stringr operators don't work very well as part of pipes so you might need to sometimes do them as their own separate lines of codes.
activities$activity <- str_squish(activities$activity)

#Test to see how many different activities there are so we can replace some that are the same but misspelled or spelled differently. We will use this as our reference table and try to reduce the number of same/similar activities by iteratively fixing things and looking back at this reference dataframe.
test1 <- activities %>% 
  group_by(activity) %>% 
  summarise(n = n())

#Using str_detect and recoding is a good tool, but it's causing us a slight issue with "walking" and "dog walking", as those will easily get overwritten with str_detect. We are going to do a somewhat janky workaround for this - you can find a more elegant solution if you want to. We will separate dog walking here and add it back at the end
dog_activities <- activities %>% 
  filter(str_detect(activity, "dog|spaniel")) %>% 
  mutate(activity = case_when(
      str_detect(activity, "dog|spaniel") ~ "dog walking",
      TRUE ~ activity)) %>% 
  count(activity)

#We'll remove dogs from the activities dataset for now
activities <- activities %>%
  filter(!str_detect(activity, "dog|spaniel"))

#Next, we'll attempt to catch some of the common activities which have multiple spellings. You can then keep adding more case_when cases as you reduce the number of activities on the test1 dataframe (remember to group_by the new recoded activity column)
activities <- activities %>% drop_na() %>% 
  mutate(
    recoded_activity = case_when(
      str_detect(activity, "walk|walks|walking") ~ "walking", 
      str_detect(activity, "beach|seaside") ~ "beach activities",
      str_detect(activity, "bird") ~ "bird watching",
      str_detect(activity, "garde|gare|grd|yard|vegetable|grow|landscap|glasshouse|planting|potting|fields|flowers|edging grass|bloom|crop|seed|soil|at home|grounds|edible|cutting grass|hanging baskets") ~ "gardening",
      str_detect(activity, "allot|alot|alllot") ~ "allotment",
      str_detect(activity, "bug") ~ "bug activity",
      str_detect(activity, "cyc|cyx|biki|off road") ~ "cycling",
      str_detect(activity, "visit|trip|site seeing|park|green|reserve|enjoying|find routes|days out|looking at nature|just looking|countryside|exploring area|excursion") ~ "visiting green spaces",
      str_detect(activity, "volun") ~ "volunteering",
      str_detect(activity, "study") ~ "studying",
      str_detect(activity, "bee") ~ "bee-keeping",
      str_detect(activity, "chicken|hens|farm|looking after|making animal|poultry|smallhold") ~ "keeping livestock",
      str_detect(activity, "hiki|hikes|hill|back-packing|pack|kennedy") ~ "hiking",
      str_detect(activity, "photo") ~ "photography",
      str_detect(activity, "feed") ~ "feeding wildlife",
      str_detect(activity, "horse|riding") ~ "horse riding",
      str_detect(activity, "conser") ~ "conservation work",
      str_detect(activity, "children|kid") ~ "taking children out",
      str_detect(activity, "living|i live") ~ "living in a rural area",
      str_detect(activity, "climb") ~ "rock climbing",
      str_detect(activity, "pond|lake|swim") ~ "swimming",
      str_detect(activity, "work|job") ~ "working",
      str_detect(activity, "holiday|travel") ~ "outdoor holidays",
      str_detect(activity, "geocaching") ~ "geocaching",
      str_detect(activity, "eating|picnic|food") ~ "eating",
      str_detect(activity, "fishing") ~ "fishing",
      str_detect(activity, "ski") ~ "skiing",
      str_detect(activity, "contemplation|meditation|tree outside my window") ~ "meditation",
      str_detect(activity, "jogging|running") ~ "running",
      str_detect(activity, "bug") ~ "bug activity",
      str_detect(activity, "sitting") ~ "sitting",
      TRUE ~ activity
    )
  )


#Check again - remember to use the new column to group by
test1 <- activities %>% 
  group_by(recoded_activity) %>% 
  summarise(n = n())

#Filter out some not plausible answers
activities <- activities %>% 
  filter(recoded_activity != "") %>% 
  filter(recoded_activity != "4") %>% 
  filter(recoded_activity != "a") %>% 
  filter(recoded_activity != "but in the current season i am only spending a few hours per week at my plot") %>% 
  filter(recoded_activity != "but lack the time") %>% 
  filter(recoded_activity != "daagse nijmegen") %>% 
  filter(recoded_activity != "diagnosed aged 35") %>% 
  filter(recoded_activity != "i try") %>% 
  filter(recoded_activity != "ing") %>% 
  filter(recoded_activity != "keeping") %>% 
  filter(recoded_activity != "mature rambles") %>% 
  filter(recoded_activity != "my activities in nature are limited by having rheumatoid arthritis") %>% 
  filter(recoded_activity != "more in summer than in winter") %>%  
  filter(recoded_activity != "n") %>%  
  filter(recoded_activity != "none") %>%  
  filter(recoded_activity != "nt") %>%  
  filter(recoded_activity != "road") %>%  
  filter(recoded_activity != "sitting in traffic") %>%  
  filter(recoded_activity != "troughs") %>%  
  filter(recoded_activity != "with intermittent periods of incapacity") %>% filter(recoded_activity != "long distance back") %>% 
  filter(recoded_activity != "wood")

#Do a new test and add some other cases that will go together - add to the str_detect code above (otherwise your cases will be overwritten)
test2 <- activities %>% 
  group_by(recoded_activity) %>% 
  summarise(n = n())

#Final list with dogs added back in
list_activities <- test2 %>% 
  rename(activity=recoded_activity) %>% 
  rbind(dog_activities)

7.5 Calculating demographics

Next, add a new code chunk to calculate the following demographics on using data_processed_1.

7.5.1 Demographics

  1. There were women, men, and transgender/non-binary participants.
data_processed_1 %>% 
  count(D_Gender)
D_Gender n
1 143
2 370
Transgender 1
non binary 1
  1. To two decimal places, the mean age of all the participants was (SD = ).
data_processed_1 %>% 
  summarise(mean_age = round(mean(D_Age), 2), 
            sd_age = round(sd(D_Age), 2))
mean_age sd_age
50.51 15.41
  1. There were allotment owners and non-allotment owners. The mean time spent in nature was hours, ranging from to hours
data_processed_1 %>% 
  count(D_allot)

data_processed_1 %>% 
  summarise(mean_hours = mean(D_hours), 
            sd_hours = sd(D_hours), 
            min = min(D_hours), 
            max = max(D_hours))
D_allot n
1 375
2 140
mean_hours sd_hours min max
12.9165 10.99606 0 80

7.6 Wrangling the data further

Rows: 515
Columns: 25
$ D_Age        <dbl> 34, 47, 56, 32, 39, 48, 44, 62, 59, 47, 51, 65, 46, 51, 5…
$ D_Gender     <chr> "2", "1", "2", "2", "2", "2", "2", "1", "1", "2", "2", "2…
$ D_hours      <dbl> 5, 12, 6, 20, 10, 6, 21, 18, 25, 10, 20, 6, 10, 12, 20, 6…
$ D_activities <chr> "Walks, gardening", "Allotment", "Gardening, allotments, …
$ D_allot      <dbl> 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, …
$ MH_1         <dbl> 4, 3, 4, 4, 4, 4, 4, 4, 5, 3, 3, 2, 3, 4, 4, 4, 4, 4, 4, …
$ MH_2         <dbl> 4, 4, 4, 4, 4, 5, 4, 4, 5, 4, 3, 4, 2, 4, 4, 4, 3, 4, 4, …
$ MH_3         <dbl> 4, 4, 4, 4, 3, 3, 4, 4, 5, 3, 4, 3, 2, 3, 4, 4, 4, 3, 4, …
$ MH_4         <dbl> 3, 4, 4, 4, 4, 4, 3, 4, 4, 4, 2, 4, 3, 4, 3, 4, 4, 5, 4, …
$ MH_5         <dbl> 4, 4, 4, 4, 3, 4, 4, 3, 5, 4, 2, 4, 3, 4, 3, 4, 4, 5, 4, …
$ MH_6         <dbl> 4, 4, 4, 4, 4, 3, 4, 3, 5, 3, 1, 4, 3, 4, 4, 3, 4, 4, 3, …
$ MH_7         <dbl> 4, 5, 4, 5, 4, 5, 4, 4, 5, 4, 3, 3, 3, 4, 4, 5, 5, 5, 3, …
$ Nat_1        <dbl> 5, 4, 4, 5, 5, 4, 5, 3, 4, 3, 5, 4, 5, 5, 5, 5, 3, 5, 4, …
$ Nat_2        <dbl> 5, 4, 5, 5, 5, 4, 5, 4, 5, 5, 5, 4, 4, 5, 4, 4, 4, 5, 4, …
$ Nat_3        <dbl> 5, 5, 5, 5, 5, 4, 5, 4, 5, 5, 5, 4, 5, 4, 4, 5, 4, 5, 4, …
$ Nat_4        <dbl> 5, 5, 0, 5, 5, 4, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 3, 5, 5, …
$ Nat_5        <dbl> 5, 4, 5, 5, 5, 3, 5, 4, 5, 5, 5, 5, 5, 5, 4, 5, 4, 5, 4, …
$ Nat_6        <dbl> 5, 5, 5, 5, 5, 4, 5, 4, 5, 5, 5, 4, 4, 5, 4, 5, 3, 5, 5, …
$ Lon_1        <dbl> 2, 3, 3, 3, 3, 3, 3, 3, 3, 3, 1, 3, 3, 3, 3, 3, 3, 3, 3, …
$ Lon_2        <dbl> 3, 1, 1, 1, 3, 2, 3, 3, 1, 3, 3, 1, 3, 1, 1, 3, 1, 2, 2, …
$ Lon_3        <dbl> 1, 2, 2, 1, 3, 3, 3, 3, 2, 3, 3, 1, 1, 1, 2, 3, 1, 3, 2, …
$ Lon_4        <dbl> 1, 1, 3, 3, 1, 3, 3, 2, 1, 3, 1, 2, 3, 1, 2, 3, 2, 3, 3, …
$ Lon_5        <dbl> 2, 1, 1, 1, 2, 1, 2, 3, 1, 1, 3, 1, 3, 1, 3, 2, 1, 2, 2, …
$ Lon_6        <dbl> 3, 3, 3, 3, 3, 3, 1, 3, 3, 3, 1, 3, 3, 3, 1, 3, 3, 3, 3, …
$ nas          <dbl> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, …
Rows: 9,785
Columns: 6
$ D_hours        <dbl> 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5…
$ D_allot        <dbl> 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1…
$ participant_id <int> 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1…
$ measure        <chr> "MH", "MH", "MH", "MH", "MH", "MH", "MH", "Nat", "Nat",…
$ item           <chr> "1", "2", "3", "4", "5", "6", "7", "1", "2", "3", "4", …
$ score          <dbl> 4, 4, 4, 3, 4, 4, 4, 5, 5, 5, 5, 5, 5, 2, 3, 1, 1, 2, 3…
TipTry this

Before we give you a task list, try and switch between the two dataframes above. Make a list of all the differences you can see between the two dataframes.

  1. Are the data in long or wide format? Which variables have been pivoted?

  2. Which variables do each of the datasets have? How would you create these variables in the further wrangled data based on the variables you have in the half wrangled data? Try to identify the new columns and figure out how they were made.

  3. Which columns contain our information about the predictors? Which column contains our outcome variable?

  4. What is the variable type of each variable?

When you get as far as you can, check the task list which explains all the steps we applied, but not how to do them. Then, you can check the solution for our code.

7.6.1 Task list

These are all the steps we applied:

  1. Get rid of columns we don’t need (demographics and the na column)

  2. Add a new column for participant number (using row numbers in wide format dataframe). We need this later.

  3. Pivot longer so we have one column for the measures and one for participants’ scores.

7.6.2 Solution

This is the code we used:

data_processed_2 <- data_processed_1 %>% 
  select(-D_Age, 
         -D_Gender, 
         -D_activities, 
         -nas) %>% 
  mutate(participant_id = row_number()) %>% 
  pivot_longer(MH_1:Lon_6, 
               names_to = "measure", 
               values_to="score") 

7.7 Final data wrangling steps

Now it’s time to do the final steps before our data is ready for analysis. At this point, we want to compute participants’ scores for each of the questionnaires they took. Remember that items 2, 3, and 5 from the loneliness questionnaire have to be reverse scored.

Rows: 9,785
Columns: 6
$ D_hours        <dbl> 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5…
$ D_allot        <dbl> 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1…
$ participant_id <int> 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1…
$ measure        <chr> "MH", "MH", "MH", "MH", "MH", "MH", "MH", "Nat", "Nat",…
$ item           <chr> "1", "2", "3", "4", "5", "6", "7", "1", "2", "3", "4", …
$ score          <dbl> 4, 4, 4, 3, 4, 4, 4, 5, 5, 5, 5, 5, 5, 2, 3, 1, 1, 2, 3…
Rows: 515
Columns: 6
$ participant_id <int> 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, …
$ D_hours        <dbl> 5, 12, 6, 20, 10, 6, 21, 18, 25, 10, 20, 6, 10, 12, 20,…
$ D_allot        <fct> 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1…
$ Lon            <dbl> 12, 11, 13, 12, 15, 15, 15, 17, 11, 16, 12, 11, 16, 10,…
$ MH             <dbl> 27, 28, 28, 29, 26, 28, 27, 26, 34, 25, 18, 24, 19, 27,…
$ Nat            <dbl> 30, 27, 24, 30, 30, 23, 30, 24, 29, 28, 30, 26, 28, 29,…
TipTry this

Before we give you a task list, try and switch between the two dataframes above. Make a list of all the differences you can see between the two dataframes.

  1. Are the data in long or wide format? Which variables have been pivoted?

  2. Which variables do each of the datasets have? How would you create these variables?

  3. Which columns have been aggregated? How was this done?

  4. What is the variable type of each variable?

When you get as far as you can, check the task list which explains all the steps we applied, but not how to do them. Then, you can check the solution for our code.

7.7.1 Task list

These are all the steps we applied:

  1. Create a new column where items 2, 3, and 5 on the loneliness scale are reverse scored and all other items keep their original score.

  2. Separate the measure column into two so you have a column for measure and for item (you need to do this to be able to group by measure next).

  3. Group by measure and participant id.

  4. Summarise (sum score) the final score column (be careful to use the one with reverse scores).

  5. Make a new dataframe where you take the info on participant id, hours in nature, and allotment ownership.

  6. Only keep unique observations of the new dataframe you made.

  7. Join this dataframe with your final scores for participants on the questionnaires.

  8. Pivot wider so each participant is represented in one row of data and your predictor variables each have their own column.

  9. Change D_allot column into a factor.

7.7.2 Solution

This is the code we used. Group_by and summarise is a great combination, but it does mean you lose some columns which have important information. There are many different ways around this - below, we show you one example (you can also aggregate scores by group_by and mutate, and then ungroup to retain all columns, but this causes a different set of issues).

final_scores <- data_processed_2 %>% 
  mutate(updated_score = ifelse(measure %in% c("Lon_2", "Lon_3", "Lon_5"), 
                                4-score, 
                                score)) %>% 
  separate(measure, 
           into = c("measure", "item"), 
           sep = "_") %>% 
  group_by(participant_id, 
           measure) %>% 
  summarise(final_score = sum(updated_score))

#Make a separate dataframe for the data you lost when grouping by and summarising. Use unique() to only keep unique data (one observation per participant)
variables <- data_processed_2 %>% 
  select(participant_id, 
         D_allot, 
         D_hours) %>% 
  unique()

#Finally, you need to pivot wider; for a multiple regression, each predictor variable needs to be represented in their own column
final_scores <- inner_join(variables, 
                           final_scores, 
                           "participant_id") %>% 
  pivot_wider(id_cols = c(participant_id, D_hours, D_allot), 
              names_from = measure, 
              values_from = final_score) %>% 
  mutate(D_allot = as.factor(D_allot))

7.8 Descriptive statistics and visualisation

Calculating descriptive statistics for regression data can be a bit tricky - if you have continuous variables, often means don’t make a lot of sense. Discuss your descriptives with your supervisor. Here for the purposes of this exercise, we’ll compute mean scores to compare well-being between the allotment and non allotment owners.

Now it’s time to actually look at the analysis. We are conducting a multiple regression, looking at the relationship of allotment ownership, hours spent in nature and well-being, and we will add the control variables of loneliness and connectedness to nature scores into our model.

Our hypotheses are:

  1. Allotment owners will have higher well-being scores in comparison to non-allotment owners.

  2. There will be a positive relationship between hours spent in nature and well-being scores.

We have a mix of continuous and categorical predictors - your final visualisations depend on the hypotheses that you are testing (although it is useful to visualise all your variables when you are exploring the data before analysis - this is good practice but you don’t need to include all your visualisations in your final results section).

Next, add a new code chunk for our descriptive statistics, calculate the cell means and SDs for allotment vs. non allotment owners, and make a visualisation with ggplot2 that shows the distribution of MH scores for these two groups. Make another visualisation to show the relationship of hours spent in nature and MH scores.

The grouped scored higher on well-being. The mean MH score for allotment owners was (SD = ). The mean MH score for non-allotment owners was (SD = ).

7.8.1 Solution

This is the code we used to compute means and suggested visualisations (you can spend some time making them more pretty and giving them more informative axis titles and labels)

descriptives <- final_scores %>% 
  group_by(D_allot) %>% 
  summarise(mean = mean(MH, na.rm=TRUE), 
            sd = sd(MH, na.rm=TRUE))

ggplot(final_scores, 
       aes(x = D_allot, y = MH, fill=D_allot)) + 
  geom_boxplot() + 
  theme_bw() + 
  scale_fill_viridis_d(option="mako", begin=0.4, end=0.65)

ggplot(final_scores, 
        aes(x = MH, y = D_hours)) + 
  geom_point() +
  geom_smooth(method="lm") + 
  theme_bw()

7.9 Inferential statistics: Multiple regression

Finally, we’ll run a multiple regression to test our hypotheses.

Run the code to conduct a multiple regression on our data and test your assumptions. We want to predict well-being (MH) scores from allotment ownership, hours spent in nature, loneliness, and connectedness to nature. You want to deviation code your allotment ownership variable and mean centre your continuous predictor variables.

Our first hypothesis of the allotment ownership being associated with higher well-being scores is .

Our second hypothesis of the positive association between well-being scores and hours spent in nature was . We should be cautious with the practical interpretation of this effect because:

The control variables of loneliness and nature connectedness show an interesting pattern. Higher loneliness scores are associated with well-being scores, and higher nature connectedness is associated with well-being scores.

The overall adjusted R squared for the model is , suggesting that the model explains % the variation in the data.

7.9.1 Solution

This is the code we used to compute the multiple regression and test the assumptions:

final_scores <- final_scores %>% 
  mutate(lon_c = Lon-mean(Lon, na.rm=TRUE), 
         nat_c = Nat-mean(Nat, na.rm=TRUE), 
         hours_c = D_hours-mean(D_hours), 
         allotment_c = ifelse(D_allot==1, 0.5, -0.5))

mod_1 <- lm(MH ~ allotment_c + hours_c + lon_c + nat_c, 
            data = final_scores)

model_summary <- tidy(summary(mod_1))

#Looks like our assumptions are fine
check_model(mod_1, 
            re_formula = NA)

7.10 Conclusion

Well done for getting through this chapter! The dataset was quite challenging and had a lot of different steps involved - hopefully, working through it has helped develop your practical skills of data analysis pipelines and how to plan your data wrangling process.