Module 1: Data Wrangling

Fill-in-the-blank coding exercises for Module 1: Data Wrangling
Author

Haley Grant

Data Manipulation

We will start by importing the data and making some updates to the variables. Then, we will turn to data summarization and visualization.

Importing the Data

Our first step is to import the data and save it as an object. The dataset is hosted on GitHub, so we will import it directly from a URL.

Fill in the blank code to import the data and call the new object cdc.

# read in CDC data and name it cdc
 <- read_csv("https://haleykgrant.github.io/tutorial_data/data/cdc.samp.csv")

# check that the object exists
exists("cdc")

If you have filled in the code correctly, the output should read:

[1] TRUE

# read in CDC data and name it cdc
cdc <- read_csv("https://haleykgrant.github.io/tutorial_data/data/cdc.samp.csv")

# check that the object exists
exists("cdc")

Data Snapshot

When you first import a dataset, it can be helpful to look at the first few rows to understand its structure.

Fill in the code to display the first 6 rows of the cdc dataset.

# view the top few rows of the cdc dataframe 
head()
# view the top few rows of the cdc dataframe 
head(cdc)

Making A New Variable

The variable hlthplan contains information about whether an individual has health insurance. It is coded as a binary variable such that a 1 indicates that an individual has a health insurance plan and a 0 indicates that the individual does not. We may prefer to have a variable show up with more readable labels. Let’s make a new variable!

Fill in the code to update the cdc dataframe to add a new column called insured holding a factor variable that reads "Yes" if the individual has health insurance and "No" if they do not.

# update dataframe to add new `insured` column
 <- cdc %>%
  mutate(insured = factor(, levels = c(0,1), labels = c(,)))
  
# take a look to make sure it works
head(cdc)

If you have filled in the code correctly, you should have a new column at the far right end of the data snapshot output showing a column with name “Insured” and either “Yes” or “No” as values.

To update the cdc. object and save the results, we are replacing the previous object called cdc with a new object also called cdc.

# update dataframe to add new `insured` column
cdc <- cdc %>%
  mutate(insured = factor(hlthplan, levels = c(0,1), labels = c("No","Yes")))
  
# take a look to make sure it works
head(cdc)

That’s all for now!