Data Manipulation
We will start by importing the data and making some updates to the variables. Then, we will turn to data summarization and visualization.
Importing the Data
Our first step is to import the data and save it as an object. The dataset is hosted on GitHub, so we will import it directly from a URL.
Fill in the blank code to import the data and call the new object cdc.
# read in CDC data and name it cdc
<- read_csv("https://haleykgrant.github.io/tutorial_data/data/cdc.samp.csv")
# check that the object exists
exists("cdc")
# read in CDC data and name it cdc
cdc <- read_csv("https://haleykgrant.github.io/tutorial_data/data/cdc.samp.csv")
# check that the object exists
exists("cdc")Data Snapshot
When you first import a dataset, it can be helpful to look at the first few rows to understand its structure.
Fill in the code to display the first 6 rows of the cdc dataset.
# view the top few rows of the cdc dataframe
head()
# view the top few rows of the cdc dataframe
head(cdc)Making A New Variable
The variable hlthplan contains information about whether an individual has health insurance. It is coded as a binary variable such that a 1 indicates that an individual has a health insurance plan and a 0 indicates that the individual does not. We may prefer to have a variable show up with more readable labels. Let’s make a new variable!
Fill in the code to update the cdc dataframe to add a new column called insured holding a factor variable that reads "Yes" if the individual has health insurance and "No" if they do not.
# update dataframe to add new `insured` column
<- cdc %>%
mutate(insured = factor(, levels = c(0,1), labels = c(,)))
# take a look to make sure it works
head(cdc)
# update dataframe to add new `insured` column
cdc <- cdc %>%
mutate(insured = factor(hlthplan, levels = c(0,1), labels = c("No","Yes")))
# take a look to make sure it works
head(cdc)That’s all for now!