Data Analysis
In the last set of exercises we imported data and made some new variables. Run the code chunk below to get started on the next set of exercises.
# read in CDC data and name it cdc
cdc <- read_csv("https://haleykgrant.github.io/tutorial_data/data/cdc.samp.csv")
# update dataframe to add new `insured` column
cdc <- cdc %>%
mutate(insured = factor(hlthplan, levels = c(0,1), labels = c("No","Yes")))
# print the column names for reference
names(cdc)
The last line of code shows the names of the columns (the variable names) from the data, which will be helpful later.
Now we are ready to start making some plots and calculating some summary statistics using our data! We will split the exercises by the number and the type of variables we are working with.
One Categorical Variable
Summary Statistics
Fill in the code to show the number of individuals who do and do not have health insurance.
# count number with and without insurance
xtabs(~ , data = cdc)
Do more people in our data have health insurance or not have health insurance?
# count number with and without insurance
xtabs(~ insured, data = cdc)insured
No Yes
10 50
There are more individuals with insurance (50) than without (10).
Plot
Now, let’s make a plot to display that information graphically.
Fill in the code to show the number of individuals who do and do not have health insurance in a plot.
# bar graph showing breakdown of insurance status
cdc %>%
ggplot(aes(x = )) +
geom_bar() +
theme_bw()
# bar graph showing breakdown of insurance status
cdc %>%
ggplot(aes(x = insured)) +
geom_bar() +
theme_bw()Adding color
Now try adding color to the plot by filling in the bars based on insurance status.
# bar graph showing breakdown of insurance status with colored bars
cdc %>%
ggplot(aes(x = insured, fill = )) +
geom_bar() +
theme_bw()
Does adding color to the plot add any additional information compared to the previous plot?
# bar graph showing breakdown of insurance status with colors bars
cdc %>%
ggplot(aes(x = insured, fill = insured)) +
geom_bar() +
theme_bw()Since we’re only plotting one variable, this doesn’t add any additional information.
One Quantitative Variable
Summary Statistics
Fill in the code to calculate some key summary statistics for age:
mean
median
standard deviation
minimum
maximum
# summary statistics for age
cdc %>%
summarise(mean = mean(),
med = median(),
sd = sd(),
min = min(),
max = max())
# summary statistics for age
cdc %>%
summarise(mean = mean(age),
med = median(age),
sd = sd(age),
min = min(age),
max = max(age))mean med sd min max
43.8 43 14.8 19 77
Plot
Now, let’s make a plot to show the full distribution of age in the data.
Fill in the code to make a histogram showing the distribution of age in the data.
# histogram showing distribution of age
cdc %>%
ggplot(aes(x = )) +
geom_histogram() +
theme_bw()
# histogram showing distribution of age
cdc %>%
ggplot(aes(x = age)) +
geom_histogram() +
theme_bw()The binwidth argument
Try a few different values of the binwidth argument.
- What happens if you make it 0.5?
- What happens if you make it 10?
# histogram showing distribution of age
cdc %>%
ggplot(aes(x = age)) +
geom_histogram(binwidth = ) +
theme_bw()
Can you describe what the binwidth argument changes in the plot?
What value would you choose if it were up to you?
# binwidth of 0.5 (small bins)
# histogram showing distribution of age
cdc %>%
ggplot(aes(x = age)) +
geom_histogram(binwidth = 0.5) +
theme_bw()
# binwidth of 10 (large bins)
cdc %>%
ggplot(aes(x = age)) +
geom_histogram(binwidth = 0.5) +
theme_bw()Binwidth changes the width (horizontal size) of the bars within a histogram. A larger value groups more individuals together (bigger intervals) and a smaller value makes narrower intervals with fewer observations. I would probably choose something like 5 or 10 in this case given the variable and the amount of age variability in the data (5-10 year age windows are probably useful and interpretable in this case).
Two Categorical Variables
Now we will turn to data summaries and visualizations for two variables.
We will start by looking at two categorical variables: health insurance and health status
Summary Statistics
Fill in the code to show breakdown of health status (genhlth) by insurance status.
# count number with and without insurance
xtabs(~ + , data = cdc)
# count number with and without insurance
xtabs(~ insured + genhlth, data = cdc)Proportions vs Counts
Remember that there were far more people with health insurance than without. To better understand the pattern here, let’s calculate the proportions of each health status within the insured and uninsured groups separately.
Fill in the code to show the proportions of each health outcome within each insurance status group.
# count number with and without insurance
xtabs(~ insured + genhlth, data = cdc) %>%
prop.table(margin = )
# count number with and without insurance
xtabs(~ insured + genhlth, data = cdc) %>%
prop.table(margin = 1)Plots
Now let’s visualize. We will start with side-by-side bar plots, then try a mosaic plot.
To start, we may want to turn the genhlth variable into a factor variable to control order of the levels in our plot. Let’s call the new variable (health).
Fill in the code to make a new factor variable called health that reorders the genhlth variable.
# make factor variable for genhlth
cdc <- cdc %>%
mutate( = factor(genhlth,
levels = c("poor", "fair", "good",
"very good", "excellent")))
Side-by-side bar plots
Fill in the code to show side-by-side bar plots with separate bars by insurance status and colors indicating health status.
# bar plot of health status by insurance
cdc %>%
ggplot(aes(x = , fill = )) +
geom_bar() +
theme_bw()
# bar plot of health status by insurance
cdc %>%
ggplot(aes(x = insured, fill = health)) +
geom_bar() +
theme_bw()Side-by-side bar plots with proportions
Notice that it is hard to compare becuase of the differnece in group sizes.
Update the code to show the proportions instead of counts.
# bar plot of health status by insurance
cdc %>%
ggplot(aes(x = insured, fill = health)) +
geom_bar(position = ) +
theme_bw()
# bar plot of health status by insurance
cdc %>%
ggplot(aes(x = insured, fill = health)) +
geom_bar(position = "fill") +
theme_bw()Mosaic plots
Now, we will make a mosaic plot, which combines the strengths of the previous two plots.
Fill in the code to create a mosaic plot showing the distribution of health status by insurance status.
# mosaic plot of health status by insurance
cdc %>%
ggplot() +
geom_mosaic(aes(x = , fill = ), alpha = 1) +
theme_bw() +
labs(title = "Health Status by Insurance",
fill = "Health Status") +
theme(axis.title.y = element_blank())
# mosaic plot of health status by insurance
cdc %>%
ggplot() +
geom_mosaic(aes(x = insured, fill = health), alpha = 1) +
theme_bw() +
labs(title = "Health Status by Insurance",
fill = "Health Status") +
theme(axis.title.y = element_blank())Ordered vs unordered variable
Try the code above using the original genhlth variable instead of the new health variable. Explain why the new variable is a better choice for plotting.
The new variable is better for plotting because the levels of the health status variable appear in a logical order.
One Categorical, One Quantitative Variable
We will now turn to summarizing data using one categorical and one quantitative variable: insurance status and age
Summary Statistics
Fill in the code calculate summary statistics for age by insurance status.
# summary statistics for age separately by insurance status
cdc %>%
summarise(mean = mean(),
med = median(),
sd = sd(),
min = min(),
max = max(), .by = )
# summary statistics for age separately by insurance status
cdc %>%
summarise(mean = mean(age),
med = median(age),
sd = sd(age),
min = min(age),
max = max(age), .by = insured)Plots
Fill in the code to make a histogram showing the distribution of age separately by insurace status.
# histogram showing distribution of by insurance status
cdc %>%
ggplot(aes(x = age, fill = )) +
geom_histogram(binwidth = 5) +
theme_bw() +
facet_wrap(~ insured, nrow = 2, scales = "free_y")
# histogram showing distribution of age by insurance status
cdc %>%
ggplot(aes(x = age, fill = insured)) +
geom_histogram(binwidth = 5) +
theme_bw() +
facet_wrap(~ insured, nrow = 2, scales = "free_y")Changing the y-axis
Try changing the scales from “free_y” to “fixed”. How does this impact the plot?
# histogram showing distribution of age by insurance status
cdc %>%
ggplot(aes(x = age, fill = insured)) +
geom_histogram(binwidth = 5) +
theme_bw() +
facet_wrap(~ insured, nrow = 2, scales = )
Two Quantitative Variables
Finally, we turn to describing the relationship between two quantitative variables: height and weight.
Summary Statistics
Fill in the code to calculate the Pearson correlation between height and weight.
# Pearson correlation between height and weight
cor(cdc$, cdc$, use = "complete.obs")
# correlation between height and weight
cor(cdc$height, cdc$weight)Spearman Correlation
Try changing from Pearson correlation (the default) to Spearman correlation. Are they similar? Different?
# Spearman correlation between height and weight
cor(cdc$height, cdc$weight, method = "", use = "complete.obs")
Try to describe what you think the scatterplot will look like for height and weight.
Plots
Fill in the code to make a scatter plot showing the relationship between height and weight
# scatter plot of height and weight
cdc %>%
ggplot(aes(x = , y = )) +
geom_point() +
theme_bw() +
labs(x = "Height (in)",
y = "Weight (lbs)")
# scatter plot of height and weight
cdc %>%
ggplot(aes(x = height, y = weight)) +
geom_point() +
theme_bw() +
labs(x = "Height (in)",
y = "Weight (lbs)")