Sampling Distributions
Recall that if we have a population with some true underlying parameter (like a mean, \(\mu\)), we use a sample statistic (like a sample mean, \(\bar{x}\)) to estimate it. Because our sample is just one of many possible samples we could have drawn, our statistic will vary from sample to sample. The sampling distribution of a statistic describes what values that statistic would take on across all possible samples of a given size, \(n\), drawn from the same population.
In this section, we’ll start from a simulated population, then repeatedly sample from it to see visualize the sampling distribution of the mean.
Our running example: length of hospital stay (LOS), in days, following a particular surgery. LOS tends to be right-skewed—most patients go home quickly, but a few stay much longer—so it’s a nice example for testing whether the Central Limit Theorem (CLT) really works even when the underlying population isn’t normal.
Simulating a Population
To explore sampling distributions, we first need a population to sample from. We’ve simulated one for you below: 100,000 patients’ length of stay, drawn from an exponential distribution with a mean of 4.5 days. The results are saved in a data frame called pop, with a column called los. (You don’t need to worry about the code that generated it—just know that we now have a full population to sample from.)
Fill in the code below to plot a histogram of the population distribution, pop$los, to see its shape.
# plot the population distribution
pop %>%
ggplot(aes(x = )) +
geom_histogram(fill = "blue", bins = 40) +
theme_bw()
# plot the population distribution
pop %>%
ggplot(aes(x = los)) +
geom_histogram(fill = "blue", bins = 40) +
theme_bw()You should see a strongly right-skewed distribution: a tall bar near 0, with a long tail stretching out to the right toward higher values. This is clearly not a normal distribution.
We can check the population’s mean and standard deviation directly:
# mean and sd of the population
mean(pop$los)
sd(pop$los)
What is the mean value of los in the population? What is the standard deviaiton?
Drawing a Single Sample
In practice, we don’t get to see the whole population—we only get to observe one sample. Let’s draw a single sample of patients and calculate the sample mean.
Draw one random sample of 30 patients (without replacement) from the population vector pop$los, and calculate the mean of that sample.
# set random seed for reproducibility
set.seed(1)
# draw one sample of size 30
one_sample = sample(pop$los, size = , replace = FALSE)
# calculate the sample mean
mean(one_sample)
# set random seed for reproducibility
set.seed(1)
# draw one sample of size 30
one_sample = sample(pop$los, size = 30, replace = FALSE)
# calculate the sample mean
mean(one_sample)Your sample mean is a single point estimate of the true population mean (4.5). If you re-ran this code with a different seed, you’d get a different sample, and therefore a different sample mean. That sample-to-sample variability is exactly what the sampling distribution describes.
Building a Sampling Distribution
To see the sampling distribution of the mean directly, we need to repeat the sampling process many times—drawing a sample, calculating its mean, and recording that mean—over and over. Let’s do this for samples of size n = 5 first.
Fill in the code below to repeat the sampling process 10,000 times, each time drawing a sample of size 5 from the population and recording the sample mean.
# set random seed for reproducibility
set.seed(2)
# number of repeated samples to draw
niter = 10000
# create a place to store the sample means
samp_means_5 = data.frame(mean_los = rep(NA, 10000))
for(i in 1:niter){
# draw one sample of size 5
samp = sample(pop$los, size = , replace = FALSE)
# store the sample mean
samp_means_5$mean_los[i] = mean(samp)
}
# plot the sampling distribution
samp_means_5 %>%
ggplot(aes(x = mean_los)) +
geom_histogram(fill = "orange", bins = 40) +
theme_bw()
The plotted distribution should still look somewhat right-skewed, though a bit less extreme than the population itself. With such a small sample size (n = 5), each sample mean is still heavily influenced by any large values that happen to get sampled.
Now let’s repeat the same process, but with a much larger sample size, n = 100.
Repeat the process above, but this time draw samples of size 100 instead of size 5.
# set random seed for reproducibility
set.seed(2)
# number of repeated samples to draw
niter = 10000
# create a place to store the sample means
samp_means_100 = data.frame(mean_los = rep(NA, niter))
for(i in 1:niter){
# draw one sample of size 100
samp = sample(pop$los, size = , replace = FALSE)
# store the sample mean
samp_means_100$mean_los[i] = mean(samp)
}
# plot the sampling distribution
samp_means_100 %>%
ggplot(aes(x = mean_los)) +
geom_histogram(fill = "orange", bins = 40) +
theme_bw()
Even though the population is strongly right-skewed, the sampling distribution of the mean for n = 100 should look approximately symmetric and bell-shaped—much closer to a normal distribution. This is the Central Limit Theorem in action: as the sample size grows, the sampling distribution of the sample mean approaches a normal distribution, regardless of the shape of the underlying population (as long as \(n\) is reasonably large).
Compare the spread (width) of the two sampling distributions you just created (n = 5 vs. n = 100). Which one is more spread out? Does this match what you’d expect based on the sample sizes used?
Standard Error
The standard deviation of a sampling distribution has a special name: the standard error (SE). The CLT tells us that if \(X\) has population mean \(\mu\) and standard deviation \(\sigma\), then for a sufficiently large sample size \(n\):
\[\bar{X} \sim Normal\left(\mu, \ \dfrac{\sigma}{\sqrt{n}}\right)\]
That is, the theoretical standard error is \(\sigma / \sqrt{n}\).
Calculate the empirical standard error from your simulated sampling distribution for n = 100 (i.e., sd() of samp_means_100$mean_los). Then calculate the theoretical standard error using the formula above, assuming the true population standard deviation is 4.5 (recall: for the exponential distribution, sd = mean) and n = 100.
# empirical standard error (from our simulation)
sd(samp_means_100$mean_los)
# theoretical standard error
/ sqrt()
# theoretical standard error
4.5 / sqrt(100)[1] 0.45
Your empirical standard error (from sd(samp_means_100$mean_los)) should be close to 0.45, though it won’t match exactly since it’s based on a finite number of simulated samples. Notice that as \(n\) increases, the standard error decreases—larger samples give us more precise (less variable) estimates of the population mean.
Applying the CLT to Calculate Probabilities
Once we know (or can assume) that a sampling distribution is approximately normal, we can use everything we learned about the normal distribution to calculate probabilities about sample means—without needing to simulate anything!
Suppose we plan to take a random sample of 100 patients. Assuming the population mean LOS is 4.5 days with a standard deviation of 4.5 days, use the CLT to find the probability that the sample mean LOS is greater than 5 days. Hint: first calculate the standard error, then use pnorm().
# standard error for n = 100
se = 4.5 / sqrt()
# probability the sample mean exceeds 5 days
pnorm(q = , mean = , sd = se, lower.tail = )
# standard error for n = 100
se = 4.5 / sqrt(100)
# probability the sample mean exceeds 5 days
pnorm(q = 5, mean = 4.5, sd = se, lower.tail = F)[1] 0.1332603
There is about a 13.3% chance that a random sample of 100 patients has a mean LOS greater than 5 days.
Now repeat the same calculation, but assume we only sampled 30 patients instead of 100. Is the probability higher, lower, or the same compared to before?
# standard error for n = 30
se = 4.5 / sqrt()
# probability the sample mean exceeds 5 days
pnorm(q = , mean = , sd = se, lower.tail = )
# standard error for n = 30
se = 4.5 / sqrt(30)
# probability the sample mean exceeds 5 days
pnorm(q = 5, mean = 4.5, sd = se, lower.tail = F)[1] 0.2714012
With a smaller sample size (n = 30), the standard error is larger, so sample means are more variable. This makes it more likely (about 27.1% vs. 13.3%) to see a sample mean as far from 4.5 as 5 days, just by chance.
We used the CLT to calculate these probabilities even though our population (length of stay) is not normally distributed. Based on what you saw in the “Building a Sampling Distribution” section, explain why this is still valid for n = 100 but might be questionable for a much smaller sample size, like n = 5. What if I wanted to calculate the probability that one individual length of stay was above some value? Could I use the CLT then?
Notice that if the underlying population itself were normally distributed, we would not need a large sample size for the CLT to apply—the sampling distribution of the mean would be exactly normal for any sample size, even n = 2 or n = 3. The “large enough n” condition is only needed to overcome skewness or other non-normal features of the population.