Chapter 6: The Normal Distribution
Shih Chien University
2026-10-08
Chapter 6 introduces the normal distribution and its role in statistical inference.
| Section | Topics |
|---|---|
| 6-1 | Normal and standard normal distributions; areas under the curve |
| 6-2 | Applications; data values from given probabilities; normality |
| 6-3 | Central limit theorem; sample means; finite population correction |
| 6-4 | Normal approximation to the binomial distribution |
After completing this chapter, you should be able to
Random variables can be discrete or continuous. Chapter 5 dealt with discrete variables; a continuous variable can assume all values between any two given values — heights of adult men, body temperatures of rats, cholesterol levels of adults.
Many continuous variables have bell-shaped distributions and are called approximately normally distributed variables. If a researcher measures the heights of 100 adult women and draws a histogram, then keeps increasing the sample size while decreasing the class width, the histogram approaches a smooth curve called a normal distribution curve (Figure 6-1). It is also called a bell curve or a Gaussian distribution curve, after Carl Friedrich Gauss (1777–1855).
No variable fits a normal distribution perfectly, because a normal distribution is a theoretical distribution. It is nevertheless useful because the deviations from it are very small.
Figure 6-1 Histograms and normal model for the distribution of heights of adult women
Normal Distribution
If a random variable has a probability distribution whose graph is continuous, bell-shaped, and symmetric, it is called a normal distribution. The graph is called a normal distribution curve.
The mathematical equation for a normal distribution is
y = \frac{e^{-(X-\mu)^2/(2\sigma^2)}}{\sigma\sqrt{2\pi}}
where e \approx 2.718, \pi \approx 3.14, \mu is the population mean and \sigma is the population standard deviation. In applied statistics, tables or technology are used instead of the equation.
The shape and position of a normal distribution curve depend on two parameters: the mean and the standard deviation (Figure 6-3).
x <- seq(-8, 10, by = 0.01)
curves <- data.frame(x, mu0_sd1 = dnorm(x, 0, 1), mu0_sd2 = dnorm(x, 0, 2),
mu2_sd2 = dnorm(x, 2, 2))
ggplot(curves, aes(x)) +
geom_line(aes(y = mu0_sd1)) +
geom_line(aes(y = mu0_sd2), linetype = "dashed") +
geom_line(aes(y = mu2_sd2), linetype = "dotted") +
labs(title = "Figure 6-3: shapes of normal distributions", x = "x", y = "Density")The solid line has \mu = 0, \sigma = 1, the dashed line \mu = 0, \sigma = 2, and the dotted line \mu = 2, \sigma = 2. Solid and dashed have the same mean; dashed and dotted have the same standard deviation.
Summary of the Properties of the Theoretical Normal Distribution
Figure 6-4 — the areas in item 8 follow the empirical rule of Section 3-2.
Symmetric and Skewed Distributions
The “tail” of the curve indicates the direction of the skewness (right is positive, left is negative).
Figure 6-5 Normal and skewed distributions
x <- seq(0, 10, by = 0.1)
shapes <- data.frame(
shape = rep(c("(a) Normal", "(b) Negatively skewed", "(c) Positively skewed"),
each = length(x)),
x = c(x, x, x),
density = c(dnorm(x, 5, 1.5), dgamma(10 - x, shape = 3), dgamma(x, shape = 3))
)
ggplot(shapes, aes(x, density)) +
geom_line() +
facet_wrap(~ shape) +
labs(y = "Density")Each normally distributed variable has its own mean and standard deviation, so a separate table of areas would be needed for every variable. To simplify this, statisticians use the standard normal distribution.
Standard Normal Distribution
The standard normal distribution is a normal distribution with a mean of 0 and a standard deviation of 1.
The equation for the standard normal distribution is
y = \frac{e^{-z^2/2}}{\sqrt{2\pi}}
Formula for the z Score (Standard Score)
z = \frac{\text{value} - \text{mean}}{\text{standard deviation}} \qquad\text{or}\qquad z = \frac{X - \mu}{\sigma}
Figure 6-6 Standard normal distribution
A two-step process is recommended:
Step 1 Draw the normal distribution curve and shade the area.
Step 2 Find the appropriate figure in the Procedure Table and follow the directions given.
Procedure Table — Finding the Area Under the Standard Normal Distribution Curve
There are three basic types of problems, and all three are summarized in the Procedure Table.
This course uses R instead of Table A-4. The three cases of the Procedure Table become three R expressions.
Key R Functions
pnorm(z) — area to the left of z1 - pnorm(z) — area to the right of zpnorm(z2) - pnorm(z1) — area between z_1 and z_2qnorm(p) — the z value whose left-tail area is p (the inverse of pnorm)Area to the Left of a z Value
Find the area under the standard normal distribution curve to the left of z = 1.72.
Area to the Right of a z Value
Find the area under the standard normal distribution curve to the right of z = -0.83.
Solution
Area Between Two z Values
Find the area under the standard normal distribution curve between z = 1.21 and z = -1.54.
Solution
Step 1 Draw the figure (Figure 6-10). The desired area lies between the two z values.
Step 2 This is the third case: look up both areas and subtract the smaller from the larger. (Do not subtract the z values.) The area for z = 1.21 is 0.8869 and the area for z = -1.54 is 0.0618, so the area between them is 0.8869 - 0.0618 = 0.8251, or 82.51\%.
The area under the standard normal distribution curve can also be thought of as a probability, or as the proportion of the population with a given characteristic. If it were possible to select a z value at random, the probability of choosing one between 0 and 2.00 would be the same as the area under the curve between 0 and 2.00, namely 0.4772.
For probabilities, a special notation is used: the probability of a z value between 0 and 2.32 is written P(0 < z < 2.32).
Caution
In a continuous distribution the probability of any exact z value is 0, since the area would be a vertical line, which in theory has no area. Hence P(a \leq z \leq b) = P(a < z < b).
Probabilities for the Standard Normal Distribution
Find the probability for each. (Assume this is a standard normal distribution.)
(a) P(0 < z < 2.14) (b) P(z < 1.35) (c) P(z > 2.27)
Solution
(a) Find the area between z = 0 and z = 2.14. The area for z = 2.14 is 0.9838 and the area for z = 0 is 0.5000: 0.9838 - 0.5000 = 0.4838, or 48.38\%.
(b) Find the area to the left of z = 1.35. It is 0.9115, so the probability is 0.9115, or 91.15\%.
(c) Find the area to the right of z = 2.27. The area for z = 2.27 is 0.9884, so 1.0000 - 0.9884 = 0.0116, or 1.16\%.
Finding a z Value for a Given Area
Find the z value such that the area under the standard normal distribution curve between 0 and the z value is 0.3078.
Solution
Draw the figure (Figure 6-14). Since Table A-4 is cumulative, add 0.5000 to the given area of 0.3078 to get the cumulative area 0.8078. Looking up 0.8078 in Table A-4 gives a row value of 0.8 and a column value of 0.07, so z = 0.87.
If the exact area cannot be found in the table, use the closest value: for a cumulative area of 0.8080 the closest table area is 0.8078, which again gives z = 0.87.
The standard normal distribution curve can be used to solve a wide variety of practical problems. The only requirement is that the variable be normally or approximately normally distributed. To solve such problems, transform the original variable to a standard normal variable:
z = \frac{\text{value} - \text{mean}}{\text{standard deviation}} \qquad\text{or}\qquad z = \frac{X - \mu}{\sigma}
Procedure Table — Finding the Area Under Any Normal Curve
Step 1 Draw a normal curve and shade the desired area.
Step 2 Convert the values of X to z values, using the formula z = \dfrac{X - \mu}{\sigma}.
Step 3 Find the corresponding area, using a table, calculator, or software.
In R the conversion is optional: pnorm(x, mean = mu, sd = sigma) works directly.
Food Delivery Times
On a food-delivery platform operating in Taipei, the time from placing an order to delivery at the door averages 27 minutes. Assume the variable is normally distributed and has a standard deviation of 4 minutes. Find the percentage of deliveries that arrive in less than 32 minutes. (hypothetical data)
Solution
Step 1 Draw a normal curve and shade the desired area (Figure 6-18).
Step 2 z = \dfrac{X - \mu}{\sigma} = \dfrac{32 - 27}{4} = \dfrac{5}{4} = 1.25
Step 3 The area to the left of z = 1.25 is 0.8944. Therefore 89.44\% of deliveries arrive in less than 32 minutes.
Order Values on an E-Commerce Platform
Completed orders on a Taiwanese e-commerce platform have an average value of 1{,}200 NT dollars. Assume the variable is approximately normally distributed and the standard deviation is 150 NT dollars. If one completed order is selected at random, find the probability that its value is
(a) between 1{,}050 and 1{,}425 NT dollars; (b) more than 1{,}380 NT dollars. (hypothetical data)
Solution (a)
Step 1 Draw a normal curve and shade the desired area (Figure 6-20).
Step 2 z_1 = \dfrac{1050 - 1200}{150} = -1.0 and z_2 = \dfrac{1425 - 1200}{150} = 1.5
Step 3 The area to the left of z_2 is 0.9332 and the area to the left of z_1 is 0.1587, so the area between them is 0.9332 - 0.1587 = 0.7745. The probability is 77.45\%.
Solution (b)
Step 1 Draw a normal curve and shade the desired area (Figure 6-22).
Step 2 z = \dfrac{1380 - 1200}{150} = \dfrac{180}{150} = 1.2
Step 3 The area to the left of z = 1.2 is 0.8849, so the area to the right is 1.0000 - 0.8849 = 0.1151. The probability that an order is worth more than 1{,}380 NT dollars is 0.1151, or 11.51\%.
Solution
Order Picking Time
At a logistics centre in Taoyuan, the time needed to pick and pack one order averages 6.4 minutes. Assume the variable is approximately normally distributed and the population standard deviation is 0.8 minute. If 60 orders are randomly selected, approximately how many will be picked and packed in less than 5.8 minutes? (hypothetical data)
Solution
Step 1 Draw a normal curve and shade the desired area (Figure 6-24).
Step 2 z = \dfrac{5.8 - 6.4}{0.8} = -0.75
Step 3 The area corresponding to z = -0.75 is 0.2266.
To find the number of orders, multiply 60 \times 0.2266 = 13.6 and round to 14. Approximately 14 of the 60 randomly selected orders are picked and packed in less than 5.8 minutes.
Note: for problems using percentages, write the percentage as a decimal before multiplying, and round the answer to the nearest whole number.
A normal distribution can also be used to find specific data values for given percentages. Starting from z = \dfrac{X-\mu}{\sigma} and solving for X gives:
Formula for Finding the Value of a Normal Variable X
X = z \cdot \sigma + \mu
Procedure Table — Finding Data Values for Specific Probabilities
Step 1 Draw a normal curve and shade the desired area that represents the probability, proportion, or percentile.
Step 2 Find the z value from the table that corresponds to the desired area.
Step 3 Calculate the X value by using the formula X = z\sigma + \mu.
Qualifying for a Rider Bonus
A delivery platform pays a quarterly bonus to the riders in the top 10\% by number of orders completed per month. Assume the number of orders completed is normally distributed with a mean of 480 and a standard deviation of 60. Find the lowest number of completed orders needed to qualify. (hypothetical data)
Solution
Step 1 Work backward. Subtract 0.1000 from 1.0000 to get the area to the left of X: 1.0000 - 0.1000 = 0.9000 (Figure 6-26).
Step 2 Look up 0.9000 in the area portion of Table A-4. The table rounds the answer to z = 1.28; computed exactly, z = 1.2816.
Step 3 X = z \cdot \sigma + \mu = 1.2816(60) + 480 = 556.89.
Because a rider cannot complete a fraction of an order, 557 completed orders should be used as the cutoff: anybody completing 557 or more qualifies for the bonus.
Daily Passengers at an MRT Station
A transport planner wishes to study the middle 60\% of weekdays at a Taipei MRT station. Assume the number of passengers entering the station on a weekday is normally distributed with a mean of 48 thousand and a standard deviation of 5 thousand.
(a) What is the passenger count for the lowest 20\% of weekdays?
(b) Find the lower and upper passenger counts for the middle 60\% of weekdays. (hypothetical data)
Solution (a)
Step 1 Draw the curve and shade the lower 20\% (Figure 6-28).
Step 2 The z value with area 0.2000 to its left is z = -0.8416 (Table A-4 rounds it to -0.84).
Step 3 X = z\sigma + \mu = (-0.8416)(5) + 48 = 43.79
The quietest 20\% of weekdays have fewer than 43.79 thousand passengers.
Solution (b)
Step 1 Two values are needed, one above and one below the mean (Figure 6-29).
Step 2 To get the area to the left of the positive z value, add 0.5000 + 0.3000 = 0.8000. The z value for a left-tail area of 0.8000 is 0.8416, and the z value for 0.2000 is -0.8416.
Step 3 X_1 = (-0.8416)(5) + 48 = 43.79 and X_2 = (0.8416)(5) + 48 = 52.21.
The middle 60\% of weekdays have between 43.79 and 52.21 thousand passengers.
A bell-shaped distribution is only one of many shapes a distribution can assume, but it is very important since many statistical methods require the distribution of values to be normally or approximately normally shaped. Statisticians check normality in several ways.
Checking a Data Set for Normality
Supplier Lead Times
A trading company recorded the number of days of finished-goods inventory held by each of 18 supplier factories. Determine if the data are approximately normally distributed. (hypothetical data)
12 21 26 33 35 38 42 45 47
49 52 56 59 63 68 74 81 93
Solution
Step 1 Construct a frequency distribution and draw a histogram (Figure 6-30).
Solution
Step 2 Check for skewness: \bar{X} = 49.67, median = 48, s = 21.17, so PC = \frac{3(49.67 - 48)}{21.17} = 0.236
[1] 49.66667
[1] 48
[1] 21.16601
[1] 0.2362278
PC is not greater than +1 nor less than -1, so the distribution is not significantly skewed.
Solution
Step 3 Check for outliers: Q_1 = 35.75 and Q_3 = 62, so IQR = 26.25. An outlier would be below 35.75 - 1.5(26.25) = -3.62 or above 62 + 1.5(26.25) = 101.38. There are no outliers.
25% 75%
35.75 62.00
[1] 26.25
[1] -3.625
[1] 101.375
Bell-shaped, not significantly skewed, no outliers: the distribution is approximately normally distributed.
Customers Served at a Night-Market Stall
The data shown are the numbers of customers served by a night-market stall on 17 randomly selected evenings. Determine if the data are approximately normally distributed. (hypothetical data)
142 95 131 152 108 145
149 25 139 118 151 135
125 60 147 78 42
Solution
Step 1 Construct a frequency distribution and draw a histogram (Figure 6-31).
(24.5,49.5] (49.5,74.5] (74.5,99.5] (99.5,124.5] (124.5,149.5]
2 1 2 2 8
(149.5,174.5]
2
Solution
Step 2 Check for skewness: \bar{X} = 114.24, median = 131, s = 40.37, so PC = \frac{3(114.24 - 131)}{40.37} = -1.25
[1] 114.2353
[1] 131
[1] 40.37098
[1] -1.245799
Since PC is less than -1, the distribution is significantly skewed to the left.
Solution
Step 3 Check for outliers: Q_1 = 95, Q_3 = 145, IQR = 50. Any value below 95 - 1.5(50) = 20 or above 145 + 1.5(50) = 220 is an outlier. There are no outliers.
25% 75%
95 145
[1] 50
[1] 20
[1] 220
In summary, the distribution is significantly negatively skewed.
Another method used to check normality is to draw a normal quantile plot. Quantiles (sometimes called fractiles) are values that separate the data set into approximately equal groups. A normal quantile plot graphs the data values against the z values of the corresponding quantiles. If the points do not lie in an approximately straight line, normality can be rejected.
Other methods include normal probability graph paper, the chi-square goodness-of-fit test (Chapter 11), the Kolmogorov-Smirnov test and the Lilliefors test.
In addition to knowing how individual data values vary about the mean for a population, statisticians want to know how the means of samples of the same size taken from the same population vary about the population mean.
Suppose an analyst selects a sample of 30 online orders and finds the mean order value to be 1{,}180 NT dollars, then selects a second sample with mean 1{,}215 NT dollars, and continues for 100 samples. The mean becomes a random variable, and the sample means 1{,}180, 1{,}215, 1{,}164, \ldots, 1{,}232 constitute a sampling distribution of sample means.
Sampling Distribution of Sample Means
A sampling distribution of sample means is a distribution using the means computed from all possible random samples of a specific size taken from a population.
If the samples are randomly selected with replacement, the sample means will for the most part be somewhat different from the population mean \mu. These differences are caused by sampling error.
Sampling Error
Sampling error is the difference between the sample measure and the corresponding population measure, due to the fact that the sample is not a perfect representation of the population.
Properties of the Distribution of Sample Means
A lecturer gave a 10-point quiz to a small class of four students; the results were 3, 5, 7 and 9. Treat the four students as the population.
[1] 6
[1] 2.236068
The graph of the original distribution is a uniform distribution (Figure 6-32): each score occurs once.
All samples of size 2 taken with replacement, and the mean of each sample:
[1] 6
[1] 1.581139
[1] 1.581139
The histogram of the sample means appears approximately normal, the mean of the sample means equals the population mean, and the standard deviation of the sample means equals \sigma/\sqrt{n}.
Standard Error of the Mean
If all possible samples of size n are taken with replacement from the same population, the mean of the sample means, denoted \mu_{\bar{X}}, equals the population mean \mu; and the standard deviation of the sample means, denoted \sigma_{\bar{X}}, equals \sigma/\sqrt{n}. The standard deviation of the sample means is called the standard error of the mean.
Mean and Standard Error of the Sample Means
\mu_{\bar{X}} = \mu \qquad\qquad \sigma_{\bar{X}} = \frac{\sigma}{\sqrt{n}}
A third property of the sampling distribution of sample means concerns the shape of the distribution, and is explained by the central limit theorem.
The Central Limit Theorem
As the sample size n increases without limit, the shape of the distribution of the sample means taken with replacement from a population with mean \mu and standard deviation \sigma will approach a normal distribution. This distribution will have a mean \mu and a standard deviation \sigma/\sqrt{n}.
z Value for the Central Limit Theorem
z = \frac{\bar{X} - \mu}{\sigma/\sqrt{n}}
The denominator must be adjusted because means are used instead of individual data values; it is the standard deviation of the sample means.
When the Central Limit Theorem Applies
In R: use pnorm(xbar, mean = mu, sd = sigma / sqrt(n)).
set.seed(42)
# rate = 0.5 gives a skewed population
simulation <- data.frame(
size = rep(c("(a) n = 2", "(b) n = 10", "(c) n = 30"), each = 2000),
sample_mean = c(replicate(2000, mean(rexp(2, rate = 0.5))),
replicate(2000, mean(rexp(10, rate = 0.5))),
replicate(2000, mean(rexp(30, rate = 0.5))))
)
ggplot(simulation, aes(sample_mean)) +
geom_histogram(bins = 40, color = "white") +
facet_wrap(~ size, scales = "free") +
labs(title = "Sample means approach normality as n increases",
x = "Sample mean", y = "Frequency")Length of Customer-Service Calls
Calls handled by the customer-service centre of an online retailer last an average of 4.5 minutes. Assume the variable is normally distributed and the standard deviation is 1.25 minutes. If 25 calls are randomly selected, find the probability that their mean length will be greater than 5.0 minutes. (hypothetical data)
Solution
Since the variable is approximately normally distributed, the distribution of sample means is approximately normal with a mean of 4.5 and standard error
\sigma_{\bar{X}} = \frac{\sigma}{\sqrt{n}} = \frac{1.25}{\sqrt{25}} = 0.25
Step 1 Draw a normal curve and shade the desired area (Figure 6-35).
Step 2 z = \dfrac{\bar{X} - \mu}{\sigma/\sqrt{n}} = \dfrac{5.0 - 4.5}{0.25} = 2.00
Step 3 The area to the right of 2.00 is 1.0000 - 0.9772 = 0.0228, or 2.28\%.
The probability that the 25 calls last more than 5.0 minutes on average is 2.28%.
Parcels Delivered per Van
A courier company’s vans deliver an average of 120 parcels per day. Assume the standard deviation is 18 parcels. If a random sample of 36 van-days is selected, find the probability that the sample mean is between 114 and 123 parcels. (hypothetical data)
Solution
Step 1 Draw a normal curve and shade the desired area (Figure 6-36). Since the sample is 30 or more, the normality assumption is not necessary.
Step 2 The standard error is 18/\sqrt{36} = 3, so z_1 = \dfrac{114 - 120}{3} = -2 and z_2 = \dfrac{123 - 120}{3} = 1.
Step 3 The area for z = -2 is 0.0228 and the area for z = 1 is 0.8413, so the area between them is 0.8413 - 0.0228 = 0.8186, or 81.86\%.
Customs Clearance Time
Imported shipments arriving at a Taiwanese port clear customs in an average of 3.2 days. Assume the distribution is approximately normal and has a standard deviation of 0.5 day.
(a) Find the probability that one randomly selected shipment clears customs in fewer than 3.3 days.
(b) If a sample of 36 shipments is randomly selected, find the probability that the mean clearance time of the sample is less than 3.3 days. (hypothetical data)
Solution (a)
Since the question concerns an individual shipment, use z = \dfrac{X - \mu}{\sigma} (Figure 6-37).
Step 2 z = \dfrac{3.3 - 3.2}{0.5} = 0.20
Step 3 The area to the left of z = 0.20 is 0.5793.
The probability of selecting a shipment that clears customs in fewer than 3.3 days is 0.5793, or 57.93\%.
Solution (b)
Since the question concerns the mean of a sample of size 36, use the central limit theorem formula z = \dfrac{\bar{X} - \mu}{\sigma/\sqrt{n}} (Figure 6-38).
Step 2 z = \dfrac{3.3 - 3.2}{0.5/\sqrt{36}} = 1.20
Step 3 The area corresponding to z = 1.20 is 0.8849, or 88.49\%.
The gap of about 30.6 percentage points between (a) and (b) arises because the distribution of sample means is much less variable than the distribution of individual data values: as the sample size increases, the standard deviation of the means decreases.
Table 6-1 Summary of Formulas and Their Uses
| Formula | Use |
|---|---|
| z = \dfrac{X - \mu}{\sigma} | Used to gain information about an individual data value when the variable is normally distributed |
| z = \dfrac{\bar{X} - \mu}{\sigma/\sqrt{n}} | Used to gain information, when applying the central limit theorem, about a sample mean when the variable is normally distributed or when the sample size is 30 or more |
Notice that the first formula contains X, an individual data value, while the second contains \bar{X}, the sample mean.
The standard error \sigma/\sqrt{n} is exact for sampling with replacement, or from a very large (infinite) population. Sampling without replacement from a finite population needs a correction factor.
Finite Population Correction Factor
\sqrt{\frac{N-n}{N-1}}
where N = population size and n = sample size. The standard error and z become
\sigma_{\bar{X}} = \frac{\sigma}{\sqrt{n}} \cdot \sqrt{\frac{N-n}{N-1}} \qquad\qquad z = \frac{\bar{X} - \mu}{\dfrac{\sigma}{\sqrt{n}} \cdot \sqrt{\dfrac{N-n}{N-1}}}
Caution
It is needed when the sample is relatively large — usually more than 5\% of the population. When the population is large and the sample small, the factor is very close to 1.00 and is generally omitted.
When n is large (say, 100), binomial calculations are too difficult by hand, so a normal distribution is used instead. Recall from Chapter 5 that a binomial distribution has these characteristics:
As n increases with p near 0.5, the binomial shape becomes similar to a normal distribution. But when p is close to 0 or 1 and n is relatively small, a normal approximation is inaccurate.
Rule of Thumb
A normal approximation should be used only when n \cdot p and n \cdot q are both greater than or equal to 5, where q = 1 - p.
Also, \mu = n \cdot p and \sigma = \sqrt{n \cdot p \cdot q}.
For example, if p = 0.3 and n = 10, then np = 3, so a normal distribution should not be used; if p = 0.5 and n = 10, then np = nq = 5 and it can be used.
binomial_shapes <- data.frame(
label = rep(c("n = 10, p = 0.3 (np = 3)", "n = 10, p = 0.5 (np = 5, nq = 5)"),
each = 11),
x = c(0:10, 0:10),
probability = c(dbinom(0:10, 10, 0.3), dbinom(0:10, 10, 0.5))
)
ggplot(binomial_shapes, aes(x, probability)) +
geom_col() +
facet_wrap(~ label) +
labs(title = "Figure 6-39: binomial distribution and a normal distribution",
x = "X", y = "P(X)")Correction for Continuity
A correction for continuity is a correction employed when a continuous distribution is used to approximate a discrete distribution.
The continuity correction means that for any specific value of X, say 8, the boundaries of X in the binomial distribution (in this case 7.5 to 8.5) must be used. For P(X = 8) the correction is P(7.5 < X < 8.5); for P(X \leq 7) it is P(X < 7.5); for P(X \geq 3) it is P(X > 2.5).
Table 6-2 Summary of the Normal Approximation to the Binomial Distribution
| When finding (binomial) | Use (normal) |
|---|---|
| 1. P(X = a) | P(a - 0.5 < X < a + 0.5) |
| 2. P(X \geq a) | P(X > a - 0.5) |
| 3. P(X > a) | P(X > a + 0.5) |
| 4. P(X \leq a) | P(X < a + 0.5) |
| 5. P(X < a) | P(X < a - 0.5) |
For all cases, \mu = n \cdot p, \sigma = \sqrt{n \cdot p \cdot q}, n \cdot p \geq 5 and n \cdot q \geq 5.
Procedure for the Normal Approximation to the Binomial Distribution
Step 1 Check to see whether the normal approximation can be used.
Step 2 Find the mean \mu and the standard deviation \sigma.
Step 3 Write the problem in probability notation, using X.
Step 4 Rewrite the problem by using the continuity correction factor, and show the corresponding area under the normal distribution.
Step 5 Find the corresponding z values.
Step 6 Find the solution.
Convenience-Store Pickup
On a cross-border e-commerce platform, 30\% of buyers choose convenience-store pickup instead of home delivery. If 150 buyers are selected at random, find the probability that exactly 52 of them choose convenience-store pickup. (hypothetical data)
Solution
Here p = 0.30, q = 0.70 and n = 150.
Step 1 np = (150)(0.30) = 45 and nq = (150)(0.70) = 105. Since np \geq 5 and nq \geq 5, the normal distribution can be used.
Step 2 \mu = np = 45 and \sigma = \sqrt{npq} = \sqrt{(150)(0.30)(0.70)} = \sqrt{31.5} \approx 5.61
Step 3 P(X = 52)
Step 4 Using approximation 1 in Table 6-2: P(51.5 < X < 52.5) (Figure 6-40).
Step 5 z_1 = \dfrac{51.5 - 45}{5.61} \approx 1.16 and z_2 = \dfrac{52.5 - 45}{5.61} \approx 1.34
Step 6 The area between z_1 and z_2 is 0.0327, or about 3.27\%.
Solution
Compute in R
[1] 45
[1] 105
[1] 5.612486
[1] 0.03268047
[1] 0.03204717
The normal approximation gives 0.0327, and the exact binomial probability is 0.032.
Returned Orders
Fifteen percent of the orders placed on an online marketplace are returned. If a random sample of 240 orders is selected, find the probability that 24 or more will be returned. (hypothetical data)
Solution
Step 1 Here p = 0.15, q = 0.85 and n = 240. Since np = (240)(0.15) = 36 and nq = (240)(0.85) = 204, the normal approximation can be used.
Step 2 \mu = np = 36 and \sigma = \sqrt{npq} = \sqrt{(240)(0.15)(0.85)} = \sqrt{30.6} \approx 5.53
Step 3 P(X \geq 24)
Step 4 Using approximation 2 in Table 6-2: P(X > 24 - 0.5) = P(X > 23.5) (Figure 6-41).
Step 5 z = \dfrac{23.5 - 36}{5.53} \approx -2.26
Step 6 The area to the right of that z value is 0.9881, or 98.81\%.
Solution
Compute in R
[1] 36
[1] 204
[1] 5.531727
[1] 0.9880798
[1] 0.9909863
In a random sample of 240 orders the probability of 24 or more being returned is 98.81%; the exact binomial probability is 0.991.
Mobile Wallet Payments
A recent in-store survey found that 45\% of shoppers at a department store pay with a mobile wallet. Find the probability that in a random sample of 120 shoppers, at most 48 will pay with a mobile wallet. (hypothetical data)
Solution
Step 1 Here p = 0.45, q = 0.55 and n = 120, so np = 54 and nq = 66. Since both values are greater than or equal to 5, the normal approximation can be used.
Step 2 \mu = np = 54 and \sigma = \sqrt{npq} = \sqrt{(120)(0.45)(0.55)} \approx 5.45
Step 3 P(X \leq 48), which with the correction factor becomes P(X < 48 + 0.5) = P(X < 48.5) (Figure 6-42).
Step 4 z = \dfrac{48.5 - 54}{5.45} \approx -1.01
Step 5 The area to the left of that z value is 0.1564, or 15.64\%.
The probability that at most 48 of the 120 shoppers pay with a mobile wallet is 15.64%.
Binomial versus Normal Approximation
When n = 12 and p = 0.5, use the binomial distribution table (Table A-2 in Appendix A) to find the probability that X = 8. Then use the normal approximation to find the probability that X = 8.
Solution
From Table A-2, for n = 12, p = 0.5 and X = 8, the probability is 0.121.
For a normal approximation, \mu = np = (12)(0.5) = 6 and \sigma = \sqrt{npq} = \sqrt{(12)(0.5)(0.5)} = \sqrt{3} \approx 1.73.
Now X = 8 is represented by the boundaries 7.5 and 8.5, so
z_1 = \frac{7.5 - 6}{1.73} \approx 0.87 \qquad z_2 = \frac{8.5 - 6}{1.73} \approx 1.44
The area between these boundaries is 0.1188 — very close to the exact binomial probability 0.1208 (Figure 6-43).
Caution
Here np = 6 and nq = 6 — only just above the minimum threshold of 5. The normal approximation may be less accurate for small n, yet 0.1188 is still close to the exact value 0.1208.
A normal distribution can also be used to approximate other distributions, such as the Poisson distribution (Table A-3 in Appendix A).
Chapter 6 Vocabulary
central limit theorem · correction for continuity · negatively or left-skewed distribution · normal distribution · positively or right-skewed distribution · sampling distribution of sample means · sampling error · standard error of the mean · standard normal distribution · symmetric distribution · z value (z score)
The Normal Distribution
Equation of a normal distribution and of the standard normal distribution: y = \frac{e^{-(X-\mu)^2/(2\sigma^2)}}{\sigma\sqrt{2\pi}} \qquad\qquad y = \frac{e^{-z^2/2}}{\sqrt{2\pi}}
Formula for the z score (or standard score), and for finding a specific data value: z = \frac{X - \mu}{\sigma} \qquad\qquad X = z \cdot \sigma + \mu
Sample Means and the Binomial Approximation
Mean of the sample means and standard error of the mean: \mu_{\bar{X}} = \mu \qquad\qquad \sigma_{\bar{X}} = \frac{\sigma}{\sqrt{n}}
z value for the central limit theorem, and the finite population correction factor: z = \frac{\bar{X} - \mu}{\sigma/\sqrt{n}} \qquad\qquad \sqrt{\frac{N-n}{N-1}}
Mean and standard deviation for the binomial distribution: \mu = n \cdot p \qquad\qquad \sigma = \sqrt{n \cdot p \cdot q}
Pearson coefficient of skewness: PC = \dfrac{3(\bar{X} - \text{median})}{s}
Key point
Key point
Copyright notice. These teaching materials follow the organization and terminology of Bluman, A. G. (2023). Elementary statistics: A step by step approach (11th ed.). McGraw Hill. All rights in the original work are reserved by its authors and publishers.
Original examples. Every worked example, data set, and R script in these slides was written for this course. The data are hypothetical unless stated otherwise.
Non-commercial use only. These materials are strictly intended for educational purposes and must not be used for commercial gain or profit.
Proper attribution. Any reproduction, distribution, or use of these materials must provide proper attribution to the original source.
Elementary Statistics: A Step by Step Approach