Statistics

Chapter 2: Frequency Distributions and Graphs

Yu-You Liou

Shih Chien University

2026-10-08

Overview

Chapter 2 shows how to organize data into frequency distributions and how to present data with charts and graphs, so that patterns which are invisible in raw data become easy to see.

Section Topics
2-1 Organizing data: categorical, grouped, and ungrouped frequency distributions; cumulative frequencies
2-2 Histograms, frequency polygons, and ogives; relative frequency graphs; distribution shapes
2-3 Other types of graphs: bar graphs, Pareto charts, time series graphs, pie graphs, dotplots, stem and leaf plots; misleading graphs

Chapter Objectives

After completing this chapter, you should be able to

  1. Organize data using a frequency distribution.
  2. Represent data in frequency distributions graphically, using histograms, frequency polygons, and ogives.
  3. Represent data using bar graphs, Pareto charts, time series graphs, pie graphs, and dotplots.
  4. Draw and interpret a stem and leaf plot.

Section 2-1: Organizing Data

Raw Data

When a researcher first gathers data for a variable, the values are in their original, unprocessed form.

Raw Data

When data are collected in original form, they are called raw data.

Suppose a researcher studies the ages of the 50 wealthiest people in the world. Little information can be obtained from looking at the raw list below, so the researcher organizes it into a frequency distribution.

45  46  64  57  85      88  45  89  67  56
92  51  71  54  48      81  58  55  62  38
27  66  76  55  69      55  56  64  81  38
54  44  54  75  46      49  68  91  56  68
61  68  78  61  83      46  47  83  71  62

Frequency Distributions

Frequency Distribution

A frequency distribution is the organization of raw data in table form, using classes and frequencies.

  • Each raw data value is placed into a quantitative or qualitative category called a class.
  • The frequency of a class is the number of data values contained in that class.

Ages of the 50 Wealthiest People

Class limits Frequency
27–35 1
36–44 3
45–53 9
54–62 15
63–71 10
72–80 3
81–89 7
90–98 2
Total 50

Now a general observation can be made: the majority of the wealthy people in the study are 45 years old or older. The values 27, 35, 36, 44, … are called class limits.

Three Types of Frequency Distributions

Type Used when
Categorical The data can be placed in specific categories — nominal- or ordinal-level data such as political affiliation, religious affiliation, or major field of study
Grouped The range of the data is large, so classes more than one unit wide are needed
Ungrouped The range of the data is small, so each class is a single data value

Categorical Frequency Distributions

The categorical frequency distribution is used for data that can be placed in specific categories.

Percentage and Relative Frequency

Find the percentage of values in each class with

\% = \frac{f}{n} \cdot 100

where f is the frequency of the class and n is the total number of values. Percentages are not normally part of a frequency distribution, but they can be added since they are used in graphs such as pie graphs. The decimal equivalent of a percent is called a relative frequency.

Example 2-1

Payment Methods at a Convenience Store

Forty customers at a campus convenience store were asked how they paid for their purchase. Construct a categorical frequency distribution for the data and summarize the results. Use these classes: C = cash, E = stored-value card, M = mobile payment, D = credit card, and G = gift voucher. (hypothetical data)

C  E  M  C  D  E  G  C  E  M
C  D  C  E  M  E  C  G  M  D
E  C  M  D  C  E  M  C  D  E
M  C  G  D  E  C  M  C  C  D

Back to Example 2-11

Example 2-1

Solution

Solve

Since the data are categorical, discrete classes are used: C, E, M, D, and G. Tally the data, count the tallies, and compute \% = \frac{f}{n}\cdot 100.

payments <- c("C","E","M","C","D","E","G","C","E","M",
              "C","D","C","E","M","E","C","G","M","D",
              "E","C","M","D","C","E","M","C","D","E",
              "M","C","G","D","E","C","M","C","C","D")

table(payments)
payments
 C  D  E  G  M 
13  7  9  3  8 
table(payments) / 40 * 100
payments
   C    D    E    G    M 
32.5 17.5 22.5  7.5 20.0 

Example 2-1

Solution

Interpretation

For the sample, most customers paid in cash (13 people, 32.5%) and the smallest number used a gift voucher (3 people, 7.5%). The stored-value card was used by 9 (22.5%), mobile payment by 8 (20%), and a credit card by 7 (17.5%). It is a good idea to add the Percent column to check that it sums to 100% — although it will not always do so because of rounding.

ggplot(data.frame(payments), aes(payments)) +
  geom_bar() +
  labs(title = "Payment Methods of 40 Customers",
       x = "Payment method", y = "Frequency")

Grouped Frequency Distributions

When the range of the data is large, the data must be grouped into classes that are more than one unit in width. The result is a grouped frequency distribution.

Class limits Class boundaries Frequency
58–64 57.5–64.5 1
65–71 64.5–71.5 6
72–78 71.5–78.5 10
79–85 78.5–85.5 14
86–92 85.5–92.5 12
93–99 92.5–99.5 5
100–106 99.5–106.5 2
Total 50

Blood glucose levels (mg/dL) for 50 randomly selected college students.

Class Limits and Class Boundaries

Limits and Boundaries

  • The lower class limit is the smallest data value that can be included in the class; the upper class limit is the largest data value that can be included in the class.
  • Class boundaries separate the classes so that there are no gaps in the frequency distribution.

Rule of Thumb for Boundaries

The class limits should have the same decimal place value as the data, but the class boundaries should have one additional place value and end in a 5.

\text{Lower limit} - 0.5 = 58 - 0.5 = 57.5 \qquad \text{Upper limit} + 0.5 = 64 + 0.5 = 64.5

If the data are in tenths, limits of 7.8–8.8 give boundaries of 7.75–8.85 (subtract and add 0.05).

Class Width and Class Midpoint

Width and Midpoint

The class width is found by subtracting the lower (or upper) class limit of one class from the lower (or upper) class limit of the next class: 65 - 58 = 7. Equivalently, subtract the lower boundary from the upper boundary of a class: 64.5 - 57.5 = 7.

The class midpoint X_m is

X_m = \frac{\text{lower boundary} + \text{upper boundary}}{2} \quad\text{or}\quad X_m = \frac{\text{lower limit} + \text{upper limit}}{2}

For the first glucose class, X_m = \dfrac{57.5 + 64.5}{2} = 61 or \dfrac{58 + 64}{2} = 61.

Class Width and Class Midpoint

Caution

Do not subtract the limits of a single class. It will result in an incorrect answer — 64 - 58 = 6, not the true width of 7.

If the class width is an even number, the midpoint is in tenths: with boundaries 5.5 and 11.5, X_m = \dfrac{5.5 + 11.5}{2} = 8.5.

Rules for Constructing a Grouped Frequency Distribution

Six Rules

  1. There should be between 5 and 20 classes — enough to present a clear description of the collected data.
  2. It is preferable but not absolutely necessary that the class width be an odd number, so the midpoint has the same place value as the data. This rule is not rigorously followed, especially when a computer is used.
  3. The classes must be mutually exclusive — nonoverlapping class limits, so data cannot be placed into two classes. Classes such as 10–20, 20–30, 30–40 are poor: where does a 40-year-old go? Use 10–20, 21–31, 32–42 instead.
  4. The classes must be continuous — even if a class has no values, it must be included. The only exception is a zero-frequency class at the very beginning or end.
  5. The classes must be exhaustive — there must be enough classes to accommodate all the data.
  6. The classes must be equal in width, which avoids a distorted view of the data. The exception is an open-ended class.

Open-Ended Distributions

Open-Ended Distribution

A frequency distribution in which the first class has no specific lower limit, or the last class has no specific upper limit, is called an open-ended distribution.

Age Frequency Minutes Frequency
10–20 3 Below 110 16
21–31 6 110–114 24
32–42 4 115–119 38
43–53 10 120–124 14
54 and above 8 125–129 5

Anybody 54 years or older is tallied in the last age class; any minute value below 110 is tallied in the first minutes class.

Procedure Table

Constructing a Grouped Frequency Distribution

Step 1 — Determine the classes.

  • Find the highest and lowest values.
  • Find the range, R = H - L.
  • Select the number of classes desired.
  • Find the width by dividing the range by the number of classes and rounding up.
  • Select a starting point (usually the lowest value or a convenient number less than it); add the width to get the lower limits.
  • Find the upper class limits.
  • Find the boundaries.

Step 2 — Tally the data.

Step 3 — Find the numerical frequencies from the tallies, and find the cumulative frequencies.

Example 2-2

Delivery Times for Online Food Orders

These data represent the delivery times, in minutes, for 50 orders handled by a food-delivery platform on one weekday evening. Construct a grouped frequency distribution for the data, using 7 classes. (hypothetical data)

31  27  35  29  22  40  30  26  33  28
25  32  37  30  43  29  36  31  27  34
38  30  28  45  32  26  35  24  31  29
18  34  41  32  27  52  30  36  25  39
23  37  21  33  42  32  47  40  35  37

Back to Example 2-4 · Example 2-5 · Example 2-6

Example 2-2

Solution

Step 1 — Determine the classes

H = 52 and L = 18, so R = 52 - 18 = 34. With 7 classes, the width is \frac{34}{7} \approx 4.9, rounded up to 5. Starting at 18 gives lower limits 18, 23, 28, …; 23 - 1 = 22 gives the first upper limit, so the classes are 18–22, 23–27, etc. Boundaries are 17.5–22.5, 22.5–27.5, and so on.

delivery_times <- c(31,27,35,29,22,40,30,26,33,28,
                    25,32,37,30,43,29,36,31,27,34,
                    38,30,28,45,32,26,35,24,31,29,
                    18,34,41,32,27,52,30,36,25,39,
                    23,37,21,33,42,32,47,40,35,37)

max(delivery_times)
[1] 52
min(delivery_times)
[1] 18
(max(delivery_times) - min(delivery_times)) / 7
[1] 4.857143

Example 2-2

Solution

Steps 2 and 3 — Tally the data and find the frequencies

class_breaks  <- seq(17.5, 52.5, by = 5)
delivery_freq <- table(cut(delivery_times, class_breaks))
delivery_freq

(17.5,22.5] (22.5,27.5] (27.5,32.5] (32.5,37.5] (37.5,42.5] (42.5,47.5] 
          3           9          16          12           6           3 
(47.5,52.5] 
          1 

Example 2-2

Solution

Interpretation

The frequency distribution shows that the class 27.5–32.5 contains the largest number of delivery times (16), followed by the class 32.5–37.5 with 12 delivery times. Hence, most of the deliveries (28) took between 28 and 37 minutes.

Analyze a frequency distribution by looking for peaks — the classes with the most data values — and for extreme values, called outliers, which are large or small relative to the other data values.

Cumulative Frequency Distributions

Cumulative Frequency Distribution

A cumulative frequency distribution is a distribution that shows the number of data values less than or equal to a specific value (usually an upper boundary).

The values are found by adding the frequencies of the classes less than or equal to the upper class boundary of a specific class. For Example 2-2 the cumulative frequency for the first class is 0 + 3 = 3; for the second, 0 + 3 + 9 = 12; for the third, 0 + 3 + 9 + 16 = 28. A shorter way is to add the frequency of the given class to the cumulative frequency of the class below: 12 + 16 = 28.

Example 2-2

Solution

Cumulative frequency distribution

cumsum(delivery_freq)
(17.5,22.5] (22.5,27.5] (27.5,32.5] (32.5,37.5] (37.5,42.5] (42.5,47.5] 
          3          12          28          40          46          49 
(47.5,52.5] 
         50 

Of the 50 delivery times, 28 are less than or equal to 32 minutes and 46 are less than or equal to 42 minutes.

Ungrouped Frequency Distributions

When the range of the data values is relatively small, a frequency distribution can be constructed using single data values for each class.

Ungrouped Frequency Distribution

An ungrouped frequency distribution uses a single data value for each class, rather than a range of values.

If the data are continuous, class boundaries can still be used: subtract 0.5 from each class value to get the lower class boundary, and add 0.5 to get the upper class boundary.

Example 2-3

Items per Shopping Basket

The data show the number of items scanned in each of 30 shopping baskets at a supermarket checkout during one lunch hour. Construct an ungrouped frequency distribution for the data and analyze the distribution. (hypothetical data)

4  9  3  5  6  4
5  3  7  4  2  6
3  5  4  8  5  3
6  4  2  5  7  4
3  6  5  4  8  7

Example 2-3

Solution

Steps 1 to 3 — Determine the classes, tally, and find the frequencies

Since the range is small (9 - 2 = 7), classes consisting of a single data value can be used: 2, 3, 4, 5, 6, 7, 8, 9.

basket_items <- c(4, 9, 3, 5, 6, 4,
                  5, 3, 7, 4, 2, 6,
                  3, 5, 4, 8, 5, 3,
                  6, 4, 2, 5, 7, 4,
                  3, 6, 5, 4, 8, 7)

table(basket_items)
basket_items
2 3 4 5 6 7 8 9 
2 5 7 6 4 3 2 1 
cumsum(table(basket_items))
 2  3  4  5  6  7  8  9 
 2  7 14 20 24 27 29 30 

Example 2-3

Solution

Interpretation

In this case, seven baskets held 4 items and this was the largest frequency. Reading the cumulative frequencies: 14 baskets held fewer than 4.5 items, 27 held fewer than 7.5 items, and all 30 held fewer than 9.5 items.

Note that several different but correct frequency distributions can be constructed for the same data by using a different class width, a different number of classes, or a different starting point. Whatever method is used, classes should be mutually exclusive, continuous, exhaustive, and of equal width.

Why Construct a Frequency Distribution?

Five Reasons

  1. To organize the data in a meaningful, intelligible way.
  2. To enable the reader to determine the nature or shape of the distribution.
  3. To facilitate computational procedures for measures of average and spread (Sections 3-1 and 3-2).
  4. To enable the researcher to draw charts and graphs for the presentation of data (Section 2-2).
  5. To enable the reader to make comparisons among different data sets.

Section 2-2: Histograms, Frequency Polygons, and Ogives

Presenting Data Graphically

The purpose of graphs in statistics is to convey the data to the viewers in pictorial form. It is easier for most people to comprehend data presented graphically than data presented numerically in tables, especially if the users have little or no statistical knowledge.

The three most commonly used graphs in research are

  1. The histogram.
  2. The frequency polygon.
  3. The cumulative frequency graph, or ogive (pronounced o-jive).

Procedure Table

Constructing a Histogram, Frequency Polygon, and Ogive

Step 1. Draw and label the x and y axes.

Step 2. On the x axis, label the class boundaries of the frequency distribution for the histogram and the ogive. Label the midpoints for the frequency polygon.

Step 3. Plot the frequencies for each class, and draw the vertical bars for the histogram and the lines for the frequency polygon and ogive.

Note: the lines for the frequency polygon begin and end on the x axis, while the lines for the ogive begin on the x axis.

The Histogram

Histogram

The histogram is a graph that displays the data by using contiguous vertical bars (unless the frequency of a class is 0) of various heights to represent the frequencies of the classes.

Karl Pearson introduced the histogram in 1891, using it to show time concepts of the various reigns of Prime Ministers.

Example 2-4

Delivery Times for Online Food Orders

Construct a histogram to represent the data shown for the delivery times of the 50 food orders (see Example 2-2).

Class boundaries Frequency
17.5–22.5 3
22.5–27.5 9
27.5–32.5 16
32.5–37.5 12
37.5–42.5 6
42.5–47.5 3
47.5–52.5 1

Back to Example 2-5 · Example 2-6

Example 2-4

Solution

Draw figure

Put the frequency on the y axis and the class boundaries on the x axis, then draw a vertical bar of that height for each class.

ggplot(data.frame(delivery_times), aes(delivery_times)) +
  geom_histogram(breaks = class_breaks, color = "white") +
  labs(title = "Delivery Times for 50 Food Orders",
       x = "Delivery time (minutes)", y = "Frequency")

The largest class (16) is 27.5–32.5, then 12 for 32.5–37.5. The graph has one peak, with the data clustering around it.

The Frequency Polygon

Frequency Polygon

The frequency polygon is a graph that displays the data by using lines that connect points plotted for the frequencies at the midpoints of the classes. The frequencies are represented by the heights of the points.

The frequency polygon and the histogram are two different ways to represent the same data set; the choice is left to the discretion of the researcher.

Example 2-5

Delivery Times for Online Food Orders

Using the frequency distribution given in Example 2-4, construct a frequency polygon.

Example 2-5

Solution

Step 1 — Find the midpoints of each class

Midpoints are found by adding the upper and lower boundaries and dividing by 2: \frac{17.5 + 22.5}{2} = 20, \frac{22.5 + 27.5}{2} = 25, and so on.

polygon_points <- data.frame(midpoint  = seq(15, 55, by = 5),
                             frequency = c(0, 3, 9, 16, 12, 6, 3, 1, 0))
polygon_points
  midpoint frequency
1       15         0
2       20         3
3       25         9
4       30        16
5       35        12
6       40         6
7       45         3
8       50         1
9       55         0

Example 2-5

Solution

Output figure

Connect adjacent points with line segments, then draw a line back to the x axis at the beginning and end of the graph, at the same distance that the previous and next midpoints would be located.

ggplot(polygon_points, aes(midpoint, frequency)) +
  geom_line() +
  geom_point() +
  labs(title = "Delivery Times for 50 Food Orders",
       x = "Delivery time (minutes)", y = "Frequency")

The Ogive

Cumulative Frequency and the Ogive

The cumulative frequency is the sum of the frequencies accumulated up to the upper boundary of a class in the distribution.

The ogive is a graph that represents the cumulative frequencies for the classes in a frequency distribution.

Example 2-6

Delivery Times for Online Food Orders

Construct an ogive for the frequency distribution described in Example 2-4.

Example 2-6

Solution

Step 1 — Find the cumulative frequency for each class

ogive_points <- data.frame(boundary   = class_breaks,
                           cumulative = cumsum(c(0, 3, 9, 16, 12, 6, 3, 1)))
ogive_points
  boundary cumulative
1     17.5          0
2     22.5          3
3     27.5         12
4     32.5         28
5     37.5         40
6     42.5         46
7     47.5         49
8     52.5         50

Plot the cumulative frequency at each upper class boundary, then extend the graph back to the first lower class boundary, 17.5, on the x axis.

Example 2-6

Solution

Output figure

ggplot(ogive_points, aes(boundary, cumulative)) +
  geom_line() +
  geom_point() +
  labs(title = "Delivery Times for 50 Food Orders",
       x = "Delivery time (minutes)", y = "Cumulative frequency")

Go up from 32.5 on the x axis to the graph, then across to the y axis: 28.

Relative Frequency Graphs

The histogram, frequency polygon, and ogive shown so far were constructed using frequencies in terms of the raw data. These distributions can be converted to proportions instead; the graphs are then called relative frequency graphs.

When to Use Relative Frequencies

Use relative frequencies when the proportion of data values falling into a class matters more than the actual number. Comparing the age distribution of Philadelphia (population 1,526,006) with Erie (population 101,786) using raw counts would make every Philadelphia bar much taller.

To convert a frequency into a relative frequency, divide the frequency for each class by the total of the frequencies. The sum of the relative frequencies is always 1.

Example 2-7

Floor Area of Retail Units

The following frequency distribution shows the floor area, in square metres, of 50 retail units in a shopping mall. Draw a relative frequency histogram, a relative frequency polygon, and a relative ogive for the data. (hypothetical data)

Class boundaries Frequency
24.5–31.5 6
31.5–38.5 13
38.5–45.5 19
45.5–52.5 8
52.5–59.5 4
Total 50

Example 2-7

Solution

Steps 1 and 2 — Find the relative and cumulative relative frequencies

For the class 24.5–31.5 the relative frequency is \frac{6}{50} = 0.12; for 31.5–38.5 it is \frac{13}{50} = 0.26; and so on. The cumulative relative frequencies are 0.12, 0.38, 0.76, 0.92, 1.00.

frequency <- c(6, 13, 19, 8, 4)

frequency / 50
[1] 0.12 0.26 0.38 0.16 0.08
cumsum(frequency / 50)
[1] 0.12 0.38 0.76 0.92 1.00

Example 2-7

Solution

Relative frequency histogram

For the histogram and ogive use the class boundaries along the x axis; for the frequency polygon use the midpoints. For the scale on the y axis, use proportions.

rel_histogram <- data.frame(midpoint = c(28, 35, 42, 49, 56),
                            rel_freq = frequency / 50)

ggplot(rel_histogram, aes(midpoint, rel_freq)) +
  geom_col(width = 7, color = "white") +
  labs(title = "Relative Frequency Histogram",
       x = "Floor area (square metres)", y = "Relative frequency")

Example 2-7

Solution

Relative frequency polygon

rel_polygon <- data.frame(midpoint = c(21, 28, 35, 42, 49, 56, 63),
                          rel_freq = c(0, frequency, 0) / 50)

ggplot(rel_polygon, aes(midpoint, rel_freq)) +
  geom_line() +
  geom_point() +
  labs(title = "Relative Frequency Polygon",
       x = "Floor area (square metres)", y = "Relative frequency")

Example 2-7

Solution

Relative ogive

rel_ogive <- data.frame(boundary = c(24.5, 31.5, 38.5, 45.5, 52.5, 59.5),
                        cum_rel  = cumsum(c(0, frequency)) / 50)

ggplot(rel_ogive, aes(boundary, cum_rel)) +
  geom_line() +
  geom_point() +
  labs(title = "Relative Ogive",
       x = "Floor area (square metres)", y = "Cumulative relative frequency")

The three graphs have exactly the same shapes as the corresponding frequency graphs; only the y axis has changed from counts to proportions.

Distribution Shapes

The shape of a distribution decides which statistical methods fit later.

Shape Description
Bell-shaped (mound) One peak, tapering at both ends; symmetric
Uniform Basically flat or rectangular
J-shaped Few values at the left, rising to the right
Reverse J-shaped The opposite of the J-shaped distribution
Right-skewed (positive) Peak at the left, tapering right
Left-skewed (negative) Clustered at the right, tapering left
Bimodal Two peaks of the same height
U-shaped High at both ends, low in the middle

Distribution Shapes

Distributions are most often not perfectly shaped, so it is not necessary to have an exact shape but rather to identify an overall pattern.

Analyzing a Histogram or Frequency Polygon

  • Does it have one peak or two? Distributions with one peak — bell-shaped, right-skewed, left-skewed — are said to be unimodal; a distribution with two peaks of the same height is bimodal.
  • Is it relatively flat, or is it U-shaped?
  • Are the data spread out, or clustered around the center?
  • Are there data values at the extreme ends? These may be outliers (Section 3-3).
  • Are there gaps in the histogram, or does the frequency polygon touch the x axis somewhere other than at the ends?
  • Are the data clustered at one end or the other, indicating a skewed distribution?

Distribution Shapes in R

Prepare data

set.seed(42)
shapes <- data.frame(
  shape = rep(c("Bell-shaped", "Uniform", "J-shaped", "Reverse J-shaped",
                "Right-skewed", "Left-skewed", "Bimodal", "U-shaped"),
              each = 500),
  value = c(rnorm(500, 50, 10), runif(500, 20, 80), 100 - rexp(500, 0.05),
            rexp(500, 0.05), rgamma(500, 2, 0.1), 80 - rgamma(500, 2, 0.1),
            rnorm(250, 30, 5), rnorm(250, 70, 5), rbeta(500, 0.4, 0.4) * 100)
)

Distribution Shapes in R

Output figure

ggplot(shapes, aes(value)) +
  geom_histogram(bins = 20, color = "white") +
  facet_wrap(~ shape, scales = "free", ncol = 4)

Section 2-3: Other Types of Graphs

Beyond the Histogram

In addition to the histogram, the frequency polygon, and the ogive, several other types of graphs are often used in statistics:

  • Bar graph — for qualitative or categorical data
  • Pareto chart — a bar graph whose bars are ordered from highest to lowest
  • Time series graph — for data collected over a period of time
  • Pie graph — for showing the relationship of the parts to the whole
  • Dotplot — each data value plotted as a point above a horizontal axis
  • Stem and leaf plot — a combination of sorting and graphing that retains the actual data

Bar Graphs

Bar Graph

A bar graph represents the data by using vertical or horizontal bars whose heights or lengths represent the frequencies of the data.

When the data are qualitative or categorical, bar graphs can be used to represent the data. A bar graph can be drawn using either horizontal or vertical bars.

Example 2-8

Monthly Student Spending

The table shows the average amount, in NT dollars, that students in an international trade programme spend each month in four categories. Draw a horizontal and a vertical bar graph for the data. (hypothetical data)

Category Average spending
Meals off campus 3,200
Commuting 1,500
Mobile data plan 750
Study materials 480

Example 2-8

Solution

Prepare data

Draw and label the x and y axes. For the horizontal bar graph place the frequency scale on the x axis; for the vertical bar graph place the frequency scale on the y axis.

monthly_spending <- data.frame(
  category = c("Meals off campus", "Commuting",
               "Mobile data plan", "Study materials"),
  amount   = c(3200, 1500, 750, 480)
)

Example 2-8

Solution

Draw figure — horizontal bar graph

ggplot(monthly_spending, aes(amount, category)) +
  geom_col() +
  labs(title = "Average Monthly Student Spending",
       x = "Average amount (NT dollars)", y = NULL)

Example 2-8

Solution

Output figure — vertical bar graph

ggplot(monthly_spending, aes(category, amount)) +
  geom_col() +
  labs(title = "Average Amount Spent",
       x = NULL, y = "Average amount (NT dollars)")

The graphs show that these students spend the most on meals off campus.

Compound Bar Graphs

Bar graphs can also be used to compare data for two or more groups; these are called compound bar graphs. It is not necessary to have equal class sizes in these types of graphs.

Store format Weekday Weekend
Convenience store 820 960
Supermarket 540 910
Hypermarket 310 720

Average number of customers visiting one store per day, by store format. (hypothetical data)

Compound Bar Graphs in R

Draw figure

store_visits <- data.frame(
  format   = rep(c("Convenience", "Supermarket", "Hypermarket"), each = 2),
  day_type = rep(c("Weekday", "Weekend"), times = 3),
  visits   = c(820, 960, 540, 910, 310, 720)
)

ggplot(store_visits, aes(format, visits, fill = day_type)) +
  geom_col(position = "dodge") +
  labs(title = "Average Daily Visitors by Store Format",
       x = "Store format", y = "Visitors per day")

The graph shows that every format is busier at the weekend, and that the weekend increase is largest for the hypermarket, whose visitor count more than doubles, and smallest for the convenience store.

Pareto Charts

Pareto Chart

A Pareto chart is used to represent a frequency distribution for a categorical variable, and the frequencies are displayed by the heights of vertical bars, which are arranged in order from highest to lowest.

Vilfredo Pareto (1848–1923) was an Italian scholar whose research on income distribution became known as Pareto’s law; the Pareto distribution is named after him.

Example 2-9

Reasons for Abandoning an Online Cart

An online retailer asked shoppers who left items in a shopping cart why they did not complete the order. The percentage of shoppers naming each reason is shown. Draw a Pareto chart for the data. (hypothetical data)

Reason Percent naming it
Shipping cost 41%
Slow delivery 26%
Checkout steps 33%
Payment options 18%
Out of stock 29%

Example 2-9

Solution

Step 1 — Arrange the data from largest value to smallest value

cart_reasons <- data.frame(
  reason  = c("Shipping cost", "Slow delivery", "Checkout steps",
              "Payment options", "Out of stock"),
  percent = c(41, 26, 33, 18, 29)
)

arrange(cart_reasons, desc(percent))
           reason percent
1   Shipping cost      41
2  Checkout steps      33
3    Out of stock      29
4   Slow delivery      26
5 Payment options      18

Example 2-9

Solution

Steps 2 and 3 — Draw and label the axes, then draw vertical bars for the percents

ggplot(cart_reasons, aes(reorder(reason, -percent), percent)) +
  geom_col() +
  labs(title = "Why Shoppers Abandon a Cart",
       x = NULL, y = "Percent of shoppers")

Shipping cost is named by 41% of the shoppers, while the payment options are named by only 18% — so the retailer should look at shipping cost first.

Suggestions for Drawing Pareto Charts

Three Suggestions

  1. Make the bars the same width.
  2. Arrange the data from largest to smallest according to frequency.
  3. Make the units that are used for the frequency equal in size.

When you analyze a Pareto chart, make comparisons by looking at the heights of the bars.

The Time Series Graph

Time Series Graph

A time series graph represents data that occur over a specific period of time.

When you analyze a time series graph, look for a trend or pattern over the time period: is the line ascending (an increase over time) or descending (a decrease)? Also look at the slope, or steepness — a steep line indicates a rapid increase or decrease over that period.

Example 2-10

Hot Spring Hotel Visitors

The data show the number of visitors, in thousands, to a hot spring hotel in each month of a one-year period. Draw a time series graph for the data and describe the results. (hypothetical data)

Jan 9.8 May 5.5 Sept 6.0
Feb 9.1 June 4.8 Oct 7.2
Mar 7.4 July 5.1 Nov 8.6
Apr 6.2 Aug 5.4 Dec 10.4

Example 2-10

Solution

Prepare data

Label the x axis for months and the y axis for thousands of visitors, then plot each point.

hotel_visitors <- data.frame(
  month    = 1:12,
  visitors = c(9.8, 9.1, 7.4, 6.2, 5.5, 4.8,
               5.1, 5.4, 6.0, 7.2, 8.6, 10.4)
)

Example 2-10

Solution

Output figure

Draw line segments connecting adjacent points; do not draw a smooth curve through the points.

ggplot(hotel_visitors, aes(month, visitors)) +
  geom_line() +
  geom_point() +
  scale_x_continuous(breaks = 1:12) +
  labs(title = "Hot Spring Hotel Visitors",
       x = "Month", y = "Visitors (thousands)")

Visitors were highest in December, January, and February and lowest in June, so the hotel’s business clearly follows the cold season.

Compound Time Series Graphs

Two or more data sets can be compared on the same graph — a compound time series graph — by drawing two or more lines.

For example, an airline records its load factor — the percentage of seats sold — for short-haul and for long-haul routes in each year from 2018 to 2024. Two lines drawn on one pair of axes show at a glance that both fell sharply in 2020, that the short-haul line recovered first, and that the long-haul line was still below its 2018 level in 2024. (hypothetical data)

The Pie Graph

Pie Graph

A pie graph is a circle that is divided into sections or wedges according to the percentage of frequencies in each category of the distribution.

Degrees for Each Section

\text{Degrees} = \frac{f}{n} \cdot 360^\circ

where f is the frequency for each class and n is the sum of the frequencies. The degrees should sum to 360^\circ, and the percentages to 100% — although rounding can make the totals slightly off.

Example 2-11

Hot Drinks Sold by a Convenience Store Chain

This frequency distribution shows the number of cups, in thousands, of each hot drink sold by a convenience store chain in one week. Construct a pie graph for the data. The percentage formula is the one introduced in Example 2-1. (hypothetical data)

Drink Cups (frequency)
Latte 12.7 thousand
Americano 9.4 thousand
Milk tea 6.1 thousand
Black tea 4.8 thousand
Hot cocoa 3.0 thousand
Total n = 36.0 thousand

Example 2-11

Solution

Steps 1 and 2 — Convert each frequency to degrees and to a percentage

cups <- c(12.7, 9.4, 6.1, 4.8, 3.0)

cups / 36 * 360
[1] 127  94  61  48  30
cups / 36 * 100
[1] 35.277778 26.111111 16.944444 13.333333  8.333333

The degrees are 127^\circ, 94^\circ, 61^\circ, 48^\circ, and 30^\circ; the percentages are 35.3%, 26.1%, 16.9%, 13.3%, and 8.3%.

Example 2-11

Solution

Step 3 — Draw the graph and label each section

pie(cups, labels = c("Latte", "Americano", "Milk tea", "Black tea", "Hot cocoa"),
    main = "Hot Drinks Sold in One Week")

Example 2-12

Packages Sorted by Shift

Construct and analyze a pie graph for the number of packages a distribution centre sorted on each of its three daily shifts last month. (hypothetical data)

Shift Frequency
1. Morning 3060
2. Afternoon 3240
3. Night 2700
Total 9000

Example 2-12

Solution

Steps 1 and 2 — Find the number of degrees and the percentages for each shift

packages <- c(3060, 3240, 2700)

packages / 9000 * 360
[1] 122.4 129.6 108.0
packages / 9000 * 100
[1] 34 36 30

Morning: \frac{3060}{9000}\cdot 360^\circ \approx 122^\circ and 34%; Afternoon: 130^\circ and 36%; Night: 108^\circ and 30%.

Example 2-12

Solution

Step 3 — Graph each section and write its name

pie(packages, labels = c("Morning", "Afternoon", "Night"),
    main = "Packages Sorted by Shift")

The numbers of packages sorted on the three shifts are about equal, although slightly more were handled on the afternoon shift.

Dotplots

Dotplot

A dotplot is a statistical graph in which each data value is plotted as a point (dot) above the horizontal axis. If the data values occur more than once, the corresponding points are plotted above one another.

Dotplots are used to show how the data values are distributed and to see whether there are any extremely high or low data values.

Example 2-13

Complaints by Product Line

The data show the number of complaints a customer service centre logged about each of 50 product lines during one month. Draw a dotplot for the data and summarize the results. (hypothetical data)

2  1  0  3  12  1  4  1  2  5
1  0  2  1   3  6  1  0  2  1
3  9  5  2   1  4  0  2  1  0
1  4  2  3   0  1  6  2  3  7
0  2  1  3   4  5  6  3  2  4

Example 2-13

Solution

Prepare data

The lowest data value is 0 and the highest is 12, so a scale from 0 to 12 is needed.

complaints <- c(2,1,0,3,12,1,4,1,2,5,
                1,0,2,1, 3,6,1,0,2,1,
                3,9,5,2, 1,4,0,2,1,0,
                1,4,2,3, 0,1,6,2,3,7,
                0,2,1,3, 4,5,6,3,2,4)

table(complaints)
complaints
 0  1  2  3  4  5  6  7  9 12 
 7 12 10  7  5  3  3  1  1  1 

Example 2-13

Solution

Output figure

Draw a horizontal scale, then plot each value above it, stacking repeated values on top of one another.

stripchart(complaints, method = "stack", pch = 19, at = 0,
           main = "Complaints by Product Line", xlab = "Number of complaints")

Most product lines drew between zero and three complaints, with 12 product lines drawing exactly one complaint — the largest frequency.

Stem and Leaf Plots

Stem and Leaf Plot

A stem and leaf plot is a data plot that uses part of the data value as the stem and part of the data value as the leaf to form groups or classes.

A data value of 34 would have 3 as the stem and 4 as the leaf; a data value of 356 would have 35 as the stem and 6 as the leaf. The stem and leaf plot has the advantage over a grouped frequency distribution of retaining the actual data while showing them in graphical form.

Example 2-14

Customers Served at a Night Market Stall

A night market food stall recorded the number of customers it served each evening for 20 evenings. Construct a stem and leaf plot for the data. (hypothetical data)

33  22  41  13  54
08  35  31  58  26
45  32  17  38  24
51  33  47  36  42

Example 2-14

Solution

Steps 1 to 3 — Arrange the data in order, separate by leading digit, and plot

evening_customers <- c(33,22,41,13,54,
                        8,35,31,58,26,
                       45,32,17,38,24,
                       51,33,47,36,42)

stem(evening_customers, scale = 2)

  The decimal point is 1 digit(s) to the right of the |

  0 | 8
  1 | 37
  2 | 246
  3 | 1233568
  4 | 1257
  5 | 148

scale = 2 gives each stem its own row. If there are no data values in a class, write the stem number and leave the leaf row blank — do not put a zero in the leaf row.

Example 2-14

Solution

Interpretation

The plot shows that the distribution peaks in the center and that there are no gaps in the data. For 7 of the 20 evenings the number of customers served was between 31 and 38. The plot also shows that the stall served from a minimum of 8 customers to a maximum of 58 customers in any one evening.

When you analyze a stem and leaf plot, look for peaks and gaps, see whether the distribution is symmetric or skewed, and check the variability by looking at the spread.

Example 2-15

Pallets Dispatched by a Distribution Centre

A distribution centre recorded the number of pallets it dispatched each day for 30 days. Construct a stem and leaf plot by using classes 50–54, 55–59, 60–64, 65–69, 70–74, and 75–79. (hypothetical data)

57  63  50  66  55
69  52  58  61  76
54  67  59  65  60
71  56  64  51  68
53  62  57  78  55
59  73  66  63  69

Example 2-15

Solution

Steps 1 to 3 — Arrange the data in order, separate by class, and plot

Because the classes are 5 units wide, each stem is repeated twice: once for the leaves 0–4 and once for the leaves 5–9.

pallets <- c(57,63,50,66,55,
             69,52,58,61,76,
             54,67,59,65,60,
             71,56,64,51,68,
             53,62,57,78,55,
             59,73,66,63,69)

stem(pallets)

  The decimal point is 1 digit(s) to the right of the |

  5 | 01234
  5 | 55677899
  6 | 012334
  6 | 5667899
  7 | 13
  7 | 68

The distribution has no gaps; the largest class, 55–59, contains 8 of the 30 days.

Back-to-Back Stem and Leaf Plots

Back-to-Back Stem and Leaf Plot

Related distributions can be compared by using a back-to-back stem and leaf plot. It uses the same digits for the stems of both distributions, but the digits used for the leaves are arranged in order out from the stems on both sides.

Stem and leaf plots are part of the techniques called exploratory data analysis; more on this topic appears in Chapter 3.

Example 2-16

Monthly Revenue of Coffee Shops in Two Cities

The monthly revenue, in hundreds of thousands of NT dollars, of two randomly selected samples of coffee shops in Taipei and Taichung is shown. Construct a back-to-back stem and leaf plot for the data and compare the distributions. (hypothetical data)

       Taipei                    Taichung
64  57  52  45  60        45  53  44  36  47
55  38  68  50  62        61  55  40  35  32
63  44  58  70  41        54  34  51  43  36
69  53  61  66  56        48  41  39  50  72
73  58  49  45  64        58  42  66  45  55

Example 2-16

Solution

Steps 1 and 2 — Arrange each data set in order with stem()

taipei   <- c(64,57,52,45,60, 55,38,68,50,62, 63,44,58,70,41,
              69,53,61,66,56, 73,58,49,45,64)
taichung <- c(45,53,44,36,47, 61,55,40,35,32, 54,34,51,43,36,
              48,41,39,50,72, 58,42,66,45,55)

stem(taipei)

  The decimal point is 1 digit(s) to the right of the |

  3 | 8
  4 | 14559
  5 | 02356788
  6 | 012344689
  7 | 03
stem(taichung)

  The decimal point is 1 digit(s) to the right of the |

  3 | 245669
  4 | 012345578
  5 | 0134558
  6 | 16
  7 | 2

Example 2-16

Solution

Step 3 — Put the two plots back to back and compare

Taipei Stem Taichung
8 3 2 4 5 6 6 9
9 5 5 4 1 4 0 1 2 3 4 5 5 7 8
8 8 7 6 5 3 2 0 5 0 1 3 4 5 5 8
9 8 6 4 4 3 2 1 0 6 1 6
3 0 7 2
  • Taipei’s coffee shops generally take in more than those in Taichung.
  • The distribution for the shops in Taipei peaks at 60 to 69, while the shops in Taichung peak at 40 to 49.
  • Taipei has 11 shops with revenue of 60 or more; Taichung has 3.

Misleading Graphs

Graphs give a visual representation that enables readers to analyze and interpret data more easily than they could simply by looking at numbers. However, inappropriately drawn graphs can misrepresent the data and lead the reader to false conclusions.

A car manufacturer’s ad stated that 98% of the vehicles it had sold in the past 10 years were still on the road, and showed a bar graph whose vertical axis ran only from 95% to 100%. Redrawn on a scale from 0 to 100%, there is hardly a noticeable difference between the manufacturer and its competitors.

Misleading Graphs in R

Prepare data

vehicles_on_road <- data.frame(
  maker   = c("Manufacturer", "Competitor I", "Competitor II"),
  percent = c(98, 97, 96.5)
)

Misleading Graphs in R

Draw figure — the truncated scale (95% to 100%)

ggplot(vehicles_on_road, aes(maker, percent)) +
  geom_col() +
  coord_cartesian(ylim = c(95, 100)) +
  labs(title = "Vehicles on the Road (scale 95 to 100)",
       x = NULL, y = "Percent of cars on road")

Misleading Graphs in R

Output figure — the full scale (0% to 100%)

ggplot(vehicles_on_road, aes(maker, percent)) +
  geom_col() +
  coord_cartesian(ylim = c(0, 100)) +
  labs(title = "Vehicles on the Road (scale 0 to 100)",
       x = NULL, y = "Percent of cars on road")

Both graphs plot exactly the same numbers. It is not wrong to truncate an axis — many times it is necessary — but the reader should be aware of it and interpret the graph accordingly.

Stretching the Scale

The average number of parcels a courier at a distribution hub delivers per hour is shown for six years; it rose from 12.4 to 13.5 parcels per hour. (hypothetical data)

Year 2019 2020 2021 2022 2023 2024
Parcels per hour 12.4 12.7 12.9 13.0 13.2 13.5

On a y axis running from 0 to 20 the increase looks slight; spreading out the scale so that it runs only from 12 to 14 in steps of 0.2 makes the same data values look like a much larger increase.

Two-Dimensional Pictures

Another misleading technique exaggerates a one-dimensional increase by showing it in two dimensions.

Suppose a logistics firm’s daily parcel volume grew from 20,000 to 60,000 — three times as large (hypothetical data). Drawn as two bars, the change is compared by the heights of the bars, which is one dimension, so the taller bar is three times the shorter one. Drawn as two circles whose diameters are in the same 1-to-3 ratio, the difference looks far larger, because the eye compares the areas of the circles, and the larger area is nine times the smaller one.

Misleading Graphs

Caution

It is not wrong to truncate a scale or to represent data by two-dimensional pictures — but when these techniques are used, the reader should be cautious of the conclusions drawn from the graph. Before accepting a graph, check:

  • Does the value axis start at zero, or has it been truncated?
  • Has the scale been stretched or compressed to exaggerate a change?
  • Are picture symbols scaled by area rather than by height?
  • Are labels or units omitted from the axes? Without numbers on the y axis, only a crude ranking can be obtained and there is no way to decide the actual magnitude of the differences.
  • Is a source given for the information? A source lets you check the reliability of the organization presenting the data.

Summary

Graph type Best used for
Histogram Grouped frequencies; contiguous vertical bars
Frequency polygon Frequencies at midpoints, joined by lines
Ogive Cumulative frequencies at upper boundaries
Bar graph Categorical data; horizontal or vertical bars
Pareto chart Categorical data, bars highest to lowest
Time series graph Data occurring over a period of time
Pie graph The relationship of the parts to the whole
Dotplot Each value as a dot; spotting extreme values
Stem and leaf plot Keeps the actual values and shows the shape

Important Terms

Chapter 2 Vocabulary

bar graph · categorical frequency distribution · class · class boundaries · class midpoint · class width · compound bar graphs · cumulative frequency · cumulative frequency distribution · dotplot · frequency · frequency distribution · frequency polygon · grouped frequency distribution · histogram · lower class limit · ogive · open-ended distribution · Pareto chart · pie graph · raw data · relative frequency graph · stem and leaf plot · time series graph · ungrouped frequency distribution · upper class limit

Key Formulas

Important Formulas

Percentage of values in each class: \displaystyle \% = \frac{f}{n} \cdot 100

Range: R = \text{highest value} - \text{lowest value}

Class width: \text{Class width} = \text{upper boundary} - \text{lower boundary}

Class midpoint: X_m = \frac{\text{lower boundary} + \text{upper boundary}}{2} \quad\text{or}\quad X_m = \frac{\text{lower limit} + \text{upper limit}}{2}

Degrees for each section of a pie graph: \displaystyle \text{Degrees} = \frac{f}{n} \cdot 360^\circ

Key Takeaways

Key point

  • Raw data — values collected in original form; little can be learned until they are organized
  • Frequency distribution — the organization of raw data in table form, using classes and frequencies
  • Three types — categorical (specific categories), grouped (large range), ungrouped (small range)
  • Class limits / boundaries / width / midpoint — limits carry the data’s decimal place; boundaries carry one more and end in 5; width is the difference between consecutive lower limits
  • Six rules — 5 to 20 classes, preferably odd width, mutually exclusive, continuous, exhaustive, equal in width (except open-ended)
  • Cumulative frequency distribution — the number of data values less than or equal to a specific value
  • Histogram / frequency polygon / ogive — boundaries, midpoints, and cumulative frequencies respectively
  • Relative frequency graphs — proportions instead of counts; needed when comparing data sets of different sizes
  • Distribution shapes — bell-shaped, uniform, J-shaped, reverse J-shaped, right-skewed, left-skewed, bimodal, U-shaped
  • Other graphs — bar graph, Pareto chart, time series graph, pie graph, dotplot, stem and leaf plot
  • Misleading graphs — truncated or stretched scales, two-dimensional pictures, and omitted labels distort the message

Acknowledgement

  • Copyright notice. These teaching materials follow the organization and terminology of Bluman, A. G. (2023). Elementary statistics: A step by step approach (11th ed.). McGraw Hill. All rights in the original work are reserved by its authors and publishers.

  • Original examples. Every worked example, data set, and R script in these slides was written for this course. The data are hypothetical unless stated otherwise.

  • Non-commercial use only. These materials are strictly intended for educational purposes and must not be used for commercial gain or profit.

  • Proper attribution. Any reproduction, distribution, or use of these materials must provide proper attribution to the original source.