0% found this document useful (0 votes)
15 views11 pages

Understanding Stem-and-Leaf Plots

The document explains stem-and-leaf plots as a quick method for displaying data distributions using integer values, allowing for easy construction in classroom settings. It includes step-by-step instructions for creating a stem-and-leaf plot using a fitness exam dataset from a PE class, as well as variations like split stems and back-to-back plots for comparisons. Additionally, it introduces scatterplots and box plots as tools for visualizing relationships and summarizing data, respectively.

Uploaded by

Shamar Francis
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
15 views11 pages

Understanding Stem-and-Leaf Plots

The document explains stem-and-leaf plots as a quick method for displaying data distributions using integer values, allowing for easy construction in classroom settings. It includes step-by-step instructions for creating a stem-and-leaf plot using a fitness exam dataset from a PE class, as well as variations like split stems and back-to-back plots for comparisons. Additionally, it introduces scatterplots and box plots as tools for visualizing relationships and summarizing data, respectively.

Uploaded by

Shamar Francis
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Version: July 2003

Stem-and-leaf plots
Description:
A stem-and-leaf plot provides an alternative to a line plot or histogram for obtaining a
picture of the data distribution when data can be represented as integer (counting)
numbers. The resulting graph looks very much like a line plot or a histogram. However,
this type of plot can be constructed very quicklyit can be used, for example, when a
teacher wants to construct a graph on a chalkboard or an overhead using data collected by
the children in the class. The children can simply call out their numbers one by one and
the teacher can enter the numbers onto the graph.

An example data set:


Eric’s third period PE class just had a fitness exam. Each child did as many sit-ups as
he/she could do in one minute. Eric’s score was 45. How well did he do relative to the
rest of the class? The data for the class are as follows:

49 58 57 43 46 35 42 56
47 45 45 43 51 36 41 50
34 42 46 47 64 40 38 41
50 48

A stem-and-leaf plot provides a convenient way to display the distribution of these data.
In a stem-and-leaf graph, we separate the digits of each data point into the right most
digit vs. the rest, so, for example, the 1s place becomes the “leaf,” and the rest of the
number, becomes the “stem.” So in the above data set, we can split the numbers into the
number of 10s (the stem) and the number of 1s (the leaf). 49 thus becomes 40 + 9 or 4
tens plus 9 ones. We are not restricted to 2-digit numbersfor example, we could have
collected forearm lengths in mm, in which case the numbers above might be 100 greater,
so 149 = 14 tens plus 9 ones, etc.

In constructing the stem-and-leaf plot we end up with something that looks very much
like a line plot or histogram. There are two differences, however. First, we pool counts
within a class of numbers, e.g. all counts 30-39, 40-49, etc. Second, we retain the actual
numbers to make the graph, rather than using dots or Xs, so that we can continue to use
the graph for determining other summaries of the data, such as the median. (Note that we
lose this information when we use a histogram, although we can still obtain such
summaries from a line plot).

⇒ Note: to get a good picture of the data, the final plot must be made quite neatly. To
encourage this, it is often convenient to use graph paper, and to write each digit in a
separate box on the graph paper.

The method of construction:


Step #1 Find the largest and smallest scores (the largest is 64 and the smallest is
34). If one person in the class is making the plot, this is easy to do orally
by asking for a small/large number and then asking for successively
smaller/larger numbers until there are no more extreme values.

6
Version: July 2003

Step #2 Because the smallest number, 34, has a 3 in the tens place, and the largest
number, 64, has a 6 in the tens place, the stems will be the digits 3 to 6.
For younger children you may want to simply write down all the tens
place digits (1-9). Write these digits vertically with a line to the right.

Stem
3 |
4 |
5 |
6 |

Step #3 Separate each score into a stem (number of tens) and a leaf (the 1s digit).
Write the leaf on the plot next to the stem. The first score is 49. The 9 is
placed next to the 4 in the stem:

Stem Leaf
3 |
4 | 9
5 |
6 |

Continue until all the scores (from left to right) in the above data set are
listed on the plot:

3 | 5 6 4 8
4 | 9 3 6 2 7 5 5 3 1 2 6 7 0 1 8
5 | 8 7 6 1 0 0
6 | 4

Step #4 On a new plot, rearrange the leaves so that they are ordered from smallest
to largest.

3 | 4 5 6 8
4 | 0 1 1 2 2 3 3 5 5 6 6 7 7 8 9
5 | 0 0 1 6 7 8
6 | 4

Step #5 Add a title and a key to the plot.

Sit-ups/min. in Period 3 PE class


3 | 4 5 6 8
4 | 0 1 1 2 2 3 3 5 5 6 6 7 7 8 9
5 | 0 0 1 6 7 8
6 | 4
3|4 represents 34 sit-ups

Interpreting the data:


The key descriptive terms, such as symmetry, skewness, mode, outliers, clusters and
gaps, can all be used with a stem-and-leaf plot, just as with a line plot or histogram. For
this example, scores are relatively symmetric with no unique mode. But since most of
7
Version: July 2003

the scores are in the 40s, we could call this the modal class or category. Where is Eric’s
score? Right in the middle of this modal category. Eric is “typical” of a student from
this class.

Variations and extensions:


For children who have mastered the concept of place value it is possible to introduce
variants on the stem-and-leaf plot.

a) Split stems:
In the above example, there were many data points in the middle categories. Spreading
the data out might reveal additional patterns. One way to do this is to split the stems into
a smaller number of equally-sized units. For example, units of 5 ones (30-34, 35-39, 40-
44, etc.). Taking the previous plot, we then get:
Sit-ups/min. in Period 3 PE class
3 | 4
• | 5 6 8
4 | 0 1 1 2 2 3 3
• | 5 5 6 6 7 7 8 9
5 | 0 0 1
• | 6 7 8
6 | 4
3 | 4 = 34 sit-ups
• | 5 = 35 sit-ups
( • is same value as stem
above it)
Can you see how this spread out the data?

b) Back-to-back stem-and-leaf plots:


Like line plots and histograms, stem-and-leaf plots can be used to compare two
situations. For example, there might be data for a second class, perhaps of older children,
or from a class which has been doing sit-ups every day for a couple of weeks:

Comparison of sit-ups/min in 2 classes

Class 2 Class 1
3 2 | 3 | 4
5 5 | • | 5 6 8
3 2 0 | 4 | 0 1 1 2 2 3 3
8 8 7 6 6 5 | • | 5 5 6 6 7 7 8 9
5 4 3 3 2 0 0 | 5 | 0 0 1
9 7 6 5 5 | • | 6 7 8
3 1 0 | 6 | 4
| 3 | 4 represents
34 sit-ups
0 | 5 | represents
50 sit-ups
Which class is generally more fit?

8
Version: July 2003

⇒ Note that back-to-back stem-and-leaf plots require writing the left-hand numbers
backwards, and therefore this method of plotting is not suitable for the younger
children.

⇒ Note also that an easy way to demonstrate the back-to-back version is to use 3
transparencies: one containing only the stem, the second, overlaid on the first for the
data from situation 1, and the third, overlaid on the first (after removing the second)
for the data on situation 2. Then you put all 3 transparencies on top of each other,
except that you flip the top one over so the numbers are written backwards. This is
suitable for an initial demonstration, or for a quick data collection later, but if you
want to keep the plot around for further discussion, you should transfer it onto a
single sheet with the digits all written in the normal fashion.

⇒ Note that for stem-and-leaf plots you retain the actual values of the numbers. The
median number of sit-ups in class 1 can be determined to be 45, while the median in
class 2 is 50 sit-ups.

9
Version: July 2003

Scatterplots and correlation


Description:
Scatterplots can be thought of as an extension of a two-way table for measurement of
(numeric) data that comes in the form of ordered pairs. An ordered pair consists of a pair
of numbers for one individual or object, such that the two numbers are two different
measurements on the same individual. Two examples might be height and weight of an
individual, or circumference and volume of a sphere. The data could be entered into a
very large two-way table with one unit of measurement for each category, but this would
not give us a very informative picture of the data. Instead, we can plot a point for each
pair of data. The composite data set then can be displayed as a scatterplot.

Scatterplots are difficult for the younger children (K-1 or K-2) to use for actual data
analysis. However, with appropriate activities (e.g., plotting number of whole peanuts
vs. peanuts inside the shell in a handful), scatterplots can be introduced to the younger
children. An earlier introduction will increase familiarity with interpretation of the plot
so that by about second or third grade it should be possible to begin to use scatterplots for
some routine data analyses.

An example:
Suppose we have measured the height (in inches) and weight (in pounds) of students in a
class. A scatterplot of these data is shown below. Notice that the plot slants upwards as
we move from lower to higher heights. This is an example of a positive correlation.
Height vs. Weight
240

220

200

180

160

Weight
140

120

100
55 60 65 70 75 80
Height

We could further emphasize this association by drawing a line that is as close as possible
to all the points; this line would have a positive slope (in more advanced classes formal
procedures for choosing this line are introduced, but for elementary and middle-school
children a line “by eye” is adequate). We might also emphasize the shape by drawing an
ellipse around the majority of the points. This ellipse would also slant upwards.

15
Version: July 2003

Height vs. Weight


240

220

200

180

160

Weight
140

120

100
55 60 65 70 75 80
Height

A second scatterplot might show number of hours of TV watched in a week vs. number
of book pages read. This might show a negative association, although with probably
more scatter than seen for the height/weight example. If we drew a line or an ellipse
around these points, it would slant downwards. Note that we would leave a few outliers
outside the ellipse.

H o u rs o f T V w a tc h e d vs . P a g e s re a d
180

160

140

120

100

Pages 80
re a d 60

40

20

0
0 2 4 6 8 10 12 14 16 18 20

T V h o u rs

Finally, a third scatterplot might show no positive or negative association; in this case
we would say there is little or no correlation between the two variables. An ellipse drawn
around these points would not slant either upwards or downwards.

16
Version: July 2003

Characteristics of our clothes


9

5
Number of 4
buttons
3

0
0 1 2 3 4 5 6 7 8 9

Number of colors

Interpreting the data:


As in all other cases, a good way to think about data interpretation is to look for patterns
and departures from patterns. In the example of height vs. weight, the overall pattern is
that weight increases with height. The line drawn through the points summarizes this
increasing trend in much the same way a mean or a median captures the center of a
distribution plotted via, e.g., a line plot.

The particular departures from the pattern may also be interesting to consider. Is there
something truly unusual about these points, or are they just errors in plotting the data?
Two examples of unusual points are the single individual in the center of the weight
distribution with very low height, and the 3 points clustered to the far right and above of
the rest of the points. Are these data-measurement errors or are they really extreme
points (outliers).

Initial preparation:
To make a scatterplot with data collected by students, start with either large sheets of
quad-ruled chart paper, or just tape a set of axes on the board or a piece of butcher paper.
Find the minimum and maximum values in the class for each variable. Mark off the axes
in even intervals to reflect these values, starting at or slightly below the minimum value.
For very young children stick with integers, but for older children the numbers can
include non-integer scales as well. Label the axes.

Graphing:
Have students come up in small groups (3-5) at a time and place a dot at the intersection
representing their pair of numbers. It is useful to model the process with one child first if
the children have not done this kind of plot before. To get appropriate placement of
dots, have the children first find the correct values for their horizontal point. Then have
them place their fingers on this point and trace upwards until they find the vertical-axis
value corresponding to their data. That is where they must place their dots.

17
Version: July 2003

If you want to bring out additional features of the data, for example, the difference
between boys and girls in their height/weight relationship, you can give the students
different colored dots.

18
Version: July 2003

Box Plots
Description:
Box plots provide a visual representation of a five-number summary of data, consisting
of the median (the midpoint of the data range), the upper and lower quartiles (the
numbers below the highest quarter of the data and above the lowest quarter, respectively)
and the largest and smallest values (the extremes). Box plots are particularly useful for
comparing distributions of the results from several experimental conditions.

Because box plots are based on simple summaries, they can be used with fairly young
childrencertainly third graders, but even, in some cases, younger children. For the
youngest children, one can “lead up” to a box plot by using only the median (middle
number) and the whiskers to the extremes.

Box plots are an important type of graph to use with children. More than any other type
of graph, they focus the user on several key statistical concepts, perhaps the most
important of which is that the data can be summarized. Because box plots focus on
representing summaries of the data, the children are not distracted by issues such as gaps
or multiple modes. In addition, the box plot is a superb way of emphasizing and
representing the variability inherent in real data in a way that is computationally
accessible to children. This is a key conceptknowing that there is a way of describing
or summarizing variability is very important, and will lay the conceptual groundwork for
other summaries that can be introduced in high school. Finally, the box plot strongly
emphasizes the idea of the center of the distribution. Again, this is a key concept,
especially as children begin to compare results from different experimental situations.

An example:
Return to the original sit-ups/min. data set from the stem-and-leaf plot section. The five-
number summary, written in the order of the lower extreme, lower quartile, median,
upper quartile, and upper extreme is:

34 41 45.5 50 64

A box plot of these numbers, plotted above the real number line for reference, would look
like the following:

-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-
30 40 50 60 70
sit-ups/min

19
Version: July 2003

Constructing the graph:


Step #1 Determine the five-number summary.

Step #2 Construct a number line that includes the extremes.

-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-
30 40 50 60 70

Step #3 Mark the position of the five numbers in the summary a little above the
number line.

* * * * *
-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-
30 40 50 60 70

Step #4 Draw a narrow box that connects the quartiles. Draw a line through the
box at the median. Draw lines from the ends of the boxes to the extremes
(“whiskers”).

* * * * *
-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-|-
30 40 50 60 70

Interpreting the data:


Half the scores will lie inside the box. Half the scores will lie outside the box on the
whiskers. Half the scores lie below the line in the box, and half lie above this line. One
quarter of the scores lie on each whisker and in each half of the box. Below are some
additional examples, without the number lines (and plotted on different scales), to
illustrate possible shapes of resulting distributions:

3rd graders’ heights

After school group heights

Ages of children in a competitive,


all-city orchestra

Years that a group of teachers has


taught

Variations and extensions:


Parallel box plots provide a convenient way to compare different situations. It is easy to
compare more than two conditions, and it is easy to determine where salient features of
the data lie. Parallel box plots are also a convenient way to compare two (or more)

20
Version: July 2003

groups with very different sample sizes (in this situation, back-to-back line plots or
histograms are hard to compare).

For example, children have counted the number of drops of clean, soapy, and sugary
water that fit on pennies:

clean

sugar

soap

One can see that there is not much of a difference in the centering of the clean vs. sugary
water, although the clean water example may be a little less variable. One can also see
the downward shift in the number of drops of soapy water, although there is still a lot of
overlap with the other two conditions.

When making box plots, it is often useful to make them after first plotting the data as a
line plot, or a stem-and-leaf plot. Then you plot the box plot next to the same set of axes:

x
x x
x x x x
x x x x x x x
x x x x x x x x x x x
| | | | | | | | | | | |
0 1 2 3 4 5 6 7 8 9 10 11
Years of teaching

This is particularly useful when you first introduce the box plot as it will relate a more
familiar type of graph to this new type of graph. If you are comparing 2 (or more)
conditions, you will find that the box plot probably will bring out the location of the
distribution better than will the link plot; by putting the box plots next to the line plots,
this point will come out quite clearly, especially if there is a slight shift in the
distribution.

21

You might also like