Business Statistics Module Guide
Business Statistics Module Guide
Module Guide
Copyright© 2022
MANCOSA
All rights reserved, no part of this book may be reproduced in any form or by any means, including photocopying machines,
without the written permission of the publisher. Please report all errors and omissions to the following email address:
modulefeedback@[Link]
This Module Guide,
Business Statistics (NQF level 6)
module guide will be used across the following programmes:
Preface.................................................................................................................................................................... 3
i
Business Statistics
List of Contents
List of Figures and Illustrations
Figure 3.7: Line Graph showing Kings Plastics Profits from 1997-2002 ........................................................... 35
Figure 3.8: Bar Graph (vertical) depicting Profits of Kings Plastics ................................................................... 36
Figure 3.9: Bar Graph (horizontal) depicting Profits of Kings Plastics ............................................................... 36
Figure 5.2: Comparison of Actual and Predicted Frequency of Student Marks ................................................. 73
1 MANCOSA
Business Statistics
Figure 7.7: Seasonal Demand history for Metro Movers ................................................................................. 109
MANCOSA 2
Business Statistics
Preface
A. Welcome
Dear Student
It is a great pleasure to welcome you to Business Statistics (BS6). To make sure that you share our passion
about this area of study, we encourage you to read this overview thoroughly. Refer to it as often as you need to
since it will certainly make studying this module a lot easier. The intention of this module is to develop both your
confidence and proficiency in this module.
The field of Business Statistics is extremely dynamic and challenging. The learning content, activities and self-
study questions contained in this guide will therefore provide you with opportunities to explore the latest
developments in this field and help you to discover the field of Business Statistics as it is practiced today.
This is a distance-learning module. Since you do not have a tutor standing next to you while you study, you need
to apply self-discipline. You will have the opportunity to collaborate with each other via social media tools. Your
study skills will include self-direction and responsibility. However, you will gain a lot from the experience! These
study skills will contribute to your life skills, which will help you to succeed in all areas of life.
Statistics is all around us. It would be difficult to go through a full week without using statistics. So, what is statistics?
Statistics is the science of designing studies, gathering data, and then classifying, summarising, interpreting, and
presenting these data to support the decisions that are needed.
The word "statistics" is used in several different senses. In the broadest sense, "statistics" refers to a range of
techniques and procedures for analysing data, interpreting data, displaying data, and making decisions based on
data. This is what courses in "statistics" generally cover.
In a second usage, a "statistic" is defined as a numerical quantity (such as the mean) calculated in a sample. Such
statistics are used to estimate parameters. A parameter is a characteristic of a population e.g. the mean. A statistic
is a measure based on sample data.
Without statistics, we could not plan our budgets, pay our taxes, enjoy games to their fullest, and evaluate
classroom performance. Are you beginning to get the picture? We need statistics.
Let us take a look at the most basic form of statistics, known as descriptive statistics.
3 MANCOSA
Business Statistics
This branch of statistics lays the foundation for all statistical knowledge (important, huh?), but it is not something
that you should learn simply so you can use it in the distant future. Descriptive statistics can be used, in a language
class, in a science class, at the football stadium, in the grocery store. You probably already know more about these
statistics than you think.
Another form of statistics is called inferential statistics, which deals with inferences (decision making, predictions,
etc.) about the process or population being studied.
This module has been divided into five Units and at the beginning of each Unit you will find a list of Learning
Outcomes. These provide you with an outline of what you should have learned by the time you have completed
the Unit, and can also be used to focus your study.
The concepts are best learned and understood through practice, and we thus suggest that you do all the exercises
recommended throughout the guide, as well as carefully going through all the worked examples. A good suggestion
is trying to work out the given examples without looking at the solution, then comparing your answers with the
model answers. In this way, you can keep track of your progress.
MANCOSA does not own or purport to own, unless explicitly stated otherwise, any intellectual property rights in or
to multimedia used or provided in this module guide. Such multimedia is copyrighted by the respective creators
thereto and used by MANCOSA for educational purposes only. Should you wish to use copyrighted material from
this guide for purposes of your own that extend beyond fair dealing/use, you must obtain permission from the
copyright owner.
MANCOSA 4
Business Statistics
B. Module Overview
This Module is a 15 credit Module at NQF level 6.
Contents and Structure
Unit 1: Introduction
In the introduction, we explain why it is important to be able to manipulate numbers and have a feel for magnitude.
We also discuss how best to carry out the exercises in the text, and what basic pre-work is needed.
Access to a Computer
Some of the exercises given to you in this module will be arduous to do without a personal computer. Before doing
the exercises on a computer, make sure you are able to apply the techniques manually, to help you better
understand the calculations carried out by the computer.
5 MANCOSA
Business Statistics
Understand the importance of statistical Statistical measures and methods of calculating and
techniques and calculations in business ascertaining them are extensively explored to provide a
management sound quantitative grounding for competence and
understanding
Perform statistical analyses and extract Principles, concepts and techniques of data gathering,
relevant information from business data classification, summarisation, interpretation and
presentation/communication are demonstrated and
critiqued to enhance understanding and appreciation of
objective statistical practices
Manipulate collected data through various Methods and techniques of data organisation and
statistical methods to generate useful processing are applied to real life scenarios to promote
information to support management understanding of the importance of producing unbiased,
decisions reliable research outcomes
Prepare and interpret reports in statistical Data processing and presentation methods and techniques
terms are examined to consolidate skills in the creation and
presentation of clear, objective and reliable statistical
reports
Assess the validity of statistical findings Statistical tests and variability measures are thoroughly
and the relevance and reliability of results examined to enhance understanding and competence in
their application and use to obtain unchallengeable
statistical results and inferences
MANCOSA 6
Business Statistics
The purpose of the Module Guide is to allow you the opportunity to integrate the theoretical concepts from the
prescribed textbook and recommended readings. We suggest that you briefly skim read through the entire guide
to get an overview of its contents. At the beginning of each Unit, you will find a list of Learning Outcomes and
Associated Assessment Criteria. This outlines the main points that you should understand when you have
completed the Unit/s. Do not attempt to read and study everything at once. Each study session should be 90
minutes without a break
This module should be studied using the prescribed and recommended textbooks/readings and the relevant
sections of this Module Guide. You must read about the topic that you intend to study in the appropriate section
before you start reading the textbook in detail. Ensure that you make your own notes as you work through both the
textbook and this module. In the event that you do not have the prescribed and recommended textbooks/readings,
you must make use of any other source that deals with the sections in this module. If you want to do further reading,
and want to obtain publications that were used as source documents when we wrote this guide, you should look at
the reference list and the bibliography at the end of the Module Guide. In addition, at the end of each Unit there
may be link to the PowerPoint presentation and other useful reading.
F. Study Material
The study material for this module includes tutorial letters, programme handbook, this Module Guide, a list of
prescribed and recommended textbooks/readings which may be supplemented by additional readings.
H. Special Features
In the Module Guide, you will find the following icons together with a description. These are designed to help you
study. It is imperative that you work through them as they also provide guidelines for examination purposes.
7 MANCOSA
Business Statistics
You may come across Activities that ask you to carry out
specific tasks. In most cases, there are no right or wrong
ACTIVITY
answers to these activities. The purpose of the activities is to
give you an opportunity to apply what you have learned.
PRACTICAL
Practical Application or Examples will be discussed to enhance
APPLICATION OR
understanding of this module.
EXAMPLES
MANCOSA 8
Business Statistics
9 MANCOSA
Business Statistics
Unit
1: Introduction to Statistics
MANCOSA 10
Business Statistics
1.1. The Principles and Methods for Understand the scope of analytical techniques
Designing Studies
1.2. Numerical and Computer Literacy Discuss the importance of displaying numerical and
computer literacy
11 MANCOSA
Business Statistics
You are probably asking yourself the question, "When and where will I use statistics?" If you read any newspaper
or watch television, or use the Internet, you will see statistical information. There are statistics about crime, sports,
education, politics, and real estate. Typically, when you read a newspaper article or watch a news program on
television, you are given sample information. With this information, you may make a decision about the correctness
of a statement, claim, or "fact." Statistical methods can help you make the "best educated guess."
Since you will undoubtedly be given statistical information at some point in your life, you need to know some
techniques to analyze the information thoughtfully. Think about buying a house or managing a budget. Think about
your chosen profession. The fields of economics, business, psychology, education, biology, law, computer science,
police science, and early childhood development require at least one course in statistics.
Think back to some of the business conversations you have had during the last week or two, and the type of
information that has been offered during these conversations. More than likely, you will have heard expressions
like:
The cost of that material is R1500-00
80% of those that responded thought it was a good idea
it took him two hours to complete the task
4500 Toyota Corollas were sold in August
You would not be blamed for feeling a little frustrated with these expressions, because they don’t really give you
good information. In each case, we are missing:
MANCOSA 12
Business Statistics
a basis of comparison
the source of the data
how it was collected
why it was collected in the first instance
If these criteria were in place, we would have some information to assist us in our decision-making.
For example, knowing that 4500 Toyota Corollas were sold in August means nothing to me if I am selling Ford
Escorts. But if I see that in July, 3000 Toyota Corollas were sold, and the sales for Ford Escorts dropped from 2000
to 500 in August, I now have a better understanding of the significance of that data. The reason for the drop in
sales of the Ford Escorts may be due to the increase in Toyota sales. I now have information on which to make a
decision on what to do next.
In this module, we will look at some of the techniques that can be applied by business managers to collect, analyse
and interpret quantitative information to make informed decisions.
13 MANCOSA
Business Statistics
Answers to Activities
Unit 1
Activity 1
(a) Quantitative, ratio
(b) Quantitative, ratio
(c) Quantitative, ratio
(d) Quantitative, ratio
(e) Qualitative, nominal
(f) Qualitative, nominal
(g) Quantitative, ratio
(h) Quantitative, ratio
(i) Qualitative, nominal
(j) Qualitative, nominal
(k) Quantitative, ratio
(l) Quantitative, ratio
(m) Quantitative, ratio
(n) Quantitative, ratio
(o) Quantitative, ratio
(p) Qualitative, nominal
MANCOSA 14
Business Statistics
Unit
2: The Nature of Data, Data Collection
and Sources
15 MANCOSA
Business Statistics
2.2. Internal and External Data Discuss internal, external, primary and secondary sources
Sources
2.3. Data Types Identify and discuss the various data types
2.4. Data Collection Methods Discuss and describe observation, interviews and experiments as
data collection methods for statistical analysis
MANCOSA 16
Business Statistics
2.1. Introduction
In this unit we examine the factors that influence the quality of data on which important management decisions are
based. Data quality is influenced by the types of data available for analysis, the sources from which data are
collected and the methods by which data are collected. The types of data available determine the appropriate types
of statistical techniques to employ, while the sources from which data are gathered and the methods of collecting
data determine the accuracy and reliability of statistical findings.
The cost of the external data depends on the source, but you may be surprised how much information is freely
available, either on the Internet or in business publications. A detailed study of the economic indicators in financial
publications will tell you a surprising amount. Statistics SA, the government’s source of statistical data, has virtually
all its data available on its home page on the internet. Many regard new motor vehicle sales as a good indicator of
the economy – and these are published for all NAAMSA members monthly.
17 MANCOSA
Business Statistics
Examples:
“Aged” market research figures
Previous financial statements
MANCOSA 18
Business Statistics
An industry market research from which you are extracting data for your company
The type of data gathered determines the type of analysis which can be performed; an incorrect application of a
statistical method to a particular data type can render the findings invalid. Data type is determined by the nature of
the random variable that the data represents.
Wegner (2007) identifies two types of random variables: qualitative and quantitative.
Quantitative data measures either how much or how many of something, i.e. a set of observations where any single
observation is a number that represents an amount or a count. Quantitative random variables yield numeric
responses, and can be meaningfully manipulated using conventional arithmetic operations. Examples are age,
distance, number of items, monetary amount, etc.
Qualitative data provide labels, or names, for categories of like items, i.e. a set of observations where any single
observation is a word or code that represents a class or category. Qualitative random variables yield categorical or
non-numeric responses. The data generated are classified into one of a number of categories.
19 MANCOSA
Business Statistics
The categories are usually represented by codes, which cannot be manipulated arithmetically. These codes are
merely used as labels. Figure 2.1 shows an example of data codes.
Each of these random variable categories can be associated with a different type of data classification.
Wegner (2007) defines two data classification types:
Data type 1 - Nominal–scaled
- Ordinal–scaled
- Interval–scaled
- Ratio–scaled
Data type 2 - Discrete
- Continuous
Nominal-scaled data: Data with no inherent order or ranking sequence, e.g. numbers used as names (group 1,
group 2), gender, etc. Nominal–scaled data is associated mainly with qualitative random variables. There is no
implied ordering between groups of the random variable, and each category is of equal importance. Figure 2.1 is
a good example
Ordinal-scaled data: Data with an ordered series, e.g. "greatly dislike, moderately dislike, indifferent, moderately
like, greatly like". Numbers assigned to such data indicate rank order only - the "distance" between the numbers
has no meaning. Ordinal-scaled data is also associated mainly with qualitative random variables. Like nominal-
scaled data, it is also assigned to one of a number of coded categories, but there is now a ranking implied between
the categories in terms of being better, bigger, longer, older, taller or stronger, etc. An example is shown in Figure
MANCOSA 20
Business Statistics
Interval-scaled data: Equally-spaced data, e.g. temperature. The difference between a temperature of 66 degrees
and 67 degrees is taken to be the same as the difference between 76 degrees and 77 degrees. Interval variables
do not have a true zero, e.g. 88 degrees is not necessarily double the temperature of 44 degrees. Interval-scaled
data is associated with quantitative random variables; differences can be measured between values. Interval-
scaled data possesses both order (implied ranking) and distance properties.
In social research studies, such as market research, the Likert Rating Scale is often used for respondents to
indicate a preference or a perception on a scale; interval-scaled properties are created for the study. An example
is shown in Figure 2.3
Example: Indicate your response to the statement “Shopping is a social experience for me.” (Tick a value from the
rating scale.)
Strongly Disagree Disagree Unsure Agree Strongly Agree
1 2 3 4 5
3Figure 2.3: Example of Interval-Scaled Data
The data does not contain an absolute origin, so the ratio of values cannot be meaningfully compared. A rating of
4 in the above example would reflect a stronger perception than a 2 (this is a property of order); it is not, however,
possible to conclude that a rating of 4 is twice as important as a rating of 2. But it is possible to conclude that the
difference in perception (or preference) between 3 and 4 is the same as between 1 and 2. This is a property of
distance.
Ratio-Scaled Data: is mainly associated with quantitative random variables; it is numeric data with a zero origin.
Examples are age, distance, time, mass, sales, units and income. Such data is the strongest form of statistical data
that can be gathered and lends itself to the widest range of statistical methods.
Ratio-scaled data is gathered through a measurement process, and can be manipulated meaningfully through
normal arithmetic operations. If ratio-scaled data is grouped into categories, that data becomes ordinal-scaled; an
example is the data for “Turnover” in Figure 2.2.
21 MANCOSA
Business Statistics
Discrete variables are usually obtained by counting. There are a finite or countable number of choices available
with discrete data. You can't have 2.63 people in the room.
Examples:
The number of students in a class
The number of cars sold in a month by a dealer
Continuous variables are usually obtained by measuring. Length, weight, and time are all examples of continuous
variables. Since continuous variables are real numbers, we usually round them. This implies a boundary depending
on the number of decimal places. For example, the measurement x = 64 is really anything in the range 63.5 < x <
64.5. Likewise, if there are two decimal places, then x = 64.03 is really anything in the range 63.025 < x < 63.035.
Boundaries always have one more decimal place than the data and end in a 5.
Examples:
time taken to travel to work daily
tensile strength of steel
speed of an aircraft
MANCOSA 22
Business Statistics
2.4.1 Observation
The use of observation as a measurement tool, assigning numerals to human behavioral acts, is discussed.
Observation has important advantages which makes it best suited for certain kinds of studies, and some limitations
which preclude its use in others. The central problems in the use of observation are:
(1) The effect of the observer on the observed, which is usually not severe and can be minimized.
(2) Observer inference, which is a crucial strength and a crucial weakness.
(3) The unit of behavior to be used, which involves the molar-molecular problem.
The considerations in planning both unstructured and structured observation studies are discussed, including what
to observe, how to record it, how to maximize validity and reliability, and how to handle the relationship between
the observer and the observed. Behavior is usually sampled using event sampling or time sampling.
Primary data can be collected by direct observation of the respondent or item in action. Examples include:
Vehicle traffic surveys
Observing the purchase behaviour of brands in a store
Quality control inspection
An advantage of direct observation is that the respondent is generally unaware of being observed and therefore
behaves naturally. This reduces the likelihood of gathering biased data.
A disadvantage is that it is a passive form of data collection, and there is little opportunity to probe for reasons or
investigate behaviour further.
Secondary data can be obtained through desk research (abstraction), from a variety of source documents. A wide
variety of organisations and individuals continually consult and use secondary data for decision making or opinion
forming.
Observation is a technique that involves systematically selecting, watching and recording the behaviour and
characteristics of living beings, objects or phenomena.
23 MANCOSA
Business Statistics
Observation of human behaviour is a much-used data collection technique. It can be undertaken in different ways:
Participant observation: The observer participates in the situation he or she observes. (For example, a
doctor hospitalised with a broken hip, who now observes hospital procedures ‘from within’.)
Non-participant observation: The observer watches the situation, openly or concealed, but does not
participate
Observations can be open (e.g., ‘shadowing’ a health worker with his/her permission during routine activities) or
concealed (e.g., ‘mystery clients’ trying to obtain antibiotics without medical prescription). They may serve different
purposes. Observations can give additional, more accurate information on behaviour of people than interviews or
questionnaires. They can also check on the information collected through interviews especially on sensitive topics
such as alcohol or drug use, or stigmatising diseases. For example, whether community members share drinks or
food with patients suffering from feared diseases (leprosy, TB, AIDS) are essential observations in a study on
stigma.
Observations of human behaviour can form part of any type of study, but as they are time consuming they are most
often used in small-scale studies.
Observations can also be made on objects. For example, the presence or absence of a latrine and its state of
cleanliness may be observed. Here observation would be the major research technique.
Observations made using a defined scale are called measurements. Measurements usually require additional tools.
For example, in nutritional surveillance weight and height are measured by using weighing scales and a measuring
board. We use thermometers for measuring body temperature.
2.4.2 Interview
An Interview is a data-collection technique that involves oral questioning of respondents, either individually or as a
group. Interviews can be conducted through direct questioning or a questionnaire. Interview data can be gathered
through personal (face-to-face) interviews, postal surveys and telephone surveys.
Answers to the questions posed during an interview can be recorded by writing them down (either during the
interview itself or immediately after the interview) or by tape-recording the responses, or by a combination of both.
Interviews can be conducted with varying degrees of flexibility. The two extremes, high and low degree of flexibility,
are described below:
MANCOSA 24
Business Statistics
These may, e.g., include how teenagers started sexual intercourse, the responsibility girls and their partners take
to prevent pregnancy (if at all), and the actions they take in the event of unwanted pregnancies. The investigator
should have an additional list of topics ready when the respondent falls silent, (e.g., when asked about abortion
methods used, who made the decision and who paid). The sequence of topics should be determined by the flow
of discussion. It is often possible to come back to a topic discussed earlier in a later stage of the interview.
The unstructured or loosely structured method of asking questions can be used for interviewing individuals as well
as groups of key informants.
A flexible method of interviewing is useful if a researcher has as yet little understanding of the problem or situation
he is investigating, or if the topic is sensitive. It is frequently applied in exploratory studies. The instrument used
may be called an interview guide or interview schedule.
Though in principle one may speak of loosely structured questionnaires, in practice the term questionnaire appears
to be so hooked to tools with pre-categorised answers that we have decided to use the term interview guide for
loosely structured tools. However, in reality there is often a mixture of open and pre-categorised answers. In that
case we will still use the term questionnaire.
A written questionnaire (also referred to as self-administered questionnaire) is a data collection tool in which written
questions are presented that are to be answered by the respondents in written form.
25 MANCOSA
Business Statistics
Personal interviews have the advantage that accurate data can be obtained immediately, and qualitative data can
be obtained by probing for reasons and observing non-verbal responses. They are, however, time consuming, and
expensive if trained interviewers are required.
Telephone interviews allow more flexibility, in that call-backs can be made if a respondent is not available initially,
and people are more willing to talk on the telephone from the security of their home or an office. It is more cost
effective, as a larger sample of respondents can be reached in a relatively short time. The main disadvantage of
telephone interviewing is that non-verbal responses cannot be observed.
Postal surveys (they can be conducted by fax or by e-mail as well) are best used when the target population is
large and/or geographically dispersed.
A larger sample of respondents can be reached, making them more cost effective. Because respondents can
answer anonymously, more honest, considered responses would be given. However, questions have to be shorter
and simpler, and the possibility of probing is limited. Data collection can take a long time, and there is no control
over who answers the questionnaire, or the possibility of check-backs on the validity of responses.
According to Wegner (2007) the response rates of postal surveys are very low (5% - 15%). The questionnaire is
the data collection instrument used to gather data in all interview situations. The design of the questionnaire is
critical to ensure that the correct research questions are addressed and that accurate and appropriate data are
collected.
2.4.3 Experimentation
Primary data can also be generated through the manipulation of variables under controlled conditions. Data on the
primary variable under study can be monitored and recorded while conscious efforts are made by the researcher
to control the effects of a number of influencing factors.
Examples:
The hardness of toughened glass can be measured for various tempering furnace temperatures
Advertising effectiveness can be measured by manipulating the frequency and choice of various media
While good quality data is collected if the experiment is correctly designed and executed, experimentation is a
costly and time consuming exercise. It may also be impossible to control certain extraneous factors that can distort
the results.
MANCOSA 26
Business Statistics
Activity 1
For each of the following variables, indicate the data type and the measurement
scale (i.e. nominal, ordinal, interval or ratio).
(a) The shelf life of milk
(b) The number of life policies issued per day
(c) The area of a shop floor
(d) The number of pages in a text book
(e) The flavours available in dog-food chunks
(f) The wood types that can be used to make a desk
(g) The size categories for shoes
(h) The voltage produced by a generator
(i) The car types in the Mercedes range
(j) The Yes/No/Sometimes response to “Do you drink Gin?”
(k) The number of loaves of bread sold daily by a bakery
(l) The income per day of a bakery
(m) The monthly birth-rate at a maternity hospital
(n) The mass of babies at birth
(o) The daily distance travelled by a courier service truck
(p) The names of teams in a cricket league.
Activity 2
Study the statistics printed in newspapers, magazines and on the Internet. The
more you study them, the more information you will start obtaining from them.
You will also start picking up trends.
2.5. Summary
In this unit we examined data as the raw material of statistics. We distinguished internal and external data sources,
and primary and secondary data. We also looked at different data types and different types of measurement scales.
Finally, we covered the methods of data collection, namely: observation, interview and experimentation.
27 MANCOSA
Business Statistics
Unit
3: Presentation of Data
MANCOSA 28
Business Statistics
3.2. Tables in Business Discuss tables as part of the data presentation process
3.3. Use of Graphs and Charts Discuss graphs and charts as part of the data presentation
process
Construct various graphs and charts
3.4. Grouped Frequency Distributions Construct grouped frequency distributions and cumulative
grouped frequency distributions
29 MANCOSA
Business Statistics
3.1. Introduction
The information obtained from a statistical analysis is meaningful to business managers only when it can be
interpreted and communicated effectively and concisely. It is customary to convey such information through the
use of summary tables and charts. Tables and charts convey information more vividly and quickly than written
reports.
Normally, the first stage of ordering is to present the data in a table. Once data is in the form of a table, we can:
Make comparisons, within the table and with other data
Perform additional calculations
Examine the component structure of the data
Using a table to list data according to category is often much clearer than writing out all the information in paragraph
form. Let's look at an example of some data first written up in a paragraph. Try to think what information is being
given and what sort of trends one could find from the information.
Example: During the 1995-1996 academic year, a survey of the holdings of university research libraries and rank
was done in the United States and Canada. It was found that Syracuse University, in New York, had 2,692,147
holdings, and was figured to rank eighty-first. Harvard University ranked first with 13,369,855 holdings. The
University of Connecticut was ranked fiftieth place, and reported 2,626,066 holdings. The Massachusetts Institute
of Technology reported 2,448,647 holdings, and was ranked in seventy-third place. (Source: Association of
Research Libraries).
As you can see, the paragraph above contains a lot of numbers and is not always easy to follow. The information
given in the paragraph would be easier to decipher if it was presented in a table. To create a table, you need to
determine the following things:
Title of the table
Label of each row and/or column
Number of rows and columns necessary
Data entry for each cell
MANCOSA 30
Business Statistics
Holdings and Rank of University Research Libraries in the U.S. And Canada--1995-1996.
U. Connecticut 50 2626066
Notice that the order of the universities in the table is different from the order they are listed in the paragraph.
When moving from the paragraph to the table, it is best to order the instances by any numerical data. In this
case, we ordered them by rank.
From the information in the table, we can see that the number one ranked university contains a lot more holdings
than the other three institutions.
We carry out these processes to convert the data into information with which we are able to make good
decisions.
A commonly used table in business is a frequency table, which shows the number of occurrences of a variable
falling into a specific range or category.
Example: Consider a group of 47 males of various ages. 12 are between 20 and 29 years of age, 13 are between
30 and 39 years of age, 7 are between 40 and 49 years of age, 8 are between 50 and 59 years of age while the
rest are between 60 and 69 years of age.
Interval (years) 20 – 29 30 - 39 40 - 49 50 - 59 60 - 69
Number of Males 12 13 7 8 7
5Figure 3.2: Grouped Frequency Distribution Showing Male Ages
31 MANCOSA
Business Statistics
Note: The sum of the number of men in each interval (i.e. 12+13+7+8+7) must equal to the total number in the
group which is 47 in this case.
Example: Consider a class consisting of 100 students. Suppose the teacher gives the entire class a statistics test
which has a maximum mark of 100. Upon marking the scripts (which are in no order whatsoever), he puts the
marks into a table as shown below. This represents raw data since there is no set order of the marks.
28 60 58 63 72 52 63 82 65 58
75 59 26 52 71 55 55 55 52 62
90 58 25 51 68 47 52 56 55 35
80 41 47 49 52 48 44 64 56 18
35 42 38 48 53 45 48 62 57 52
12 44 85 47 45 41 75 51 51 48
65 46 76 46 46 40 25 52 48 56
50 54 74 36 32 50 66 53 46 48
55 51 72 24 8 51 50 48 42 47
45 35 65 56 44 60 55 49 40 50
Since the highest possible mark in this case is 100, and the lowest mark is 0, an interval size (width) of 10 is easy
to work with. Hence, we may choose to use the following intervals:
MANCOSA 32
Business Statistics
Interval
0 9
10 19
20 29
30 39
40 49
50 59
60 69
70 79
80 89
90 99
Now that we have decided on the intervals, we can do a frequency count, which we do by registering each mark
in the correct interval. For example, we would register the first value, 28, in the interval 20 - 29; we would tally the
rest as shown in figure 3.4.
Interval Tally
0–9 /
10 – 19 //
20 – 29 /////
30 – 39 //////
40 – 49 /////////////////////////////
50 – 59 //////////////////////////////////
60 – 69 ////////////
70 – 79 ///////
80 – 89 ///
90 – 99 /
7Figure 3.4: Tally of Student Marks
Note: This function is easily done on a spreadsheet – if you are unsure how to do it, look up “frequency distribution”
in the HELP facility. When the tallying is complete, you can construct a frequency table as shown in figure 3.5.
33 MANCOSA
Business Statistics
Cumulative
Interval Frequency
Frequency
0–9 1 1
10 – 19 2 3
20 – 29 5 8
30 – 39 6 14
40 – 49 29 43
50 – 59 34 77
60 – 69 12 89
70 – 79 7 96
80 – 89 3 99
90 – 99 1 100
8Figure 3.5: Frequency table of Student Marks
The column Cumulative Frequency refers to the total number of observations/measurements encountered up
until a particular interval. This frequency table could be modified to show ratios as percentages if this is preferred.
Example: The table in figure 3.6 shows the total profit made by Kings Plastics for six consecutive years, from 1997
– 2002.
MANCOSA 34
Business Statistics
1997 78
1998 65
1999 53
2000 49
2001 38
2002 16
9Figure 3.6: Profit of Kings Plastic from 1997-2002
A study of the table shows that the profits have dropped considerably. This trend can be better depicted using a
line graph (Figure 3.7) as shown on the following page.
100
80
Profit (R'000)
60
40
20
0
1996 1997 1998 1999 2000 2001 2002 2003
Year
10Figure 3.7: Line Graph showing Kings Plastics Profits from 1997-2002
The line graph clearly shows the drop in profits from 1997 to 2002.
Note the features of a line graph:
1. The vertical (y) and horizontal (x) axes are perpendicular to each other.
2. An appropriate scale is used such that the data points are reasonably spaced.
3. The data points are clearly marked (using small square points in this case).
4. The points are joined by lines (usually straight), which indicate a clear trend of the variables
concerned.
Note: In this case, the yearly rate at which the profits decrease changes, hence the slope (or gradient) of the graph
changes from year to year.
35 MANCOSA
Business Statistics
Example: If we take the previous example of the profits of Kings Plastics, we my plot the bar chart as follows:
8
6
4
2
0
15 25 35 45 55 65
Year
65
55
45
Year
35
25
15
0 5 10 15
Profit (R'000)
MANCOSA 36
Business Statistics
Example: Consider a father who gives spending money to each of his three sons. Josh, the eldest gets R 120; Matt
gets R 80 and David, the youngest gets R 50. This data can be expressed in a pie chart as follows:
Firstly, we calculate the percentage (in terms of the total amount the father gave out) that each son receives. The
total in this case is R 120 + R 80 + R 50 = R 250. The percentages are:
Josh: (120/250) x 100 = 48 %
Matt: (80/250) x 100 = 32 %
David: (50/250) x 100 = 20 %
Note: The sum of the three percentages must be 100 % (the full pie!)
37 MANCOSA
Business Statistics
3. Each segment must be fully labelled with the percentage it represents and what context it is used (in this case,
the sons’ names).
The advantage of pie charts and bar charts is the visual impact that they have in conveying information. Pie charts
are limited to a relatively small amount of data; with more data you will need to resort to a bar chart.
3.3.4 Histogram
A histogram is a graphic display of a frequency distribution, using a bar-like graph. Earlier, we constructed
a frequency table for the ages of 47 males. We can illustrate this information on a histogram as follows:
14
12
Number of Males
10
8
6
4
2
0
10 20 30 40 50 60 70 80
Ages (years)
The advantage of representing information in a histogram is the visual impact it has. It is quicker and easier to
see how ages are distributed.
Exercise: Use the grouped frequency distribution corresponding to the student marks and draw the
corresponding histogram. What conclusions can you draw from the histogram?
MANCOSA 38
Business Statistics
Cumulative
Interval Frequency
Frequency
0–9 1 1
10 – 19 2 3
20 – 29 5 8
30 – 39 6 14
40 – 49 29 43
50 – 59 34 77
60 – 69 12 89
70 - 79 7 96
80 – 89 3 99
90 – 99 1 100
15Figure 3.13: Cumulative Frequencies of Student Marks
Note: The width (interval size) of all intervals is 10. This is one of the main features of a grouped frequency
distribution. Also, intervals must never overlap. Once the grouped frequency table has been constructed, an OGIVE
curve can be drawn. An OGIVE curve is simply a line graph depicting the upper limit of each interval on the
horizontal (x) axis (for example, the upper limit of the interval 40 - 50 is just 49) against the cumulative frequency
on the vertical (y) axis.
120
100
Cumulative Number of Students
80
60
40
20
0
0 10 20 30 40 50 60 70 80 90 100 110
Student Marks
39 MANCOSA
Business Statistics
We can see from the table, or more easily from the graph, that 43 students (or 43%, since we have exactly 100
observations) had marks of 49% or less – so 57% of the class managed to achieve at least 50%.
Evident from Figure 3.14 is the very steep line between marks of 40 and 59; 14% of students had a mark of 40 or
less; 77% of students had a mark of 59 or less. Thus 63% of students marks thus fall in the interval of 40 to 59.
Again, one may question whether this is a good distribution of marks.
Another term associated with this type of analysis is the percentile, which, in our example, would refer to a mark
that a percentage of students have not achieved. For example, one can read from figure 3.14:
The 90th percentile is a mark of 69; thus 90% of students received a mark of 69% or less; or 10% of students
received a mark of at least 70
The 25th percentile is 45; 25% of students received 45 or less; or 75% of students received a mark of more
than 45
Also associated with this type of analysis is quartiles, which divide an ordered data-set into quarters.
The lower quartile (or 25th percentile) is that observation which separates the lower 25 percent of observation from
the top 75 percent of ordered observation.
The middle quartile (or 50th percentile) is the median. It divides an ordered data set into two equal halves.
The upper quartile (or 75th percentile) is that observation which separates the top 25 percent of observations from
the bottom 75 percent of ordered observations.
Activity 1
1. The table below gives the number of ocean-going vessels that arrived in
South African ports over a certain period.
MANCOSA 40
Business Statistics
Represent the data in the form of a frequency table, using three classes of
equal width.
3.5. Summary
In this unit we outlined different modes of displaying data and conveying the information from statistical analyses.
Charts such as the pie and bar charts vividly display data associated with qualitative (categorical) random variables’
examples, how data may be presented using bar charts, pie charts, histograms, line graphs, frequency tables and
ogive curves.
41 MANCOSA
Business Statistics
Answers to Activities
Unit 3
Activity 1
1.
3000 2346
2000 1602
802
1000 306
109 14
0
RB DBN EL PE MB CT SB
Port
1602
4127
9306
109
2346 802
14
306
RB DBN EL PE MB CT SB Total
2.
Interval Frequency
5 – 19 5
20 – 34 4
35 –49 1
3.
MANCOSA 42
Business Statistics
20
15
10
Y
0
0 2 4 6 8 10
X
43 MANCOSA
Business Statistics
Unit
4: Management Statistics
MANCOSA 44
Business Statistics
4.2. Measures of Central Location Understand, calculate and interpret the measures of
central location
Discuss the concept of skewness and measures of
central tendency
4.3. Other Measures of Central Location Discuss other measures of central location
45 MANCOSA
Business Statistics
4.1. Introduction
We saw in the previous chapter that graphical displays of statistical data are useful as a means of communicating
broad overviews of the behaviour of a random variable. However, there is a need for numerical measures (statistics)
about the behaviour pattern of a random variable.
After discussing these behaviour patterns, we will look at frequently used probability distribution functions in
business, the binomial and normal distributions. Commonly used index numbers will be demonstrated, after which
we will look at Sampling and Sampling Methods.
A measure of central tendency is a typical or representative score. If the mayor is asked to provide a single value
which best describes the income level of the city, he or she would answer with a measure of central tendency.
A central location statistic represents a typical value or middle data point of a set of observations and is useful for
comparing data sets.
The computation of each of these measures differs for ungrouped (or raw) data and grouped data (data
summarised into a frequency distribution). The latter is of utmost importance.
MANCOSA 46
Business Statistics
where:
n = sample size.
𝑥𝑖 = 𝑖 𝑡ℎ observation of random variable X.
n
n
i.e. x
i 1
i x1 x2 ...... xn
If we take our previous example of student marks (the raw data is shown in Figure 3.3), we obtain the following
value for the mean using the equation shown above:
𝑥̅ ≈ 51.3
Exercise: Show the full calculation and verify the above result.
An easier method of calculating the mean is to use the grouped data shown in Figure 4.1 below, where the midpoint
and frequency of observations for each interval is tabled.
10 – 19 2 14.5 29.0
20 – 29 5 24.5 122.5
30 – 39 6 34.5 207.0
40 – 49 29 44.5 1290.5
50 – 59 34 54.5 1853.0
60 – 69 12 64.5 774.0
47 MANCOSA
Business Statistics
70 – 79 7 74.5 521.5
80 – 89 3 84.5 253.5
90 – 99 1 94.5 94.5
∑ =100 ∑ =5150.0
17Figure 4.1: Grouped Data for Student Marks
With this grouped data, we can calculate the mean using the formula:
𝑘
1
𝑥̅ = ∑ 𝑓𝑖 𝑚𝑖
𝑛
𝑖=1
where:
𝑚𝑖 is the midpoint of the 𝑖 𝑡ℎ interval of frequency 𝑓𝑖 .
The average is:
5150
𝑥̅ = = 51.5
100
If we compare this to the value calculated from the raw data (51.3), we see that this method can give a very close
approximation.
4.2.2 Median
The median is the value of a random variable that divides an ordered data-set into two equal parts, i.e. half the
observations will fall below the median value and the other half above it.
For ungrouped or raw data set of n observations arranged in ascending order, there are two possibilities:
1. If n is an odd number, then the median will be the middle value of the ordered data set. i.e. the [(n+1)/2]
value is the median.
2. If n is an even number, then the median is the average of the middle two values of the ordered data set,
i.e. the average of the [n/2] and [(n/2)+1] values.
For grouped data, the median can be calculated using the following formula:
MANCOSA 48
Business Statistics
n
C Fbelow
Median L
2
f med
where:
L = lower bound of median class
C = class width
4.2.3 Mode
The mode is the most frequently occurring value in a set of data. It is seldom computed for ungrouped data, since
it is simply the value occurring the most number of times.
We can also calculate the mode for grouped data, using the formula:
𝐶(𝑓𝑚𝑜 − 𝑓𝑎𝑏𝑜𝑣𝑒 )
𝑀𝑜𝑑𝑒 = 𝐿 +
2𝑓𝑚𝑜 − 𝑓𝑎𝑏𝑜𝑣𝑒 − 𝑓𝑏𝑒𝑙𝑜𝑤
Where:
L = the lower limit of the modal class.
𝐶 = the class width.
𝑓𝑚𝑜 = the frequency of the modal class.
𝑓𝑏𝑒𝑙𝑜𝑤 = the frequency of the class before (below) the modal class.
𝑓𝑎𝑏𝑜𝑣𝑒 = the frequency of the class after (above) the modal class.
In our student marks example, the interval 50 < 60 has the most observations (34), and thus qualifies as the “modal
class”. We note 𝐿𝑚𝑜 = 50, 𝑓𝑚𝑜 = 34, 𝑓𝑏𝑒𝑙𝑜𝑤 = 29, 𝑓𝑎𝑏𝑜𝑣𝑒 = 12. Therefore:
10(34−29)
Mode = 50 + 2(34)−29−12 ≈ 51.9
49 MANCOSA
Business Statistics
An exception to this is the case of a bi-modal symmetrical distribution. In this case the mean and the median fall
at the same point, while the two modes correspond to the two highest points of the distribution.
Mode Mode
Mean = Median
18Figure 4.4: A bimodal frequency distribution
A positively skewed distribution is asymmetrical and points in the positive direction. If a test was very difficult and
almost everyone in the class did very poorly on it, the resulting distribution would most likely
be positively skewed.
MANCOSA 50
Business Statistics
In the case of a positively skewed distribution, the mode is smaller than the median, which is smaller than the
mean. This relationship exists because the mode is the point on the x-axis corresponding to the highest point, that
is the score with greatest value, or frequency. The median is the point on the x-axis that cuts the distribution in half,
such that 50% of the area falls on each side.
The mean is the balance point of the distribution. Because points further away from the balance point change the
centre of balance, the mean is pulled in the direction the distribution is skewed. For example, if the distribution is
positively skewed, the mean would be pulled in the direction of the skewness, or be pulled toward larger numbers.
One way to remember the order of the mean, median and mode in a skewed distribution is to remember that the
mean is pulled in the direction of the extreme scores. In a positively skewed distribution, the extreme scores are
larger, thus the mean is larger than the median.
A negatively skewed distribution is asymmetrical and points in the negative direction, such as would result with a
very easy test. On an easy test, almost all students would perform well and only a few would do poorly.
51 MANCOSA
Business Statistics
The order of the measures of central tendency would be the opposite of the positively skewed distribution, with the
mean being smaller than the median, which is smaller than the mode.
The choice of a representative central location value depends on the shape of the frequency distribution. If a
distribution is distorted by extreme values (i.e. skewed), then the median or the mode is more representative of the
distribution than the mean.
For a skewed distribution, the median may be the best measure of central location as it is not pulled by extreme
values (as the mean is), nor is it as highly influenced by the frequency of occurrence of a single value (as the mode
is).
MANCOSA 52
Business Statistics
Measures of dispersion provide useful information with which the reliability of the central value may be judged.
Widely dispersed observations indicate that the central value has low reliability, and does not represent the
observations very well. Conversely, a high concentration of observations about the central value indicates higher
reliability, with the central value being more representative.
4.4.1 Range
The range is the difference between the highest and lowest observed values in a data set. It is simply the largest
score minus the smallest score. It is a quick and dirty measure of variability, although when a test is given back to
students they very often wish to know the range of scores.
Because the range is greatly affected by extreme scores, it may give a distorted picture of the scores. The following
two distributions have the same range, 13, yet appear to differ greatly in the amount of variability.
Distribution 1 32 35 36 36 37 38 40 42 42 43 43 45
Distribution 2 32 32 33 33 33 34 34 34 34 34 35 45
For this reason, among others, the range is not the most important measure of variability.
Referring again to the example of student exam marks (see figure 3.3), we can use the ungrouped data to
calculate the range:
Maximum value = 90
Minimum value = 8
53 MANCOSA
Business Statistics
Range = 90 – 8 = 82
Obviously the range calculated from the grouped data is not as accurate a measure as that calculated with the raw
data; in this case, it is also a poor estimate. For larger sets of data, it is normally a much closer estimate.
The range is a crude estimate of spread. It is easily calculated, but is distorted by extreme values (“out-liers”) An
“out-lier” would be the minimum or maximum value. It is thus a volatile and unstable measure of dispersion as it
can vary greatly between samples taken from the same population. It also provides no information on the clustering
of observations within the data set about a central value as it uses only two observations (i.e. the maximum and
minimum) in its computation.
Figure 3.14 showed the cumulative frequency polygon (or OGIVE) for our student marks as an example.
From the figure, we can easily read off the 75th and 25th percentile. We obtain
75th percentile = 58
25th percentile = 45
Inter-quartile range = 58 – 45 = 13
This measure of dispersion removes much of the instability inherent in the range by excluding “out-liers”, but it
excludes 50 percent of all observations from further analysis. It also provides no information on the clustering of
observations within the data set as it uses only two observations (Q1 & Q3) in its calculation.
MANCOSA 54
Business Statistics
The quartile deviation is useful as a measure of dispersion if the sample of observations contains excessive
“outliers”, as it ignores the top 25% and bottom 25% of the ranked observations.
As with the inter-quartile range, the quartile deviation does not use all the observations and therefore gives no
indication of the spread of values between the upper and lower quartiles.
4.4.4 Variance
The variance has become the most used measure of dispersion, because it:
Takes every observation into account, and
Is based on an average deviation from a central value
Note that the variance could almost be the average squared deviation around the mean if the expression were
divided by n rather than n-1. It is divided by n-1, called the degrees of freedom, for theoretical reasons. If the mean
is known, as it must be to compute the numerator of the expression, then only n-1 scores that are free to vary. That
is if the mean and n-1 scores are known, then it is possible to figure out the nth score.
The formula for the variance presented above is a definitional formula, it defines what the variance means. The
variance may be computed from this formula, but in practice this is rarely done. It is done here to better describe
what the formula means. The computation is performed in a number of steps, which are presented below:
Step One - Find the mean of the scores.
55 MANCOSA
Business Statistics
Example: Consider the following simple example showing the ages of 7 second-hand cars:
13 7 10 15 12 18 9
13+7+10+15+12+18+9
𝑥̅ =
7
84
=
7
= 12 years
The calculation of the squared deviation of each observation from the sample mean is shown in Figure 4.7:
Car Ages, x ̅)
(𝒙 − 𝒙 ̅)𝟐
(𝒙 − 𝒙
13 1 1
7 -5 25
10 -2 4
15 3 9
12 0 0
18 6 36
9 -3 9
∑=0 ∑ = 84
21Figure 4.7: Calculation of Squared Deviation
MANCOSA 56
Business Statistics
Continuing with the car ages example, we can calculate the variance as follows:
Car age, x x2
13 169
7 49
10 100
15 225
12 144
18 324
9 81
∑ = 84 ∑= 1092
22Figure 4.8: Car ages –Variance Calculation
where
𝑚𝑖 = midpoint of i interval
𝑘
1
𝑥̅ = ∑(𝑓𝑖 𝑚𝑖 )
𝑛
𝑖=1
57 MANCOSA
Business Statistics
For student marks problem, the intermediate calculations for the variance are shown in Figure 4.9.
𝑆 = √𝑆 2
The standard deviation measures variability in units of measurement, while the variance does so in units of
measurement squared. For example, if one measured height in inches, then the standard deviation would be in
inches, while the variance would be in inches squared. For this reason, the standard deviation is usually the
preferred measure when describing the variability of distributions.
In our examples of exam marks, we can easily calculate the standard deviation.
Ungrouped data: 𝑆 = √210.1 ≃ 14.5
MANCOSA 58
Business Statistics
Coefficient of Variation
It is sometimes necessary to compare samples of data from different random variables to establish which sample
data shows greater variability. A direct comparison of their respective standard deviations would be misleading as
the random variables may be measured in different units.
The comparison would be more meaningful if the measures of variability were expressed in the same units. This
can be achieved by producing a measure of relative variability, i.e. relative to their mean, expressed in percentage
terms.
A statistic that shows this relative dispersion about a mean for a random variable is called the coefficient of variation,
𝑆
and is defined as: 𝐶𝑉 = 𝑥̅
× 100%.
A coefficient close to zero indicates low variability and a tight clustering of observations about the mean.
Conversely, a large coefficient of variation indicates that the observations are more spread about the mean value.
The coefficient of variation for the student marks, using the ungrouped data gives:
14.5
𝐶𝑉 = 51.3
× 100% = 28.3%
This low value indicates that, in spite of the large range of the data, the marks are generally tightly clustered around
the mean.
38 24 35 17 56
29 45 19 46 28
33 34 27 31 52
41 51 32 44 22
59 MANCOSA
Business Statistics
2. Group the data in a frequency distribution. Let the lower limit of the initial class be 10 faulty ATMs and use a
class width of 10 faulty ATMs.
3. From the grouped frequency distribution, determine each of the following for the 20-day period.
(i) mean number of faulty ATMs.
(ii) median. number of faulty ATMs
(iii) modal number of faulty ATMs
4. Determine the standard deviation and interpret its value.
5. Draw an ogive curve and use it to estimate the median. How does your estimate compare with result of 3 (ii)
above?
6. What type of data (discrete or continuous) is portrayed in the table? Explain.
2.
MANCOSA 60
Business Statistics
5.
20
15
Median ≈ 34 ATMs
10
0
0 10 20 30 40 50 60
Faulty ATMs
Activity 1
4.1 The net annual salary (in R’000s) for 20 clerks at an auditing firm is given
below.
4.1.1 Using the raw data, determine the range.
4.1.2 Group the data into a grouped frequency distribution with a lowest
class lower limit of R 120 000 and a class width of R10 000.
4.1.3 Determine the mean and mode using the raw data.
4.1.4 Draw an OGIVE curve corresponding to the data.
4.2 The number of hernia repair surgeries performed on patients of
different ages by a general surgeon over a six-month period is given in
the table below.
4.2.1 For these patients, determine:
(a) the mean age for hernia repair surgery.
(b) the median age for hernia surgery.
(c) the modal age for hernia surgery.
4.2.2 Determine the standard deviation.
61 MANCOSA
Business Statistics
4.6. Summary
Sample statistics serve to estimate population parameters and describe the data characteristics. Two categories
of statistics were described in this chapter, namely: measures of central tendency and measures of variability. In
the former category were the mean, median, and mode. In the latter were the range, interquartile range and
standard deviation. Measures of central tendency describe a typical or representative score, while measures of
variability describe the spread or dispersion of scores about a central measure.
MANCOSA 62
Business Statistics
Answers to Activities
Unit 4
Activity 1
1. Range = max – min = 176 – 121 = 55 ≡ R55000
2.
Class
Freq, f F
(R’000)
120 - 130 4 4
130 - 140 7 11
140 - 150 3 14
150 - 160 3 17
160 - 170 2 19
170 - 180 1 20
4.
20
15
10
0
110 120 130 140 150 160 170 180 190
Annual Salary (R'000)
4.1.
Class Freq, f F midpt, m fm fx2
10 < 20 4 4 15 60 900
63 MANCOSA
Business Statistics
MANCOSA 64
Business Statistics
Unit
5: Probability Distribution Functions
65 MANCOSA
Business Statistics
5.2. Binomial and Normal Probability Distinguish between the various distributions and calculate
distributions the associated probabilities
5.3. Sampling and Sampling Understand appropriate sampling techniques for obtaining
Distributions statistical data
MANCOSA 66
Business Statistics
5.1. Introduction
This unit focuses on three topics: probability distributions, sampling and index number.
A probability distribution is a list of all the possible outcomes of a random variable and their associated probabilities
of occurrence. There are numerous problem situations in practice where the outcomes of a specific random variable
follow known probability patterns. If the behaviour of a random variable can be matched to a known probability
pattern, then probabilities for the random variable can be found directly by applying an appropriate theoretical
probability distribution function.
We examine probability distributions for discrete and continuous random variables, specifically the binomial
distribution (for discrete variable) and the normal distribution (for continuous variables). These distributions are
described explicitly by theoretical functions that enable one to calculate the probability of occurrences of events.
Index numbers play an important role in economic activities. An index number is a summary measure of the
change in the level of activity of a single item or collection (often referred to as basket) of related items from one
time period to another. We look at the calculation and application of these numbers.
Sampling is the process of selecting a representative subset of observations from a population to determine
characteristics of the random variable under study.
Sampling methods may be classified as non-probability and probability methods. We also introduce concept of
sampling distribution illustrate sampling distribution of the mean by means of an example.
67 MANCOSA
Business Statistics
These can be summarized as an experiment with a fixed number of independent trials, each of which can only
have two possible outcomes. The fact that each trial is independent actually means that the probabilities remain
constant.
The binomial formula calculates the probability of r successes, and is stated as follows:
𝑛!
𝑃(𝑟) = 𝑝𝑟 𝑞 (𝑛−𝑟)
𝑟! (𝑛 − 𝑟)!
where
n = number of trials (observations0
r = 0, 1, 2… n = number of success outcomes in n trials
p = probability of success outcomes
MANCOSA 68
Business Statistics
Example: What is the probability of rolling exactly two sixes in 6 rolls of a die?
There are five things you need to do to work a binomial story problem.
1. Define Success first. Success must be for a single trial. Success = "Rolling a 6 on a single die"
2. Define the probability of success (p): p = 1/6
3. Find the probability of failure: q = 5/6
4. Define the number of trials: n = 6
5. Define the number of successes out of those trials: r = 2
Anytime a six appears, it is a success (denoted S) and anytime something else appears, it is a failure (denoted F).
The ways you can get exactly 2 successes in 6 trials are given below. The probability of each is written to the right
of the way it could occur. Because the trials are independent, the probability of the event (all six dice) is the product
of each probability of each outcome (die).
5 5 5 5 1 1 1 2 5 4
1. 𝐹𝐹𝐹𝐹𝑆𝑆 × × × × × =( ) ( )
6 6 6 6 6 6 6 6
5 5 5 1 5 1 1 2 5 4
2. 𝐹𝐹𝐹𝑆𝐹𝑆 × × × × × =( ) ( )
6 6 6 6 6 6 6 6
5 5 5 1 1 5 1 2 5 4
3. 𝐹𝐹𝐹𝑆𝑆𝐹 × × × × × =( ) ( )
6 6 6 6 6 6 6 6
5 5 1 5 5 1 1 2 5 4
4. 𝐹𝐹𝑆𝐹𝐹𝑆 × × × × × =( ) ( )
6 6 6 6 6 6 6 6
5 5 1 5 1 5 1 2 5 4
5. 𝐹𝐹𝑆𝐹𝑆𝐹 × × × × × =( ) ( )
6 6 6 6 6 6 6 6
5 5 1 1 5 5 1 2 5 4
6. 𝐹𝐹𝑆𝑆𝐹𝐹 × × × × × =( ) ( )
6 6 6 6 6 6 6 6
5 1 5 5 5 1 1 2 5 4
7. 𝐹𝑆𝐹𝐹𝐹𝑆 × × × × × =( ) ( )
6 6 6 6 6 6 6 6
5 1 5 5 1 5 1 2 5 4
8. 𝐹𝑆𝐹𝐹𝑆𝐹 × × × × × =( ) ( )
6 6 6 6 6 6 6 6
5 1 5 1 5 5 1 2 5 4
9. 𝐹𝑆𝐹𝑆𝐹𝐹 × × × × × =( ) ( )
6 6 6 6 6 6 6 6
4
5 1 1 5 5 5 1 2 5
10. 𝐹𝑆𝑆𝐹𝐹𝐹 × × × × × =( ) ( )
6 6 6 6 6 6 6 6
4
1 5 5 5 5 1 1 2 5
11. 𝑆𝐹𝐹𝐹𝐹𝑆 × × × × × =( ) ( )
6 6 6 6 6 6 6 6
69 MANCOSA
Business Statistics
1 5 5 5 1 5 1 2 5 4
12. 𝑆𝐹𝐹𝐹𝑆𝐹 × × × × × =( ) ( )
6 6 6 6 6 6 6 6
1 5 5 1 5 5 1 2 5 4
13. 𝑆𝐹𝐹𝑆𝐹𝐹 × × × × × =( ) ( )
6 6 6 6 6 6 6 6
1 5 1 5 5 5 1 2 5 4
14 𝑆𝐹𝑆𝐹𝐹𝐹 × × × × × =( ) ( )
6 6 6 6 6 6 6 6
4
1 1 5 5 5 5 1 2 5
15. 𝑆𝑆𝐹𝐹𝐹𝐹 × × × × × =( ) ( )
6 6 6 6 6 6 6 6
Further note that there are fifteen ways this can occur. This is the number of ways 2 successes can be occur in 6
trials without repetition and order not being important, or a combination of 6 things, 2 at a time. 6Hence, the
probability is 15 x 0.0134 = 0.201 (20.1 %).
With larger values of n and r, these calculations can become more elaborate.
Fortunately, most spreadsheets have the formula built in; for the above example, we could use the following formula
in Microsoft Excel: BINOMDIST(2, 5, 0.25, FALSE) ≃ 0.264.
MANCOSA 70
Business Statistics
There are three areas on a standard normal curve that all statistics students should know. The first is that the total
area below 0.0 is 0.50, as the standard normal curve is symmetrical like all normal curves. This result generalizes
to all normal curves in that the total area below the value of μ is 0.50 on any member of the family of normal curves.
The second area that should be memorized is between Z-scores of -1.00 and +1.00. It is 0.68 or 68%.
71 MANCOSA
Business Statistics
The total area between plus and minus one σ unit on any member of the family of normal curves is also 0.68.
The third area is between Z-scores of -2.00 and +2.00 and is 0.95 or 95%.
This area (0.95) also generalizes to plus and minus two σ units on any normal curve.
Knowing these areas allow computation of additional areas. For example, the area between a Z-score of 0.0 and
1.0 may be found by taking 1/2 the area between Z-scores of -1.0 and 1.0, because the distribution is symmetrical
between those two points. The answer in this case is 0.34 or 34%. A similar logic and answer is found for the area
between 0.0 and -1.0 because the standard normal distribution is symmetrical around the value of 0.0.
The area below a Z-score of 1.0 may be computed by adding 0.34 and 0.50 to get .84. The area above a Z-score
of 1.0 may now be computed by subtracting the area just obtained from the total area under the distribution (1.00),
giving a result of 1.00 - 0.84 or 0.16 or 16%.
The area between -2.0 and -1.0 requires additional computation. First, the area between 0.0 and -2.0 is 1/2 of 0.95
or 0.475. Because the 0.475 includes too much area, the area between 0.0 and -1.0 (0.34) must be subtracted i*n
order to obtain the desired result. The correct answer is 0.475 - 0.34 or 0.135.
MANCOSA 72
Business Statistics
Using a similar kind of logic to find the area between Z-scores of .5 and 1.0 will result in an incorrect answer
because the curve is not symmetrical around 0.5. The correct answer must be something less than 0.17, because
the desired area is on the smaller side of the total divided area.
If a data set follows a normal distribution, we can predict the frequency of data in intervals defined by the mean
and standard deviation, as follows:
𝜇−𝜎 < 𝑥 < 𝜇+𝜎 ∶ 68.2%
𝜇 − 2𝜎 < 𝑥 < 𝜇 + 2𝜎 ∶ 95.5%
𝜇 − 3𝜎 < 𝑥 < 𝜇 + 3𝜎 ∶ 99.7%
𝜇 = population mean
𝜎 = population standard deviation
If we look at our student marks, we can calculate the intervals and compare the count with the normal distribution
calculation.
Recall: For the student marks problem, mean = 51.3 and standard deviation = 14.5 (for the ungrouped data).
The table in figure 5.2 shows how the values predicted by the normal distribution compares with the actual
distribution (as determined from the ogive curve in figure.
Actual % by Normal
Interval
Distribution (%) Distribution
We observe that except for the one standard deviation range, the predictions are quite accurate – certainly
accurate enough for practical applications.
73 MANCOSA
Business Statistics
The values in the table (figure 5.3) lie between zero (mean) and given z-scores, i.e. P (0 < Z < z-score).
z 0.00 0.01 0.02 0.03 0.04 0.05 0.06 0.07 0.08 0.09
0.0 0.0000 0.0040 0.0080 0.0120 0.0160 0.0199 0.0239 0.0279 0.0319 0.0359
0.1 0.0398 0.0438 0.0478 0.0517 0.0557 0.0596 0.0636 0.0675 0.0714 0.0753
0.2 0.0793 0.0832 0.0871 0.0910 0.0948 0.0987 0.1026 0.1064 0.1103 0.1141
0.3 0.1179 0.1217 0.1255 0.1293 0.1331 0.1368 0.1406 0.1443 0.1480 0.1517
0.4 0.1554 0.1591 0.1628 0.1664 0.1700 0.1736 0.1772 0.1808 0.1844 0.1879
0.5 0.1915 0.1950 0.1985 0.2019 0.2054 0.2088 0.2123 0.2157 0.2190 0.2224
0.6 0.2257 0.2291 0.2324 0.2357 0.2389 0.2422 0.2454 0.2486 0.2517 0.2549
0.7 0.2580 0.2611 0.2642 0.2673 0.2704 0.2734 0.2764 0.2794 0.2823 0.2852
0.8 0.2881 0.2910 0.2939 0.2967 0.2995 0.3023 0.3051 0.3078 0.3106 0.3133
0.9 0.3159 0.3186 0.3212 0.3238 0.3264 0.3289 0.3315 0.3340 0.3365 0.3389
1,0 0.3413 0.3438 0.3461 0.3485 0.3508 0.3531 0.3554 0.3577 0.3599 0.3621
1.1 0.3643 0.3665 0.3686 0.3708 0.3729 0.3749 0.3770 0.3790 0.3810 0.3830
1.2 0.3849 0.3869 0.3888 0.3907 0.3925 0.3944 0.3962 0.3980 0.3997 0.4015
1.3 0.4032 0.4049 0.4066 0.4082 0.4099 0.4115 0.4131 0.4147 0.4162 0.4177
1.4 0.4192 0.4207 0.4222 0.4236 0.4251 0.4265 0.4279 0.4292 0.4306 0.4319
1.5 0.4332 0.4345 0.4357 0.4370 0.4382 0.4394 0.4406 0.4418 0.4429 0.4441
1.6 0.4452 0.4463 0.4474 0.4484 0.4495 0.4505 0.4515 0.4525 0.4535 0.4545
1.7 0.4554 0.4564 0.4573 0.4582 0.4591 0.4599 0.4608 0.4616 0.4625 0.4633
1.8 0.4641 0.4649 0.4656 0.4664 0.4671 0.4678 0.4686 0.4693 0.4699 0.4706
1.9 0.4713 0.4719 0.4726 0.4732 0.4738 0.4744 0.4750 0.4756 0.4761 0.4767
2,0 0.4772 0.4778 0.4783 0.4788 0.4793 0.4798 0.4803 0.4808 0.4812 0.4817
2.1 0.4821 0.4826 0.4830 0.4834 0.4838 0.4842 0.4846 0.4850 0.4854 0.4857
2.2 0.4861 0.4864 0.4868 0.4871 0.4875 0.4878 0.4881 0.4884 0.4887 0.4890
2.3 0.4893 0.4896 0.4898 0.4901 0.4904 0.4906 0.4909 0.4911 0.4913 0.4916
2.4 0.4918 0.4920 0.4922 0.4925 0.4927 0.4929 0.4931 0.4932 0.4934 0.4936
2.5 0.4938 0.4940 0.4941 0.4943 0.4945 0.4946 0.4948 0.4949 0.4951 0.4952
2.6 0.4953 0.4955 0.4956 0.4957 0.4959 0.4960 0.4961 0.4962 0.4963 0.4964
2.7 0.4965 0.4966 0.4967 0.4968 0.4969 0.4970 0.4971 0.4972 0.4973 0.4974
2.8 0.4974 0.4975 0.4976 0.4977 0.4977 0.4978 0.4979 0.4979 0.4980 0.4981
2.9 0.4981 0.4982 0.4982 0.4983 0.4984 0.4984 0.4985 0.4985 0.4986 0.4986
3.0 0.4987 0.4987 0.4987 0.4988 0.4988 0.4989 0.4989 0.4989 0.4990 0.4990
26Figure 5.3: Standard Normal Distribution Table
MANCOSA 74
Business Statistics
Some examples that can be read off from the tables are:
For z = 0, P (z > 0) = 0.5
For z = 1, P (0 < z < 1) = 0.3413 → P (z > 1) = 0.5 – 0.3413 = 0.1587
Negative z-values: Because the table is symmetrical, the area to the left of the negative value will be the same as
that to the right of the positive value. Thus P (0 < Z < -1) 0.3413; P (Z < -1) = P (Z > 1) = 0.1587
How do we determine the area (probability) between any two values z1 and z2 of z, i.e P (z1 < Z < z2) Since P(z1
< Z < z2) = P(z1 < Z < 0) + P(0 < Z < z2). We read off the areas for z1 and z2 from the z table and then take the
sum.
For example P (-1 < Z < 1) = P (-1 < Z < 0) + P (0< Z < 1) = 0.3413 + 0.3413 = 0.6826
where:
= population arithmetic mean
= population standard deviation
Let’s take another look at our example of student marks.
In this problem we have mean = 51.3, and standard deviation =14.5. Assuming the marks to be normally distributed,
then for any given value of x, the corresponding z score can be estimated using the transformation
𝑥 − 51.3
𝑧=
14.5
We can then use the z table to estimate the probabilities in Figure 5.2.
For example, for one standard deviation on either side of the mean, we have
(51.3 – 14.5) < x < (51.3 + 14.5) = 36.8 < x < 65.8. This can be converted to:
36.8−51.3 65.8−51.3
14.5
<𝑧< 14.5
→ −1 < 𝑧 < 1
Similarly, we can read off the areas for two and three standard deviations on both sides of the mean.
Two standard deviations:
22.3 < x < 80.3 or –2 < z < 2
This gives us an area of 2 × 0.4772 = 0.9544 (95.4%)
75 MANCOSA
Business Statistics
The Standard Normal Distribution table is most useful for determining the probable frequency of a variable between
any two limits, and can be used for any data set that follows a normal distribution. and with a reasonably good
accuracy an approximately normal distribution.
Example:
A test is normally distributed with a mean of 60 and a standard deviation of 10. What proportion of the scores is
above 85?
Solution:
This problem is very similar to figuring out the percentile rank of a person scoring 85. The first step is to figure out
the proportion of scores less than or equal to 85. This is done by figuring out how many standard deviations above
the mean 85 is. Since 85 is 85-60 = 25 points above the mean and since the standard deviation is 10, a score of
85 is 25/10 = 2.5 standard deviations above the mean. Or, in terms of the formula,
𝑥−𝜇 85−60
𝑧= 𝜎
= 10
= 2.5
From the z table, P (0 < z < 2.5) = 0.4938. Thus P(x > 85) = P (z > 2.5) = 0.5 – 0.4938 = .0062 (0.62%)
MANCOSA 76
Business Statistics
Situation Instructions
Between zero and any number. Look up the area in the table.
Between two positives, or between two negatives. Look up both areas and subtract smaller from larger.
Between a negative and a positive. Look up both areas and add them together.
Less than a negative or greater than a positive. Look up the area and subtract from 0.5000.
Greater than a negative or less than a positive. Look up the area and add to 0.5000.
Situation Instruction
Area between 0 and a value. Look up the area and make negative if in left tail.
Area including one complete half. (Less than Subtract 0.5000 from the area. Look up difference and make
a positive or greater than a negative.) negative if on the left side.
Two tails with equal area. (More than z units Subtract the area from 0.5000. Look up difference and use
from mean.) negative and positive scores.
77 MANCOSA
Business Statistics
A population comprises all possible observations of the random variable under study. Examples are:
All the residents in a suburb, town or city under study
The entire population of cell phon7e owners in the country
According to Wegner (2007), a measure that is found from analysing sample data is called a statistic, while a
measure describing a population is called a parameter. The various notations used for these measures are
shown in figure 5.6.
size n N
mean 𝑥̅ 𝜇
Standard Deviation 𝑠 𝜎
Proportion p 𝜋
28Figure 5.6: Notations for Samples and Populations
Most of the data used in managerial decision making is derived from a sample of observations. However, a
manager’s need is to know about the population parameter values of a random variable, not its sample statistics.
MANCOSA 78
Business Statistics
For example, a quality controller in a beer bottling process is more interested to know the population mean volume
of all bottles filled, rather than the sample mean volume of the few filled bottles drawn regularly from the production
line and tested.
Inferential statistics is that area of statistics that aims to estimate the true population parameters, with the following
process:
Draw a sample of observations on the random variable under study
Produce the appropriate sample statistics
Derive estimates of the values of the corresponding population parameters based on these simple statistics
The major disadvantage of non-probability sampling methods is the unrepresentative nature of the sample with
respect to the population from which it is drawn. Consequently, results from any statistical inference would probably
be invalid.
79 MANCOSA
Business Statistics
However, non-probability samples can be used in exploratory research to obtain initial impressions of the
characteristics of a random variable under study.
There are four types of probability sampling methods: Random, Systematic, Cluster and Stratified sampling.
Random sampling is analogous to putting everyone's name into a hat and drawing out several names.
Each element in the population has an equal chance of occurring. While this is the preferred way of
sampling, it is often difficult to do. It requires that a complete list of every element in the population be
obtained. Computer generated lists are often used with random sampling
Systematic sampling is easier to do than random sampling. In systematic sampling, the list of elements is
"counted off". That is, every kth element is taken. This is similar to lining everyone up and numbering off
"1,2….k; 1,2…k; etc.". When done numbering, all persons numbered k would be chosen
Cluster sampling is accomplished by dividing the population into groups -- usually geographically. These
groups are called clusters or blocks. The clusters are randomly selected, and each element in the selected
clusters are used
Stratified sampling also divides the population into groups called strata. However, this time it is by some
characteristic, not geographically. For instance, the population might be separated into males and
females. A sample is taken from each of these strata using either random, systematic, or convenience
sampling
Measures of sample statistics whose behaviour is generally described with respect to their corresponding
population parameters are:
Mean
Proportion
Difference between two means
Difference between two proportions
MANCOSA 80
Business Statistics
Let us look at the sampling distribution of a single sample mean through an example from Wegner (2007). We will
find the probability that a single sample mean lies within a certain distance of its unknown population.
Example: Assume that typing speed, measured in words per minute, is normally distributed. A random sample of
100 typists is selected and their typing speeds measured. Assume that the population deviation of typing speed is
8 words per minute.
What is the probability that the sample mean differs from the unknown population mean of typing speeds by no
more than one word per minute in either direction?
We are studying the behaviour of the sample mean with respect to its population mean. We can use the sampling
distribution of the sample means to find the required probability.
The standard deviation of sample means, also called the standard error, is calculated using the formula:
𝜎
𝜎𝑥̅ =
√𝑛
Irrespective of the population distribution, the distribution of sample means will always be normal, so we can use
the properties of the normal distribution to predict behaviour. The sampling distribution of the sample mean is
related to the standard normal probability distribution (the z-distribution) through the following transformation
formula:
𝑥̅ − 𝜇
𝑧=
𝜎𝑥̅
81 MANCOSA
Business Statistics
Reading between the tails from the table, the area between the two values of z is 0.7888.
Thus there is a 78.9% chance that a single sample mean of typing speeds will lie within 1 word per minute of the
true (but unknown) population mean of typing speeds. This is based on a sample size of 100 typists, and drawn
from a normal population of typing speeds, with a standard deviation of 8 words per minute. Alternatively, there is
a 21.1% chance that it will be outside one word per minute.
There are two major categories of index numbers – price and quantity. In both cases, a single or composite index
may be used.
A price index measures the percentage change in price between any two periods of time.
For a single item, the relative price change from one-time period to another is found by computing its price relative:
𝑝1
𝑃𝑟𝑖𝑐𝑒 𝑅𝑒𝑙𝑎𝑡𝑖𝑣𝑒 = × 100%
𝑝0
where
𝑝1 = current period price
𝑝0 = base period price
A quantity index measures the percentage change in consumption level of either an individual item or a basket of
items from one-time period to another.
MANCOSA 82
Business Statistics
For a single item, the relative quantity changes from one-time period to another is found by computing its quantity
relative.
𝑞1
𝑄𝑢𝑎𝑛𝑡𝑖𝑡𝑦 𝑅𝑒𝑙𝑎𝑡𝑖𝑣𝑒 = × 100%
𝑞0
where
𝑞1 = current period quantity
𝑞0 = base period quantity
Example: In the following share portfolio problem, the Laspeyres composite index is calculated for price and
quantity. The base year is 1986.
83 MANCOSA
Business Statistics
3. An executive usually replies to his e-mails fairly quickly. The mean time he takes to reply to his e-mails is 30
minutes, with a standard deviation of 6 minutes. Determine the following probabilities.
i. he takes between 18 and 42 minutes to answer an e-mail.
ii. he takes between 24 and 36 minutes to answer an e-mail.
Solutions :
1. (i) n = 7, r = 3, n – r = 4, p = 0.34, q = 1 - 0.34 = 0.66
7!
𝑃(3) = (0.34)3 (0.66)4 = 0.261 (26.1%)
3! 4!
(ii) n = 7, r = 4, n – r = 3, p = 0.38, q = 1 - 0.38 = 0.62
MANCOSA 84
Business Statistics
7!
𝑃(4) = (0.38)4 (0.62)3 = 0.174 (17.4%)
4! 3!
2. n = 7, r = 0, n – r = 7, p = 0.28, q = 0.72
7!
𝑃(0) = (0.28)3 (0.72)4 = 0.101
0! 7!
𝑃(𝑟 ≥ 1) = 1 − 𝑃(0) = 1 − 0.101 = 0.899 (89.9%)
3. The outcomes are mutually exclusive and collective exhaustive. Even though there are three possible
outcomes, the binomial distribution is applicable. This can be seen as follows. One of the 3 outcomes is the
desired outcome (successful outcome) with probability p. The remaining outcomes together constitute the
failure outcome with probability q which is the sum of the probabilities of the unsuccessful outcomes. Thus, p
+ q = 1 as required.
4. = 30 minutes, = 6 minutes,
18−30 42−30
(i) 𝑃(18 < 𝑥 < 42) = 𝑃( 6
< z< 6
) = 𝑃(−2 < 𝑧 < 2)
Activity 1
1. According to a survey, four out of ten South African drivers have
outstanding traffic fines. For a randomly-selected group comprising eight
South African drivers, what is the probability that at least six drivers have
outstanding traffic fines?
2. The mean monthly electricity bill for a complex of apartments is R1800.
Assuming that the electricity bills are normally distributed with a standard
deviation of R 250, approximately what percentage of these apartments
have monthly electricity bills in excess of R2000?
3. An automatic machine fills jars of jam with a mean net weight of 340 grams.
Assume a normal distribution with a standard deviation of 8 grams. What is
the probability that a randomly- selected jar of jam weights between 338
grams and 344 grams?
4. The life span of a particular brand of squash balls has a normal distribution
with a mean of 48 months and a standard deviation of 6 months. What
percentage of these squash balls last between 40 and 50 months?
85 MANCOSA
Business Statistics
5. The data in the table below shows the price (in Rand) and quantity of three
food items in 2011 and 2012
Using 2011 as a base year, calculate the Laspeyres price and quantity indices
5.6 Summary
In this unit we covered three important topics: probability distributions, index numbers and sampling.
We firstly examined the properties and applications of two theoretical probability distributions, namely the binomial
and normal distributions. The binomial distribution enables us to calculate the probability for any given value of a
binomial random variable using the binomial formula, while the normal distribution enables calculating the
probabilities for any given range of a continuous variable by using the standard normal distribution table. The
applications of these distributions were illustrated by means of examples.
We then looked at the construction and application of simple and composite index numbers, specifically the
Laspeyres price and quantity index numbers.
Finally, we introduced the concept of sampling. We outlined the different types of probability and non-probability
sampling methods. We also illustrated the concept of random distribution of the mean through an example.
MANCOSA 86
Business Statistics
Answers to Activities
Unit 5
Activity 1
1. p =4/10 = 0.4, q = 0.6
P (r ≥6) = P (6) + P (7) +P (8)
8! 8! 8!
= 6!(8−6)! (0.4)6 (0.6)2 + 7!(8−7)! (0.4)7 (0.6)1 + 8!(8−8)! (0.4)8 (0.6)0
3. x 1 = 338 g, x2 = 344 g
z1 = (338 – 340)/8 = - 0.25 z2 = (344 – 340)/8 = 0.50
P (338 < X < 344)) = P (-0.25 < Z < 0.50) = 0.0987 + 0.1915 = 0.2902
4. z = (x -µ)/σ
z1 = (40 – 48)/6 = -8/6 = -1.33
z2 = (50 – 48)/6 = 1/3 = 0.33
→ P (40 < X < 50) = P (-2.33 < Z < 0.33) = 0.4082 + 0.1293 = 0.5375 (≈53.8%)
5.
𝑝0 𝑞0 𝑝1 𝑞1 𝑝0 𝑞0 𝑝1 𝑞0 𝑝0 𝑞1
8.50 50 10.50 60 425 525 510
13.00 35 14.00 25 455 490 325
9.00 120 9.50 138 1080 1140 1242
∑=1960 ∑=2155 ∑=2077
87 MANCOSA
Business Statistics
Unit
6: Prediction
(Correlation and Regression)
MANCOSA 88
Business Statistics
6.4. Linear Regression and Perform linear regression and correlation analysis
Correlation Analysis
89 MANCOSA
Business Statistics
6.1. Introduction
When two variables are related, it is possible to predict the values on one variable from the values on the other
variable with better than chance accuracy. This Unit describes how these predictions are made and what can be
learned about the relationship between the variables by developing a prediction equation. It will be assumed that
the relationship between the two variables is linear. Although there are methods for making predictions when the
relationship is nonlinear, these methods are beyond the scope of this module. Regression and correlation analyses
are statistical methods that attempt to quantify and describe possible relationships between variables. This
relationship can assist with the prediction of unknown values of certain variables from known values of the related
variables. Regression analysis quantifies the underlying structural relationship between variables. Correlation
analysis determines the strength of this identified association.
The other random variable is termed the dependent variable (y). Values are not readily known and need to be
estimated from values of the independent variable (x). Given that the relationship is linear, the prediction problem
becomes one of finding the straight line that best fits the data. Since the terms "regression" and "prediction" are
synonymous, this line is called the regression line.
The table in figure 6.1 shows pairs of random variables, between which possible relationships exist.
MANCOSA 90
Business Statistics
Regression analysis aims to find a linear function i.e., a straight line that best fits the actual observations. A
straight-line graph is defined as follows:
𝑦̂ = 𝑎 + 𝑏𝑥
where:
𝑦̂ =estimated value of dependent variable
𝑥 = value of independent variable
𝑎 = y-intercept (where regression line cuts the y-axis)
𝑏 = slope of the regression line (for every unit change in x, y changes by b units).
Graphically the straight line (y = a + bx) may look as shown in figure 6.2:
It is however uncommon to find such a perfect straight-line relationship shown in figure 6.2.
We usually talk about the “best-fit” straight line, i.e. a line passing through as many of the data points as possible.
6.3 Scatterplot
A scatterplot is a graphical plot of the values of the independent and dependent variables. The independent
variables x is recorded along the horizontal axis and the dependent values y along the vertical axis. Pairs of x
and y observations are plotted in space.
A visual inspection of the likely relationship between the two variables x and y, as provided by a scatterplot, will
provide an initial insight into the likely regression and correlation analysis results.
91 MANCOSA
Business Statistics
If for example the data points are widely scattered and the range of y values is large for any given x value, then a
linear regression function will be of little value as an estimation function for y, and the correlation measure will
show almost no association. Examples of various scatterplots are shown in figure 6.3 below:
MANCOSA 92
Business Statistics
No Linear Relationship
The regression line is that line which minimises the sum of the squared deviations of the observations from the
fitted line. Without providing the derivation, the coefficients a and b that result from this “method of least squares”
are as follows:
n n n
n xi yi xi yi
b i 1 i 1 i 1
2
n
n
n x xi
2
i
i 1 i 1
n n
yi b xi
a i 1 i 1
The correlation coefficient most commonly used is Pearson’s correlation coefficient (r), which is calculated as
follows:
n n n
n xi yi xi yi
r i 1 i 1 i 1
n 2 n 2
n 2 n 2
n xi xi n yi yi
i 1 i 1 i 1 i 1
93 MANCOSA
Business Statistics
r is always between -1 and 1 inclusive. -1 means perfect negative linear correlation and +1 means
perfect positive linear correlation. 0 means a poor (or no) correlation
r has the same sign as the slope of the regression (best fit) line
r does not change if the independent (x) and dependent (y) variables are interchanged
r does not change if the scale on either variable is changed. You may multiply, divide, add, or subtract a
value to/from all the x-values or y-values without changing the value of r
A correlation does not necessarily imply a cause and effect relationship, merely an observed association.
Example: Most of South Africa’s power stations are coal fired. Assume a random sample of 10 power stations
was selected and their coal usage and electricity generated for 1992 was obtained. The data are shown in figure
6.4.
15 35
6 18
10 24
18 32
9 24
7 20
14 32
11 29
5 14
8 22
33Figure 6.4: Coal Usage and Electricity Generated
MANCOSA 94
Business Statistics
In this case, electricity generated is the dependent variable y and the coal usage the independent variable x.
A scatterplot of the data is shown below in Figure 6.5, along with the best fit line (dashed).
35
30
25
20
15
10
3 5 7 9 11 13 15 17 19
Coal Usage (megatons)
From the scatterplot, we can already see a strong linear (direct) relationship between coal usage x and electricity
generated. There is little dispersion, since the points lie near the best line fit.
When carrying out linear regression calculations it is useful to construct the table shown in figure 6.6.
x y x2 xy y2
6 18 36 108 324
9 24 81 216 576
7 20 49 140 400
5 14 25 70 196
8 22 64 176 484
95 MANCOSA
Business Statistics
Using the above formulae, we obtain the following values for b and a:
10 × 2818 − 103 × 250
𝑏= ≈ 1.52
10 × 1221 − 1032
250 − 1.52 × 103
𝑎= ≈ 9.37
10
We can therefore define the estimated regression line as:
𝑦̂ = 9.37 + 1.52𝑥
(5 ≤ 𝑥 ≤ 18)
This correlation coefficient is close to +1, hence the association between x and y is very strong and positive.
Values of x can therefore confidently be used to estimate values of y.
The regression line can be used to estimate values of y from known values of x, by substituting the given x value
into the regression equation.
For example, estimate the level of electricity that would be generated for 12 million tons of coal:
𝑦̂ = 9.37 + 1.52 × 12 = 27.61
Thus with 12 million tons of coal, 27.61 million kilowatt hours of electricity can be expected to be generated.
Note on extrapolation.
Extrapolation is the process of estimating values of y, using values of x which lie outside the domain x values
used in the construction of the regression line. In our example, valid estimates of y are produced only from values
within the interval 5 ≤ x ≤ 18.
If values of y are estimated outside the domain of x, the estimates can be unreliable as the relationship between
x and y outside these limits is unknown and may in fact be quite different to that which is defined within the
domain.
MANCOSA 96
Business Statistics
32 8
23 6
37 9
11 3
60 14
45 11
Solution:
1.
97 MANCOSA
Business Statistics
70
60
50
Bonus (R'000)
40
30
20
10
0
0 2 4 6 8 10 12 14 16
Service Period (years)
2.
x y xy x2 y2
8 32 256 64 1024
6 23 138 36 529
9 37 333 81 1369
3 11 33 9 121
6×2095−51×208
2. 𝑟 = ≈ 0.999
√(6×507−512 )(6×8668−2082 )
This value of r indicates a strong, direct linear relationship between number of years of service and bonus.
6×2095−51×208
3. 𝑏 = 6×507−512
≈ 4.4
208−4.45×51
𝑎= 6
≈ −3.16
MANCOSA 98
Business Statistics
(3 ≤ 𝑥 ≤ 14)
For y = 70, we can use to regression equation to find x.
3.16+70
𝑥= 4.45
= 16 years.
Activity 1
The monthly salary (in thousands of Rand) of 5 employees at ABC agencies
as well as the number of years of experience of each
Experience Monthly salary
(years) (R’000)
2 3
6 9
11 13
13 16
15 20
6.6 Summary
This unit focussed on linear regression and correlation. We distinguished between dependent and independent
variables and how they connected by the linear regression equation. We also looked how to estimate the strength
of the linear regression by means of the scatterplot and to calculate the strength using Pearson’s correlation
coefficient. Finally illustrated the application of linear regression and correlation by means of an practical example.
99 MANCOSA
Business Statistics
Answers to Activities
Unit 6
Activity 1
1.
16
14
12
10
8
6
4
2
0
0 5 10 15 20
Experience (years)
𝑛 ∑ 𝑥𝑦−∑ 𝑥 ∑ 𝑦
2. r =
√[𝑛 ∑ 𝑥 2 −(∑ 𝑥)2 ][𝑛 ∑ 𝑦 2 −(∑ 𝑦)2 ]
5×711−47×61
= = 0.99
√(5×555−47×47)(5×915−61×61)
3. y (x) = a + bx
𝑛 ∑ 𝑥𝑦−∑ 𝑥 ∑ 𝑦 5×711−47×61
b= = = 1.216
𝑛 ∑ 𝑥 2 −(∑ 𝑥)2 5×555−47×47
61 47
a= 5
− 0.774 × 5
= 0.774
MANCOSA 100
Business Statistics
Unit
7: Forecasting Methods Using
Time Series Analysis
101 MANCOSA
Business Statistics
7.2. Components of a Time Series Understand and state the principles of forecasting and time
series
State the principles of seasonality and trend
7.3. Trend Analysis using Moving Calculate a trend using moving averages, and illustrate it
averages on a graph
MANCOSA 102
Business Statistics
7.1. Introduction
Forecasting is an integral part of business management. The better the forecast, the better management will be
able to plan for the future. Although there are many methods for making forecasts, some are better suited than
others for particular situations. Forecasting is a critical function that needs to be done by businesses. It is needed
to assist us in financial planning determining staff levels and ordinary raw materials for production and other
business functions. The most common tool used for forecasting is time series analysis. It assumes that the actual
values of a random variable in a time series are influenced by a variety of environmental forces operating over
time. Time series analysis attempts to isolate and quantify the influence of these different environmental forces
operating on the time series into a number of different components.
Time series analysis assumes that four underlying forces individually and collectively determine the random
variables value in a time series in any time period. They are
Trend (T)
Cyclical Variations ( C )*
Seasonal Variations (S)
Random (irregular) variation (R )
103 MANCOSA
Business Statistics
The most common form of cycle is the business cycle between periods of relatively good economic activity to poor
economic activity. The causes of these are difficult to determine. Action by government, trade unions and world
organisations induce levels of pessimism and optimism into the economy which are reflected in changes in the
time series levels. Index numbers are used to describe cyclical fluctuations. An illustration of cycles is shown in
figure 7.2.
MANCOSA 104
Business Statistics
Sales
Summer
Winter
Spring Fall
Time (Quarterly)
The trend component is expressed in the active units of the variable we are looking at. The seasonal and cyclical
indices are, by definition, index numbers and expressed relative to the trend.
Mathematically, this is expressed as:
y T C S I
Statistical analysis can be used effectively to isolate the trend (T) and the seasonal (S) components, but is of less
value in quantifying the cyclical movements, and of no value in isolating irregular components.
105 MANCOSA
Business Statistics
We will examine statistical approaches to quantify Trend and Seasonal variation only. More sophisticated models
would be needed to isolate the other two components.
To illustrate the method, let’s say we sold 30 widgets during the month of June. We want to estimate what our sales
will be for July. Our best guess might be that we will sell 30 widgets during July – we have used a “one month
moving average” as our forecast.
When we want to forecast for August, we may want to take into account what happened during June and July. Let’s
say we had sales of 40 during July. If we took a two-month moving average, our forecast for August would be
( Actual) June ( Actual) July 30 40
( F / C ) August 35
2 2
What do we do for September? Let’s say the sales for August were 30. We now have a choice between 3 forecasts:
1-Month moving average:
( F / C ) Sept ( Actual ) Aug 30
The table in the figure 7.4 below demonstrates the calculation for each forecast, while figure 7.5 shows the results
graphically.
MANCOSA 106
Business Statistics
1 2 3 4 5 6 7 8 9 10 11 12
Month
107 MANCOSA
Business Statistics
The graph of past quarterly demand for Metro Movers is shown in figure 7.7. It is clear that summer demand is by
far the highest in every year and autumn demand is generally the lowest. The seasonal index measures how much
higher and how much lower. Figure 7.8 shows calculations of seasonal indices for the 16 available past demands.
(Note. Besides seasonality, it looks like there is a slight upward trend over the 16 quarters. We shall ignore the
trend for now.).
MANCOSA 108
Business Statistics
Spring 90 - -
Winter 120 - -
41Figure 7.8: Seasonal Index Calculations
109 MANCOSA
Business Statistics
The mean seasonal demand is a four-period moving average centred in the middle of a given season, that is a
month and a half into the season. It includes demands going back six months and forward six months from that
point. Thus, the first figure in column 3 is based on demands for the last one and a half months of spring 1997; and
all of summer, autumn and winter 1997; and the first one and a half months of spring 1998. So,
(90 / 2) 160 70 120 (130 / 2)
115
4
This is a bit cumbersome, but it ensures that no one season is weighted more heavily than any other. The seasonal
indices are shown rearranged by year and season in figure 7.9. The three values for each season need to somehow
be reduced to a single index. The index for autumn is steadily rising, from 0.61 to 0.73 to 0.96. That is not sufficient
reason to expect it to continue to rise, however, especially since the other seasons do not show trends. Thus, the
projections of the seasonal indices for 2001 are the means of each column.
(The future seasonal index is obtained by calculating the mean of corresponding indices for past years.)
Metro movers may now use the seasonal indices in fine-tuning its demand forecasts for each coming season. For
example, suppose that they expect to move 480 vans of goods next year based on projection of the mean of past
years’ demands. It would be naïve to divide 480 by 4 and project 120 vans in each season. Instead,
Divide 480/4 = 120 = average number of vans per season. This average is now multiplied by the mean seasonal
index.
Spring 2001: 120 x 0.85 = 102 vans
Summer 2001: 120 x 1.46 = 175 vans
Autumn 2001: 120 x 0.76 = 91 vans
Winter 2001: 120 x 0.93 = 112 vans
MANCOSA 110
Business Statistics
Activity 1
The number of tennis racquets sold by a sports store is recorded per quarter,
for the past three years. The results are presented in the table below.
Year Q1 Q2 Q3 Q4
Year 1 200 220 405 300
Year 2 190 240 540 298
Year 3 180 198 680 307
7.5 Summary
In this unit we considered the ratio-to-moving-average method of analysing a time series. We identified and
described the nature of the trend, cyclical, seasonal and irregular influences on the time series. Using the
multiplicative model, we used the technique of time series analysis to decompose the time series into its constituent
components. We examined seasonal components were considered. Trend component can be described by linear
regression and seasonal component by finding seasonal indexes using the method of centred moving averages.
The method was illustrated by means of an example.
111 MANCOSA
Business Statistics
Answers to Activities
Activity 1
1.
(Year, Quarter) Data Mean Seasonal Seasonal
Demand Index
(1, Q1) 200 - -
(1, Q2) 220 - -
(1, Q3) 405 280.0 1.45
(1, Q4) 300 281.3 1.07
(2, Q1) 190 300.6 0.63
(2, Q2) 240 317.3 0.76
(2, Q3) 540 315.8 1.71
(2, Q4) 298 309.3 0.96
(3, Q1) 180 321.5 0.56
(3, Q2) 198 340.1 0.58
(3, Q3) 680 - -
(3, Q4) 307 - -
Summary Table
Q1 Q2 Q3 Q4
year 1 - - 1.45 1.07
year 2 0.63 0.76 1.71 0.96
year 3 0.56 0.58 - -
Mean SI 0.60 0.67 1.58 1.03
2.
Year-4 quarterly forecasts:
Q1 225
Q2 251
Q3 593
Q4 386
MANCOSA 112
Business Statistics
Bibliography
Glyn Davis, Branko Pecar and Leonard Santana. Business Statistics using Excel: A first Course for
South African Students.; Oxford University Press Southern Africa (2017)
Mann, Prem S. (2004). Introductory Statistics. 5th Ed. John Wiley & Sons, Inc
Ross, Sheldon M. (2005). Introductory Statistics. 2nd Ed. Elsevier Academic Press
Sanders, Donald S. (1995). Statistics : A first course. 5th Ed. McGraw-Hill, Inc
Schonberger, Richard J. & KNOD, Edward M., Jr. (1985). Operations Management.
Serving the Customer. 3rd Ed. Homewood, Illinois: BPI Irwin
Stevenson, William J. (1999). Production/Operations Management. 6th Ed. Irwin: McGraw Hill
Wegner, Trevor. (2007). Applied Business Statistics. Methods and Applications. Kenwyn: Juta & Co, Ltd
Weiers, Ronald M (2005) Essentials of Business Statistics. Thomson Learning, Inc
Wisniewski, M. and STEAD R. (1996). Foundation Quantitative Methods for Business. London: Prentice
Hall (Chapter 13)
113 MANCOSA
Business Statistics
MANCOSA 114
Tables and graphs are effective for structuring and presenting data in an understandable way. Tables work well for displaying precise values and detailed comparisons, while graphs are valuable for visualizing trends and distributions quickly. The choice between them depends on the need for clarity in understanding complex relationships or providing detailed quantitative insights .
Understanding the nature of data and data collection methods is crucial in establishing accuracy and reliability in statistical findings because it ensures the data sourced is relevant, unbiased, and accurately represents the population or process being studied. Different data types and appropriate collection techniques, such as surveys or observational studies, help avoid sampling errors and increase confidence in the analytical results .
Time series analysis in forecasting provides the advantage of identifying patterns over time, such as trends and seasonality, which can improve the accuracy of future predictions . However, challenges include the need for large datasets to identify patterns accurately and the difficulty of accounting for unexpected events or changes in underlying processes that may disrupt established patterns .
Measures of central tendency, such as mean, median, and mode, simplify complex datasets by providing a single value representing a typical data point, which aids in summarizing and comparing data sets effectively . The choice between them involves considering data characteristics such as skewness and the presence of outliers; the mean is informative for symmetric distributions, the median is robust to skew and outliers, and the mode reflects the most frequent observation .
In business statistics, the binomial distribution is a discrete distribution used when an experiment or process has two potential outcomes, like success or failure, with fixed probabilities across trials . The normal distribution, in contrast, is a continuous distribution representing data that clusters around a mean with symmetrical decay towards the extremes, suitable for naturally occurring datasets like heights or test scores, where a myriad of small influences govern outcomes .
Learning outcomes in an educational module, such as business statistics, guide the instructional design by outlining the knowledge and skills students should acquire and be able to demonstrate. They facilitate focused learning and self-assessment, ensuring that educational objectives are met across various cognitive levels .
The least squares method is utilized to derive a best-fit straight line through a set of data points, which can then be used to make future predictions about the data . The Pearson correlation coefficient helps in determining the strength and direction of the linear relationship between two variables, providing insight into how closely the prediction model might fit the actual data .
Descriptive statistics focus on summarizing and organizing data to describe the characteristics of a dataset, such as mean, median, and mode, which are useful in situations like evaluating classroom performance or analyzing game statistics . In contrast, inferential statistics go beyond immediate data characteristics to make predictions or decisions about a broader population based on sample data, useful in scenarios like market research or policy formulation .
The standard normal distribution, characterized by a mean of 0 and a standard deviation of 1, is used in statistical analysis to calculate the probability of occurrences within a dataset. Probabilities are derived from the area under the curve, which is symmetrical and standardized, allowing for the use of Z-scores to determine the likelihood of a particular data point appearing within a certain range .
Variance measures the average squared deviation from the mean, providing a sense of data spread or variability. It is crucial for comparing data sets with different units or scales, though it is expressed in squared units which can be abstract . The standard deviation, being the square root of variance, provides a measure of spread in the original unit of data, offering more intuitive insights into how much data values deviate from the mean, which is vital for understanding variability .