100% found this document useful (2 votes)
406 views117 pages

Business Statistics Module Guide

The Business Statistics Module Guide provides a comprehensive overview of statistical techniques and their applications in business management. It covers essential topics such as data collection, presentation, management statistics, probability distributions, correlation, regression, and forecasting methods. The guide is structured into seven units, each with specific learning outcomes and assessment criteria to facilitate effective self-directed study.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
100% found this document useful (2 votes)
406 views117 pages

Business Statistics Module Guide

The Business Statistics Module Guide provides a comprehensive overview of statistical techniques and their applications in business management. It covers essential topics such as data collection, presentation, management statistics, probability distributions, correlation, regression, and forecasting methods. The guide is structured into seven units, each with specific learning outcomes and assessment criteria to facilitate effective self-directed study.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

BUSINESS STATISTICS

Module Guide

Copyright© 2022
MANCOSA
All rights reserved, no part of this book may be reproduced in any form or by any means, including photocopying machines,
without the written permission of the publisher. Please report all errors and omissions to the following email address:
modulefeedback@[Link]
This Module Guide,
Business Statistics (NQF level 6)
module guide will be used across the following programmes:

 Advanced Certificate in Management Studies


 Bachelor of Commerce in International Business
 Bachelor of Business Administration
 Bachelor of Commerce in Human Resource Management
 Bachelor of Commerce in Marketing Management
 Bachelor of Commerce in Supply Chain Management
BUSINESS STATISTICS

List of Contents ....................................................................................................................................................... 1

Preface.................................................................................................................................................................... 3

Unit 1: Introduction to Statistics............................................................................................................................. 10

Unit 2: he Nature of Data, Data Collection and Sources ....................................................................................... 15

Unit 3: Presentation of Data .................................................................................................................................. 28

Unit 4: Management Statistics .............................................................................................................................. 44

Unit 5: Probability Distribution Functions .............................................................................................................. 65

Unit 6: Prediction (Correlation and Regression) .................................................................................................... 88

Unit 7: Forecasting Methods Using Time Series Analysis................................................................................... 101

Bibliography ........................................................................................................................................................ 113

i
Business Statistics

List of Contents
List of Figures and Illustrations

Figure 2.1: Data Codes for Qualitative Data ..................................................................................................... 20

Figure 2.2: Example of ordinal-scaled data ....................................................................................................... 21

Figure 2.3: Example of Interval-Scaled Data ..................................................................................................... 21

Figure 3.1: Table depicting rank of University research libraries ....................................................................... 31

Figure 3.2: Grouped Frequency Distribution Showing Male Ages ..................................................................... 31

Figure 3.3: Student Marks ................................................................................................................................. 32

Figure 3.4: Tally of Student Marks .................................................................................................................... 33

Figure 3.5: Frequency table of Student Marks .................................................................................................. 34

Figure 3.6: Profit of Kings Plastic from 1997-2002 ............................................................................................ 35

Figure 3.7: Line Graph showing Kings Plastics Profits from 1997-2002 ........................................................... 35

Figure 3.8: Bar Graph (vertical) depicting Profits of Kings Plastics ................................................................... 36

Figure 3.9: Bar Graph (horizontal) depicting Profits of Kings Plastics ............................................................... 36

Figure 3.10: Pie Chart Showing Amounts Given to Sons .................................................................................. 37

Figure 3.12: Ages of 47 Males .......................................................................................................................... 38

Figure 3.13: Cumulative Frequencies of Student Marks ................................................................................... 39

Figure 3.14: Graph of Cumulative Frequencies of Student Marks .................................................................... 39

Figure 4.1: Grouped Data for Student Marks .................................................................................................... 48

Figure 4.4: A bimodal frequency distribution ..................................................................................................... 50

Figure 4.5: A Positively Skewed Distribution ..................................................................................................... 50

Figure 4.6: A Negatively Skewed Distribution ................................................................................................... 52

Figure 4.7: Calculation of Squared Deviation .................................................................................................... 56

Figure 4.8: Car ages –Variance Calculation...................................................................................................... 57

Figure 4.9: Intermediate Calculations for Variance ........................................................................................... 58

Figure 5.1: A Typical Normal Distribution .......................................................................................................... 71

Figure 5.2: Comparison of Actual and Predicted Frequency of Student Marks ................................................. 73

1 MANCOSA
Business Statistics

Figure 5.3: Standard Normal Distribution Table ................................................................................................ 74

Figure 5.4: Percentage of values on a normal distribution ................................................................................ 76

Figure 5.6: Notations for Samples and Populations .......................................................................................... 78

Figure 5.5: Share Portfolio Performance ........................................................................................................... 83

Figure 6.1: Relationships................................................................................................................................... 90

Figure 6.2: Straight-line Relationship ................................................................................................................ 91

Figure 6.3: Illustrations of Scatterplots .............................................................................................................. 93

Figure 6.4: Coal Usage and Electricity Generated ............................................................................................ 94

Figure 6.5: Scatterplot of Coal Usage and Electricity Generation ..................................................................... 95

Figure 6.6: Calculations for Linear Regression ................................................................................................. 96

Figure 7.1: Illustration of Trend ....................................................................................................................... 103

Figure 7.2: Illustration of Cycles ...................................................................................................................... 104

Figure 7.3: Illustration of Seasonal Fluctuations.............................................................................................. 105

Figure 7.5: Moving Average Plots ................................................................................................................... 107

Figure 7.7: Seasonal Demand history for Metro Movers ................................................................................. 109

Figure 7.8: Seasonal Index Calculations ......................................................................................................... 109

Figure 7.9: Summary and Projection of Seasonal Indices............................................................................... 110

MANCOSA 2
Business Statistics

Preface
A. Welcome
Dear Student
It is a great pleasure to welcome you to Business Statistics (BS6). To make sure that you share our passion
about this area of study, we encourage you to read this overview thoroughly. Refer to it as often as you need to
since it will certainly make studying this module a lot easier. The intention of this module is to develop both your
confidence and proficiency in this module.

The field of Business Statistics is extremely dynamic and challenging. The learning content, activities and self-
study questions contained in this guide will therefore provide you with opportunities to explore the latest
developments in this field and help you to discover the field of Business Statistics as it is practiced today.

This is a distance-learning module. Since you do not have a tutor standing next to you while you study, you need
to apply self-discipline. You will have the opportunity to collaborate with each other via social media tools. Your
study skills will include self-direction and responsibility. However, you will gain a lot from the experience! These
study skills will contribute to your life skills, which will help you to succeed in all areas of life.

Statistics is all around us. It would be difficult to go through a full week without using statistics. So, what is statistics?
Statistics is the science of designing studies, gathering data, and then classifying, summarising, interpreting, and
presenting these data to support the decisions that are needed.

The word "statistics" is used in several different senses. In the broadest sense, "statistics" refers to a range of
techniques and procedures for analysing data, interpreting data, displaying data, and making decisions based on
data. This is what courses in "statistics" generally cover.

In a second usage, a "statistic" is defined as a numerical quantity (such as the mean) calculated in a sample. Such
statistics are used to estimate parameters. A parameter is a characteristic of a population e.g. the mean. A statistic
is a measure based on sample data.

Without statistics, we could not plan our budgets, pay our taxes, enjoy games to their fullest, and evaluate
classroom performance. Are you beginning to get the picture? We need statistics.

Let us take a look at the most basic form of statistics, known as descriptive statistics.

3 MANCOSA
Business Statistics

This branch of statistics lays the foundation for all statistical knowledge (important, huh?), but it is not something
that you should learn simply so you can use it in the distant future. Descriptive statistics can be used, in a language
class, in a science class, at the football stadium, in the grocery store. You probably already know more about these
statistics than you think.

Another form of statistics is called inferential statistics, which deals with inferences (decision making, predictions,
etc.) about the process or population being studied.

This module has been divided into five Units and at the beginning of each Unit you will find a list of Learning
Outcomes. These provide you with an outline of what you should have learned by the time you have completed
the Unit, and can also be used to focus your study.

The concepts are best learned and understood through practice, and we thus suggest that you do all the exercises
recommended throughout the guide, as well as carefully going through all the worked examples. A good suggestion
is trying to work out the given examples without looking at the solution, then comparing your answers with the
model answers. In this way, you can keep track of your progress.

We hope you enjoy the module.

MANCOSA does not own or purport to own, unless explicitly stated otherwise, any intellectual property rights in or
to multimedia used or provided in this module guide. Such multimedia is copyrighted by the respective creators
thereto and used by MANCOSA for educational purposes only. Should you wish to use copyrighted material from
this guide for purposes of your own that extend beyond fair dealing/use, you must obtain permission from the
copyright owner.

MANCOSA 4
Business Statistics

B. Module Overview
This Module is a 15 credit Module at NQF level 6.
Contents and Structure
Unit 1: Introduction
In the introduction, we explain why it is important to be able to manipulate numbers and have a feel for magnitude.
We also discuss how best to carry out the exercises in the text, and what basic pre-work is needed.

Unit 2: The Nature of Data and Data Collection


This Unit deals with “data”, which is raw information. We discuss the different types of data that is available to us,
and where to source it, as well as techniques for data collection.

Unit 3: Presentation of Data


We discuss how to structure and present data in an understandable format, and which formats are appropriate to
use. In particular, the use of tables and graphs are discussed.

Unit 4: Management Statistics


In this Unit, statistical measures and behaviours of random variables are presented and analysed.

Unit 5: Probability distributions


This Unit of the module deals with the Binomial (discrete) and Normal (continuous) distribution functions. The
difference between each distribution is discussed and probabilities are calculated using these two distribution
functions.

Unit 6: Prediction (Correlation and Regression)


Here we discuss the least squares method which is essentially deriving a straight line graph to fit a set of data
points and use the regression line to make predictions. In addition, we discuss the strength of the relationship
between two variables using the Pearson correlation co-efficient.

Unit 7: Forecasting Methods using Time Series Analysis


Different forecasting techniques and their merits are demonstrated, as well as moving averages.

Access to a Computer
Some of the exercises given to you in this module will be arduous to do without a personal computer. Before doing
the exercises on a computer, make sure you are able to apply the techniques manually, to help you better
understand the calculations carried out by the computer.

5 MANCOSA
Business Statistics

C. Learning Outcomes and Associated Assessment Criteria of the Module

LEARNING OUTCOMES OF THE MODULE ASSOCIATED ASSESSMENT CRITERIA OF THE MODULE

 Understand the importance of statistical  Statistical measures and methods of calculating and
techniques and calculations in business ascertaining them are extensively explored to provide a
management sound quantitative grounding for competence and
understanding

 Perform statistical analyses and extract  Principles, concepts and techniques of data gathering,
relevant information from business data classification, summarisation, interpretation and
presentation/communication are demonstrated and
critiqued to enhance understanding and appreciation of
objective statistical practices

 Manipulate collected data through various  Methods and techniques of data organisation and
statistical methods to generate useful processing are applied to real life scenarios to promote
information to support management understanding of the importance of producing unbiased,
decisions reliable research outcomes

 Prepare and interpret reports in statistical  Data processing and presentation methods and techniques
terms are examined to consolidate skills in the creation and
presentation of clear, objective and reliable statistical
reports

 Assess the validity of statistical findings  Statistical tests and variability measures are thoroughly
and the relevance and reliability of results examined to enhance understanding and competence in
their application and use to obtain unchallengeable
statistical results and inferences

D. Learning Outcomes and the of the Units


You will find the Unit Learning Outcomes on the introductory pages of each Unit in the Module Guide. The Unit
Learning Outcomes lists is an overview of the areas you must demonstrate knowledge in and the practical skills
you must be able to achieve at the end of each Unit lesson in the Module Guide.

E. How to Use this Module


This Module Guide was compiled to help you work through your units and textbook for this module, by breaking
your studies into manageable parts. The Module Guide gives you extra theory and explanations where necessary,
and so enables you to get the most from your module.

MANCOSA 6
Business Statistics

The purpose of the Module Guide is to allow you the opportunity to integrate the theoretical concepts from the
prescribed textbook and recommended readings. We suggest that you briefly skim read through the entire guide
to get an overview of its contents. At the beginning of each Unit, you will find a list of Learning Outcomes and
Associated Assessment Criteria. This outlines the main points that you should understand when you have
completed the Unit/s. Do not attempt to read and study everything at once. Each study session should be 90
minutes without a break

This module should be studied using the prescribed and recommended textbooks/readings and the relevant
sections of this Module Guide. You must read about the topic that you intend to study in the appropriate section
before you start reading the textbook in detail. Ensure that you make your own notes as you work through both the
textbook and this module. In the event that you do not have the prescribed and recommended textbooks/readings,
you must make use of any other source that deals with the sections in this module. If you want to do further reading,
and want to obtain publications that were used as source documents when we wrote this guide, you should look at
the reference list and the bibliography at the end of the Module Guide. In addition, at the end of each Unit there
may be link to the PowerPoint presentation and other useful reading.

F. Study Material
The study material for this module includes tutorial letters, programme handbook, this Module Guide, a list of
prescribed and recommended textbooks/readings which may be supplemented by additional readings.

G. Prescribed and Recommended Textbook/Readings


There is at least one prescribed and recommended textbooks/readings allocated for the module.

The prescribed and recommended textbooks/readings for this module is:


 Business Statistics using Excel: A first Course for South African Students
 Glyn Davis, Branko Pecar and Leonard Santana; Oxford University Press Southern Africa (2017)

H. Special Features
In the Module Guide, you will find the following icons together with a description. These are designed to help you
study. It is imperative that you work through them as they also provide guidelines for examination purposes.

Special Feature Icon Explanation

LEARNING The Learning Outcomes indicate aspects of the particular Unit


OUTCOMES you have to master.

7 MANCOSA
Business Statistics

The Associated Assessment Criteria is the evaluation of the


ASSOCIATED students’ understanding which are aligned to the outcomes.
ASSESSMENT The Associated Assessment Criteria sets the standard for the
CRITERIA successful demonstration of the understanding of a concept or
skill.

A Think Point asks you to stop and think about an issue.


THINK POINT Sometimes you are asked to apply a concept to your own
experience or to think of an example.

You may come across Activities that ask you to carry out
specific tasks. In most cases, there are no right or wrong
ACTIVITY
answers to these activities. The purpose of the activities is to
give you an opportunity to apply what you have learned.

At this point, you should read the references supplied. If you


are unable to acquire the suggested readings, then you are
READINGS
welcome to consult any current source that deals with the
subject.

PRACTICAL
Practical Application or Examples will be discussed to enhance
APPLICATION OR
understanding of this module.
EXAMPLES

You may come across Knowledge Check Questions at the end


KNOWLEDGE
of each Unit in the form of Knowledge Check Questions
CHECK
(KCQ’s) that will test your knowledge. You should refer to the
QUESTIONS
Module Guide or your textbook(s) for the answers.

You may come across Revision Questions that test your


REVISION understanding of what you have learned so far. These may be
QUESTIONS attempted with the aid of your textbooks, journal articles and
Module Guide.

Case Studies are included in different sections in this Module


CASE STUDY Guide. This activity provides students with the opportunity to
apply theory to practice.

MANCOSA 8
Business Statistics

You may come across links to Videos Activities as well as


VIDEO ACTIVITY
instructions on activities to attend to after watching the video.

9 MANCOSA
Business Statistics

Unit
1: Introduction to Statistics

MANCOSA 10
Business Statistics

Unit Learning Outcomes

CONTENT LIST LEARNING OUTCOMES OF THIS UNIT:

1.1. The Principles and Methods for  Understand the scope of analytical techniques
Designing Studies

1.2. Numerical and Computer Literacy  Discuss the importance of displaying numerical and
computer literacy

Prescribed and Recommended Textbooks/Readings


 Business Statistics using Excel: A first Course for South African
Students
 Glyn Davis, Branko Pecar and Leonard Santana; Oxford University
Press Southern Africa (2017)

11 MANCOSA
Business Statistics

1.1. Principles and Methods for Designing Studies


A statistic is an algebraic expression combining scores into a single number. Statistics serve two functions: they
estimate parameters in population models and they describe the data.
1. Collecting data
2. Presenting and analysing data
3. Interpreting the results

Statistics has been described as:


 Turning data into information
 Data-based decision making
 The technology of the "Scientific Method"

The scientific approach to decision making can be summarised as:


Hypothesis Data Conclusion

You are probably asking yourself the question, "When and where will I use statistics?" If you read any newspaper
or watch television, or use the Internet, you will see statistical information. There are statistics about crime, sports,
education, politics, and real estate. Typically, when you read a newspaper article or watch a news program on
television, you are given sample information. With this information, you may make a decision about the correctness
of a statement, claim, or "fact." Statistical methods can help you make the "best educated guess."

Since you will undoubtedly be given statistical information at some point in your life, you need to know some
techniques to analyze the information thoughtfully. Think about buying a house or managing a budget. Think about
your chosen profession. The fields of economics, business, psychology, education, biology, law, computer science,
police science, and early childhood development require at least one course in statistics.

Think back to some of the business conversations you have had during the last week or two, and the type of
information that has been offered during these conversations. More than likely, you will have heard expressions
like:
 The cost of that material is R1500-00
 80% of those that responded thought it was a good idea
 it took him two hours to complete the task
 4500 Toyota Corollas were sold in August

You would not be blamed for feeling a little frustrated with these expressions, because they don’t really give you
good information. In each case, we are missing:

MANCOSA 12
Business Statistics

 a basis of comparison
 the source of the data
 how it was collected
 why it was collected in the first instance

If these criteria were in place, we would have some information to assist us in our decision-making.
For example, knowing that 4500 Toyota Corollas were sold in August means nothing to me if I am selling Ford
Escorts. But if I see that in July, 3000 Toyota Corollas were sold, and the sales for Ford Escorts dropped from 2000
to 500 in August, I now have a better understanding of the significance of that data. The reason for the drop in
sales of the Ford Escorts may be due to the increase in Toyota sales. I now have information on which to make a
decision on what to do next.

In this module, we will look at some of the techniques that can be applied by business managers to collect, analyse
and interpret quantitative information to make informed decisions.

1.2. Numerical and Computer Literacy


Any manager who does not have numerate proficiency and computer literacy will be at a disadvantage in today’s
business environment. It is important that you are able to do basic arithmetic manipulation of numbers and that you
have a feel for magnitude. What do we mean by a feel for magnitude? It means you are able to recognise a number
that is obviously wrong. Imagine if through a finger error, you calculated that your budget to replace ten new
personal computers in your department was R8000-00 instead of R80000-00. If you didn’t notice the one zero
missing before submitting your budget, you will be very embarrassed when the time to effect the purchase arrives.

13 MANCOSA
Business Statistics

Answers to Activities

Unit 1
Activity 1
(a) Quantitative, ratio
(b) Quantitative, ratio
(c) Quantitative, ratio
(d) Quantitative, ratio
(e) Qualitative, nominal
(f) Qualitative, nominal
(g) Quantitative, ratio
(h) Quantitative, ratio
(i) Qualitative, nominal
(j) Qualitative, nominal
(k) Quantitative, ratio
(l) Quantitative, ratio
(m) Quantitative, ratio
(n) Quantitative, ratio
(o) Quantitative, ratio
(p) Qualitative, nominal

MANCOSA 14
Business Statistics

Unit
2: The Nature of Data, Data Collection
and Sources

15 MANCOSA
Business Statistics

Unit Learning Outcomes

CONTENT LIST LEARNING OUTCOMES OF THIS UNIT:

2.1. Introduction  Introduce topic areas for the unit

2.2. Internal and External Data  Discuss internal, external, primary and secondary sources
Sources

2.3. Data Types  Identify and discuss the various data types

2.4. Data Collection Methods  Discuss and describe observation, interviews and experiments as
data collection methods for statistical analysis

2.5. Summary  Summarises topic areas in the unit

Prescribed and Recommended Textbooks/Readings


Business Statistics using Excel: A first Course for South African Students.
Glyn Davis, Branko Pecar and Leonard Santana; Oxford University Press
Southern Africa (2017).

MANCOSA 16
Business Statistics

2.1. Introduction
In this unit we examine the factors that influence the quality of data on which important management decisions are
based. Data quality is influenced by the types of data available for analysis, the sources from which data are
collected and the methods by which data are collected. The types of data available determine the appropriate types
of statistical techniques to employ, while the sources from which data are gathered and the methods of collecting
data determine the accuracy and reliability of statistical findings.

2.2. Internal and External Data Sources


When one is confronted with a situation in which you need data, it is often surprising how much is actually available
to us. Newspapers, magazines and the Internet, as well as internal company records have a wealth of data that
can be put to good use. Data sources may be classified as:
 Internal
 External
 Primary
 Secondary

2.2.1 Internal Data Sources


Within any organisation, internal data is generated during the course of normal business activities. Examples
include:
Financial Data – sales vouchers, credit notes, accounts receivable.
Production Data – monthly production, defect rates, WIP levels.
Human Resource Data – time sheets, staff demographics, wage and salary schedules.
Marketing Data – monthly sales, advertising expenditure, customer profiles.

2.2.2 External Data Sources


Sources for data external to an organisation may be private institutions, trade/ employer/ employee associations,
profit motivated organisations and government bodies.

The cost of the external data depends on the source, but you may be surprised how much information is freely
available, either on the Internet or in business publications. A detailed study of the economic indicators in financial
publications will tell you a surprising amount. Statistics SA, the government’s source of statistical data, has virtually
all its data available on its home page on the internet. Many regard new motor vehicle sales as a good indicator of
the economy – and these are published for all NAAMSA members monthly.

Private sources of information include:


 South African Chamber of Business (SACOB)

17 MANCOSA
Business Statistics

 Business Partners (previously Small Business Development Corporation)


 Industrial Development Corporation (IDC)
 Bureau of Economic Research
 Bureau of Market Research
 Bureau of Financial Analysis
 SA Labour Development Research Unit

Public Domain sources include:


 Newspapers, journals, trade magazines
 Reference libraries
 Bank economic reports
 Human Sciences Research Council (HSRC)
 Council for Scientific and Industrial Research (CSIR)

2.2.3 Primary Data Sources


Primary data - collected by the researcher himself, and is captured at the point where it is generated for the first
time, normally with a specific purpose in mind. Examples include salary surveys and market research surveys.
Advantages of primary data:
 Directly relevant to the problem at hand
 Generally, offers greater control over data accuracy

Disadvantages of primary data:


 Time consuming to collect
 Generally, more expensive to collect (ask a market research company for a quote!)

2.2.4 Secondary Data Sources


Secondary data is data which has been collected by individuals or agencies for purposes other than those of our
particular research study. For example, if a government department has conducted a survey of, say, family food
expenditures, and then a food manufacturer might use this data in the organization's evaluations of the total
potential market for a new product. Similarly, statistics prepared by a ministry on agricultural production will prove
useful to a whole host of people and organizations, including those marketing agricultural supplies. Such data is
already in existence either within or outside an organisation.

Examples:
 “Aged” market research figures
 Previous financial statements

MANCOSA 18
Business Statistics

 An industry market research from which you are extracting data for your company

Advantages of secondary data:


 Data already in existence
 Access time is relatively short
 Generally, less expensive to acquire

Disadvantages of secondary data:


 May not be problem specific or entirely relevant to your situation
 May be dated and hence inappropriate
 May be difficult to determine the data accuracy
 May not be suitable for further manipulation
 Previously manipulated by the time you receive it; you have no way of knowing how reliable the
data is what has been omitted or what has been extrapolated, which could lead you to wrong
conclusions. i.e. has the data been "massaged"

2.3. Data Types


Wegner (2007) suggests two reasons why an understanding of the nature of data is necessary:
 To assess data quality
 To select the appropriate statistical method to use to analyse the data

The type of data gathered determines the type of analysis which can be performed; an incorrect application of a
statistical method to a particular data type can render the findings invalid. Data type is determined by the nature of
the random variable that the data represents.

Wegner (2007) identifies two types of random variables: qualitative and quantitative.

Quantitative data measures either how much or how many of something, i.e. a set of observations where any single
observation is a number that represents an amount or a count. Quantitative random variables yield numeric
responses, and can be meaningfully manipulated using conventional arithmetic operations. Examples are age,
distance, number of items, monetary amount, etc.

Qualitative data provide labels, or names, for categories of like items, i.e. a set of observations where any single
observation is a word or code that represents a class or category. Qualitative random variables yield categorical or
non-numeric responses. The data generated are classified into one of a number of categories.

19 MANCOSA
Business Statistics

The categories are usually represented by codes, which cannot be manipulated arithmetically. These codes are
merely used as labels. Figure 2.1 shows an example of data codes.

Random Variable Response Categories Data Codes

Management level Supervisor 1


Unit Head 2
Department Head 3
General Manager 4

Wine Preference Yes 2


(Do you like red wine?) No 1

1Figure 2.1: Data Codes for Qualitative Data

Each of these random variable categories can be associated with a different type of data classification.
Wegner (2007) defines two data classification types:
 Data type 1 - Nominal–scaled
- Ordinal–scaled
- Interval–scaled
- Ratio–scaled
 Data type 2 - Discrete
- Continuous

Nominal-scaled data: Data with no inherent order or ranking sequence, e.g. numbers used as names (group 1,
group 2), gender, etc. Nominal–scaled data is associated mainly with qualitative random variables. There is no
implied ordering between groups of the random variable, and each category is of equal importance. Figure 2.1 is
a good example

Ordinal-scaled data: Data with an ordered series, e.g. "greatly dislike, moderately dislike, indifferent, moderately
like, greatly like". Numbers assigned to such data indicate rank order only - the "distance" between the numbers
has no meaning. Ordinal-scaled data is also associated mainly with qualitative random variables. Like nominal-
scaled data, it is also assigned to one of a number of coded categories, but there is now a ranking implied between
the categories in terms of being better, bigger, longer, older, taller or stronger, etc. An example is shown in Figure

MANCOSA 20
Business Statistics

Random Variable Response Categories Data Codes

T-shirt size Small / medium / large 1, 2, 3

Turnover <5m / 5-10m / >10m 1, 2, 3


2Figure 2.2: Example of ordinal-scaled data

Interval-scaled data: Equally-spaced data, e.g. temperature. The difference between a temperature of 66 degrees
and 67 degrees is taken to be the same as the difference between 76 degrees and 77 degrees. Interval variables
do not have a true zero, e.g. 88 degrees is not necessarily double the temperature of 44 degrees. Interval-scaled
data is associated with quantitative random variables; differences can be measured between values. Interval-
scaled data possesses both order (implied ranking) and distance properties.

In social research studies, such as market research, the Likert Rating Scale is often used for respondents to
indicate a preference or a perception on a scale; interval-scaled properties are created for the study. An example
is shown in Figure 2.3

Example: Indicate your response to the statement “Shopping is a social experience for me.” (Tick a value from the
rating scale.)
Strongly Disagree Disagree Unsure Agree Strongly Agree
1 2 3 4 5
3Figure 2.3: Example of Interval-Scaled Data

The data does not contain an absolute origin, so the ratio of values cannot be meaningfully compared. A rating of
4 in the above example would reflect a stronger perception than a 2 (this is a property of order); it is not, however,
possible to conclude that a rating of 4 is twice as important as a rating of 2. But it is possible to conclude that the
difference in perception (or preference) between 3 and 4 is the same as between 1 and 2. This is a property of
distance.

Ratio-Scaled Data: is mainly associated with quantitative random variables; it is numeric data with a zero origin.
Examples are age, distance, time, mass, sales, units and income. Such data is the strongest form of statistical data
that can be gathered and lends itself to the widest range of statistical methods.

Ratio-scaled data is gathered through a measurement process, and can be manipulated meaningfully through
normal arithmetic operations. If ratio-scaled data is grouped into categories, that data becomes ordinal-scaled; an
example is the data for “Turnover” in Figure 2.2.

21 MANCOSA
Business Statistics

2.3.1 Discrete Data


Variables whose observations can take on only specific values, usually only integer (whole numbers) values, are
referred to as discrete.

Discrete variables are usually obtained by counting. There are a finite or countable number of choices available
with discrete data. You can't have 2.63 people in the room.
Examples:
 The number of students in a class
 The number of cars sold in a month by a dealer

2.3.2 Continuous Data


A random variable whose observation can take on any value in an interval is said to generate continuous data, and
that any value between a lower and upper limit is valid.

Continuous variables are usually obtained by measuring. Length, weight, and time are all examples of continuous
variables. Since continuous variables are real numbers, we usually round them. This implies a boundary depending
on the number of decimal places. For example, the measurement x = 64 is really anything in the range 63.5 < x <
64.5. Likewise, if there are two decimal places, then x = 64.03 is really anything in the range 63.025 < x < 63.035.
Boundaries always have one more decimal place than the data and end in a 5.

Examples:
 time taken to travel to work daily
 tensile strength of steel
 speed of an aircraft

MANCOSA 22
Business Statistics

2.4. Data Collection Methods


Wegner (2007) suggests three approaches to gathering data for statistical analyses:
 Observation
 Interview
 Experimentation

2.4.1 Observation
The use of observation as a measurement tool, assigning numerals to human behavioral acts, is discussed.
Observation has important advantages which makes it best suited for certain kinds of studies, and some limitations
which preclude its use in others. The central problems in the use of observation are:
(1) The effect of the observer on the observed, which is usually not severe and can be minimized.
(2) Observer inference, which is a crucial strength and a crucial weakness.
(3) The unit of behavior to be used, which involves the molar-molecular problem.

The considerations in planning both unstructured and structured observation studies are discussed, including what
to observe, how to record it, how to maximize validity and reliability, and how to handle the relationship between
the observer and the observed. Behavior is usually sampled using event sampling or time sampling.

Primary data can be collected by direct observation of the respondent or item in action. Examples include:
 Vehicle traffic surveys
 Observing the purchase behaviour of brands in a store
 Quality control inspection

An advantage of direct observation is that the respondent is generally unaware of being observed and therefore
behaves naturally. This reduces the likelihood of gathering biased data.

A disadvantage is that it is a passive form of data collection, and there is little opportunity to probe for reasons or
investigate behaviour further.

Secondary data can be obtained through desk research (abstraction), from a variety of source documents. A wide
variety of organisations and individuals continually consult and use secondary data for decision making or opinion
forming.

Observation is a technique that involves systematically selecting, watching and recording the behaviour and
characteristics of living beings, objects or phenomena.

23 MANCOSA
Business Statistics

Observation of human behaviour is a much-used data collection technique. It can be undertaken in different ways:
 Participant observation: The observer participates in the situation he or she observes. (For example, a
doctor hospitalised with a broken hip, who now observes hospital procedures ‘from within’.)
 Non-participant observation: The observer watches the situation, openly or concealed, but does not
participate

Observations can be open (e.g., ‘shadowing’ a health worker with his/her permission during routine activities) or
concealed (e.g., ‘mystery clients’ trying to obtain antibiotics without medical prescription). They may serve different
purposes. Observations can give additional, more accurate information on behaviour of people than interviews or
questionnaires. They can also check on the information collected through interviews especially on sensitive topics
such as alcohol or drug use, or stigmatising diseases. For example, whether community members share drinks or
food with patients suffering from feared diseases (leprosy, TB, AIDS) are essential observations in a study on
stigma.

Observations of human behaviour can form part of any type of study, but as they are time consuming they are most
often used in small-scale studies.

Observations can also be made on objects. For example, the presence or absence of a latrine and its state of
cleanliness may be observed. Here observation would be the major research technique.

Observations made using a defined scale are called measurements. Measurements usually require additional tools.
For example, in nutritional surveillance weight and height are measured by using weighing scales and a measuring
board. We use thermometers for measuring body temperature.

2.4.2 Interview
An Interview is a data-collection technique that involves oral questioning of respondents, either individually or as a
group. Interviews can be conducted through direct questioning or a questionnaire. Interview data can be gathered
through personal (face-to-face) interviews, postal surveys and telephone surveys.

Answers to the questions posed during an interview can be recorded by writing them down (either during the
interview itself or immediately after the interview) or by tape-recording the responses, or by a combination of both.
Interviews can be conducted with varying degrees of flexibility. The two extremes, high and low degree of flexibility,
are described below:

High degree of flexibility:


When studying sensitive issues such as teenage pregnancy and abortions, the investigator may use a list of topics
rather than fixed questions.

MANCOSA 24
Business Statistics

These may, e.g., include how teenagers started sexual intercourse, the responsibility girls and their partners take
to prevent pregnancy (if at all), and the actions they take in the event of unwanted pregnancies. The investigator
should have an additional list of topics ready when the respondent falls silent, (e.g., when asked about abortion
methods used, who made the decision and who paid). The sequence of topics should be determined by the flow
of discussion. It is often possible to come back to a topic discussed earlier in a later stage of the interview.

The unstructured or loosely structured method of asking questions can be used for interviewing individuals as well
as groups of key informants.

A flexible method of interviewing is useful if a researcher has as yet little understanding of the problem or situation
he is investigating, or if the topic is sensitive. It is frequently applied in exploratory studies. The instrument used
may be called an interview guide or interview schedule.

Low degree of flexibility:


Less flexible methods of interviewing are useful when the researcher is relatively knowledgeable about expected
answers or when the number of respondents being interviewed is relatively large. Then questionnaires may be
used with a fixed list of questions in a standard sequence, which have mainly fixed or pre-categorised answers.
Example: After a number of observations on the (hygienic) behaviour of women drawing water at a well and
some key informant interviews on the use and maintenance of the wells, one may conduct a larger survey on
water use and satisfaction with the quantity and quality of the water.

Though in principle one may speak of loosely structured questionnaires, in practice the term questionnaire appears
to be so hooked to tools with pre-categorised answers that we have decided to use the term interview guide for
loosely structured tools. However, in reality there is often a mixture of open and pre-categorised answers. In that
case we will still use the term questionnaire.

A written questionnaire (also referred to as self-administered questionnaire) is a data collection tool in which written
questions are presented that are to be answered by the respondents in written form.

A written questionnaire can be administered in different ways, such as by:


 Sending questionnaires by mail with clear instructions on how to answer the questions and asking for
mailed responses
 Gathering all or part of the respondents in one place at one time, giving oral or written instructions, and
letting the respondents fill out the questionnaires; or
 Hand-delivering questionnaires to respondents and collecting them later

25 MANCOSA
Business Statistics

Personal interviews have the advantage that accurate data can be obtained immediately, and qualitative data can
be obtained by probing for reasons and observing non-verbal responses. They are, however, time consuming, and
expensive if trained interviewers are required.

Telephone interviews allow more flexibility, in that call-backs can be made if a respondent is not available initially,
and people are more willing to talk on the telephone from the security of their home or an office. It is more cost
effective, as a larger sample of respondents can be reached in a relatively short time. The main disadvantage of
telephone interviewing is that non-verbal responses cannot be observed.

Postal surveys (they can be conducted by fax or by e-mail as well) are best used when the target population is
large and/or geographically dispersed.

A larger sample of respondents can be reached, making them more cost effective. Because respondents can
answer anonymously, more honest, considered responses would be given. However, questions have to be shorter
and simpler, and the possibility of probing is limited. Data collection can take a long time, and there is no control
over who answers the questionnaire, or the possibility of check-backs on the validity of responses.

According to Wegner (2007) the response rates of postal surveys are very low (5% - 15%). The questionnaire is
the data collection instrument used to gather data in all interview situations. The design of the questionnaire is
critical to ensure that the correct research questions are addressed and that accurate and appropriate data are
collected.

2.4.3 Experimentation
Primary data can also be generated through the manipulation of variables under controlled conditions. Data on the
primary variable under study can be monitored and recorded while conscious efforts are made by the researcher
to control the effects of a number of influencing factors.

Examples:
 The hardness of toughened glass can be measured for various tempering furnace temperatures
 Advertising effectiveness can be measured by manipulating the frequency and choice of various media

While good quality data is collected if the experiment is correctly designed and executed, experimentation is a
costly and time consuming exercise. It may also be impossible to control certain extraneous factors that can distort
the results.

MANCOSA 26
Business Statistics

Activity 1
For each of the following variables, indicate the data type and the measurement
scale (i.e. nominal, ordinal, interval or ratio).
(a) The shelf life of milk
(b) The number of life policies issued per day
(c) The area of a shop floor
(d) The number of pages in a text book
(e) The flavours available in dog-food chunks
(f) The wood types that can be used to make a desk
(g) The size categories for shoes
(h) The voltage produced by a generator
(i) The car types in the Mercedes range
(j) The Yes/No/Sometimes response to “Do you drink Gin?”
(k) The number of loaves of bread sold daily by a bakery
(l) The income per day of a bakery
(m) The monthly birth-rate at a maternity hospital
(n) The mass of babies at birth
(o) The daily distance travelled by a courier service truck
(p) The names of teams in a cricket league.

Activity 2
Study the statistics printed in newspapers, magazines and on the Internet. The
more you study them, the more information you will start obtaining from them.
You will also start picking up trends.

2.5. Summary
In this unit we examined data as the raw material of statistics. We distinguished internal and external data sources,
and primary and secondary data. We also looked at different data types and different types of measurement scales.
Finally, we covered the methods of data collection, namely: observation, interview and experimentation.

27 MANCOSA
Business Statistics

Unit
3: Presentation of Data

MANCOSA 28
Business Statistics

Unit Learning Outcomes

CONTENT LIST LEARNING OUTCOMES OF THIS UNIT:

3.1. Introduction  Introduce topic areas for the unit

3.2. Tables in Business  Discuss tables as part of the data presentation process

 Construct various tables for given data sets

3.3. Use of Graphs and Charts  Discuss graphs and charts as part of the data presentation
process
 Construct various graphs and charts

3.4. Grouped Frequency Distributions  Construct grouped frequency distributions and cumulative
grouped frequency distributions

3.5. Summary  Summarises topic areas in the unit

Prescribed and Recommended Textbooks/Readings:


 Business Statistics using Excel: A first Course for South African
Students
 Glyn Davis, Branko Pecar and Leonard Santana; Oxford University
Press Southern Africa (2017)

29 MANCOSA
Business Statistics

3.1. Introduction
The information obtained from a statistical analysis is meaningful to business managers only when it can be
interpreted and communicated effectively and concisely. It is customary to convey such information through the
use of summary tables and charts. Tables and charts convey information more vividly and quickly than written
reports.

3.2. Tables in Business


The major difference between data and information is that the former is presented in an ordered format. One such
compact and efficient way of presenting data is in the form of a table.

Normally, the first stage of ordering is to present the data in a table. Once data is in the form of a table, we can:
 Make comparisons, within the table and with other data
 Perform additional calculations
 Examine the component structure of the data

Using a table to list data according to category is often much clearer than writing out all the information in paragraph
form. Let's look at an example of some data first written up in a paragraph. Try to think what information is being
given and what sort of trends one could find from the information.

Example: During the 1995-1996 academic year, a survey of the holdings of university research libraries and rank
was done in the United States and Canada. It was found that Syracuse University, in New York, had 2,692,147
holdings, and was figured to rank eighty-first. Harvard University ranked first with 13,369,855 holdings. The
University of Connecticut was ranked fiftieth place, and reported 2,626,066 holdings. The Massachusetts Institute
of Technology reported 2,448,647 holdings, and was ranked in seventy-third place. (Source: Association of
Research Libraries).

As you can see, the paragraph above contains a lot of numbers and is not always easy to follow. The information
given in the paragraph would be easier to decipher if it was presented in a table. To create a table, you need to
determine the following things:
 Title of the table
 Label of each row and/or column
 Number of rows and columns necessary
 Data entry for each cell

MANCOSA 30
Business Statistics

Holdings and Rank of University Research Libraries in the U.S. And Canada--1995-1996.

Institution Rank Holdings

Harvard Univ. 1 13369855

U. Connecticut 50 2626066

Mass. Inst. Tech. 73 2448647

Syracuse Univ. 81 2692147


4Figure 3.1: Table depicting rank of University research libraries

Now the statistics are talking to us!

Notice that the order of the universities in the table is different from the order they are listed in the paragraph.
When moving from the paragraph to the table, it is best to order the instances by any numerical data. In this
case, we ordered them by rank.

From the information in the table, we can see that the number one ranked university contains a lot more holdings
than the other three institutions.

We carry out these processes to convert the data into information with which we are able to make good
decisions.

A commonly used table in business is a frequency table, which shows the number of occurrences of a variable
falling into a specific range or category.

Example: Consider a group of 47 males of various ages. 12 are between 20 and 29 years of age, 13 are between
30 and 39 years of age, 7 are between 40 and 49 years of age, 8 are between 50 and 59 years of age while the
rest are between 60 and 69 years of age.

This data can be presented in a frequency table as follows:

Interval (years) 20 – 29 30 - 39 40 - 49 50 - 59 60 - 69

Number of Males 12 13 7 8 7
5Figure 3.2: Grouped Frequency Distribution Showing Male Ages

31 MANCOSA
Business Statistics

Immediately we can make the following conclusions:


1. Majority of the males are “young”, i.e. below 40 years of age.
2. Most lie between 30 and 39 years of age.

Note: The sum of the number of men in each interval (i.e. 12+13+7+8+7) must equal to the total number in the
group which is 47 in this case.

Example: Consider a class consisting of 100 students. Suppose the teacher gives the entire class a statistics test
which has a maximum mark of 100. Upon marking the scripts (which are in no order whatsoever), he puts the
marks into a table as shown below. This represents raw data since there is no set order of the marks.

28 60 58 63 72 52 63 82 65 58

75 59 26 52 71 55 55 55 52 62

90 58 25 51 68 47 52 56 55 35

80 41 47 49 52 48 44 64 56 18

35 42 38 48 53 45 48 62 57 52

12 44 85 47 45 41 75 51 51 48

65 46 76 46 46 40 25 52 48 56

50 54 74 36 32 50 66 53 46 48

55 51 72 24 8 51 50 48 42 47

45 35 65 56 44 60 55 49 40 50

6Figure 3.3: Student Marks

On examining the data, we see the following:


Maximum value = 90
Minimum value =8
Range (max – min) = 82

Since the highest possible mark in this case is 100, and the lowest mark is 0, an interval size (width) of 10 is easy
to work with. Hence, we may choose to use the following intervals:

MANCOSA 32
Business Statistics

Interval

Lower limit Upper limit

0 9

10 19

20 29

30 39

40 49

50 59

60 69

70 79

80 89

90 99

Now that we have decided on the intervals, we can do a frequency count, which we do by registering each mark
in the correct interval. For example, we would register the first value, 28, in the interval 20 - 29; we would tally the
rest as shown in figure 3.4.

Interval Tally
0–9 /
10 – 19 //
20 – 29 /////
30 – 39 //////
40 – 49 /////////////////////////////
50 – 59 //////////////////////////////////
60 – 69 ////////////
70 – 79 ///////
80 – 89 ///
90 – 99 /
7Figure 3.4: Tally of Student Marks

Note: This function is easily done on a spreadsheet – if you are unsure how to do it, look up “frequency distribution”
in the HELP facility. When the tallying is complete, you can construct a frequency table as shown in figure 3.5.

33 MANCOSA
Business Statistics

Cumulative
Interval Frequency
Frequency

0–9 1 1

10 – 19 2 3

20 – 29 5 8

30 – 39 6 14

40 – 49 29 43

50 – 59 34 77

60 – 69 12 89

70 – 79 7 96

80 – 89 3 99

90 – 99 1 100
8Figure 3.5: Frequency table of Student Marks

The column Cumulative Frequency refers to the total number of observations/measurements encountered up
until a particular interval. This frequency table could be modified to show ratios as percentages if this is preferred.

3.3. Use of Graphs and Charts


A table is a useful way to present detailed data, but a picture or diagram is more powerful if we want to focus
attention on a particular aspect, such as a dominant feature or a trend. There are many different types of graphs
we can use to portray data. The important graphs we will discuss are the line graph, bar graph and pie graph.

3.3.1 Line Graphs


One method of showing trends or comparative trends is to use a line graph. A good example is one that is used to
show the growth in terms of profits of a particular organization.

Example: The table in figure 3.6 shows the total profit made by Kings Plastics for six consecutive years, from 1997
– 2002.

MANCOSA 34
Business Statistics

Year Profit (R’000)

1997 78

1998 65

1999 53

2000 49

2001 38

2002 16
9Figure 3.6: Profit of Kings Plastic from 1997-2002

A study of the table shows that the profits have dropped considerably. This trend can be better depicted using a
line graph (Figure 3.7) as shown on the following page.

100
80
Profit (R'000)

60
40
20
0
1996 1997 1998 1999 2000 2001 2002 2003
Year

10Figure 3.7: Line Graph showing Kings Plastics Profits from 1997-2002

The line graph clearly shows the drop in profits from 1997 to 2002.
Note the features of a line graph:
1. The vertical (y) and horizontal (x) axes are perpendicular to each other.
2. An appropriate scale is used such that the data points are reasonably spaced.
3. The data points are clearly marked (using small square points in this case).
4. The points are joined by lines (usually straight), which indicate a clear trend of the variables
concerned.

Note: In this case, the yearly rate at which the profits decrease changes, hence the slope (or gradient) of the graph
changes from year to year.

35 MANCOSA
Business Statistics

3.3.2 Bar Chart


The two most commonly used charts for business presentations are bar charts and pie charts – both of these very
clearly and simply convey a large amount of information. We will look at the bar chart first. A bar chart consists of
a series of bars, the length of each bar representing the value of the variable being plotted. The bars can be drawn
either vertically or horizontally.

Example: If we take the previous example of the profits of Kings Plastics, we my plot the bar chart as follows:

Profits of Kings Plastics (1997 - 2002)


14
12
10
Profit (R'000)

8
6
4
2
0
15 25 35 45 55 65
Year

11Figure 3.8: Bar Graph (vertical) depicting Profits of Kings Plastics

The corresponding horizontal bar chart would look like:

65

55

45
Year

35

25

15

0 5 10 15
Profit (R'000)

12Figure 3.9: Bar Graph (horizontal) depicting Profits of Kings Plastics

MANCOSA 36
Business Statistics

Note the features of a bar chart:


1. The width of each bar must be the same (we use length to represent magnitude of the value, not width).
2. An appropriate scale must be used on each axis such that the lengths of the bars are reasonable.
3. The distance between each bar must be kept constant to give the bar chart uniformity.
4. The bars may or may not be coloured (this is purely up to the person drawing the bar chart).

3.3.3 Pie Chart:


A pie chart is usually used when proportions are to be depicted relative to a whole. In essence, it is a circle divided
into segments, with the size of each segment proportional to the value of the variable, relative to the whole, and is
usually expressed in percentage terms. To illustrate the use of a pie cart, consider the following example.

Example: Consider a father who gives spending money to each of his three sons. Josh, the eldest gets R 120; Matt
gets R 80 and David, the youngest gets R 50. This data can be expressed in a pie chart as follows:
Firstly, we calculate the percentage (in terms of the total amount the father gave out) that each son receives. The
total in this case is R 120 + R 80 + R 50 = R 250. The percentages are:
Josh: (120/250) x 100 = 48 %
Matt: (80/250) x 100 = 32 %
David: (50/250) x 100 = 20 %
Note: The sum of the three percentages must be 100 % (the full pie!)

13Figure 3.10: Pie Chart Showing Amounts Given to Sons

Note the features of a pie chart:


1. The segments (or slices) of the pie must be an accurate reflection of the percentage values. To be totally
accurate, you can use angles (with 3600) as the total and express each slice in terms of the angle.
2. The slices may or may not be coloured (although colouring each segment in a different colour looks nice!).

37 MANCOSA
Business Statistics

3. Each segment must be fully labelled with the percentage it represents and what context it is used (in this case,
the sons’ names).

The advantage of pie charts and bar charts is the visual impact that they have in conveying information. Pie charts
are limited to a relatively small amount of data; with more data you will need to resort to a bar chart.

3.3.4 Histogram
A histogram is a graphic display of a frequency distribution, using a bar-like graph. Earlier, we constructed
a frequency table for the ages of 47 males. We can illustrate this information on a histogram as follows:

14
12
Number of Males

10
8
6
4
2
0
10 20 30 40 50 60 70 80
Ages (years)

14Figure 3.12: Ages of 47 Males

The advantage of representing information in a histogram is the visual impact it has. It is quicker and easier to
see how ages are distributed.

Exercise: Use the grouped frequency distribution corresponding to the student marks and draw the
corresponding histogram. What conclusions can you draw from the histogram?

3.4. Grouped Frequency Distribution


Cumulative frequencies are useful for determining the portion or percentage of observations that fall below or above
a given value. In our example of student marks, we could, for example, ask what percentage of students achieved
more than 50 %. A table of the cumulative frequencies is shown in figure 3.13, while the information is shown
graphically in figure 3.14.

MANCOSA 38
Business Statistics

Cumulative
Interval Frequency
Frequency

0–9 1 1

10 – 19 2 3

20 – 29 5 8

30 – 39 6 14

40 – 49 29 43

50 – 59 34 77

60 – 69 12 89

70 - 79 7 96

80 – 89 3 99

90 – 99 1 100
15Figure 3.13: Cumulative Frequencies of Student Marks

Note: The width (interval size) of all intervals is 10. This is one of the main features of a grouped frequency
distribution. Also, intervals must never overlap. Once the grouped frequency table has been constructed, an OGIVE
curve can be drawn. An OGIVE curve is simply a line graph depicting the upper limit of each interval on the
horizontal (x) axis (for example, the upper limit of the interval 40 - 50 is just 49) against the cumulative frequency
on the vertical (y) axis.

120

100
Cumulative Number of Students

80

60

40

20

0
0 10 20 30 40 50 60 70 80 90 100 110
Student Marks

16Figure 3.14: Graph of Cumulative Frequencies of Student Marks

39 MANCOSA
Business Statistics

We can see from the table, or more easily from the graph, that 43 students (or 43%, since we have exactly 100
observations) had marks of 49% or less – so 57% of the class managed to achieve at least 50%.

Evident from Figure 3.14 is the very steep line between marks of 40 and 59; 14% of students had a mark of 40 or
less; 77% of students had a mark of 59 or less. Thus 63% of students marks thus fall in the interval of 40 to 59.
Again, one may question whether this is a good distribution of marks.

Another term associated with this type of analysis is the percentile, which, in our example, would refer to a mark
that a percentage of students have not achieved. For example, one can read from figure 3.14:
 The 90th percentile is a mark of 69; thus 90% of students received a mark of 69% or less; or 10% of students
received a mark of at least 70
 The 25th percentile is 45; 25% of students received 45 or less; or 75% of students received a mark of more
than 45

Also associated with this type of analysis is quartiles, which divide an ordered data-set into quarters.
The lower quartile (or 25th percentile) is that observation which separates the lower 25 percent of observation from
the top 75 percent of ordered observation.

The middle quartile (or 50th percentile) is the median. It divides an ordered data set into two equal halves.
The upper quartile (or 75th percentile) is that observation which separates the top 25 percent of observations from
the bottom 75 percent of ordered observations.

Activity 1
1. The table below gives the number of ocean-going vessels that arrived in
South African ports over a certain period.

Richards Bay 1602


Durban 4127
East London 109
Port Elizabeth 802
Mossel Bay 14
Cape Town 2346
Saldanha Bay 306
∑ = 9306

MANCOSA 40
Business Statistics

2. Consider the following raw data:


12 9 18 22 5 13 32 49 25 28

Represent the data in the form of a frequency table, using three classes of
equal width.

3. Draw a line graph depicting the following data:


x y
2 5
4 8
6 14
8 18

3.5. Summary
In this unit we outlined different modes of displaying data and conveying the information from statistical analyses.
Charts such as the pie and bar charts vividly display data associated with qualitative (categorical) random variables’
examples, how data may be presented using bar charts, pie charts, histograms, line graphs, frequency tables and
ogive curves.

41 MANCOSA
Business Statistics

Answers to Activities

Unit 3
Activity 1
1.

Ocean-Going Vessels in SA Ports


5000
4127
4000
Number of Vessels

3000 2346

2000 1602

802
1000 306
109 14
0
RB DBN EL PE MB CT SB
Port

Ocean-Going Vessels in SA Ports

1602

4127
9306

109
2346 802
14

306

RB DBN EL PE MB CT SB Total

2.
Interval Frequency
5 – 19 5
20 – 34 4
35 –49 1

3.

MANCOSA 42
Business Statistics

20

15

10
Y

0
0 2 4 6 8 10
X

43 MANCOSA
Business Statistics

Unit
4: Management Statistics

MANCOSA 44
Business Statistics

Unit Learning Outcomes

CONTENT LIST LEARNING OUTCOMES OF THIS UNIT:

4.1. Introduction  Introduce topic areas for the unit

4.2. Measures of Central Location  Understand, calculate and interpret the measures of
central location
 Discuss the concept of skewness and measures of
central tendency

4.3. Other Measures of Central Location  Discuss other measures of central location

4.4. Measures of dispersion/variability  Understand, calculate and interpret the measures of


dispersion for given data sets

4.5. Worked Examples  Discuss calculations of examples

4.6. Summary  Summarises topic areas in the unit

Prescribed and Recommended Textbooks/Readings


 Business Statistics using Excel: A first Course for South African
Students
 Glyn Davis, Branko Pecar and Leonard Santana; Oxford University
Press Southern Africa (2017)

45 MANCOSA
Business Statistics

4.1. Introduction
We saw in the previous chapter that graphical displays of statistical data are useful as a means of communicating
broad overviews of the behaviour of a random variable. However, there is a need for numerical measures (statistics)
about the behaviour pattern of a random variable.

The behaviour pattern of any random variable can be described by:


 A measure of central location, and
 A measure of spread of observations about this central value

After discussing these behaviour patterns, we will look at frequently used probability distribution functions in
business, the binomial and normal distributions. Commonly used index numbers will be demonstrated, after which
we will look at Sampling and Sampling Methods.

4.2. Measures of Central Location/Tendency


Observations of a random variable tend to group about some central value. The statistical measures that quantify
where the majority of observations are concentrated are referred to as measures of central location/tendency.

A measure of central tendency is a typical or representative score. If the mayor is asked to provide a single value
which best describes the income level of the city, he or she would answer with a measure of central tendency.
A central location statistic represents a typical value or middle data point of a set of observations and is useful for
comparing data sets.

There are three main measures of central location:


 Arithmetic mean (or average)
 Mode and
 Median (also called second quartile or the 50th percentile)

The computation of each of these measures differs for ungrouped (or raw) data and grouped data (data
summarised into a frequency distribution). The latter is of utmost importance.

4.2.1 Arithmetic Mean


The arithmetic mean (average) for raw (ungrouped) data is calculated is given by:
𝑠𝑢𝑚 𝑜𝑓 𝑎𝑙𝑙 𝑜𝑏𝑠𝑒𝑤𝑟𝑣𝑎𝑡𝑖𝑜𝑛𝑠
𝑥̅ =
𝑡𝑜𝑡𝑎𝑙 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑜𝑏𝑠𝑒𝑟𝑣𝑎𝑡𝑖𝑜𝑛𝑠
In equation form:
𝑛
1
𝑥̅ = ∑ 𝑥𝑖
𝑛
𝑖=1

MANCOSA 46
Business Statistics

where:
n = sample size.
𝑥𝑖 = 𝑖 𝑡ℎ observation of random variable X.
n

  shorthand notation for sum of n individual observations.


i 1

n
i.e. x
i 1
i  x1  x2  ......  xn

Example: Calculate the mean for the following data:


32 35 36 37 38 38 39 39 39 40 40 42 45
Solution:
n
500
x   xi   38.5
i 1 13

If we take our previous example of student marks (the raw data is shown in Figure 3.3), we obtain the following
value for the mean using the equation shown above:
𝑥̅ ≈ 51.3

Exercise: Show the full calculation and verify the above result.
An easier method of calculating the mean is to use the grouped data shown in Figure 4.1 below, where the midpoint
and frequency of observations for each interval is tabled.

Interval Frequency, f Midpoint, m fm

0–9 1 4.5 4.5

10 – 19 2 14.5 29.0

20 – 29 5 24.5 122.5

30 – 39 6 34.5 207.0

40 – 49 29 44.5 1290.5

50 – 59 34 54.5 1853.0

60 – 69 12 64.5 774.0

47 MANCOSA
Business Statistics

70 – 79 7 74.5 521.5

80 – 89 3 84.5 253.5

90 – 99 1 94.5 94.5

∑ =100 ∑ =5150.0
17Figure 4.1: Grouped Data for Student Marks

With this grouped data, we can calculate the mean using the formula:
𝑘
1
𝑥̅ = ∑ 𝑓𝑖 𝑚𝑖
𝑛
𝑖=1

where:
𝑚𝑖 is the midpoint of the 𝑖 𝑡ℎ interval of frequency 𝑓𝑖 .
The average is:
5150
𝑥̅ = = 51.5
100

If we compare this to the value calculated from the raw data (51.3), we see that this method can give a very close
approximation.

4.2.2 Median
The median is the value of a random variable that divides an ordered data-set into two equal parts, i.e. half the
observations will fall below the median value and the other half above it.

For ungrouped or raw data set of n observations arranged in ascending order, there are two possibilities:
1. If n is an odd number, then the median will be the middle value of the ordered data set. i.e. the [(n+1)/2]
value is the median.
2. If n is an even number, then the median is the average of the middle two values of the ordered data set,
i.e. the average of the [n/2] and [(n/2)+1] values.

Example: Determine the median for the following data:


32 35 36 37 38 38 39 39 39 40 40 42 45
Solution: The data is already arranged in ascending order. In this case there are 13 values. Hence the median is
the middle value which is 39.

For grouped data, the median can be calculated using the following formula:

MANCOSA 48
Business Statistics

n 
C   Fbelow 
Median  L   
2
f med
where:
L = lower bound of median class
C = class width

f med = frequency of median class


Fbelow = cumulative frequency below median class, i.e. the total number of observations below the lower bound
of median class.
n = sample size
The median interval is that interval into which the n/2 observation falls.
For the student marks problem we have n = 100, n/2 = 50 which falls in the interval 50 to 60. Thus
L= 50, 𝑓𝑚𝑒𝑑 = 34, 𝐹𝑏𝑒𝑙𝑜𝑤 = 43. Therefore:
10(50 − 43)
𝑀𝑒𝑑𝑖𝑎𝑛 = 50 + ≈ 52.1
34

4.2.3 Mode
The mode is the most frequently occurring value in a set of data. It is seldom computed for ungrouped data, since
it is simply the value occurring the most number of times.

Example: Determine the mode for the following data: 32 35 36 37 38 38 39 39 39 40 40 42 45


Solution: The value appearing the most times (three times in this case) is 39. Hence the mode is 39.

We can also calculate the mode for grouped data, using the formula:
𝐶(𝑓𝑚𝑜 − 𝑓𝑎𝑏𝑜𝑣𝑒 )
𝑀𝑜𝑑𝑒 = 𝐿 +
2𝑓𝑚𝑜 − 𝑓𝑎𝑏𝑜𝑣𝑒 − 𝑓𝑏𝑒𝑙𝑜𝑤
Where:
L = the lower limit of the modal class.
𝐶 = the class width.
𝑓𝑚𝑜 = the frequency of the modal class.
𝑓𝑏𝑒𝑙𝑜𝑤 = the frequency of the class before (below) the modal class.
𝑓𝑎𝑏𝑜𝑣𝑒 = the frequency of the class after (above) the modal class.
In our student marks example, the interval 50 < 60 has the most observations (34), and thus qualifies as the “modal
class”. We note 𝐿𝑚𝑜 = 50, 𝑓𝑚𝑜 = 34, 𝑓𝑏𝑒𝑙𝑜𝑤 = 29, 𝑓𝑎𝑏𝑜𝑣𝑒 = 12. Therefore:
10(34−29)
Mode = 50 + 2(34)−29−12 ≈ 51.9

49 MANCOSA
Business Statistics

4.2.4 Skewed Distributions and Measures of Central Tendency


Skewness refers to the asymmetry of the distribution, such that a symmetrical distribution exhibits no skewness.
In a symmetrical distribution the mean, median, and mode all fall at the same point, as in the following
distribution.

An exception to this is the case of a bi-modal symmetrical distribution. In this case the mean and the median fall
at the same point, while the two modes correspond to the two highest points of the distribution.

Mode Mode

Mean = Median
18Figure 4.4: A bimodal frequency distribution

A positively skewed distribution is asymmetrical and points in the positive direction. If a test was very difficult and
almost everyone in the class did very poorly on it, the resulting distribution would most likely
be positively skewed.

19Figure 4.5: A Positively Skewed Distribution

MANCOSA 50
Business Statistics

In the case of a positively skewed distribution, the mode is smaller than the median, which is smaller than the
mean. This relationship exists because the mode is the point on the x-axis corresponding to the highest point, that
is the score with greatest value, or frequency. The median is the point on the x-axis that cuts the distribution in half,
such that 50% of the area falls on each side.

The mean is the balance point of the distribution. Because points further away from the balance point change the
centre of balance, the mean is pulled in the direction the distribution is skewed. For example, if the distribution is
positively skewed, the mean would be pulled in the direction of the skewness, or be pulled toward larger numbers.

One way to remember the order of the mean, median and mode in a skewed distribution is to remember that the
mean is pulled in the direction of the extreme scores. In a positively skewed distribution, the extreme scores are
larger, thus the mean is larger than the median.

A negatively skewed distribution is asymmetrical and points in the negative direction, such as would result with a
very easy test. On an easy test, almost all students would perform well and only a few would do poorly.

51 MANCOSA
Business Statistics

The order of the measures of central tendency would be the opposite of the positively skewed distribution, with the
mean being smaller than the median, which is smaller than the mode.

20Figure 4.6: A Negatively Skewed Distribution

The choice of a representative central location value depends on the shape of the frequency distribution. If a
distribution is distorted by extreme values (i.e. skewed), then the median or the mode is more representative of the
distribution than the mean.

For a skewed distribution, the median may be the best measure of central location as it is not pulled by extreme
values (as the mean is), nor is it as highly influenced by the frequency of occurrence of a single value (as the mode
is).

4.3. Other Measures of Central Location


Other measures you will come across, but not used as frequently as the mean, median and mode are:
Geometric mean – used for percentage changes or growth rates.
Harmonic mean – used when a data set represents rates of change.
Weighted arithmetic mean – used if the importance (weight) of each observation is different.

MANCOSA 52
Business Statistics

4.4. Measures of Dispersion / Variability


Variability refers to the spread or dispersion of scores. A distribution of scores is said to be highly variable if the
scores differ widely from one another, so the classical definition of dispersion (or spread) is the extent by which the
observations of random variable are scattered about the central value.

Measures of dispersion provide useful information with which the reliability of the central value may be judged.
Widely dispersed observations indicate that the central value has low reliability, and does not represent the
observations very well. Conversely, a high concentration of observations about the central value indicates higher
reliability, with the central value being more representative.

The measures that are used to describe dispersion are:


 Range
 Inter-quartile range
 Quartile deviation
 Variance
 Standard deviation

4.4.1 Range
The range is the difference between the highest and lowest observed values in a data set. It is simply the largest
score minus the smallest score. It is a quick and dirty measure of variability, although when a test is given back to
students they very often wish to know the range of scores.

Range = Maximum value – Minimum value, for ungrouped data


= Upper limit (highest class) – Lower limit (lowest class), for grouped data

Because the range is greatly affected by extreme scores, it may give a distorted picture of the scores. The following
two distributions have the same range, 13, yet appear to differ greatly in the amount of variability.

Distribution 1 32 35 36 36 37 38 40 42 42 43 43 45

Distribution 2 32 32 33 33 33 34 34 34 34 34 35 45

For this reason, among others, the range is not the most important measure of variability.
Referring again to the example of student exam marks (see figure 3.3), we can use the ungrouped data to
calculate the range:
Maximum value = 90
Minimum value = 8

53 MANCOSA
Business Statistics

Range = 90 – 8 = 82

If we use the grouped data from Figure 4.2:


Upper limit (highest class) = 99
Lower limit (lowest class) = 0
Range = 99 – 0 = 99

Obviously the range calculated from the grouped data is not as accurate a measure as that calculated with the raw
data; in this case, it is also a poor estimate. For larger sets of data, it is normally a much closer estimate.

The range is a crude estimate of spread. It is easily calculated, but is distorted by extreme values (“out-liers”) An
“out-lier” would be the minimum or maximum value. It is thus a volatile and unstable measure of dispersion as it
can vary greatly between samples taken from the same population. It also provides no information on the clustering
of observations within the data set about a central value as it uses only two observations (i.e. the maximum and
minimum) in its computation.

4.4.2 Inter-Quartile Range


Because the range can be distorted by extreme values (“out-liers”), a modified range that excludes these is often
calculated. The inter-quartile range considers the viability shown by only the middle 50 percent of observations,
and is the difference between the upper and lower quartiles.

Inter-quartile range = Q3 – Q1 = 75th percentile – 25th percentile

Figure 3.14 showed the cumulative frequency polygon (or OGIVE) for our student marks as an example.

From the figure, we can easily read off the 75th and 25th percentile. We obtain
75th percentile = 58
25th percentile = 45
Inter-quartile range = 58 – 45 = 13

This measure of dispersion removes much of the instability inherent in the range by excluding “out-liers”, but it
excludes 50 percent of all observations from further analysis. It also provides no information on the clustering of
observations within the data set as it uses only two observations (Q1 & Q3) in its calculation.

4.4.3 Quartile Deviation


This measure of variation is simply the inter-quartile range divided by 2:
Quartile Deviation (QD) = (Q3 – Q1)/2

MANCOSA 54
Business Statistics

Continuing with our student marks example,


QD = 13/2 = 6.5
Remembering the median of 52.1, we interpret this as follows:
50% of all observations are expected to lie within 6.5 marks either side of 52.1, i.e. between 45.6 and 58.6.
Alternatively, 25% of the marks are expected to be within 6.5 marks below the median (i.e. from 45.6 to 52.1), and
25% of the marks are expected to lie within 6.5 marks above the median, (i.e. from 52.1 to 58.6).

The quartile deviation is useful as a measure of dispersion if the sample of observations contains excessive
“outliers”, as it ignores the top 25% and bottom 25% of the ranked observations.

As with the inter-quartile range, the quartile deviation does not use all the observations and therefore gives no
indication of the spread of values between the upper and lower quartiles.

4.4.4 Variance
The variance has become the most used measure of dispersion, because it:
 Takes every observation into account, and
 Is based on an average deviation from a central value

The calculation depends on whether the data is ungrouped or grouped.


Formula for ungrouped data:
𝑠𝑢𝑚 𝑜𝑓 𝑠𝑞𝑢𝑎𝑟𝑒𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛𝑠
𝑉𝑎𝑟𝑖𝑎𝑛𝑐𝑒 =
𝑠𝑎𝑚𝑝𝑙𝑒 𝑠𝑖𝑧𝑒 − 1

Expressed mathematically, the variance is given by the formula


𝑛
2
1
𝑆 = ∑(𝑥𝑖 − 𝑥̅ )2
𝑛−1
𝑖=1

Note that the variance could almost be the average squared deviation around the mean if the expression were
divided by n rather than n-1. It is divided by n-1, called the degrees of freedom, for theoretical reasons. If the mean
is known, as it must be to compute the numerator of the expression, then only n-1 scores that are free to vary. That
is if the mean and n-1 scores are known, then it is possible to figure out the nth score.

The formula for the variance presented above is a definitional formula, it defines what the variance means. The
variance may be computed from this formula, but in practice this is rarely done. It is done here to better describe
what the formula means. The computation is performed in a number of steps, which are presented below:
Step One - Find the mean of the scores.

55 MANCOSA
Business Statistics

Step Two - Subtract the mean from every score.


Step three - Square the results of step two.
Step Four - Sum the results of step three.
Step Five - Divide the results of step four by n-1.

Example: Consider the following simple example showing the ages of 7 second-hand cars:
13 7 10 15 12 18 9
13+7+10+15+12+18+9
𝑥̅ =
7
84
=
7

= 12 years

The calculation of the squared deviation of each observation from the sample mean is shown in Figure 4.7:

Car Ages, x ̅)
(𝒙 − 𝒙 ̅)𝟐
(𝒙 − 𝒙

13 1 1

7 -5 25

10 -2 4

15 3 9

12 0 0

18 6 36

9 -3 9

∑=0 ∑ = 84
21Figure 4.7: Calculation of Squared Deviation

Note: The sum of deviations (2nd column) must be zero.


Applying the formula on page 56,
84
𝑆2 = 6
= 14 years2.

An alternative formula that is easier to use is:


𝑛
1
2
𝑆 = ∑(𝑥𝑖2 − 𝑛𝑥̅ 2 )
𝑛−1
𝑖=1

MANCOSA 56
Business Statistics

Continuing with the car ages example, we can calculate the variance as follows:

Car age, x x2

13 169

7 49

10 100

15 225

12 144

18 324

9 81

∑ = 84 ∑= 1092
22Figure 4.8: Car ages –Variance Calculation

Using the alternative formula:


1092 − 7 × 122
𝑆2 = = 14
6

The variance can also be calculated for grouped data.

Calculation for Grouped Data


For grouped data, the variance can be calculated as:
𝑘
1
2
𝑆 = ∑(𝑓𝑖 𝑚𝑖2 − 𝑛𝑥̅ 2 )
𝑛−1
𝑖=1

where
𝑚𝑖 = midpoint of i interval
𝑘
1
𝑥̅ = ∑(𝑓𝑖 𝑚𝑖 )
𝑛
𝑖=1

𝑚𝑖 = midpoint of ith interval


𝑓𝑖 = frequency of ith interval
k = number of intervals

57 MANCOSA
Business Statistics

For student marks problem, the intermediate calculations for the variance are shown in Figure 4.9.

Interval Frequency, f Midpoint, m fm fm2


0–9 1 4.5 4.5 20.25
10 – 19 2 14.5 29.0 420.50
20 – 29 5 24.5 122.5 3001.25
30 – 39 6 34.5 207.0 7141.50
40 – 49 29 44.5 1290.5 57427.25
50 – 59 34 54.5 1853.0 100988.50
60 – 69 12 64.5 774.0 49923.00
70 – 79 7 74.5 521.5 38851.75
80 – 89 3 84.5 253.5 21420.75
90 – 99 1 94.5 94.5 8930.25
∑ =100 ∑ =5150.0 ∑ =288125.00
23Figure 4.9: Intermediate Calculations for Variance

Using the formula above,


288125 − 100 (51.5)2
𝑆2 = = 231.3
99
This is fairly close to the value of 210.10 calculated for the raw data earlier.
The variance is a measure of average squared deviation about the arithmetic mean. It is expressed in squared
units. Consequently, its meaning in a practical sense is obscure. To provide meaning, the dispersion measure
should be expressed in the original unit of measure of the random variable.

4.4.5 Standard Deviation


The standard deviation is a statistical measure, which expresses the average deviation about the mean in the
original units of the random variable. It is the square root of the variance and is written as

𝑆 = √𝑆 2

The standard deviation measures variability in units of measurement, while the variance does so in units of
measurement squared. For example, if one measured height in inches, then the standard deviation would be in
inches, while the variance would be in inches squared. For this reason, the standard deviation is usually the
preferred measure when describing the variability of distributions.

In our examples of exam marks, we can easily calculate the standard deviation.
Ungrouped data: 𝑆 = √210.1 ≃ 14.5

MANCOSA 58
Business Statistics

Grouped data: 𝑆 = √231.3 ≃ 15.2


The standard deviation is a relatively stable measure of dispersion across different samples of the same random
variable. It is therefore a powerful statistic, which describes how the observations are spread across the mean.

Coefficient of Variation
It is sometimes necessary to compare samples of data from different random variables to establish which sample
data shows greater variability. A direct comparison of their respective standard deviations would be misleading as
the random variables may be measured in different units.

The comparison would be more meaningful if the measures of variability were expressed in the same units. This
can be achieved by producing a measure of relative variability, i.e. relative to their mean, expressed in percentage
terms.

A statistic that shows this relative dispersion about a mean for a random variable is called the coefficient of variation,
𝑆
and is defined as: 𝐶𝑉 = 𝑥̅
× 100%.

A coefficient close to zero indicates low variability and a tight clustering of observations about the mean.
Conversely, a large coefficient of variation indicates that the observations are more spread about the mean value.
The coefficient of variation for the student marks, using the ungrouped data gives:
14.5
𝐶𝑉 = 51.3
× 100% = 28.3%

This low value indicates that, in spite of the large range of the data, the marks are generally tightly clustered around
the mean.

4.5. Worked Example


(Solution on next page)
Faulty ATMs are a major problem plaguing a leading bank in South Africa. The accompanying table gives the
number of faulty ATMs of this bank reported around the country over a 20-day period.

38 24 35 17 56
29 45 19 46 28
33 34 27 31 52
41 51 32 44 22

1. Determine the range from the data in the table.

59 MANCOSA
Business Statistics

2. Group the data in a frequency distribution. Let the lower limit of the initial class be 10 faulty ATMs and use a
class width of 10 faulty ATMs.
3. From the grouped frequency distribution, determine each of the following for the 20-day period.
(i) mean number of faulty ATMs.
(ii) median. number of faulty ATMs
(iii) modal number of faulty ATMs
4. Determine the standard deviation and interpret its value.
5. Draw an ogive curve and use it to estimate the median. How does your estimate compare with result of 3 (ii)
above?
6. What type of data (discrete or continuous) is portrayed in the table? Explain.

Solutions to worked example:


1. Range = xmax – xmin = 56 – 17 = 39 faulty ATMs

2.

Interval Frequency, f Midpoint, m fm fm2 F

10 – 19 2 14.5 29.0 420.50 2

20 – 29 5 24.5 122.5 3001.25 7

30 – 39 6 34.5 207.0 7141.50 13

40 – 49 4 44.5 178.0 7921.00 17

50 – 59 3 54.5 163.5 8910.75 20

∑ =20 ∑ =700.0 ∑ =27395.00

3. (i) Mean = 700 / 20 = 35 faulty ATMs


(ii) Median = 30 + 10[10 – 7] / 6 = 35 faulty ATMs
(iii) Mode = 30 + 10[6 – 5] / (2 x 6 – 5 – 4) = 33 faulty ATMs
4. Variance = (27395 – 20x352)/ (20-1) = 152.37
Standard deviation = 152.371/2 = 12.34 faulty ATMs.
This value is small, showing consistency of the ATMs.

MANCOSA 60
Business Statistics

5.

Cumulative Number of Days 25

20

15
Median ≈ 34 ATMs

10

0
0 10 20 30 40 50 60
Faulty ATMs

6. Discrete data, since number of faulty ATMs has to be a whole number.

Activity 1
4.1 The net annual salary (in R’000s) for 20 clerks at an auditing firm is given
below.
4.1.1 Using the raw data, determine the range.
4.1.2 Group the data into a grouped frequency distribution with a lowest
class lower limit of R 120 000 and a class width of R10 000.
4.1.3 Determine the mean and mode using the raw data.
4.1.4 Draw an OGIVE curve corresponding to the data.
4.2 The number of hernia repair surgeries performed on patients of
different ages by a general surgeon over a six-month period is given in
the table below.
4.2.1 For these patients, determine:
(a) the mean age for hernia repair surgery.
(b) the median age for hernia surgery.
(c) the modal age for hernia surgery.
4.2.2 Determine the standard deviation.

61 MANCOSA
Business Statistics

4.6. Summary
Sample statistics serve to estimate population parameters and describe the data characteristics. Two categories
of statistics were described in this chapter, namely: measures of central tendency and measures of variability. In
the former category were the mean, median, and mode. In the latter were the range, interquartile range and
standard deviation. Measures of central tendency describe a typical or representative score, while measures of
variability describe the spread or dispersion of scores about a central measure.

MANCOSA 62
Business Statistics

Answers to Activities

Unit 4
Activity 1
1. Range = max – min = 176 – 121 = 55 ≡ R55000
2.
Class
Freq, f F
(R’000)
120 - 130 4 4
130 - 140 7 11
140 - 150 3 14
150 - 160 3 17
160 - 170 2 19
170 - 180 1 20

3. Mean (raw data) = 2840/20 =142 ≡ R142000


Mode (raw data) = R133000

4.

Annual Salary of Clerks


25
Cumulative Number of Employees

20

15

10

0
110 120 130 140 150 160 170 180 190
Annual Salary (R'000)

4.1.
Class Freq, f F midpt, m fm fx2
10 < 20 4 4 15 60 900

63 MANCOSA
Business Statistics

20 < 30 9 13 25 225 5625


30 < 40 18 31 35 630 22050
40 < 50 26 57 45 1170 52650
50 < 60 17 74 55 935 51425
60 < 70 6 80 65 390 25350
∑=3410 ∑=158000

(a) mean age = 3410/80 = 42.6 years


(b) median age = 40 + 10(40 – 31)/26 = 43.5 years
(c) modal age = 40 + 10(26 – 18)/(52 – 18 – 17) = 44.7 years

4.2. s = √(158000 − 80 × 42.6 × 42.6)/79 = 12.65 years

MANCOSA 64
Business Statistics

Unit
5: Probability Distribution Functions

65 MANCOSA
Business Statistics

Unit Learning Outcomes

CONTENT LIST LEARNING OUTCOMES OF THIS UNIT:

5.1. Introduction  Introduce topic areas for the unit

5.2. Binomial and Normal Probability  Distinguish between the various distributions and calculate
distributions the associated probabilities

5.3. Sampling and Sampling  Understand appropriate sampling techniques for obtaining
Distributions statistical data

5.4. Index Numbers  Calculate and use index numbers

5.5. Worked Examples  Discuss calculations of examples

5.6. Summary  Summarises topic areas of the unit

Prescribed and Recommended Textbooks/Readings


Business Statistics using Excel: A first Course for South African Students.
Glyn Davis, Branko Pecar and Leonard Santana; Oxford University Press
Southern Africa (2017).

MANCOSA 66
Business Statistics

5.1. Introduction
This unit focuses on three topics: probability distributions, sampling and index number.
A probability distribution is a list of all the possible outcomes of a random variable and their associated probabilities
of occurrence. There are numerous problem situations in practice where the outcomes of a specific random variable
follow known probability patterns. If the behaviour of a random variable can be matched to a known probability
pattern, then probabilities for the random variable can be found directly by applying an appropriate theoretical
probability distribution function.

We examine probability distributions for discrete and continuous random variables, specifically the binomial
distribution (for discrete variable) and the normal distribution (for continuous variables). These distributions are
described explicitly by theoretical functions that enable one to calculate the probability of occurrences of events.

Index numbers play an important role in economic activities. An index number is a summary measure of the
change in the level of activity of a single item or collection (often referred to as basket) of related items from one
time period to another. We look at the calculation and application of these numbers.

Sampling is the process of selecting a representative subset of observations from a population to determine
characteristics of the random variable under study.
Sampling methods may be classified as non-probability and probability methods. We also introduce concept of
sampling distribution illustrate sampling distribution of the mean by means of an example.

5.2. Binomial and Normal Probability Distributions


A probability distribution is a list of all the possible outcomes of a random variable and their associated probabilities
of occurrence. There are numerous problem situations in practice where the outcomes of a specific random variable
follow known probability patterns. If the behaviour of a random variable can be matched to a known probability
pattern, then probabilities for the random variable can be found directly by applying an appropriate theoretical
probability distribution function.

Probability distribution functions can be classified as:


 Discrete, which assumes that the outcomes of a random variable can take on only specific (usually integer)
values. An example is the Binomial Probability Distribution
 Continuous, where the variable can take on any value (as opposed to only discrete values) in an interval. They
are used to find probabilities associated with intervals of x values. An example is the Normal Distribution

67 MANCOSA
Business Statistics

5.2.1 Binomial Probability Distribution


A discrete random variable can be described by the Binomial distribution if it satisfies the following four
conditions:
(i) There are only two mutually exclusive and collectively exhaustive outcomes of the random variable.
Generally, these two outcomes are referred to as success or failure. Each outcome has an associated probability:
 The probability of the success outcome is denoted by p
 The probability of the failure is denoted by q
 p + q = 1; hence q = (1 – p)
(ii) The random variable is observed n times. Each observation of the random variable in its problem setting is
called a trial. Each trial generates either a success or failure outcome. Thus, n outcomes are observed.
(iii) The trials are assumed to be independent of one another. Thus, the outcome on any trial is in no way
influenced by the outcome on any other trial. This means that p and q remain constant for each trial of the
process under study.
(iv) The binomial question is “What is the probability that r successes will occur in n trials of the process under
study?”

These can be summarized as an experiment with a fixed number of independent trials, each of which can only
have two possible outcomes. The fact that each trial is independent actually means that the probabilities remain
constant.

Examples of binomial experiments:


 Tossing a coin 20 times to see how many tails occur
 Asking 200 people if they watch ABC news
 Rolling a die to see if a 5 appears

Examples which aren't binomial experiments:


 Rolling a die until a 6 appears (not a fixed number of trials)
 Asking 20 people how old they are (not two outcomes)
 Drawing 5 cards from a deck for a poker hand (done without replacement, so not independent)

The binomial formula calculates the probability of r successes, and is stated as follows:
𝑛!
𝑃(𝑟) = 𝑝𝑟 𝑞 (𝑛−𝑟)
𝑟! (𝑛 − 𝑟)!
where
n = number of trials (observations0
r = 0, 1, 2… n = number of success outcomes in n trials
p = probability of success outcomes

MANCOSA 68
Business Statistics

q = 1 – p = probability of failure outcomes


! is the mathematical symbol for factorial, where n! = n x (n – 1) x (n – 2) x …x2 x 1
It has the property 0! = 1

Example: What is the probability of rolling exactly two sixes in 6 rolls of a die?
There are five things you need to do to work a binomial story problem.
1. Define Success first. Success must be for a single trial. Success = "Rolling a 6 on a single die"
2. Define the probability of success (p): p = 1/6
3. Find the probability of failure: q = 5/6
4. Define the number of trials: n = 6
5. Define the number of successes out of those trials: r = 2

Anytime a six appears, it is a success (denoted S) and anytime something else appears, it is a failure (denoted F).
The ways you can get exactly 2 successes in 6 trials are given below. The probability of each is written to the right
of the way it could occur. Because the trials are independent, the probability of the event (all six dice) is the product
of each probability of each outcome (die).
5 5 5 5 1 1 1 2 5 4
1. 𝐹𝐹𝐹𝐹𝑆𝑆 × × × × × =( ) ( )
6 6 6 6 6 6 6 6
5 5 5 1 5 1 1 2 5 4
2. 𝐹𝐹𝐹𝑆𝐹𝑆 × × × × × =( ) ( )
6 6 6 6 6 6 6 6
5 5 5 1 1 5 1 2 5 4
3. 𝐹𝐹𝐹𝑆𝑆𝐹 × × × × × =( ) ( )
6 6 6 6 6 6 6 6
5 5 1 5 5 1 1 2 5 4
4. 𝐹𝐹𝑆𝐹𝐹𝑆 × × × × × =( ) ( )
6 6 6 6 6 6 6 6
5 5 1 5 1 5 1 2 5 4
5. 𝐹𝐹𝑆𝐹𝑆𝐹 × × × × × =( ) ( )
6 6 6 6 6 6 6 6
5 5 1 1 5 5 1 2 5 4
6. 𝐹𝐹𝑆𝑆𝐹𝐹 × × × × × =( ) ( )
6 6 6 6 6 6 6 6
5 1 5 5 5 1 1 2 5 4
7. 𝐹𝑆𝐹𝐹𝐹𝑆 × × × × × =( ) ( )
6 6 6 6 6 6 6 6
5 1 5 5 1 5 1 2 5 4
8. 𝐹𝑆𝐹𝐹𝑆𝐹 × × × × × =( ) ( )
6 6 6 6 6 6 6 6
5 1 5 1 5 5 1 2 5 4
9. 𝐹𝑆𝐹𝑆𝐹𝐹 × × × × × =( ) ( )
6 6 6 6 6 6 6 6
4
5 1 1 5 5 5 1 2 5
10. 𝐹𝑆𝑆𝐹𝐹𝐹 × × × × × =( ) ( )
6 6 6 6 6 6 6 6
4
1 5 5 5 5 1 1 2 5
11. 𝑆𝐹𝐹𝐹𝐹𝑆 × × × × × =( ) ( )
6 6 6 6 6 6 6 6

69 MANCOSA
Business Statistics

1 5 5 5 1 5 1 2 5 4
12. 𝑆𝐹𝐹𝐹𝑆𝐹 × × × × × =( ) ( )
6 6 6 6 6 6 6 6
1 5 5 1 5 5 1 2 5 4
13. 𝑆𝐹𝐹𝑆𝐹𝐹 × × × × × =( ) ( )
6 6 6 6 6 6 6 6
1 5 1 5 5 5 1 2 5 4
14 𝑆𝐹𝑆𝐹𝐹𝐹 × × × × × =( ) ( )
6 6 6 6 6 6 6 6
4
1 1 5 5 5 5 1 2 5
15. 𝑆𝑆𝐹𝐹𝐹𝐹 × × × × × =( ) ( )
6 6 6 6 6 6 6 6

Notice that each of the 15 probabilities are exactly the same:


1 2 5 4
( ) ( ) = 0.0134
6 6
1 5
Also, note that 6
is the probability of success and you needed 2 successes, and that 6
is the probability of failure,

and if 2 of the 6 trials were success, then 4 of the 6 must be failures.


Note that 2 is the value of r and 4 is the value of (n-r).

Further note that there are fifteen ways this can occur. This is the number of ways 2 successes can be occur in 6
trials without repetition and order not being important, or a combination of 6 things, 2 at a time. 6Hence, the
probability is 15 x 0.0134 = 0.201 (20.1 %).

Example: Car-Hire Problem


A car hire firm rents out only BMW and Mazda cars. Experience has shown that one in four clients choose a BMW.
If 5 reservations are randomly selected from today’s bookings, what is the probability that 2 will have requested a
BMW?
One in four clients hires a BMW, so:
1
Probability of a success outcome (BMW hired): 𝑝 = = 0.25
4
3
Probability of a failure outcome (BMW not hired): 𝑞 = 4 = 0.75

We want to know the probability of 2 success outcomes,


i.e. we require P (2) and we have 5 observations, i.e. n = 5.

Hence, using the binomial formula,


5!
𝑃(2) = (0.25)2 (0.75)(5−2) ≈ 0.264
2! (5 − 2)!

With larger values of n and r, these calculations can become more elaborate.
Fortunately, most spreadsheets have the formula built in; for the above example, we could use the following formula
in Microsoft Excel: BINOMDIST(2, 5, 0.25, FALSE) ≃ 0.264.

MANCOSA 70
Business Statistics

5.2.2 The Normal Probability Distribution


A normal probability distribution function finds the probabilities for a continuous random variable and has the
following characteristics:
 It is bell-shaped
 It is symmetrical about a central value
 The tails of the distribution never touch the x-axis, i.e. asymptotic
 A normally distributed random variable is described by two parameters – the mean (μ) and standard deviation
(σ)
 The area under the curve equals one
 The probability associated with a particular range of x values is described by the area under the curve
between the limits of the x range; for example: 𝑥1 < 𝑥 < 𝑥2

Figure 5.1 below illustrates the features of the normal curve.

24Figure 5.1: A Typical Normal Distribution

There are three areas on a standard normal curve that all statistics students should know. The first is that the total
area below 0.0 is 0.50, as the standard normal curve is symmetrical like all normal curves. This result generalizes
to all normal curves in that the total area below the value of μ is 0.50 on any member of the family of normal curves.

The second area that should be memorized is between Z-scores of -1.00 and +1.00. It is 0.68 or 68%.

71 MANCOSA
Business Statistics

The total area between plus and minus one σ unit on any member of the family of normal curves is also 0.68.
The third area is between Z-scores of -2.00 and +2.00 and is 0.95 or 95%.

This area (0.95) also generalizes to plus and minus two σ units on any normal curve.
Knowing these areas allow computation of additional areas. For example, the area between a Z-score of 0.0 and
1.0 may be found by taking 1/2 the area between Z-scores of -1.0 and 1.0, because the distribution is symmetrical
between those two points. The answer in this case is 0.34 or 34%. A similar logic and answer is found for the area
between 0.0 and -1.0 because the standard normal distribution is symmetrical around the value of 0.0.

The area below a Z-score of 1.0 may be computed by adding 0.34 and 0.50 to get .84. The area above a Z-score
of 1.0 may now be computed by subtracting the area just obtained from the total area under the distribution (1.00),
giving a result of 1.00 - 0.84 or 0.16 or 16%.

The area between -2.0 and -1.0 requires additional computation. First, the area between 0.0 and -2.0 is 1/2 of 0.95
or 0.475. Because the 0.475 includes too much area, the area between 0.0 and -1.0 (0.34) must be subtracted i*n
order to obtain the desired result. The correct answer is 0.475 - 0.34 or 0.135.

MANCOSA 72
Business Statistics

Using a similar kind of logic to find the area between Z-scores of .5 and 1.0 will result in an incorrect answer
because the curve is not symmetrical around 0.5. The correct answer must be something less than 0.17, because
the desired area is on the smaller side of the total divided area.

If a data set follows a normal distribution, we can predict the frequency of data in intervals defined by the mean
and standard deviation, as follows:
𝜇−𝜎 < 𝑥 < 𝜇+𝜎 ∶ 68.2%
𝜇 − 2𝜎 < 𝑥 < 𝜇 + 2𝜎 ∶ 95.5%
𝜇 − 3𝜎 < 𝑥 < 𝜇 + 3𝜎 ∶ 99.7%
𝜇 = population mean
𝜎 = population standard deviation

If we look at our student marks, we can calculate the intervals and compare the count with the normal distribution
calculation.

Recall: For the student marks problem, mean = 51.3 and standard deviation = 14.5 (for the ungrouped data).
The table in figure 5.2 shows how the values predicted by the normal distribution compares with the actual
distribution (as determined from the ogive curve in figure.

Actual % by Normal
Interval
Distribution (%) Distribution

36.8 – 65.8 74 68.3

22.3 – 80.3 94 95.4

7.8 – 94.8 100 99.7


25Figure 5.2: Comparison of Actual and Predicted Frequency of Student Marks

We observe that except for the one standard deviation range, the predictions are quite accurate – certainly
accurate enough for practical applications.

5.2.3 The Standard Normal Distribution (Z-Distribution)


It is more common to use a Standard Normal Distribution since the probabilities (equal to the area under the
curve) are worked out and presented in tables. The random variable is defined in terms of a parameter Z, and
hence Standard Normal Distribution is also referred to as the Z-distribution. It has the following characteristics:
 Mean (µ) = 0
 Standard Deviation () = 1

73 MANCOSA
Business Statistics

The values in the table (figure 5.3) lie between zero (mean) and given z-scores, i.e. P (0 < Z < z-score).
z 0.00 0.01 0.02 0.03 0.04 0.05 0.06 0.07 0.08 0.09
0.0 0.0000 0.0040 0.0080 0.0120 0.0160 0.0199 0.0239 0.0279 0.0319 0.0359
0.1 0.0398 0.0438 0.0478 0.0517 0.0557 0.0596 0.0636 0.0675 0.0714 0.0753
0.2 0.0793 0.0832 0.0871 0.0910 0.0948 0.0987 0.1026 0.1064 0.1103 0.1141
0.3 0.1179 0.1217 0.1255 0.1293 0.1331 0.1368 0.1406 0.1443 0.1480 0.1517
0.4 0.1554 0.1591 0.1628 0.1664 0.1700 0.1736 0.1772 0.1808 0.1844 0.1879
0.5 0.1915 0.1950 0.1985 0.2019 0.2054 0.2088 0.2123 0.2157 0.2190 0.2224
0.6 0.2257 0.2291 0.2324 0.2357 0.2389 0.2422 0.2454 0.2486 0.2517 0.2549
0.7 0.2580 0.2611 0.2642 0.2673 0.2704 0.2734 0.2764 0.2794 0.2823 0.2852
0.8 0.2881 0.2910 0.2939 0.2967 0.2995 0.3023 0.3051 0.3078 0.3106 0.3133
0.9 0.3159 0.3186 0.3212 0.3238 0.3264 0.3289 0.3315 0.3340 0.3365 0.3389
1,0 0.3413 0.3438 0.3461 0.3485 0.3508 0.3531 0.3554 0.3577 0.3599 0.3621
1.1 0.3643 0.3665 0.3686 0.3708 0.3729 0.3749 0.3770 0.3790 0.3810 0.3830
1.2 0.3849 0.3869 0.3888 0.3907 0.3925 0.3944 0.3962 0.3980 0.3997 0.4015
1.3 0.4032 0.4049 0.4066 0.4082 0.4099 0.4115 0.4131 0.4147 0.4162 0.4177
1.4 0.4192 0.4207 0.4222 0.4236 0.4251 0.4265 0.4279 0.4292 0.4306 0.4319
1.5 0.4332 0.4345 0.4357 0.4370 0.4382 0.4394 0.4406 0.4418 0.4429 0.4441
1.6 0.4452 0.4463 0.4474 0.4484 0.4495 0.4505 0.4515 0.4525 0.4535 0.4545
1.7 0.4554 0.4564 0.4573 0.4582 0.4591 0.4599 0.4608 0.4616 0.4625 0.4633
1.8 0.4641 0.4649 0.4656 0.4664 0.4671 0.4678 0.4686 0.4693 0.4699 0.4706
1.9 0.4713 0.4719 0.4726 0.4732 0.4738 0.4744 0.4750 0.4756 0.4761 0.4767
2,0 0.4772 0.4778 0.4783 0.4788 0.4793 0.4798 0.4803 0.4808 0.4812 0.4817
2.1 0.4821 0.4826 0.4830 0.4834 0.4838 0.4842 0.4846 0.4850 0.4854 0.4857
2.2 0.4861 0.4864 0.4868 0.4871 0.4875 0.4878 0.4881 0.4884 0.4887 0.4890
2.3 0.4893 0.4896 0.4898 0.4901 0.4904 0.4906 0.4909 0.4911 0.4913 0.4916
2.4 0.4918 0.4920 0.4922 0.4925 0.4927 0.4929 0.4931 0.4932 0.4934 0.4936
2.5 0.4938 0.4940 0.4941 0.4943 0.4945 0.4946 0.4948 0.4949 0.4951 0.4952
2.6 0.4953 0.4955 0.4956 0.4957 0.4959 0.4960 0.4961 0.4962 0.4963 0.4964
2.7 0.4965 0.4966 0.4967 0.4968 0.4969 0.4970 0.4971 0.4972 0.4973 0.4974
2.8 0.4974 0.4975 0.4976 0.4977 0.4977 0.4978 0.4979 0.4979 0.4980 0.4981
2.9 0.4981 0.4982 0.4982 0.4983 0.4984 0.4984 0.4985 0.4985 0.4986 0.4986
3.0 0.4987 0.4987 0.4987 0.4988 0.4988 0.4989 0.4989 0.4989 0.4990 0.4990
26Figure 5.3: Standard Normal Distribution Table

MANCOSA 74
Business Statistics

Some examples that can be read off from the tables are:
For z = 0, P (z > 0) = 0.5
For z = 1, P (0 < z < 1) = 0.3413 → P (z > 1) = 0.5 – 0.3413 = 0.1587

Negative z-values: Because the table is symmetrical, the area to the left of the negative value will be the same as
that to the right of the positive value. Thus P (0 < Z < -1) 0.3413; P (Z < -1) = P (Z > 1) = 0.1587
How do we determine the area (probability) between any two values z1 and z2 of z, i.e P (z1 < Z < z2) Since P(z1
< Z < z2) = P(z1 < Z < 0) + P(0 < Z < z2). We read off the areas for z1 and z2 from the z table and then take the
sum.

For example P (-1 < Z < 1) = P (-1 < Z < 0) + P (0< Z < 1) = 0.3413 + 0.3413 = 0.6826

Using the z-distribution to find probabilities for a range of x-values:


Values of x associated with any normally distributed random variable can be converted to their corresponding z-
values via the transformation equation:
𝑥−𝜇
𝑧=
𝜎

where:
= population arithmetic mean
= population standard deviation
Let’s take another look at our example of student marks.
In this problem we have mean = 51.3, and standard deviation =14.5. Assuming the marks to be normally distributed,
then for any given value of x, the corresponding z score can be estimated using the transformation
𝑥 − 51.3
𝑧=
14.5

We can then use the z table to estimate the probabilities in Figure 5.2.
For example, for one standard deviation on either side of the mean, we have
(51.3 – 14.5) < x < (51.3 + 14.5) = 36.8 < x < 65.8. This can be converted to:
36.8−51.3 65.8−51.3
14.5
<𝑧< 14.5
→ −1 < 𝑧 < 1

From the z tables, we obtain an area of 2 × 0.3413 = 0.6826 (68.3%)

Similarly, we can read off the areas for two and three standard deviations on both sides of the mean.
Two standard deviations:
22.3 < x < 80.3 or –2 < z < 2
This gives us an area of 2 × 0.4772 = 0.9544 (95.4%)

75 MANCOSA
Business Statistics

Three standard deviations:


7.9 < x < 94.8 or –3 < z < 3
This gives us an area of 2 × 0.4986 = 0.9972 (99.7%)
These values are illustrated in Figure 5.4 below.

27Figure 5.4: Percentage of values on a normal distribution

The Standard Normal Distribution table is most useful for determining the probable frequency of a variable between
any two limits, and can be used for any data set that follows a normal distribution. and with a reasonably good
accuracy an approximately normal distribution.

Example:
A test is normally distributed with a mean of 60 and a standard deviation of 10. What proportion of the scores is
above 85?

Solution:
This problem is very similar to figuring out the percentile rank of a person scoring 85. The first step is to figure out
the proportion of scores less than or equal to 85. This is done by figuring out how many standard deviations above
the mean 85 is. Since 85 is 85-60 = 25 points above the mean and since the standard deviation is 10, a score of
85 is 25/10 = 2.5 standard deviations above the mean. Or, in terms of the formula,
𝑥−𝜇 85−60
𝑧= 𝜎
= 10
= 2.5

From the z table, P (0 < z < 2.5) = 0.4938. Thus P(x > 85) = P (z > 2.5) = 0.5 – 0.4938 = .0062 (0.62%)

MANCOSA 76
Business Statistics

5.2.4 Computing Normal Probabilities


There are several different situations that can arise when asked to find normal probabilities.

Situation Instructions

Between zero and any number. Look up the area in the table.

Between two positives, or between two negatives. Look up both areas and subtract smaller from larger.

Between a negative and a positive. Look up both areas and add them together.

Less than a negative or greater than a positive. Look up the area and subtract from 0.5000.

Greater than a negative or less than a positive. Look up the area and add to 0.5000.

This can be shortened into two rules.


1. If there is only one z score given, use 0.5000 for the second area, otherwise look up both z-scores in the table.
2. If the two numbers are the same sign, then subtract; if they are different signs, and then add. If there is only
one z score, then use the inequality to determine the second sign (< is negative, and > is positive).

5.2.5 Finding z-scores from Probabilities


This is more difficult, and requires you to use the table inversely. You must look up the area between zero and the
value on the inside part of the table, and then read the z score from the outside. Finally, decide if the z score should
be positive or negative, based on whether it was on the left side or the right side of the mean. Remember, z-scores
can be negative, but areas or probabilities cannot be.

Situation Instruction

Area between 0 and a value. Look up the area and make negative if in left tail.

Subtract the area from 0.5000. Look up difference and


Area in one tail.
make negative if in left tail.

Area including one complete half. (Less than Subtract 0.5000 from the area. Look up difference and make
a positive or greater than a negative.) negative if on the left side.

Look up half the area. Use both positive and negative z-


Within z units of mean.
scores.

Two tails with equal area. (More than z units Subtract the area from 0.5000. Look up difference and use
from mean.) negative and positive scores.

77 MANCOSA
Business Statistics

5.3. Sampling and Sampling Distributions


It is seldom possible to gather all the data on a random variable under study for analysis purposes; usually only a
subset, or sample, is collected.
Any statistical analysis performed on the sample is valid for that sample. But the behaviour of the whole population
can be inferred from the behaviour of the sample.
It is important to distinguish between a population and a sample.

A population comprises all possible observations of the random variable under study. Examples are:
 All the residents in a suburb, town or city under study
 The entire population of cell phon7e owners in the country

Gathering data on all possible observations in a population is called a census.


It is seldom practical or necessary to gather data on every possible observation in the population. More commonly,
a subset of all observations, or a sample, is gathered on the random variable, analysed and used as the basis for
decision making. Sampling is generally preferred, for the following reasons:
 It is more cost effective to gather sample data
 Sample data can be collected more timeously
 Some data requires destructive testing (e.g. battery life, impact resistance of items, shelf life), and a census
would not be appropriate
 The data collection for a sample is easier to control, and will thus be more accurate

According to Wegner (2007), a measure that is found from analysing sample data is called a statistic, while a
measure describing a population is called a parameter. The various notations used for these measures are
shown in figure 5.6.

Measure Sample Population

size n N

mean 𝑥̅ 𝜇

Standard Deviation 𝑠 𝜎

Proportion p 𝜋
28Figure 5.6: Notations for Samples and Populations

Most of the data used in managerial decision making is derived from a sample of observations. However, a
manager’s need is to know about the population parameter values of a random variable, not its sample statistics.

MANCOSA 78
Business Statistics

For example, a quality controller in a beer bottling process is more interested to know the population mean volume
of all bottles filled, rather than the sample mean volume of the few filled bottles drawn regularly from the production
line and tested.

Inferential statistics is that area of statistics that aims to estimate the true population parameters, with the following
process:
 Draw a sample of observations on the random variable under study
 Produce the appropriate sample statistics
 Derive estimates of the values of the corresponding population parameters based on these simple statistics

Statistical inference is performed in two ways:


 Through estimation where the sample statistic is used to estimate likely values of the corresponding
population parameter. Point estimation and confidence interval estimation are techniques used for this
 Through hypothesis testing, where a claim, statement or hypothesis is made about the population parameter,
and tested using sample evidence

5.3.1 Sampling Methods


Sampling is the process of selecting a representative subset of observations from a population to determine
characteristics (i.e. the population parameters) of the random variable under study.

There are two basic methods of sampling:


 Non-probability sampling methods
 Probability sampling methods

Non-Probability Sampling Methods


Any sampling method in which the observations are not selected randomly is called non-probability sampling. There
are 3 types:
 Convenience sampling, in which a sample is drawn to suit the convenience of the researcher
 Judgement sampling, where the researcher uses his or her judgement to select the sample
 Quota sampling, in which the population is divided into segments, and a quota of observations is collected
from each segment

The major disadvantage of non-probability sampling methods is the unrepresentative nature of the sample with
respect to the population from which it is drawn. Consequently, results from any statistical inference would probably
be invalid.

79 MANCOSA
Business Statistics

However, non-probability samples can be used in exploratory research to obtain initial impressions of the
characteristics of a random variable under study.

Probability Sampling Methods


Probability sampling includes all selection methods where the observations to be included in a sample have been
selected on a purely random basis from the population.

There are four types of probability sampling methods: Random, Systematic, Cluster and Stratified sampling.
 Random sampling is analogous to putting everyone's name into a hat and drawing out several names.
Each element in the population has an equal chance of occurring. While this is the preferred way of
sampling, it is often difficult to do. It requires that a complete list of every element in the population be
obtained. Computer generated lists are often used with random sampling
 Systematic sampling is easier to do than random sampling. In systematic sampling, the list of elements is
"counted off". That is, every kth element is taken. This is similar to lining everyone up and numbering off
"1,2….k; 1,2…k; etc.". When done numbering, all persons numbered k would be chosen
 Cluster sampling is accomplished by dividing the population into groups -- usually geographically. These
groups are called clusters or blocks. The clusters are randomly selected, and each element in the selected
clusters are used
 Stratified sampling also divides the population into groups called strata. However, this time it is by some
characteristic, not geographically. For instance, the population might be separated into males and
females. A sample is taken from each of these strata using either random, systematic, or convenience
sampling

5.3.2 The Sampling Distribution


A sampling distribution shows the relationship between a sample statistic and its corresponding population
parameter. It describes how a particular sample statistic varies about the true population parameter. From this
relationship, the level of confidence in estimating the population parameter from a single sample statistic can be
established.

Measures of sample statistics whose behaviour is generally described with respect to their corresponding
population parameters are:
 Mean
 Proportion
 Difference between two means
 Difference between two proportions

MANCOSA 80
Business Statistics

Let us look at the sampling distribution of a single sample mean through an example from Wegner (2007). We will
find the probability that a single sample mean lies within a certain distance of its unknown population.

Example: Assume that typing speed, measured in words per minute, is normally distributed. A random sample of
100 typists is selected and their typing speeds measured. Assume that the population deviation of typing speed is
8 words per minute.

What is the probability that the sample mean differs from the unknown population mean of typing speeds by no
more than one word per minute in either direction?

Expressed mathematically, we need to find:


𝑃(−1 < 𝑥̅ − 𝜇 < +1)

We are studying the behaviour of the sample mean with respect to its population mean. We can use the sampling
distribution of the sample means to find the required probability.

The standard deviation of sample means, also called the standard error, is calculated using the formula:
𝜎
𝜎𝑥̅ =
√𝑛

Substituting n = 100 and 𝜎 = 8, we obtain


8
𝜎𝑥̅ = = 0.8
10

Irrespective of the population distribution, the distribution of sample means will always be normal, so we can use
the properties of the normal distribution to predict behaviour. The sampling distribution of the sample mean is
related to the standard normal probability distribution (the z-distribution) through the following transformation
formula:
𝑥̅ − 𝜇
𝑧=
𝜎𝑥̅

Which can be rewritten as


𝑧𝜎𝑥̅ = 𝑥̅ − 𝜇

The original equation can therefore be written as


𝑃(−1 < 𝑧𝜎𝑥̅ < 1)
or
𝑃(−1 < 𝑧(0.8) < 1)

81 MANCOSA
Business Statistics

Dividing throughout by 0.8, gives:


𝑃(−1.25 < 𝑧 < 1.25)

Reading between the tails from the table, the area between the two values of z is 0.7888.
Thus there is a 78.9% chance that a single sample mean of typing speeds will lie within 1 word per minute of the
true (but unknown) population mean of typing speeds. This is based on a sample size of 100 typists, and drawn
from a normal population of typing speeds, with a standard deviation of 8 words per minute. Alternatively, there is
a 21.1% chance that it will be outside one word per minute.

5.4 Index Numbers


An index number is a summary measure of the change in the level of activity of a single item or collection (often
referred to as basket) of related items from one-time period to another. It is constructed by expressing the value of
an item in the current period as a ratio of its value in the base period. In percentage terms,
𝐶𝑢𝑟𝑟𝑒𝑛𝑡 𝑃𝑒𝑟𝑖𝑜𝑑 𝑉𝑎𝑙𝑢𝑒
𝐼𝑛𝑑𝑒𝑥 𝑁𝑢𝑚𝑏𝑒𝑟 = × 100%
𝐵𝑎𝑠𝑒 𝑃𝑒𝑟𝑖𝑜𝑑 𝑉𝑎𝑙𝑢𝑒

The base period is normally given a value of 100.


Some of the better known index numbers in South Africa are:
JSE Actuaries Indices – all share index, gold index, industrial index.
CPI – Consumer Price Index (1985 = 100)
PPI – Production Price Index (1980 = 100)

There are two major categories of index numbers – price and quantity. In both cases, a single or composite index
may be used.

A price index measures the percentage change in price between any two periods of time.
For a single item, the relative price change from one-time period to another is found by computing its price relative:
𝑝1
𝑃𝑟𝑖𝑐𝑒 𝑅𝑒𝑙𝑎𝑡𝑖𝑣𝑒 = × 100%
𝑝0

where
𝑝1 = current period price
𝑝0 = base period price
A quantity index measures the percentage change in consumption level of either an individual item or a basket of
items from one-time period to another.

MANCOSA 82
Business Statistics

For a single item, the relative quantity changes from one-time period to another is found by computing its quantity
relative.
𝑞1
𝑄𝑢𝑎𝑛𝑡𝑖𝑡𝑦 𝑅𝑒𝑙𝑎𝑡𝑖𝑣𝑒 = × 100%
𝑞0

where
𝑞1 = current period quantity
𝑞0 = base period quantity

A composite index combines the relative prices and quantities.


A commonly used composite index is the Laspeyres index.

The Laspeyres price index is given by


 p1q0
Lp  100%
 p0 q0
where quantities at base period levels are held constant.

The Laspeyres quantity index is given by:


 p0 q1
Lq  100%
 p0 q0 .
where prices at base period levels are held constant

Example: In the following share portfolio problem, the Laspeyres composite index is calculated for price and
quantity. The base year is 1986.

Share Base Year 1992

p0 q0 p1 q1 p0q0 p1q0 p0q1

A 65 350 115 300 22750 40250 19500

B 200 240 120 60 48000 28800 12000

C 1260 50 1890 100 63000 94500 126000

∑ =133750 ∑ =163550 ∑ =157500


29Figure 5.5: Share Portfolio Performance

83 MANCOSA
Business Statistics

Laspeyres quantity index:


163550
𝐿𝑝 = × 100% = 122.3%
133750

Value of share units increased, on average, by 22.3 %.

Laspeyres quantity index:


157500
𝐿𝑝 = × 100% = 117.8%
133750

Number of share units increased, on average, by 17.8 %.


The price index indicates the increase in the value of the portfolio if all quantities of shares remain the same.
Conversely, the quantity index indicates the increase in the number of shares bought if all prices are held constant.
Index numbers are generally based on samples of items. Hence sampling errors are introduced. Furthermore,
technological changes, product quality changes and changes in consumer purchasing patterns can individually and
collectively make comparisons over time unreliable.

5.5 Worked Examples


(Solutions on next page.)
1. According to a local newspaper, 34% of the employees of a leading bank resign because of poor pay, 28%
resign because of a career change, while the remainder resign due to family commitments. Determine the
following probabilities (for a group of seven randomly-selected employees who resigned):
(i) Three resigned due to poor pay.
(ii) Four resigned due to family commitments.
(iii) At least one resigned due to career change.

2. Explain why the binomial distribution is relevant is this situation.

3. An executive usually replies to his e-mails fairly quickly. The mean time he takes to reply to his e-mails is 30
minutes, with a standard deviation of 6 minutes. Determine the following probabilities.
i. he takes between 18 and 42 minutes to answer an e-mail.
ii. he takes between 24 and 36 minutes to answer an e-mail.

Solutions :
1. (i) n = 7, r = 3, n – r = 4, p = 0.34, q = 1 - 0.34 = 0.66
7!
𝑃(3) = (0.34)3 (0.66)4 = 0.261 (26.1%)
3! 4!
(ii) n = 7, r = 4, n – r = 3, p = 0.38, q = 1 - 0.38 = 0.62

MANCOSA 84
Business Statistics

7!
𝑃(4) = (0.38)4 (0.62)3 = 0.174 (17.4%)
4! 3!
2. n = 7, r = 0, n – r = 7, p = 0.28, q = 0.72
7!
𝑃(0) = (0.28)3 (0.72)4 = 0.101
0! 7!
𝑃(𝑟 ≥ 1) = 1 − 𝑃(0) = 1 − 0.101 = 0.899 (89.9%)

3. The outcomes are mutually exclusive and collective exhaustive. Even though there are three possible
outcomes, the binomial distribution is applicable. This can be seen as follows. One of the 3 outcomes is the
desired outcome (successful outcome) with probability p. The remaining outcomes together constitute the
failure outcome with probability q which is the sum of the probabilities of the unsuccessful outcomes. Thus, p
+ q = 1 as required.

4.  = 30 minutes,  = 6 minutes,
18−30 42−30
(i) 𝑃(18 < 𝑥 < 42) = 𝑃( 6
< z< 6
) = 𝑃(−2 < 𝑧 < 2)

= 2 × 0.4775 = 0.955 (95.5%)


24−30 36−30
(ii) 𝑃(24 < 𝑥 < 36) = 𝑃( 6
< z< 6
) = 𝑃(−1 < 𝑧 < 1)

= 2 × 0.3413 = 0.6826 (68.3%)

Activity 1
1. According to a survey, four out of ten South African drivers have
outstanding traffic fines. For a randomly-selected group comprising eight
South African drivers, what is the probability that at least six drivers have
outstanding traffic fines?
2. The mean monthly electricity bill for a complex of apartments is R1800.
Assuming that the electricity bills are normally distributed with a standard
deviation of R 250, approximately what percentage of these apartments
have monthly electricity bills in excess of R2000?
3. An automatic machine fills jars of jam with a mean net weight of 340 grams.
Assume a normal distribution with a standard deviation of 8 grams. What is
the probability that a randomly- selected jar of jam weights between 338
grams and 344 grams?
4. The life span of a particular brand of squash balls has a normal distribution
with a mean of 48 months and a standard deviation of 6 months. What
percentage of these squash balls last between 40 and 50 months?

85 MANCOSA
Business Statistics

5. The data in the table below shows the price (in Rand) and quantity of three
food items in 2011 and 2012

Price Quantity Price Quantity


Item
(2011) (2011) (2012) (2012)

Bread (loaf) 8.50 50 10.50 60

Eggs (dozen) 13.00 35 14.00 25

Milk (litre) 9.00 120 9.50 138

Using 2011 as a base year, calculate the Laspeyres price and quantity indices

5.6 Summary
In this unit we covered three important topics: probability distributions, index numbers and sampling.
We firstly examined the properties and applications of two theoretical probability distributions, namely the binomial
and normal distributions. The binomial distribution enables us to calculate the probability for any given value of a
binomial random variable using the binomial formula, while the normal distribution enables calculating the
probabilities for any given range of a continuous variable by using the standard normal distribution table. The
applications of these distributions were illustrated by means of examples.

We then looked at the construction and application of simple and composite index numbers, specifically the
Laspeyres price and quantity index numbers.

Finally, we introduced the concept of sampling. We outlined the different types of probability and non-probability
sampling methods. We also illustrated the concept of random distribution of the mean through an example.

MANCOSA 86
Business Statistics

Answers to Activities

Unit 5
Activity 1
1. p =4/10 = 0.4, q = 0.6
P (r ≥6) = P (6) + P (7) +P (8)
8! 8! 8!
= 6!(8−6)! (0.4)6 (0.6)2 + 7!(8−7)! (0.4)7 (0.6)1 + 8!(8−8)! (0.4)8 (0.6)0

= 0.124 + 0.041+ 0.008 = 0.172

2. z = (2000 – 1800)/250 = 0.8


P (X > R2000) = P (Z > 0.8) = 0.5 – 0.2881 = 0.1119 (11.19%)

3. x 1 = 338 g, x2 = 344 g
z1 = (338 – 340)/8 = - 0.25 z2 = (344 – 340)/8 = 0.50
P (338 < X < 344)) = P (-0.25 < Z < 0.50) = 0.0987 + 0.1915 = 0.2902

4. z = (x -µ)/σ
z1 = (40 – 48)/6 = -8/6 = -1.33
z2 = (50 – 48)/6 = 1/3 = 0.33
→ P (40 < X < 50) = P (-2.33 < Z < 0.33) = 0.4082 + 0.1293 = 0.5375 (≈53.8%)

5.
𝑝0 𝑞0 𝑝1 𝑞1 𝑝0 𝑞0 𝑝1 𝑞0 𝑝0 𝑞1
8.50 50 10.50 60 425 525 510
13.00 35 14.00 25 455 490 325
9.00 120 9.50 138 1080 1140 1242
∑=1960 ∑=2155 ∑=2077

𝑝0 , 𝑞0 = price, quantity respectively in 2011


𝑝1 , 𝑞1 = price, quantity respectively in 2012
 p1q0 2155
Lp  100%  100%  109.9%
 p0 q0 1960
 p0 q1 2077
Lq   100%   100%  106.0%
 p0 q0 1960

87 MANCOSA
Business Statistics

Unit
6: Prediction
(Correlation and Regression)

MANCOSA 88
Business Statistics

Unit Learning Outcomes

CONTENT LIST LEARNING OUTCOMES OF THIS UNIT:

6.1. Introduction  Introduce topic areas for the unit

6.2. Simple Linear Regression  Identify independent and dependent variables

6.3. Scatterplot  Define and discuss the concept scatterplot

6.4. Linear Regression and  Perform linear regression and correlation analysis
Correlation Analysis

6.5. Worked example  Discuss calculations of examples

6.6. Summary  Summarises topic areas of the unit

Prescribed and Recommended Textbooks/Readings


Business Statistics using Excel: A first Course for South African Students.
Glyn Davis, Branko Pecar and Leonard Santana; Oxford University Press
Southern Africa (2017).

89 MANCOSA
Business Statistics

6.1. Introduction
When two variables are related, it is possible to predict the values on one variable from the values on the other
variable with better than chance accuracy. This Unit describes how these predictions are made and what can be
learned about the relationship between the variables by developing a prediction equation. It will be assumed that
the relationship between the two variables is linear. Although there are methods for making predictions when the
relationship is nonlinear, these methods are beyond the scope of this module. Regression and correlation analyses
are statistical methods that attempt to quantify and describe possible relationships between variables. This
relationship can assist with the prediction of unknown values of certain variables from known values of the related
variables. Regression analysis quantifies the underlying structural relationship between variables. Correlation
analysis determines the strength of this identified association.

6.2. Simple Linear Regression


Simple linear regression analysis aims to find a linear relationship between the values of two random variables
only. One of the random variables is termed the independent variable (x). It is the variable for which values are
known or easily determined, and in certain instances can be controlled.

The other random variable is termed the dependent variable (y). Values are not readily known and need to be
estimated from values of the independent variable (x). Given that the relationship is linear, the prediction problem
becomes one of finding the straight line that best fits the data. Since the terms "regression" and "prediction" are
synonymous, this line is called the regression line.

The table in figure 6.1 shows pairs of random variables, between which possible relationships exist.

Independent Variable (x);Potential predictors of Dependant Variable (y); Variable to be


y. estimated.
Advertising Company Turnover

Training Labour Productivity

Speed Fuel Consumption

Hours Worked Machine Output

Daily Temperature Electricity Demand

Hours Studied Examination Results

Product price Product Sales Level

Bond Interest Rate Number of Bond Defaulters

Cost of living Poverty

30Figure 6.1: Relationships

MANCOSA 90
Business Statistics

Regression analysis aims to find a linear function i.e., a straight line that best fits the actual observations. A
straight-line graph is defined as follows:
𝑦̂ = 𝑎 + 𝑏𝑥
where:
𝑦̂ =estimated value of dependent variable
𝑥 = value of independent variable
𝑎 = y-intercept (where regression line cuts the y-axis)
𝑏 = slope of the regression line (for every unit change in x, y changes by b units).
Graphically the straight line (y = a + bx) may look as shown in figure 6.2:

31Figure 6.2: Straight-line Relationship

It is however uncommon to find such a perfect straight-line relationship shown in figure 6.2.
We usually talk about the “best-fit” straight line, i.e. a line passing through as many of the data points as possible.

6.3 Scatterplot
A scatterplot is a graphical plot of the values of the independent and dependent variables. The independent
variables x is recorded along the horizontal axis and the dependent values y along the vertical axis. Pairs of x
and y observations are plotted in space.

A visual inspection of the likely relationship between the two variables x and y, as provided by a scatterplot, will
provide an initial insight into the likely regression and correlation analysis results.

91 MANCOSA
Business Statistics

If for example the data points are widely scattered and the range of y values is large for any given x value, then a
linear regression function will be of little value as an estimation function for y, and the correlation measure will
show almost no association. Examples of various scatterplots are shown in figure 6.3 below:

Direct linear relationship


with small dispersion

Inverse linear relationship


with small dispersion

Direct linear relationship


with greater dispersion

Inverse Linear Relationship


with greater Dispersion

MANCOSA 92
Business Statistics

No Linear Relationship

32Figure 6.3: Illustrations of Scatterplots

6.4 Linear Regression and Correlation Analysis


When we do linear regression analysis we try to find the “best fit” line. The strength of the fit is indicated by the
correlation coefficient.

The regression line is that line which minimises the sum of the squared deviations of the observations from the
fitted line. Without providing the derivation, the coefficients a and b that result from this “method of least squares”
are as follows:
n n n
n xi yi   xi  yi
b i 1 i 1 i 1
2

n
 n
n x    xi 
2
i
i 1  i 1 
n n

 yi  b xi
a i 1 i 1

The correlation coefficient most commonly used is Pearson’s correlation coefficient (r), which is calculated as
follows:
n n n
n xi yi   xi  yi
r i 1 i 1 i 1

 n 2  n  2
  n 2  n 2 
n xi    xi   n yi    yi  
 i 1  i 1    i 1  i 1  

Here are some properties of r


 r only measures the strength of a linear relationship. There are other kinds of relationships besides
linear

93 MANCOSA
Business Statistics

 r is always between -1 and 1 inclusive. -1 means perfect negative linear correlation and +1 means
perfect positive linear correlation. 0 means a poor (or no) correlation
 r has the same sign as the slope of the regression (best fit) line
 r does not change if the independent (x) and dependent (y) variables are interchanged
 r does not change if the scale on either variable is changed. You may multiply, divide, add, or subtract a
value to/from all the x-values or y-values without changing the value of r

The correlation coefficient is a dimensionless number since it is a proportion.


A low correlation does not necessarily imply that the variables are unrelated, but simply that a straight line poorly
describes the relationship. A nonlinear relationship may well exist. Pearson’s correlation coefficient does not
identify non-linear association.

A correlation does not necessarily imply a cause and effect relationship, merely an observed association.
Example: Most of South Africa’s power stations are coal fired. Assume a random sample of 10 power stations
was selected and their coal usage and electricity generated for 1992 was obtained. The data are shown in figure
6.4.

Coal Usage in 1992 Electricity Generated


(Mega tons)
(million KW hours)

15 35

6 18

10 24

18 32

9 24

7 20

14 32

11 29

5 14

8 22
33Figure 6.4: Coal Usage and Electricity Generated

MANCOSA 94
Business Statistics

In this case, electricity generated is the dependent variable y and the coal usage the independent variable x.
A scatterplot of the data is shown below in Figure 6.5, along with the best fit line (dashed).

Electricity Generated (mega KWh) 40

35

30

25

20

15

10
3 5 7 9 11 13 15 17 19
Coal Usage (megatons)

34Figure 6.5: Scatterplot of Coal Usage and Electricity Generation

From the scatterplot, we can already see a strong linear (direct) relationship between coal usage x and electricity
generated. There is little dispersion, since the points lie near the best line fit.

When carrying out linear regression calculations it is useful to construct the table shown in figure 6.6.

x y x2 xy y2

15 35 225 525 1225

6 18 36 108 324

10 24 100 240 576

18 32 324 576 1024

9 24 81 216 576

7 20 49 140 400

14 32 196 448 1024

11 29 121 319 841

5 14 25 70 196

8 22 64 176 484

95 MANCOSA
Business Statistics

103 250 1221 2818 6670


35Figure 6.6: Calculations for Linear Regression

Using the above formulae, we obtain the following values for b and a:
10 × 2818 − 103 × 250
𝑏= ≈ 1.52
10 × 1221 − 1032
250 − 1.52 × 103
𝑎= ≈ 9.37
10
We can therefore define the estimated regression line as:
𝑦̂ = 9.37 + 1.52𝑥
(5 ≤ 𝑥 ≤ 18)

Pearson’s correlation coefficient can be calculated as follows:


10 × 2818 − 103 × 250
𝑟= ≈ 0.94
√(10 × 1221 − 1032 )(10 × 6670 − 2502 )

This correlation coefficient is close to +1, hence the association between x and y is very strong and positive.
Values of x can therefore confidently be used to estimate values of y.

The regression line can be used to estimate values of y from known values of x, by substituting the given x value
into the regression equation.

For example, estimate the level of electricity that would be generated for 12 million tons of coal:
𝑦̂ = 9.37 + 1.52 × 12 = 27.61

Thus with 12 million tons of coal, 27.61 million kilowatt hours of electricity can be expected to be generated.
Note on extrapolation.

Extrapolation is the process of estimating values of y, using values of x which lie outside the domain x values
used in the construction of the regression line. In our example, valid estimates of y are produced only from values
within the interval 5 ≤ x ≤ 18.

If values of y are estimated outside the domain of x, the estimates can be unreliable as the relationship between
x and y outside these limits is unknown and may in fact be quite different to that which is defined within the
domain.

For example, if we substitute x = 0 in our regression equation, we obtain.

MANCOSA 96
Business Statistics

𝑦 = 9.37 + 1.52(0) = 9.37


We could interpret this as meaning that 9.37 million kilowatt-hours of electricity will be generated if no coal is
used. This is clearly nonsensical!

6.5 Worked Example


(Solution on next page)
The cashiers of a leading bank get rewarded with annual bonuses, depending on the number of years of service
given to the bank. The accompanying table shows the bonus (in R 000 s) and the number of years of service of
the employees.

Bonus Number of years


of service
(R 000’s)

32 8

23 6

37 9

11 3

60 14

45 11

1. Portray the above information in a scatter-graph.


2. Determine the co-efficient of correlation and interpret its value.
3. Calculate the linear regression equation and use it to predict the number of years of Service of an
employee who received a bonus of R 70 000.
4. Is this value reliable? Explain.

Solution:
1.

97 MANCOSA
Business Statistics

70

60

50
Bonus (R'000)

40

30

20

10

0
0 2 4 6 8 10 12 14 16
Service Period (years)

2.
x y xy x2 y2

8 32 256 64 1024

6 23 138 36 529

9 37 333 81 1369

3 11 33 9 121

14 60 840 196 3600

11 45 495 121 2025

∑ = 51 ∑ = 208 ∑ = 2095 ∑ = 507 ∑ = 8668

6×2095−51×208
2. 𝑟 = ≈ 0.999
√(6×507−512 )(6×8668−2082 )

This value of r indicates a strong, direct linear relationship between number of years of service and bonus.

6×2095−51×208
3. 𝑏 = 6×507−512
≈ 4.4
208−4.45×51
𝑎= 6
≈ −3.16

We can therefore define the estimated regression line as:


𝑦̂ = −3.16 + 4.45𝑥

MANCOSA 98
Business Statistics

(3 ≤ 𝑥 ≤ 14)
For y = 70, we can use to regression equation to find x.
3.16+70
𝑥= 4.45
=  16 years.

5. No. The line is valid only for 3 ≤ 𝑥 ≤ 14.

Activity 1
The monthly salary (in thousands of Rand) of 5 employees at ABC agencies
as well as the number of years of experience of each
Experience Monthly salary
(years) (R’000)

2 3

6 9

11 13

13 16

15 20

1. Draw a scatter diagram depicting the data.


2. Calculate Pearson’s correlation coefficient.
3. Determine the linear regression equation and use it to predict the monthly
salary of an employee with 8 years experience. Is the result valid? Explain

6.6 Summary
This unit focussed on linear regression and correlation. We distinguished between dependent and independent
variables and how they connected by the linear regression equation. We also looked how to estimate the strength
of the linear regression by means of the scatterplot and to calculate the strength using Pearson’s correlation
coefficient. Finally illustrated the application of linear regression and correlation by means of an practical example.

99 MANCOSA
Business Statistics

Answers to Activities

Unit 6
Activity 1
1.

Monthly salary (R’000)


22
20
18
Monthly Salary (R'000)

16
14
12
10
8
6
4
2
0
0 5 10 15 20
Experience (years)

𝑛 ∑ 𝑥𝑦−∑ 𝑥 ∑ 𝑦
2. r =
√[𝑛 ∑ 𝑥 2 −(∑ 𝑥)2 ][𝑛 ∑ 𝑦 2 −(∑ 𝑦)2 ]

5×711−47×61
= = 0.99
√(5×555−47×47)(5×915−61×61)

3. y (x) = a + bx
𝑛 ∑ 𝑥𝑦−∑ 𝑥 ∑ 𝑦 5×711−47×61
b= = = 1.216
𝑛 ∑ 𝑥 2 −(∑ 𝑥)2 5×555−47×47
61 47
a= 5
− 0.774 × 5
= 0.774

→ y (x) = 0.774 + 1.216x (2 ≤ x ≤ 15)


Y (18) = 0.774 + 1.216x18 = 22.662 ≡ R22662
No, Regression equation valid in the domain 2 ≤ x ≤ 15

MANCOSA 100
Business Statistics

Unit
7: Forecasting Methods Using
Time Series Analysis

101 MANCOSA
Business Statistics

Unit Learning Outcomes

CONTENT LIST LEARNING OUTCOMES OF THIS UNIT:

7.1. Introduction  Introduce topic areas for the unit

7.2. Components of a Time Series  Understand and state the principles of forecasting and time
series
 State the principles of seasonality and trend

7.3. Trend Analysis using Moving  Calculate a trend using moving averages, and illustrate it
averages on a graph

7.4. Worked example  Discuss calculations of examples

7.5. Summary  Summarise the topic areas of the unit

Prescribed and Recommended Textbooks/Readings


Business Statistics using Excel: A first Course for South African Students.
Glyn Davis, Branko Pecar and Leonard Santana; Oxford University Press
Southern Africa (2017).

MANCOSA 102
Business Statistics

7.1. Introduction
Forecasting is an integral part of business management. The better the forecast, the better management will be
able to plan for the future. Although there are many methods for making forecasts, some are better suited than
others for particular situations. Forecasting is a critical function that needs to be done by businesses. It is needed
to assist us in financial planning determining staff levels and ordinary raw materials for production and other
business functions. The most common tool used for forecasting is time series analysis. It assumes that the actual
values of a random variable in a time series are influenced by a variety of environmental forces operating over
time. Time series analysis attempts to isolate and quantify the influence of these different environmental forces
operating on the time series into a number of different components.

7.2. Components of a Time Series


A time series is a set of numeric data of a random variable gathered over time at regular intervals and arranged in
chronological (time) order.

Time series analysis assumes that four underlying forces individually and collectively determine the random
variables value in a time series in any time period. They are
 Trend (T)
 Cyclical Variations ( C )*
 Seasonal Variations (S)
 Random (irregular) variation (R )

7.2.1 Trend (T)


Trend is defined as a long-term smooth underlying movement in time series. It describes the effect that long-term
factors have on the series. These long-term factors tend to operate fairly gradually and in one direction for a long
period of time. Thus, a smooth curve or a straight line, such as shown in figure 7.1 usually describes the trend
component:

36Figure 7.1: Illustration of Trend

103 MANCOSA
Business Statistics

Examples of long-term trends are:


 Population growth
 Urbanisation
 Technological improvements
 Shifts in habits and attitudes

7.2.2 Cyclical Fluctuation (C)


Cycles are medium to long term deviations from the trend. They reflect alternating periods of relative expansion
and contraction. They are wave like movements in a time series that can vary greatly in duration and amplitude.
They are difficult to measure statistically and their use in statistical forecasting is limited.

The most common form of cycle is the business cycle between periods of relatively good economic activity to poor
economic activity. The causes of these are difficult to determine. Action by government, trade unions and world
organisations induce levels of pessimism and optimism into the economy which are reflected in changes in the
time series levels. Index numbers are used to describe cyclical fluctuations. An illustration of cycles is shown in
figure 7.2.

37Figure 7.2: Illustration of Cycles

7.2.3 Seasonal Variations (S)


Seasonal variations are fluctuations that are repeated periodically, usually within a year (i.e. daily, weekly, monthly
or quarterly). They are readily isolated through statistical analysis. Seasonal fluctuations are caused by re-occurring
events such as climatic conditions, special occurring events (e.g. Easter, Christmas) and religious, public and
school holidays. An example is shown in figure 7.3. The regular patterns of seasonal fluctuation are measured by
seasonal indices.

MANCOSA 104
Business Statistics

Sales

Summer
Winter

Spring Fall

Time (Quarterly)

38Figure 7.3: Illustration of Seasonal Fluctuations

7.2.4 Random Fluctuation (I)


These are caused by unpredictable occurrences, which may be evident or sometimes not so evident. Examples of
evident events are natural disasters such as floods, droughts or fires and man-made disasters such as strikes or
boycotts. These variations follow no specific pattern, and cannot be analysed statistically, and thus cannot be
incorporated into forecasts.

7.3. Trend Analysis Using Moving Averages


7.3.1 Decomposition of a Time Series
By using time series analysis, we try to isolate the influence of each of the four components on the series. We do
this through the Multiplicative Time Series Model, which states that the actual values of a time series, y, can be
found by multiplying the trend component by each of the following
o Cyclical index: C
o Seasonal index: S
o Irregular measure

The trend component is expressed in the active units of the variable we are looking at. The seasonal and cyclical
indices are, by definition, index numbers and expressed relative to the trend.
Mathematically, this is expressed as:
y  T C  S  I

Statistical analysis can be used effectively to isolate the trend (T) and the seasonal (S) components, but is of less
value in quantifying the cyclical movements, and of no value in isolating irregular components.

105 MANCOSA
Business Statistics

We will examine statistical approaches to quantify Trend and Seasonal variation only. More sophisticated models
would be needed to isolate the other two components.

7.3.2 Moving Averages


The most common methods used for trend isolation are:
 The moving average which produce a smooth curve
 Regression analysis which involves fitting a straight line

Regression analysis was discussed in the last Unit.


The moving average removes the short-term fluctuations in a time series by taking successive averages of groups
of observations.

To illustrate the method, let’s say we sold 30 widgets during the month of June. We want to estimate what our sales
will be for July. Our best guess might be that we will sell 30 widgets during July – we have used a “one month
moving average” as our forecast.

When we want to forecast for August, we may want to take into account what happened during June and July. Let’s
say we had sales of 40 during July. If we took a two-month moving average, our forecast for August would be
( Actual) June  ( Actual) July 30  40
( F / C ) August    35
2 2

What do we do for September? Let’s say the sales for August were 30. We now have a choice between 3 forecasts:
1-Month moving average:
( F / C ) Sept  ( Actual ) Aug  30

2-Month moving average:


( Actual) July  ( Actual) Aug 40  30
( F / C ) Sept    35
2 2

3-Month moving average


( Actual) June  ( Actual) July  ( Actual) Aug 30  40  30
( F / C ) Sept    33.3
3 3

The table in the figure 7.4 below demonstrates the calculation for each forecast, while figure 7.5 shows the results
graphically.

MANCOSA 106
Business Statistics

Moving Average Forecast


Month Actual 1 month 3 month 5 month
1 30 - -
2 40 30 - -
3 30 40 - -
4 35 30 33 -
5 32 35 35 -
6 45 32 32 33
7 52 45 37 36
8 35 52 43 39
9 36 35 44 40
10 38 36 41 40
11 60 38 36 41
12 50 60 45 44
13 45 50 49 44
14 50 45 52 46
15 55 50 48 49
Sales

1 2 3 4 5 6 7 8 9 10 11 12
Month

39Figure 7.5: Moving Average Plots

107 MANCOSA
Business Statistics

From the graph plots, we can see that


 No plot is available in the early months for the 3-month and the 5-month moving averages, as enough data is
not available. For example, the first month in which a 5 month moving average forecast can be used is month
6
 The higher the number of months used in a moving average, the smoother the curve – compare the 5 month
to the 3 month and the 1 month curves. Another way of seeing it is that the higher the number of months, the
less the curve is influenced by variations in trend
 The 1-month moving average replicates the actual exactly, but “ lags” by one month. This is characteristic of
all simple moving average curves – they will lag the actual figures, i.e. upward or downward shifts in trend will
only be detected after the event. The more months used to calculate the moving average, the longer it takes
for the change to register

7.3.4 Seasonal Analysis


A seasonal index is the ratio of the demand for a particular season to the average seasonal demand. For example,
if the average seasonal demand is 100 units and the demand for the summer season is 80, the summer season
index is 80 / 100 = 0.8. An averaging process used to arrive at the average seasonal value is illustrated in the
following example.

7.4 Worked Example: (Metro Movers and Seasonal Indices)


There is no need to look at actual demand data to know that a moving company has a seasonal demand – it is
common sense. Even so, a good starting point in seasonal analysis is scrutiny of the demand graph.

The graph of past quarterly demand for Metro Movers is shown in figure 7.7. It is clear that summer demand is by
far the highest in every year and autumn demand is generally the lowest. The seasonal index measures how much
higher and how much lower. Figure 7.8 shows calculations of seasonal indices for the 16 available past demands.
(Note. Besides seasonality, it looks like there is a slight upward trend over the 16 quarters. We shall ignore the
trend for now.).

MANCOSA 108
Business Statistics

40Figure 7.7: Seasonal Demand history for Metro Movers

Year Season Actual Demand Mean Seasonal Demand Seasonal Index

Spring 90 - -

1997 Summer 160 - -

Autumn 70 115 0.61

Winter 120 125 0.96

Spring 130 133 0.98

1998 Summer 200 133 1.50

Autumn 90 124 0.73

Winter 100 114 0.88

Spring 80 115 0.70

1999 Summer 170 125 1.36

Autumn 130 136 0.96

Winter 140 148 0.95

Spring 130 146 0.89

Summer 210 138 1.52


2000
Autumn 80 - -

Winter 120 - -
41Figure 7.8: Seasonal Index Calculations

109 MANCOSA
Business Statistics

The mean seasonal demand is a four-period moving average centred in the middle of a given season, that is a
month and a half into the season. It includes demands going back six months and forward six months from that
point. Thus, the first figure in column 3 is based on demands for the last one and a half months of spring 1997; and
all of summer, autumn and winter 1997; and the first one and a half months of spring 1998. So,
(90 / 2)  160  70  120  (130 / 2)
 115
4

This is a bit cumbersome, but it ensures that no one season is weighted more heavily than any other. The seasonal
indices are shown rearranged by year and season in figure 7.9. The three values for each season need to somehow
be reduced to a single index. The index for autumn is steadily rising, from 0.61 to 0.73 to 0.96. That is not sufficient
reason to expect it to continue to rise, however, especially since the other seasons do not show trends. Thus, the
projections of the seasonal indices for 2001 are the means of each column.

Year Spring Summer Autumn Winter

1997 - - 0.61 0.96

1998 0.98 1.50 0.73 0.88

1999 0.70 1.36 0.96 0.95

2000 0.89 1.52 - -

Mean SI 0.85 1.46 0.76 0.93

42Figure 7.9: Summary and Projection of Seasonal Indices

(The future seasonal index is obtained by calculating the mean of corresponding indices for past years.)

Metro movers may now use the seasonal indices in fine-tuning its demand forecasts for each coming season. For
example, suppose that they expect to move 480 vans of goods next year based on projection of the mean of past
years’ demands. It would be naïve to divide 480 by 4 and project 120 vans in each season. Instead,

Divide 480/4 = 120 = average number of vans per season. This average is now multiplied by the mean seasonal
index.
Spring 2001: 120 x 0.85 = 102 vans
Summer 2001: 120 x 1.46 = 175 vans
Autumn 2001: 120 x 0.76 = 91 vans
Winter 2001: 120 x 0.93 = 112 vans

MANCOSA 110
Business Statistics

Yearly Total = 102 + 175 + 91 + 112 = 480 vans

Activity 1
The number of tennis racquets sold by a sports store is recorded per quarter,
for the past three years. The results are presented in the table below.
Year Q1 Q2 Q3 Q4
Year 1 200 220 405 300
Year 2 190 240 540 298
Year 3 180 198 680 307

1. Calculate the mean seasonal indices for each quarter of year 4.


2. If the store expects to sell 1500 tennis racquets in year 4, how many
racquets is it likely to sell in each quarter of year 4?

7.5 Summary
In this unit we considered the ratio-to-moving-average method of analysing a time series. We identified and
described the nature of the trend, cyclical, seasonal and irregular influences on the time series. Using the
multiplicative model, we used the technique of time series analysis to decompose the time series into its constituent
components. We examined seasonal components were considered. Trend component can be described by linear
regression and seasonal component by finding seasonal indexes using the method of centred moving averages.
The method was illustrated by means of an example.

111 MANCOSA
Business Statistics

Answers to Activities

Activity 1
1.
(Year, Quarter) Data Mean Seasonal Seasonal
Demand Index
(1, Q1) 200 - -
(1, Q2) 220 - -
(1, Q3) 405 280.0 1.45
(1, Q4) 300 281.3 1.07
(2, Q1) 190 300.6 0.63
(2, Q2) 240 317.3 0.76
(2, Q3) 540 315.8 1.71
(2, Q4) 298 309.3 0.96
(3, Q1) 180 321.5 0.56
(3, Q2) 198 340.1 0.58
(3, Q3) 680 - -
(3, Q4) 307 - -

Summary Table
Q1 Q2 Q3 Q4
year 1 - - 1.45 1.07
year 2 0.63 0.76 1.71 0.96
year 3 0.56 0.58 - -
Mean SI 0.60 0.67 1.58 1.03

2.
Year-4 quarterly forecasts:
Q1 225
Q2 251
Q3 593
Q4 386

MANCOSA 112
Business Statistics

Bibliography

 Glyn Davis, Branko Pecar and Leonard Santana. Business Statistics using Excel: A first Course for
South African Students.; Oxford University Press Southern Africa (2017)
 Mann, Prem S. (2004). Introductory Statistics. 5th Ed. John Wiley & Sons, Inc
 Ross, Sheldon M. (2005). Introductory Statistics. 2nd Ed. Elsevier Academic Press
 Sanders, Donald S. (1995). Statistics : A first course. 5th Ed. McGraw-Hill, Inc
 Schonberger, Richard J. & KNOD, Edward M., Jr. (1985). Operations Management.
 Serving the Customer. 3rd Ed. Homewood, Illinois: BPI Irwin
 Stevenson, William J. (1999). Production/Operations Management. 6th Ed. Irwin: McGraw Hill
 Wegner, Trevor. (2007). Applied Business Statistics. Methods and Applications. Kenwyn: Juta & Co, Ltd
 Weiers, Ronald M (2005) Essentials of Business Statistics. Thomson Learning, Inc
 Wisniewski, M. and STEAD R. (1996). Foundation Quantitative Methods for Business. London: Prentice
Hall (Chapter 13)

113 MANCOSA
Business Statistics

MANCOSA 114

Common questions

Powered by AI

Tables and graphs are effective for structuring and presenting data in an understandable way. Tables work well for displaying precise values and detailed comparisons, while graphs are valuable for visualizing trends and distributions quickly. The choice between them depends on the need for clarity in understanding complex relationships or providing detailed quantitative insights .

Understanding the nature of data and data collection methods is crucial in establishing accuracy and reliability in statistical findings because it ensures the data sourced is relevant, unbiased, and accurately represents the population or process being studied. Different data types and appropriate collection techniques, such as surveys or observational studies, help avoid sampling errors and increase confidence in the analytical results .

Time series analysis in forecasting provides the advantage of identifying patterns over time, such as trends and seasonality, which can improve the accuracy of future predictions . However, challenges include the need for large datasets to identify patterns accurately and the difficulty of accounting for unexpected events or changes in underlying processes that may disrupt established patterns .

Measures of central tendency, such as mean, median, and mode, simplify complex datasets by providing a single value representing a typical data point, which aids in summarizing and comparing data sets effectively . The choice between them involves considering data characteristics such as skewness and the presence of outliers; the mean is informative for symmetric distributions, the median is robust to skew and outliers, and the mode reflects the most frequent observation .

In business statistics, the binomial distribution is a discrete distribution used when an experiment or process has two potential outcomes, like success or failure, with fixed probabilities across trials . The normal distribution, in contrast, is a continuous distribution representing data that clusters around a mean with symmetrical decay towards the extremes, suitable for naturally occurring datasets like heights or test scores, where a myriad of small influences govern outcomes .

Learning outcomes in an educational module, such as business statistics, guide the instructional design by outlining the knowledge and skills students should acquire and be able to demonstrate. They facilitate focused learning and self-assessment, ensuring that educational objectives are met across various cognitive levels .

The least squares method is utilized to derive a best-fit straight line through a set of data points, which can then be used to make future predictions about the data . The Pearson correlation coefficient helps in determining the strength and direction of the linear relationship between two variables, providing insight into how closely the prediction model might fit the actual data .

Descriptive statistics focus on summarizing and organizing data to describe the characteristics of a dataset, such as mean, median, and mode, which are useful in situations like evaluating classroom performance or analyzing game statistics . In contrast, inferential statistics go beyond immediate data characteristics to make predictions or decisions about a broader population based on sample data, useful in scenarios like market research or policy formulation .

The standard normal distribution, characterized by a mean of 0 and a standard deviation of 1, is used in statistical analysis to calculate the probability of occurrences within a dataset. Probabilities are derived from the area under the curve, which is symmetrical and standardized, allowing for the use of Z-scores to determine the likelihood of a particular data point appearing within a certain range .

Variance measures the average squared deviation from the mean, providing a sense of data spread or variability. It is crucial for comparing data sets with different units or scales, though it is expressed in squared units which can be abstract . The standard deviation, being the square root of variance, provides a measure of spread in the original unit of data, offering more intuitive insights into how much data values deviate from the mean, which is vital for understanding variability .

You might also like