Statistics in Management
Chapter-1
March 31, 2021
Statistics
Statistics is defined as a set of mathematically based tools and techniques to
transform raw (unprocessed) data into a few summary measures that represent useful
and usable information to support effective decision making.
These summary measures are used to describe profiles (patterns) of data, produce
estimates, test relationships between sets of data and identify trends in data over
time.
Statistics
Transformation process from data to information
The Language of Statistics
Important terms and concepts used in Statistics.
a random variable and its data
a sampling unit
a population and its characteristics, called population parameters
a sample and its characteristics, called sample statistics
Random Variable & Data
Random variable
A random variable is any attribute of interest on which data is collected and
analysed.
Data
Data is the actual values (numbers) or outcomes recorded on a random variable.
Examples:
the travel distances of delivery vehicles (data: 34 km, 13 km, 21 km)
the daily occupancy rates of hotels in Cape Town (data: 45%, 72%, 54%)
brand of coffee preferred (data: Nescafé, Ricoffy, Frisco)
Sampling Unit
A sampling unit is the object being measured, counted or observed with respect to
the random variable under study.
Sampling unit could be a consumer, an employee, a household, a company or a
product. More than one random variable can be defined for a given sampling unit.
For example, an employee could be measured in terms of age, qualification and
gender.
Population & Population Parameter
Population
A population is the collection of all possible data values that exist for the random
variable under study.
Examples
for a study on hotel occupancy levels (the random variable) in Cape Town only,
all hotels in Cape Town would represent the target population
to research the age, gender and savings levels of banking clients (three random
variables being studied), the population would be all savings account holders at
all banks.
Population Parameter
A population parameter is a measure that describes a characteristic of a population.
Population average & Population proportion are population parameters.
Sample & Sample Statistic
Sample
A sample is a subset of data values drawn from a population.
Samples are used because it is often not possible to record every data value of the
population, mainly because of cost, time and possibly item destruction.
Examples:
A sample of 25 hotels in Cape Town is selected to study hotel occupancy levels
a sample of 50 savings account holders from each of four national banks is
selected to study the profile of their age, gender and savings account balances.
Sample Statistic
A sample statistic is a measure that describes a characteristic of a sample.
The sample average and a sample proportion are two typical sample statistics.
Examples:
the average hotel occupancy level for the sample of 25 hotels surveyed
the average age of savers, the proportion of savers who are female and the
average savings account balances of the total sample of 200 surveyed clients.
Examples of populations and associated samples
Symbolic notation for sample and population measures
Components of Statistics
Statistics consists of three major components: Descriptive Statistics, Inferential
Statistics and Statistical Modelling.
Descriptive Statistics
Descriptive statistics condenses sample data into a few summary descriptive
measures.
Inferential Statistics
Inferential statistics generalises sample findings to the broader population.
Statistical Modelling
Statistical modelling builds models of relationships between random variables.
Components of Statistics
Data and Data Quality
An understanding of the nature of data is necessary for two reasons. It enables a user
to assess data quality
to select the most appropriate statistical method to apply to the data.
Both factors affect the validity and reliability of statistical findings.
Data Quality
Data is the raw material of statistical [Link] quality is influenced by four
factors:
data type
data source
methods of data collection
appropriate data preparation
Selection of Statistical Method
The choice of the most appropriate statistical method to use depends firstly on the
management problem to be addressed and secondly on the type of data available.
Data Types and Measurement Scales
The type of data available for analysis is determined by the nature of its random
variable. A random variable is either qualitative (categorical) or quantitative
(numeric) in nature.
Qualitative random variables
Qualitative random variables generate categorical (non-numeric) response data.
The data is represented by categories only.
Examples:
The gender of a consumer is either male or female.
An employee’s highest qualification is either a matric, a diploma or a degree.
Quantitative random variables
Quantitative random variables generate numeric response data.
These are real numbers that can be manipulated using arithmetic operations (add,
subtract, multiply and divide).
Examples:
the age of an employee (e.g. 46 years; 28 years; 32 years)
the price of a product in different stores (e.g. R6.75; R7.45; R7.20; R6.99)
Discrete Data & Continuous Data
Numeric data can be classified as either discrete or continuous.
Discrete data is whole number (or integer) data.
Example:
the number of students in a class (e.g. 24; 37; 41; 46)
the number of cars sold by a dealer in a month (e.g. 14; 27; 21; 16)
Continuous data is any number that can occur in an interval.
Example:
a passenger’s hand luggage can have a mass between 0.5 kg and 10 kg (e.g. 2.4
kg)
the volume of fuel in a car tank can be between 0 litres and 55 litres (e.g. 42.38
litres)
Measurement Scales
Data can also be classified in terms of its scale of measurement. This indicates the
’strength’ of the data in terms of how much arithmetic manipulation on the data is
possible.
There are four types of measurement scales
Nominal data
Ordinal data
Interval data
Ratio data
Nominal data & Ordinal data
Nominal and Ordinal data are associated with categorical data.
Nominal data
For Nominal data, all the categories of a qualitative random variable are of equal
importance.
Examples:
gender (1 = male; 2 = female)
city of residence (1 = Pretoria; 2 = Durban; 3 = Cape Town; 4 = Bloemfontein)
home language (1 = Xhosa; 2 = Zulu; 3 = English; 4 = Afrikaans; 5 = Sotho)
Ordinal data
Ordinal data has an implied ranking between the different categories of the
qualitative random variable. Each consecutive category possesses either more or less
than the previous category of a given characteristic.
Examples:
size of clothing (1 = small; 2 = medium; 3 = large; 4 = extra large)
income category (1 = lower; 2 = middle; 3 = upper)
Interval data & Ratio data
Interval data and Ratio data are associated with numeric data and quantitative
random variables.
Interval data
Interval data is generated mainly from rating scales, which are used in survey
questionnaires to measure respondents’ attitudes, motivations, preferences and
perceptions.
Example:
How would you rate your chances of promotion after the next performance
appraisal?
Very poor Poor Unsure Good Very good
1 2 3 4 5
The performance appraisal system is biased in favour of technically oriented
employees.
Strongly disagree Disagree Unsure Agree Strongly agree
1 2 3 4 5
Ratio data
Ratio data
Ratio data consists of all real numbers associated with quantitative random
variables. Ratio data has all the properties of numbers (order, distance and an
absolute origin of zero) that allow such data to be manipulated using all arithmetic
operations. It is the strongest data for statistical analysis.
Example:
employee ages (years)
distance travelled (km)
Customer income (R)
machine speed
tyre pressure
product prices (R),
More statistical methods can be applied to ratio data than to any other data type.
Classification of data types