0% found this document useful (0 votes)
13 views87 pages

Introduction to Business Statistics

The document provides an introduction to statistics, defining it as the art and science of collecting, analyzing, presenting, and interpreting data. It discusses the significance of statistics in business, including its applications in decision-making, forecasting, and quality control. Additionally, it covers types of data, scales of measurement, and the distinction between descriptive and inferential statistics.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
13 views87 pages

Introduction to Business Statistics

The document provides an introduction to statistics, defining it as the art and science of collecting, analyzing, presenting, and interpreting data. It discusses the significance of statistics in business, including its applications in decision-making, forecasting, and quality control. Additionally, it covers types of data, scales of measurement, and the distinction between descriptive and inferential statistics.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd

INTRODUCTION

TO STATISTICS
What is Statistics?
■ Look at these statements
– For the past 30 days, the benchmark index of BSE, SENSEX
has been growing by an average of 250 points daily (Personal
observation)
– According to a survey by Statista, an average 62% of
generation Z millennials, either own a business or want to
start a venture of their own ([Link])
– According to World Bank, 70% of Indian population have bank
accounts, but 50% of these people do not access these
accounts (World Bank’s LDB on Financial Inclusion)
– India's foreign exchange reserves stood at US$ 575.29 billion
as of November 20, 2020, according to data from RBI.
([Link])
■ All these statements are using statistics to either make a point or
What is Statistics?
■ From this view, the term statistics refers to numerical facts
such as averages, medians, percentages, and index numbers
that help us understand a variety of business and economic
situations (Anderson, Sweeney and Williams, 2014)
■ A broader definition can be
■ “The art and science of collecting, analyzing, presenting, and
interpreting data”. (Anderson, Sweeney and Williams, 2014)
Definition:
Croxton and Cowden –
“Statistics may be defined as the collection, presentation,
analysis and interpretation of numerical data.“

Seligman –
“Statistics is the science which deals with the methods of
collecting, classifying, presenting, comparing and
interpreting numerical data collected to throw some light
on any sphere of enquiry."
Are the following statements statistical
data?
i) Weekly wages of 100 workers of a factory.

ii) Height of Ram is six feet.

iii) Mohan's weight is 70 Kgs, Sohan's height is 6.2 feet, and


Ram's monthly income is Rs. 1,500.

iv) Sales of a company during the past 10 years.


Significance of statistics in
Business
■ Business managers continually strive to get a better knowledge
about the business and economic environment
■ The understanding of statistics in business is intended to
– Give an idea of presentation and description of
information
– Understand the process of deriving inferences about a
population based on study of a sample
– Introduce the utilization of data to make more informed and
better decisions for business improvement and policy
formulation
– Guide how to use the data for forecasting purposes
■ Such information is provided by collecting, organizing, analyzing,
presenting, and interpreting data
Applications of Statistics in
business
■ Sampling of accounts for audit
■ Statistical information related to stock returns for
forecasting
■ Information collected by market research firms for test
marketing
■ Data collected for quality control purposes
■ Statistical information about economy such as index
numbers
Data and its components
■ Data - Facts and figures collected for statistical analysis and
interpretation.
■ All the data collected in a particular study are referred to as the data
set for the study. For e.g. a data set containing prices of shares over a
period of time
■ Elements are the entities on which data are collected. For e.g. the
shares for which price data is collected
■ A variable is a characteristic of interest for the elements. For e.g. for
the data set mentioned above the prices of shares are variable
■ Measurements collected on each variable for every element in a study
provide the data.
■ The set of measurements obtained for a particular element is called
an observation. For e.g. the price of share of RIL on a date is Rs 1925
Data, Data Sets,
Elements, Variables, and Observations
Observation Variables

Element
Names Stock Annual Earn/
Company Exchange Sales(Rs.M) Share(Rs.)

Dataram BSE 73.10 0.86


EnergySouth NSE 74.00 1.67
Keystone Nasdaq 365.70 0.86
LandCare MCX 111.40 0.33
Psychemedics N 17.60 0.13

Data Set
Measurement????

Can’t be measured can’t be improved???


Scales of Measurement

Data

Categorical Quantitativ
e

Nomina
Nomina Ordina Interval Ratio
ll l
Scales of Measurement

Scales
Scales of
of measurement
measurement include:
include:
Nominal Interval

Ordinal Ratio

The
The scale
scale determines
determines thethe amount
amount of
of information
information
contained
contained in
in the
the data.
data.

The
The scale
scale indicates
indicates the
the data
data summarization
summarization and
and
statistical
statistical analyses
analyses that
that are
are most
most appropriate.
appropriate.
Scales of Measurement
■ Nominal

Data
Data are
are labels
labels or
or names
names used
used to
to identify
identify an
an
attribute
attribute of
of the
the element.
element.

A
A nonnumeric
nonnumeric label
label or
or numeric
numeric code
code may
may be
be used.
used.

Example: Employment Classification


1 for Educator
2 for Construction Worker
3 for Manufacturing Worker
Example: Ethnicity
1 for African-American
2 for Anglo-American
3 for Hispanic-American
Scales of Measurement
 Nominal

Example:
Example:
Students
Students of
of aa university
university are
are classified
classified by
by the
the
school
school in
in which
which they
they are
are enrolled
enrolled using
using aa
nonnumeric
nonnumeric label
label such
such as
as Business,
Business, Humanities,
Humanities,
Education,
Education, and
and soso on.
on.
Alternatively,
Alternatively, aa numeric
numeric code
code could
could be
be used
used for
for
the
the school
school variable
variable (e.g.
(e.g. 1
1 denotes
denotes Business,
Business,
2
2 denotes
denotes Humanities,
Humanities, 3 3 denotes
denotes Education,
Education, and
and
so
so on).
on).
Scales of Measurement
■ Ordinal

The
The data
data have
have the
the properties
properties of
of nominal
nominal data
data and
and
the
the order
order or
or rank
rank of
of the
the data
data is
is meaningful.
meaningful.

A
A nonnumeric
nonnumeric label
label or
or numeric
numeric code
code may
may be
be used.
used.
Example: Ranking productivity of employees
Example: Taste test ranking of three brands of soft drink
Example: Position within an organization
1 for President
2 for Vice President
3 for Plant Manager
4 for Department Supervisor
5 for Employee
Scales of Measurement
■ Ordinal

Example:
Example:
Students
Students of
of aa university
university are
are classified
classified by
by their
their
class
class standing
standing using
using aa nonnumeric
nonnumeric label
label such
such asas
Freshman,
Freshman, Sophomore,
Sophomore, Junior,
Junior, or
or Senior.
Senior.
Alternatively,
Alternatively, aa numeric
numeric code
code could
could be
be used
used for
for
the
the class
class standing
standing variable
variable (e.g.
(e.g. 1
1 denotes
denotes
Freshman,
Freshman, 2 2 denotes
denotes Sophomore,
Sophomore, and and so
so on).
on).
Ordinal Data

Faculty and staff should receive preferential


treatment for parking space.

Strongly Agree Neutral Disagree Strongly


Agree Disagree

1 2 3 4 5
Scales of Measurement
■ Interval

The
The data
data have
have the
the properties
properties of
of ordinal
ordinal data,
data, and
and
the
the interval
interval between
between observations
observations is
is expressed
expressed in
in
terms
terms of
of aa fixed
fixed unit
unit of
of measure.
measure.

Interval
Interval data
data are
are always
always numeric.
numeric.
Scales of Measurement
■ Interval

Example:
Example:
Tushar
Tushar has
has an
an SAT
SAT score
score of
of 1205,
1205, while
while Priya
Priya
has
has an
an SAT
SAT score
score of
of 1090.
1090. Tushar
Tushar scored
scored 115
115
points
points more
more than
than Priya.
Priya.

Example: Fahrenheit Temperature


Example: Calendar Time
Scales of Measurement
■ Ratio

The
The data
data have
have all
all the
the properties
properties of
of interval
interval data
data
and
and the
the ratio
ratio of
of two
two values
values is
is meaningful.
meaningful.

Variables
Variables such
such as
as distance,
distance, height,
height, weight,
weight, and
and time
time
use
use the
the ratio
ratio scale.
scale.

This
This scale
scale must
must contain
contain aa zero
zero value
value that
that indicates
indicates
that
that nothing
nothing exists
exists for
for the
the variable
variable at
at the
the zero
zero point.
point.

Example: Monetary Variables, such as Profit and Loss, Revenues, and


Expenses
Example: Financial ratios, such as P/E Ratio, Inventory Turnover, and
Quick Ratio.
Scales of Measurement
■ Ratio

Example:
Example:
If
If we
we compare
compare the
the cost
cost of
of Rs.
Rs. 30000
30000 for
for one
one
automobile
automobile toto the
the cost
cost of
of Rs.
Rs. 15000
15000 for
for aa second
second
automobile,
automobile, the
the ratio
ratio property
property shows
shows that
that the
the first
first
automobile
automobile isis Rs
Rs 30000/Rs
30000/Rs 15000=
15000= 2 2 times
times the
the cost
cost
of
of the
the second
second one.
one.
Tushar’s
Tushar’s college
college record
record shows
shows 3636 credit
credit hours
hours
earned,
earned, while
while Priya’s
Priya’s record
record shows
shows 72
72 credit
credit
hours
hours earned.
earned. Priya
Priya has
has twice
twice as
as many
many credit
credit
hours
hours earned
earned as
as Tushar.
Tushar.
Scales of Measurement
Type of Type of data Example
Scale
Nominal Labels or names used to identify Name of students in a class,
Scale an attribute of the element their roll numbers or their
gender

Ordinal Rank order data or ratings. Ranking of students on the basis


Scale Distance between variables of marks, rating given to
cannot be calculated products on e - commerce app

Interval Data have all the properties of Total marks in senior secondary
Scale ordinal data and the interval of a group of students which can
between values is expressed in be ranked in a decreasing order
terms of a fixed unit of measure. and where the difference
between the marks of two
students is on a common scale
Primary scales of measurement
Nominal Numbers Finish
Assigned
7 8 3
to Runners

Ordinal Rank Order Finish


of Winners
Third Second First
place place place

Interval
Performance 8.2 9.1 9.6
Rating on a

0 to 10 Scale
15.2 14.1 13.4

Ratio Time to
Illustration of primary scales of measurement

Nominal Ordinal Interval


Ratio
Scale
Scale Scale
Preference
Scale Ratings
Preference 1-7 11-17 $
spent last No. 7Store 79 5 Rankings
15 0
3 months
2 25 7 17 200
8 82 4 14 0
[Link] & Taylor 3 30 6 16 100
[Link]’s 1 10 7 17 250
[Link] 5 53 5 15 35
[Link]’s 9 95 4 14 0
5.J.C. Penney 6 61 5 15 100
[Link] Marcus 4 45 6 16 0
[Link] 10 115 2 12 10
8. Saks Fifth Avenue
9. Sears
Basic Statistical concepts
■ Population and sample
– Population is a collection of all the elements under statistical
study
– Sample is a portion of population which is selected in a way
so as to be a representative of the population
■ Descriptive and inferential statistics
– Summated measures arrived from the data are called
descriptive statistics while inferring the characteristics of
population using sample statistics is inferential statistics
Descriptive vs. Inferential Statistics

■ Descriptive Statistics — using data gathered on a group to


describe or reach conclusions about that same group only

■ Inferential Statistics — using sample data to reach conclusions


about the population from which the sample was taken
Classification of Data based on
sources
■ Primary Data
– Data that has been observed, experienced or recorded
close to the event are called primary data
– Data collected for the first time to address a specific
research problem
■ Secondary Data
– Sometimes the data required to address a particular
research problem is already available
– Such data is easy to search nowadays, given the
technological advancement
– Such data is available in form of various publications
Methods of primary data
collection
■ Survey using Questionnaires
■ Conducting interviews
■ Observing without getting involved
■ Participative observation/immersing oneself in a situation
■ Doing experiments
Survey using a questionnaire
■ Questionnaires are a particularly suitable tool for gaining quantitative data but can
also be used for qualitative data.
■ Useful when the researcher requires structured data. Some advantages of this
method are
– Structured format, can be customized easily according to the data requirements
– Easy and convenient for respondents
– Cheap and quick to administer to a large number of cases covering large
geographical areas
■ There are three methods of delivering questionnaires, personally, by post or through
the Internet
■ Questions can be
– Close ended – Single choice, true/false, multiple choice, scale
– Open ended
■ Research in social sciences, politics, business, healthcare etc. often needs to gain
the opinions, feelings and reactions of a large number of people, most easily done
with a survey.
Conducting interviews
■ Interviews are more suitable for questions that require probing to obtain
adequate information.
■ Structured interview – standardized questions read out by the interviewer
according to an interview schedule. Answers may be closed format.
■ Unstructured interview – a flexible format, usually based on a question
guide but where the format remains the choice of the interviewer. No closed
format questions.
■ Semi-structured interview – one that contains structured and unstructured
sections with standardized and open type questions
■ Face-to-face interviews – Allows the interviewer to probe even deeper and
offer clarity wherever required
■ Focus Group – Group interview which tends to concentrate in depth on a
particular theme or topic with an element of interaction.
■ Telephonic interviews - Avoid the necessity of travelling to the respondents
and can therefore be carried out more quickly than face-to-face
Observation and participation
■ Observation
– The aim is to take a detached view of the phenomena, and
be ‘invisible’, either in fact or in effect
– When studying humans or animals, this detachment assumes
an absence of involvement in the group even if the subjects
are aware that the observation is taking place.
– Use of all five senses to collect data
■ Participation
– The researcher usually disguises himself/herself to study a
particular phenomenon
– Method mostly used in social sciences
Experiments
■ Usually conducted in controlled atmosphere to find the causal relationships
■ Most expensive method of data collection
■ Used to find the interaction between
– One thing and another thing
– People and things
– People and people
■ The design of experiments and models depends very much on the type of
event investigated, the sort of variables involved, and the level of accuracy
and reliability aimed at practical issues such as time and resources
available.
Secondary data
■ The advantage of using sets of secondary data is that it has been produced
by teams of expert researchers, often with large budgets and extensive
resources, so it cuts out the need for time consuming fieldwork.
■ Secondary data can also be used to compare with primary data you may
have collected, in order to triangulate the findings and put your data into a
larger context
■ Before using data from secondary sources the following issues must be
addressed
– Locating and accessing the data
– Authenticating the sources
– Assessing credibility
– Gauging how representative they are
– Selecting methods to interpret them.
Types of Secondary Data
■ Written materials – organizational records such as internal reports,
annual reports, production records, personnel data, committee
reports and minutes of meetings; communications such as emails,
letters, notes; publications, such as books, journals, newspapers,
advertising copy, government publications of all kinds etc.
■ Non-written materials – television programmes, radio programmes,
tape recordings, video tapes, films of all types, including
documentary, live reporting, interviews, etc. works of art, historical
artefacts etc.
■ Survey data – government census of population, employment,
household surveys, economic data, organizational surveys of
markets, sales, economic forecasts, employee attitudes. These may
be carried out on a periodic basis, with frequent regularity or
continuously, or ad hoc or one-off occasions. They may also be
limited to sector, time, area.
Sources of Secondary Data
■ Online data sets
■ Cultural texts
■ Libraries, museums and other archives
■ Commercial and professional bodies
■ Government publications
Selection of appropriate method
of data collection
■ Nature, object and scope of enquiry
■ Level of accuracy desired
■ Availability of funds
■ Time available
GRAPHICAL
PRESENTATION
OF DATA
Introduction
• Graphical presentation of frequency distribution facilitate easy understanding
of data presentation and interpretation.
• The shape of graph offers easy answers to several questions.
• It offers an easy technique for quick and effective comparison between two
or more frequency distributions.
• Advantages-
- Diagrams given an effective and elegant presentations.
- Diagrams leave good visual impact.
- Diagrams facilitate comparison.
- Diagrams save time.
- Diagrams simplify complexity and depict the characteristics of the data.
• Limitations-
-They provide only an approximate picture of the data.
- They cannot be used as alternative to tabulation of data.
- Problems to select a suitable method.
- Loss of accuracy of data- Sometimes Illusionary data effect
creates a wrong impression on the minds of the viewer.
Types of Diagrams
Bar charts

Types
Frequenc of Frequenc
y Curve
Diagra y Polygon

ms

Cumulativ
e
Frequenc
y
Distributi
on
(Ogive)
Simple Bar Charts

• Bar charts are used to represent only one characteristics


of data and there will be a many bars as number of
observations.
• The bars are of the same width and only the length
varies, the relationship among them can be easily
established.
• Data is presented via vertical or horizontal columns.
•Example- The data on the production of oil seeds in a particular year is
presented:
Oil seeds Yield (Million tones)
Ground nut 5.80
Rapeseed 3.30
Coconut 1.18
Cotton 2.20
Soya bean 1.00

Yield (Million tones)


7
6
5.8
5
4
3 3.3
2 2.2
1 1.18 1
0
Ground nut Rapeseed Coconut Cotton Soya bean
The following data gives the information of the number of children involved in different activities.
Activities No. of children
Dance 30
Music 40
Art 25
Cricket 20
Football 53
Histograms
• Also known as Area Diagram.
• A histogram is the most commonly used graph to show frequency distributions.
• Value of variables (the characteristics to be measured) are scaled along the horizontal
axis and the number of observations (or frequencies) along the vertical axis of the
graphs.
• Convert class limits to class boundaries. (Basically it should be an exclusive series)
Height Range (Class Interval) Number of Trees
(Frequencies)
60-65 3
66-70 3
71-75 8
76-80 10
81-85 5
86-90 1
Multiple bar charts

• A multiple bar chart is also known as groups (or compound) bar


charts.
• Such charts are useful for direct comparison between two or more
sets of data.
• The technique of drawing such a chart is same as that of a single bar
chart with difference that each set of data is represented in
different shades or colors on the same scale.
• An index explaining shades or colors must be given.
■ Example- The data on fund flow (Rs. crore) of an International Airport
Authority during financial years 2001-02 to 2003-04 are given below:

2001-2002 2002-2003 2003-2004


Non- traffic 40.00 50.75 70.25
revenue
Traffic revenue 70.25 80.75 110.00
Profit before tax 40.15 50.50 80.25
120

100

80

60

40

20

0
Non- traffic revenue Traffic revenue Profit before tax

2001-2002 2002-2003 2003-2004


Populations Over Time (millions)
Country 1980 1990 2000
France 55 56 65
United Kingdom 50 53 63
Mexico 65 78 80
Nigeria 60 82 85
Pakistan 57 65 74
Sub-divided bar chart

• Sub-divided bar charts are suitable for expressing information in terms of


ratios or percentage.
• Different shades must be used to represent various ratio values but the
shade of each component should remain the same in all the other bars.
• An index of the shades should be given with the diagram.
• Sub-divided bar chart is used to represent the data in which total magnitude is
divided into different components.
• Used to present the data having two or more components.
■ Example- The data on sales (Rs. million) of a company are given below:

2005 2006 2007


Export 1.4 1.8 2.29
Home 1.6 2.7 2.9
Total 3.0 4.5 5.18
6

0
2005 2006 2007

Export Home
•Example- The data shows number of students in college A and college B that use
mobile phones from Samsung, Oppo and Apple.

Products College A College B


Samsung 590 800
Oppo 880 750
Apple 100 150
Total Magnitude 1570 1700
Frequency Polygon
• The frequency polygon is formed by marking the mid-points at the top of the
horizontal bars and then joining these dots by a series of straight line.

STEPS TO BE FOLLOWED:
Step 1: Draw Histogram
Step 2: Calculate the mid-points = (Lower limit+ Upper limit)/2
Step 3: Mark all the mid points on the horizontal axis.
Step 4: Join all the plotted points using a line segment.
Step 5: This resulting curve is called the frequency polygon.
Height Range (Class Interval) Number of Trees
(Frequencies)
60-65 3
66-70 3
71-75 8
76-80 10
81-85 5
86-90 1
Cumulative frequency distribution (ogive)

• A cumulative frequency curve popularly known as Ogive is another form of


graphic presentation of a cumulative frequency distribution.
• Y- axis= Total frequencies, X-axis= Upper limits in case of less than ogive; Lower
limits in case of more than ogive.
• Example-

Class Upper Frequency CF (Less CF (More


Interval class than) than)
10-15 15 6 6 40
15-20 20 11 6+11= 17 40-6= 34
20-25 25 9 17+9= 26 34-11=23
25-30 30 7 26+7= 33 23-9= 14
30-35 35 5 33+5= 38 14-7= 7
35-40 40 2 38+2= 40 7-5=2
More than Less than
type type

Ogive
45

40 40 40
38
35 34 33
30

25 26
23
20
17
15 14
10
6 7
5
2
0
15 20 25 30 35 40

CF (Less than) CF (More than)


Pie Chart
• These diagrams are normally used to show the total number of
observation of different types in the data set on a percentage basic
rather on an absolute basis through circle.
• Also known as circle chart.
• Uses of Pie Chart-
- Within a business, it is used to compare areas of growth, such as
turnover, profit, and exposure.
- To represent categorical data.
Example- The distribution of sales of the laptop industry between five
companies:
Company % market share
HP 22
Dell 33
Lenovo 13
Apple 15
Acer 17
INTRODUCTION
TO FREQUENCY
DISTRIBUTIONS
Meaning
• Process of arranging data in groups/classes on the basis of certain
properties.
• Purpose-
1. Condense the raw data into a form suitable for statistical analysis.
2. Removes complexities and highlights the feature of the data.
3. Facilities comparisons and drawing inference from the data.
4. Provides information about the mutual relationships among elements of
a data set.
5. Helps in statistical analysis by separating the elements of the data set
into homogeneous groups and hence brings out the point of similarity
and dissimilarity.
Basis of classification
• Generally data are classified on the basis of following four bases:
1. Geographical Classification:
- Data is classified on the basis of geographical or locational difference such as
cities, districts, or villages between various elements of data set.
- Example-

City Mumbai Kolkata Delhi Chennai


Population 654 685 423 205
Density (per
square km)
- Elements can be classified alphabetically or based on the frequency size.
■ 2. Chronological Classification:
■ - Data classified on the basis of time is known as chronological classification.
■ - Also known as time series.
■ - In such a classification, data are classified either in ascending or in
descending order with reference to time such as years, quarters, months, weeks,
etc.
■ - Example
Year 1941 1951 1961 1971 1981 1991 2001
Populatio 31.9 36.9 43.9 54.7 75.6 85.9 98.6
n (crore)
■ 3. Qualitative Classification:
■ - Data are classified on the basis of descriptive characteristics or on the basis of
attributes like gender, literacy, region, caste or education, which cannot be quantified.
■ - Done in two ways:
■ 1. Simple Classification-
• Each class is sub divided into two sub-classes and only one attribute is studied -
Dichotomous.
• Example- Male and Female, Blind and Not blind, Educated and Uneducated.
■ 2. Manifold Classification-
• A class is subdivided into more than two sub-classes which may be sub-divided
further.
• Example- Population in a country can be classified in terms of gender as male and
female. These two sub-classes may be further classified in terms of literacy as literate
and illiterate.
A manifold classification

Populatio
n

Male Female

Uneducat Uneducat
Educated Educated
ed ed

Rural Urban Rural Urban Rural Urban Rural Urban


■ 4. Quantitative Classification:
• Data are classified on the basis of some characteristics which can be measured on
numerical scale. Example- Height, weight, income, expenditure, production, sales.
• Two types of Quantitative classification-
1. Discrete (or discontinuous) data -
• Discrete Data can only take certain values. It is a data obtained by counting.
• Example: the number of students in a class, the results of rolling 2 dice. Number
of children in a family.
Discrete Series
Number of Children Number of
Families
0 10
1 30
2 60
3 90
4 110
5 20
■ 2. Continuous data-
• Continuous Data can take any value (within a range). It is a data obtained by measuring.
• Example- A person’s weight: could be any value (within the range of human weights),
not just certain fixed weight; Time in a race: you could even measure it to fractions of a
second.

Continuous Series
Weights (kgs) Number of
Persons
40-50 10
50-60 20
60-70 25
70-80 35
80-90 50
Identify the type of distribution

■ No. of accidents on certain days of a ■ Discrete


month –
■ Discrete
■ No. of shares sold on a day in market
■ Continuous
■ Lengths of 1,000 bolts produced at a
■ Discrete
factory
■ Continuous
■ No. of books on a library shelf
■ Discrete
■ Speed of an automobile in Kph
■ No. of individuals in a family
Difference between Discrete and Continuous Data

Discrete Data Continuous Data

Expressed in whole numbers and not Can take any numeric value including
fractions. fractions and decimals.

Countable. Measureable.

Broken or separate data. Unbroken data.

When plotted on a graph, points are When points are plotted on a graph, a
isolated and as such may or may not pattern is formed either of a straight
form a pattern. line or a curvy line.
Example- number of months in a year, Example- time, weight, height,
number of days in a week. temperature.
Frequency Distribution
■ An effort is made to distribute the data into classes or categories and to
determine the number of elements belonging to each class or category to
derive meaningful information out of the data collected
■ According to Lawrence Lapin “Frequency of occurrence lies at the heart of
statistical analysis. It allows us to create order from confusion.
Meaningfully arranged numbers can tell us a story and help us choose
methods and procedures for generalizing about populations.”
Meaning of Frequency
Distribution
• Frequency- It is the number of times the observation occurred/recorded in a
study.
• Example- Sam played football on:
■ - Saturday Morning,
■ - Saturday Afternoon
■ - Thursday Afternoon
The frequency was 2 on Saturday, 1 on Thursday and 3 for the whole week.
• Frequency Distribution- It is a list, table or graph that displays the frequency
of various outcomes in a sample.
• Thus, frequency distribution refers to a table that shows an item and its
frequency.
Some Definitions
■ Morris Humburg
■ “A frequency distribution or frequency table is simply a table in which the
data grouped into classes and the number of cases which fall in each
class are recorded. The numbers in each class are referred to as
frequencies.”
■ Croxton and Cowden
■ “Frequency distribution is a statistical table in which different values of
variable are shown in the sequence of magnitude along with
corresponding frequencies.”
■ Murray. R. Spiegal
■ “A tabular arrangement of data by class together with the corresponding
class frequencies is called a frequency distribution or frequency table”
• Examples-
- A school conducted a blood donation camp. The blood groups of 30 students
were recorded as follows.
- A, B, O, O, AB, O, A, O B, A, O, B, A, O, O, A, AB, O, A, A, O, O, AB, B, A, O, B, A, B,
O.
- We can present this data in Group
Blood a tabular form. Number of Students

Total

- This table is known as a frequency distribution table.


- Example (Business Situation)- A marketing manger wants to know how many
units of each product sells in a particular region during a given period.
Advantages of frequency
distribution
• Present raw data in an organized and easy-to-read format.
• Data are expressed in a more compact form.
• One can quickly note the pattern of distribution of observations falling in various
classes.
• Permits the use of more complex statistical techniques which helps reveal
certain hidden characteristics of the data.
Disadvantages of frequency
distribution

• As we group the data, individual observations lose their


identity.
• Statistically calculations are based only on the values of
the class mark and not on the values of the observation
in that class.
Organizing the data
• To examine large set of numerical data is first to organize and present it in an
appropriate tabular and graphical format.
Raw Data Pertaining to total time hours worked by Machinists
94, 89, 88, 89, 90, 94, 92, 88, 87, 85, 88, 93,
94, 93, 94,
93, 92, 88, 94, 90, 93, 84, 93, 84, 91, 93, 85,
• This raw89,
91, data95
do not highlight any characteristics/trend, such as highest,
lowest, average weekly hours.
• Unless these data are reorganized, no meaningful inferences can be drawn.
• The raw data can be reorganized in a data array and frequency distribution.
• When a raw data set is arranged in rank order, from smallest to largest
observation or vice versa, the ordered sequence obtained is called an ordered
array.
Raw Data Pertaining to total time hours worked by Machinists

■ 1. Provides quick look at the highest and lowest observation in the data within
which individual values vary.
■ 2. Helps in dividing the data into various sections or parts.
■ 3. Helps us to know the degree of concentration around a particular observation.
■ 4. Helps to identify whether any values appear more than once in the array.
Array and Tallies
Number of Total time Tally Number of Workers
Hours (Frequency)
84
85
86
87
88
89
90
91
92
93
94
95
Types of frequency distribution
Exclusive
FD
Inclusive FD

Types
Open End
FD
Cumulative
FD
Mid Values
FD
Exclusive Method
• When the data are classified in such a way that the upper limit of a class
interval is the lower limit of the succeeding class interval, then it is said to
be the exclusive method of classifying data.
Table: Exclusive
Dividends Declared method of
in % (Class data classification
Number of companies
Interval) (Frequencies)
0-10 5
10-20 7
20-30 15
30-40 10
• Such classification ensures continuity of data as the upper limit of one class is the
lowerDividends
limit of succeeding
Declared class.
in % (Class Number of companies
Interval) (Frequencies)
0 but less than 10 5
10 but less than 20 7
20 but less than 30 15
Inclusive method
• When the data are classified in such a way that both lower and upper limits of
a class interval are included in the interval itself, then it is said to be the
inclusive method of classifying data.

Number Table: Inclusive


of Accidents MethodNumber
(Class of Data of
Classification
Weeks
Interval) (Frequencies)
40-44 5
45-59 22
60-74 13
75-89 8
90-104 2
• Certain adjustment in the class interval is needed to obtain the continuity.
• To ensure continuity, first calculate correction factor:
x= Upper limit of a class- Lower limit of the next higher class/ 2
• Subtract the correction factor (x) from the lower limit and add to the upper limit
of all the classes.
• For Example X=
Number of Accidents (Class Number of Weeks
Interval) (Frequencies)
Cumulative Frequency
Distribution
• Cumulative Frequency series is that series in which the frequencies are
continuously added corresponding to each class interval in the series.
• Distribution which shows the cumulative number of observations below
the upper limit of each frequency distribution.
• A cumulative frequency distribution is of two types:
1. Less than type (Upper limit)
2. More than type (Lower limit)
• Less than cumulative frequency distribution-
- The frequencies of each class interval are added successively from top
to bottom.
- Represents the cumulative number of observations less than or equal to
the class frequency to which it relates.
■ Example:
■ Cumulative Frequency Distribution
Number of Number of Weeks Cumulative
Accidents Frequency
(Less than)
0-4 5
5-9 22
10-14 13
15-19 8
20-24 2
• More than Cumulative Frequency Distribution-
- The frequencies of each class interval are added successively from bottom to top.
- Represents the cumulative number of observations greater than or equal to class
frequency to which the relates.
- Example-
Cumulative Frequency Distribution
Number of Number of Weeks Cumulative
Accidents Frequency
(More than)
0-4 5
5-9 22
10-14 13
15-19 8
20-24 2
Number of Cumulative
Accidents Frequency
(Upper limits) (Less than)
Less than 4
Less than 9
Less than 14
Less than 19
Less than 24

Number of Accidents Cumulative Frequency


(Lower limits) (More Than)
More than 0
More than 5
More than 10
More than 15
More than 20
Thank You

You might also like