0% found this document useful (0 votes)
22 views15 pages

Business Data Analysis: Statistics Overview

The document provides an extensive overview of statistics, its history, definitions, stages, and importance across various fields. It discusses methods of data collection, including census and sampling techniques, as well as the classification and presentation of data. Additionally, it outlines the limitations of statistics, the types of data, and the design of questionnaires for effective data gathering.

Uploaded by

Vrushali P.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
22 views15 pages

Business Data Analysis: Statistics Overview

The document provides an extensive overview of statistics, its history, definitions, stages, and importance across various fields. It discusses methods of data collection, including census and sampling techniques, as well as the classification and presentation of data. Additionally, it outlines the limitations of statistics, the types of data, and the design of questionnaires for effective data gathering.

Uploaded by

Vrushali P.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

BUSINESS DATA ANALYSIS

Module I: Introduction to Statistics


The subject “Statistics” is not a new discipline, it has originated as a science of statehood and
found applications slowly and steadily in Agriculture, Economics, Commerce, Biology,
Medicine, Industry, planning, education and so on. The word is derived from the Latin word
Status, means a political state. And, Sir Ronald A. Fisher is called as the Father of Statistics.

According to Spiegal , “Statistics is concerned with scientific methods for collecting,


organizing, summarizing, presenting and analysing data as well as drawing valid conclusions
and making reasonable decisions on the basis of this analysis”.
According to Secrist, “Aggregate of facts, affected to a marked extent by multiplicity of
causes, numerically expressed, enumerated or estimated according to reasonable standards of
accuracy, collected in a systematic manner for a predetermined purpose and placed in relation
to each other”.
Thus, Statistics mainly involves four stages:
a) Collection of data
b) Presentation of data
c) Analysis of data
d) Interpretation of data

Nature of a Statistical Study


1) Formulation of the Problem
2) Objectives of the Study
3) Determining the sources of Data
4) Designing Data Collection Forms (Observational or Survey method)
5) Conducting the Field Survey (Censes or Sample survey)
6) Organising the Data (Tables/Charts/Graphs)
7) Analysing the Data (Statistical Techniques)
8) Reaching Statistical Findings
9) Presentation of Findings (Oral or Written)

Importance and Scope of Statistics


In ancient times statistics was used only as the science of Statecraft for devising military and
fiscal policies. But today the scope of statistics has widened to include the social as well as
economic phenomenon. Statistics is been used in various disciplines such as Economics,
Accountancy, Industry, Physical sciences, Social sciences, Business & Management, etc.
The various disciplines where statistics is widely used are:
Planning State
Economics Business & Management
Accountancy & Auditing Industry
Physical Sciences Social Sciences
Biology and Medical Sciences Psychology and Education

Statistics in Business
The planning of The setting up of The function of
operations standards control

Statistical Quality Control Methods Personnel Management


Seasonal Behaviour Export Marketing
Maintenance of Cost Records Management of Inventory
Expenditure on Advertising & Sales Mutual Funds
Relevance in BFSI Institutions

Limitations of Statistics
 Statistics is not suitable to the study of qualitative phenomenon (intelligence, bravery,
taste, etc.)
 Statistics studies include Sampling more than Census.
 Statistical reveals the average behaviour that may not be applied to particular
situations or individuals.
 Statistics is not 100% precise as in Maths or Accounting.
 Statistics is only one of the methods of studying a problem
Subdivisions within Statistics
Statistics
Descriptive Statistics

Inferential Statistics

Descriptive Statistics: A collection of methods that enable us to organise, display and


describe data using various devices. E.g. Charts, Tables or Graphs, frequency distributions,
Averages, etc.
Inferential Statistics: A collection of methods that enable us in making decisions about a
population based on sample results. E.g. Correlation, Regression, Probability, etc.

COLLECTION OF DATA
One of the main functions of statistics is to provide information which will help in making
decisions. Thus, collection of data is the first step for any statistical investigation. For any
statistical enquiry, the crucial part is that of collecting facts and figures for the study
conducted. The person who conducts the statistical enquiry is known as the Investigator. The
persons from whom the information is collected are known as Respondents. The items on
which the measurements are taken are called as the Statistical Units. The Process of counting
or measurement together with the systematic recording of results is called the Collection of
Statistical Data.
Before collection of data for any statistical enquiry, it is essential to examine carefully the
following points.
a) Objectives & scope of the enquiry
b) Statistical units to be used
c) Sources of information (data)
d) Method of data collection
e) Degree of accuracy aimed at in the final results
f) Type of enquiry

Method of Data Collection


This step is unnecessary if we are selecting secondary data as our source of information.
But, if primary data are to be collected, then deciding the method of collecting data becomes
important. There are two techniques available for this purpose:
(a) Census Method (b) Sampling technique
Census Method:
Under this method, 100 % inspection of the population is conducted and each and every unit
of the population is enumerated. In other words, every element of the population is included
in the investigation. For example, if we study the average annual income of the families of a
particular city or area, and if there are 1000 families in that area, we must study the income of
all 1000 families.
Merits of Census Method:
 The data are collected from each and every item of the population.
 The results are more accurate and reliable, because every item of the universe is required.
 Intensive study is possible.
 The data collected may be used for various surveys, analyses etc.
Limitations of Census Method:
 It requires a large number of enumerators and it is a costly method.
 It requires more money, labour, time energy etc.
 It is not possible in some circumstances where the universe is infinite.

Sampling Technique
Under this method, only a selected representative and adequate fraction of the population is
inspected or studied. The word Sample is used to describe a portion chosen from the
population. It is a finite subset of statistical individuals defined in a population. The number
of units in a sample is called the sample size.
In a statistical enquiry, all the items which fall within the purview of enquiry, are known as
Population or Universe. A sample is a portion of the population selected for analysis. A
parameter is a numerical measure that describes a characteristic of a population. A statistic is
a numerical measure that describes a characteristic of a sample.
Merits of Sampling Technique
 Sampling saves time and labour.
 It results in reduction of cost in terms of money and man-hour.
 Sampling ends up with greater accuracy of results.
 It has greater scope and greater adaptability.
Demerits of Sampling Technique
 Sampling is to be done by qualified and experienced persons. Otherwise, the information
will be unbelievable.
 Sample method may give the extreme values sometimes instead of the mixed values.
 There is the possibility of sampling errors. Census survey is free from sampling error.
Types of Data
The data collection required for any statistical survey can be collected from the following
sources: (a) Primary Data (b) Secondary Data

Primary Data:
The data which are originally collected by an investigator or agency for the first time for any
statistical investigation and used by them in the statistical analysis are termed as “Primary
Data.” In other words, it is the one, which is collected by the investigator himself for the
purpose of a specific study. Such data is original in character and is generated by survey
conducted by individuals, research institution or any organization.
Secondary Data:
The data (published or unpublished) which have already been collected and processed by
some agency or person and taken over from there and used by any other agency for their
statistical work are termed as “Secondary Data”.
In other words, it is that data which has been already collected and analysed by some earlier
agency for its own use; and later the same data is used by a different agency.

Methods of Collecting Primary Data


a) Direct personal investigation / interviews:
This method consists of collection of data personally by the investigator from the sources
concerned. Here, the investigator has to go to the field personally for making enquiries
and asking for information from respondents. This method should be used only of the
investigation is generally local (confined to a locality, region, or area).
b) Indirect Oral interviews:
When “Direct Personal Interview” is not suitable or possible, an indirect oral
investigation is carried out. In this method, the information is obtained by interviewing
the respondent’s friends’, relatives, etc. who know him thoroughly well. In this method,
the data is collected through enumerators, who are specially appointed for this purpose.
c) Information from correspondents/ through local agencies:
In this method the information is neither collected by the investigator nor the
enumerators. Here, it is collected by the local agents (called Correspondents) These
agents are appointed by the investigator, in different parts of the field of enquiry. This
technique is usually employed by newspaper or periodical agencies who require
information in different fields, like sports, accidents, share market, etc.
d) Mailed questionnaire method:
This method consists of preparing a questionnaire, which is mailed to the respondents
with a request for quick response within the specified time. A Questionnaire is a list of
questions relating to the field of enquiry and providing space for the answers to be filled
by the respondents. This method is usually used by the research workers, private
individuals, non-official agencies and sometimes even by government.
e) Schedules sent through enumerators:
A schedule is a device of obtaining answers to the list of questions in a form. In this
method, the schedule is filled by the enumerators or interviewers in a face to face
situation with the respondents. This method is generally used by big business houses,
large public enterprises and research institutions.

Sources of Secondary Data


a) Published Sources
There are a number of national and international agencies which collect statistical data
relating to business, trade, prices, industries, population, income, etc., and publish
their findings in statistical reports on a regular basis.
Some of the published may include: Official Publications of Central Government,
Publications of Semi-Government Statistical Organizations, Publications of Research
Institutions, Publications of Commercial and Financial Institutions, Reports of
Various Committees & commissions appointed by the Government, Newspapers and
Periodicals, or International Publications.
b) Unpublished sources:
There are various sources of unpublished statistical data that are maintained by the
private firms or business enterprises. These data are not readily available.

Questionnaire
A questionnaire is a research instrument consisting of a series of questions for the purpose of
gathering information from respondents. They can be carried out face to face, by telephone,
computer or post.
Designing a Questionnaire
While designing a questionnaire, the following points have to be borne in mind:
a) Type of Information to be collected
b) Types of questions (Open-ended, Dichotomous & multiple-choice)
c) Phrasing of the Questions
d) Order of questions
e) How many questions to be asked
f) Layout of the questionnaire
Editing: Process by which data are prepared for subsequent coding process. Here errors and
omissions are examined in the collected data for & necessary corrections are made.
Coding: is the procedure of classifying the responses in a questionnaire into meaningful
groups with proper symbols indicating codes for each response.

CLASSIFICATION OF DATA
After the survey, the investigator will have huge data at hand. But, this data is known as raw
data, because it is generally voluminous, huge, clumsy and incomprehensible. Thus the next
step after data collection is to organize and present it in a proper reasonable state. The
presentation of data is broadly divided into the following two categories: Tabular
Presentation and Diagrammatic or Graphical Presentation.

Classification:
Before tabulating, the arrangement of the raw data into different homogenous groups is
necessary. The process of arranging the data into groups or classes according to resemblances
and similarities is called as Classification. Classification is the process of arranging data into
sequences and groups according to their common characteristics or separating them into
different but related parts. The technique of dividing the given data into different classes
w.r.t. more than one basis simultaneously is called Cross-Classification.
Functions of Classification:
 It condenses the data.
 It facilitates comparisons.
 It helps to study the relationships.
 It facilitates the statistical treatment of the data .
Rules for Classification
 It should be unambiguous.
 It should be exhaustive and mutually exclusive.
 It should be stable.
 It should be suitable for the purpose.
 It should be flexible.

Bases of Classification
 Geographical (area wise or regional):
Here the basis of classification is the geographical or locational differences between the
various items in the data like States, Cities, Regions, Zones, Areas, etc. E.g. Population of
Belgaum city, Yield of Wheat in Punjab, etc.
 Chronological (according to time)
Here the data are classified on the basis of differences in time. E.g. The profits of a
company for various years, the birth rate in the country for different years, etc.
 Qualitative (character or attribute)
Here the data are classified according to some qualitative characteristic which cannot be
measured in quantitative terms like intelligence, customer satisfaction, etc.
o If the data are classified into only two classes of the given attribute, like Yes or
No, the classification is termed as Simple or Dichotomous.
o If the given population is classified into more than two classes relating to an
attribute, then it is called Manifold classification.
 Quantitative (numerical values)
Here the data is classified on the basis of criteria which are capable of quantitative
measurement like age, height, income, etc.

Attribute: is a characteristic for which numerical measurements cannot be made. In simple


words, the criteria that cannot be measured or quantified are called as Attributes.
Variable: is a characteristic, number, or quantity that increases or decreases over time, or
takes different values in different situations. In simple words, a characteristic that can be
measured or quantified is called as Variable.

Kinds of Variables (a) Continuous variable (b) Discrete Variable


Continuous variable: those variables that can take all the possible values (integrals as well as
fractions) in a given range are termed as continuous variable. E.g. Weight, Time, etc.
Discrete Variable: The variables which cannot take all the possible values within a given
specified range are termed as discrete variables. E.g. Marks in a test, students in a class, etc.

Frequency Distribution
The arrangement and display of data in a form with the values of the variable paired with its
frequency is called a frequency distribution. The types of arrangement may be in any of the
following ways:
a) Individual observations (raw data):
When data is collected it is voluminous and incomprehensible, thus it is called raw data.
It is data that has not been processed for use. The data in raw form does not give any
useful information and is confusing. A better presentation of the data will be to arrange
them in an ascending or descending order of magnitude which is called Arraying of the
data.
b) Discrete or ungrouped frequency distribution:
This is a way of arranging the available data where the occurrence of each of the values
of the variable in the data is counted i.e. counting how many number of times each value
of the variable occurs in the data. This is facilitated through the technique of Tally Marks
or Tally Bars. This is better than arraying but still it does not condense the data much and
is quite cumbersome to grasp and understand and it will be suitable only if:
 The values of the variable are largely repeated.
 The variable under consideration takes only a few values.
c) Grouped frequency distribution:
This method consists of classifying the data into different classes (or class intervals) by
dividing all the variables into classes (a suitable number of groups) and then recording the
frequencies (number of observations in each group/class). E.g. Age: 20-25, 25-30, 30-35,
etc. The various groups into which the values of the variable are classified are known as
classes or class intervals. The length of the class interval is called width or magnitude of
the class. The two values specifying the class are class limits. The larger value is called
the upper class limit and the smaller value is called the lower class limit.
d) Continuous frequency distribution:
While dealing with a continuous variable it is not desirable to present the data into
grouped frequency distribution. Under this method, we form continuous class intervals
(without any gaps) of the following type: Below 10; 10 or more but less than 20; 20 or
more but less than 30; etc. The presentation of the data into continuous classes along with
the corresponding frequencies is known as Continuous frequency distribution.

Cumulative Frequency Distribution: is a modification of the given frequency distribution


and is obtained on successively adding the frequencies of the values of the variable according
to a certain law. The new frequencies obtained are called as cumulative frequencies (c.f.).
There are two types: Less Than c.f. and More Than c.f.
The frequency distributions relating to only single variable are called as “Univariate
frequency distributions”. The distribution obtained by taking into consideration two
variables is called as “Bivariate frequency distribution”.

Bivariate Frequency Distribution


“Bi” means “Two” and “variate” means “variable”, thus, Bivariate is a commonly used term
that describes a data which consists of observations showing two attributes or two variables.
Thus, when the data set are classified (or grouped in a frequency distribution) on the basis of
two variables, the distribution so formed is known as Bivariate Frequency Distribution.
Age (X)
Below20 20-30 30-40 Above 40 F (Y)
Income (Y)
Less than 10,000 8 5 2 2 17
10,000-20,000 3 7 12 20 42
20,000-30,000 5 8 15 25 53
More than 30,000 8 10 20 20 58
F (X) 24 30 49 67 N = 170
Basic Principles for forming a Grouped Frequency Distribution
1) Type of Classes: Classes should be properly defined. They must be mutually exclusive
(No data value should fall into 2 different classes) and must be all inclusive or exhaustive
(All data values must be included). The classes must be continuous with no gaps in a
frequency distribution.
2) Number of Classes: there is no specific rule for the number of classes in a frequency
distribution. But it is ideally said that there should be between 5 and 20 classes.
There is a general rule of thumb for calculating the number of classes (Sturges formula):
k = 1 + 3.322 log10N
3) Size of Class Intervals: it is inversely proportional to the number of classes. The class
width should be an odd number (So that midpoints are integers not decimals) and the
classes must be equal in width. Again using the Sturges rule we can calculate the size of
the class interval as:
i = Range / Number of Classes
4) Types of Class Intervals:
 Inclusive Type Classes: The type of classes in which both the upper and the lower
limits are included in the class are called “Inclusive Classes”. For example: 10-19; 20-
29; 30-39; etc. are inclusive.
 Exclusive Type Classes: The type of classes in which upper limits are excluded from
the respective classes and are included in the immediate next class are termed as
“Exclusive Classes”. For example: 10-15; 15-20; 20-25; etc. are exclusive.
 Open Ended Classes: The classification is termed as open-ended if the lower limit of
the first class or the upper limit of the last class are not specified. And the classes in
which one of the limits is missing are called as ‘open end classes’. For example: Less
than 50; More than 20; etc.

Converting a grouped frequency distribution into a continuous distribution:


A grouped frequency distribution needs to be converted into a continuous distribution to fill
the gaps between the upper limit of any class and lower limit of the succeeding class by
applying a correction factor. (called converting inclusive classes to exclusive classes). The
upper limits and the lower limits of this new “exclusive type” classes are called “Class
Boundaries”. Thus, correction factor is the means by which a grouped frequency
distribution is converted into a continuous frequency distribution. The Mid-Value/Class
Mark is the value of the variable which is exactly at the middle of the class. It is obtained by
dividing the sum of the upper and lower class limits by two.
Mid-Value = [Lower Class Limit + Upper Class Limit] / 2

TABULATION OF DATA
A statistical table is an orderly and logical arrangement of data into rows and columns. It
attempts to present the voluminous and heterogeneous (diversified) data in a condensed and
homogeneous (identical) form. Tabulation refers to the systematic representation of the
information contained in the data, in rows and columns according to certain characteristics. It
is a device of presenting data in a condensed and understandable form to furnish maximum
information contained in the data in minimum possible space.

Parts of a Table
 Table number (if more than one is prepared)
 Title (brief for the contents of the table)
 Head notes or Prefatory notes (Explanatory notes)
 Captions and Stubs (headings which explain the contents along the columns are captions
and for the rows are known as stubs)
 Body of the table (numerical information to be presented)
 Foot-note (used for further elaboration of some feature)
 Source note( Medium from which data is obtained)
DIAGRAMMATIC AND GRAPHIC PRESENTATION OF DATA
One of the important, convincing and easily understood method of presenting the statistical
data is the use of diagrams and graphs. They are geometrical figures like points, lines, bars,
squares, rectangles, circles, cubes, etc. pictures, maps or charts.
Benefits of Diagrams and Graphs
 Visual aids that present data in simple, readily comprehensible form.
 More attractive & fascinating than a set of numerical data.
 They have universal applicability & extensively used to present statistical figures & facts
 They save a lot of time & help draw meaningful inferences compared to statistical data.
 They facilitate easy comparisons and relationship study between data.
 Graphs reveal trends better than tables.

Types of Diagrams
(a) Line Diagrams (b) Bar Diagrams
Line Diagram: This is the simplest of all the diagrams. It consists of drawing vertical lines,
each vertical line being equal to the frequency. The variable (‘X’) values are presented on
suitable scale along the X-axis and the corresponding frequencies along Y-axis. Line
diagrams help comparisons.
Bar Diagrams: Bar diagrams are one of the easiest and the most commonly used devices of
presenting most of the business and economic data. They include a group of equally spaced
rectangles, one for each group/category of the data in which the values are represented by the
length or height of the rectangles. The different types of Bar diagrams are:
 Simple Bar Diagram: Simple bar diagram is the simplest of the bar diagrams and is used
frequently for the comparative study of two or more values of a single variable.

 Sub-divided / Component Bar Diagram: Simple bar diagram studies only one
characteristics or classification at a time. To over come this limitation, sub-divided bar
diagram is used. These diagrams are useful for presenting several items of a variable and
also help for comparative study of different parts/components .
 Percentage Bar Diagrams: A sub-divided / component bar diagram presented
graphically on percentage basis give percentage bar diagrams. It is used to highlight the
relative importance of the various components parts to the whole.
 Multiple Bar Diagrams: When two or more sets of inter-related variables are to be
presented graphically, multiple bar diagrams are used.
 Deviation Bar Diagram: Deviation bars are specially useful for graphic presentation of
net quantities like, net profit or loss, net exports or imports, etc. which have both positive
and negative values. The positive deviations (like profits) are presented by bars above the
base line while negative deviations (like loss) are represented by bars below the base line.

Types of Graphs:
 Histogram: is a popular device for data representation and is most suitable when the data
is represented using continuous frequency distribution. Here a series of adjacent vertical
rectangles are constructed, on the sections of the horizontal axis (X-axis), (with the bases
equal to the width of class intervals) and heights representing the frequencies of the
corresponding classes.
 Frequency Polygon: This is another device of graphic presentation of a frequency
distribution (Continuous, grouped or discrete). In case of discrete frequency distribution,
frequency polygon is obtained on plotting the frequencies on the vertical axis (Y-axis)
against the corresponding values of the variable on the horizontal axis (X-axis) and
joining the points so obtained by straight lines.
 Frequency Curve: A frequency curve is a smooth free hand curve drawn through the
vertices of a frequency polygon. It is same as that of the histogram or frequency polygon.
The only difference between the two is that histogram and frequency polygon have sharp
edges and under frequency curve
 “Ogives” or Cumulative Frequency Curves: Ogives are graphic presentations of the
cumulative frequency (c.f.) distribution of continuous variable. It consists of plotting the
c.f. (along the Y-axis) against the class boundaries (along X-axis). Since there are two
types of cumulative frequency distributions, there are two types of Ogives:
 ‘Less Than’ Ogive: consists of plotting the ‘Less Than’ cumulative frequencies
against the upper class boundaries of the respective classes. The points obtained
should be joined by a smooth free hand curve which is called ‘less than ogive curve’.
 ‘More Than’ Ogive: consists of plotting the ‘More Than’ cumulative frequencies
against the lower class boundaries of the respective classes. The points so obtained
are joined to give the ‘more than ogive curve’.

Limitations of Diagrams and Graphs


 They are supplements for classification and tabulation but they cannot substitute them.
 They give a clear general idea of the data, furnish only limited and approximate
information.
 They are subjective in character and may be interpreted differently by different people.
 All the diagrams and graphs are not easy to construct.
 In case of large figures (observations) they fail to reveal small differences in them.
 The choice of a particular diagram or graph to present the given data requires a great deal
of expertise.
 Diagrams and graphs can be used only for comparative analysis of different sets of data.
They are not useful if absolute information is to be represented.

You might also like