INTRODUCTION
TO STATISTICS
What is Statistics?
■ Look at these statements
– For the past 30 days, the benchmark index of BSE, SENSEX has been
growing by an average of 250 points daily (Personal observation)
– According to a survey by Statista, an average 62% of generation Z
millennials, either own a business or want to start a venture of their own
([Link])
– According to World Bank, 70% of Indian population have bank accounts, but
50% of these people do not access these accounts (World Bank’s LDB on
Financial Inclusion)
– India's foreign exchange reserves stood at US$ 575.29 billion as of
November 20, 2020, according to data from RBI. ([Link])
■ All these statements are using statistics to either make a point or enhance
credibility of the statement
What is Statistics?
■ From this view, the term statistics refers to numerical facts such as averages,
medians, percentages, and index numbers that help us understand a variety of
business and economic situations (Anderson, Sweeney and Williams, 2014)
■ A broader definition can be
■ “The art and science of collecting, analyzing, presenting, and interpreting
data”. (Anderson, Sweeney and Williams, 2014)
Definition:
Croxton and Cowden –
“Statistics may be defined as the collection, presentation, analysis and
interpretation of numerical data.“
Seligman –
“Statistics is the science which deals with the methods of collecting,
classifying, presenting, comparing and interpreting numerical data
collected to throw some light on any sphere of enquiry."
Are the following statements statistical data?
i) Weekly wages of 100 workers of a factory.
ii) Height of Ram is six feet.
iii) Mohan's weight is 70 Kgs, Sohan's height is 6.2 feet, and Ram's monthly
income is Rs. 1,500.
iv) Sales of a company during the past 10 years.
Significance of statistics in Business
■ Business managers continually strive to get a better knowledge about the
business and economic environment
■ The understanding of statistics in business is intended to
– Give an idea of presentation and description of information
– Understand the process of deriving inferences about a population based on
study of a sample
– Introduce the utilization of data to make more informed and better
decisions for business improvement and policy formulation
– Guide how to use the data for forecasting purposes
■ Such information is provided by collecting, organizing, analyzing, presenting,
and interpreting data
Applications of Statistics in business
■ Sampling of accounts for audit
■ Statistical information related to stock returns for forecasting
■ Information collected by market research firms for test marketing
■ Data collected for quality control purposes
■ Statistical information about economy such as index numbers
Data and its components
■ Data - Facts and figures collected for statistical analysis and interpretation.
■ All the data collected in a particular study are referred to as the data set for the study.
For e.g. a data set containing prices of shares over a period of time
■ Elements are the entities on which data are collected. For e.g. the shares for which
price data is collected
■ A variable is a characteristic of interest for the elements. For e.g. for the data set
mentioned above the prices of shares are variable.
Data, Data Sets,
Elements, Variables, and Observations
Observation Variables
Element
Names Stock Annual Earn/
Exchange Sales(Rs.M)
Company Share(Rs.)
Dataram BSE 73.10 0.86
EnergySouth NSE 74.00 1.67
Keystone Nasdaq 365.70 0.86
LandCare MCX 111.40 0.33
Psychemedics N 17.60 0.13
Data Set
Basic Statistical concepts
■ Population and sample
– Population is a collection of all the elements under statistical study
– Sample is a portion of population which is selected in a way so as to be a
representative of the population
■ Descriptive and inferential statistics
– Summated measures arrived from the data are called descriptive statistics
while inferring the characteristics of population using sample statistics is
inferential statistics
Descriptive vs. Inferential Statistics
■ Descriptive Statistics — using data gathered on a group to describe or reach
conclusions about that same group only
■ Inferential Statistics — using sample data to reach conclusions about the
population from which the sample was taken
Classification of Data based on sources
■ Primary Data
– Data that has been observed, experienced or recorded close to the
event are called primary data
– Data collected for the first time to address a specific research
problem
■ Secondary Data
– Sometimes the data required to address a particular research
problem is already available
– Such data is easy to search nowadays, given the technological
advancement
– Such data is available in form of various publications
Methods of primary data collection
■ Survey using Questionnaires
■ Conducting interviews
■ Observing without getting involved
■ Participative observation/immersing oneself in a situation
■ Doing experiments
Sources of Secondary Data
■ Online data sets
■ Cultural texts
■ Libraries, museums and other archives
■ Commercial and professional bodies
■ Government publications
FREQUENCY
DISTRIBUTIONS
Meaning
• Process of arranging data in groups/classes on the basis of certain properties.
• Purpose-
1. Condense the raw data into a form suitable for statistical analysis.
2. Removes complexities and highlights the feature of the data.
3. Facilities comparisons and drawing inference from the data.
4. Provides information about the mutual relationships among elements of a data set.
5. Helps in statistical analysis by separating the elements of the data set into
homogeneous groups and hence brings out the point of similarity and dissimilarity.
Meaning of Frequency Distribution
• Frequency- It is the number of times the observation occurred/recorded in a study.
• Example- Sam played football on:
■ - Saturday Morning,
■ - Saturday Afternoon
■ - Thursday Afternoon
The frequency was 2 on Saturday, 1 on Thursday and 3 for the whole week.
• Frequency Distribution- It is a list, table or graph that displays the frequency of various outcomes in
a sample.
• Thus, frequency distribution refers to a table that shows an item and its frequency.
Some Definitions
■ Morris Humburg
■ “A frequency distribution or frequency table is simply a table in which the data grouped
into classes and the number of cases which fall in each class are recorded. The numbers
in each class are referred to as frequencies.”
■ Croxton and Cowden
■ “Frequency distribution is a statistical table in which different values of variable are
shown in the sequence of magnitude along with corresponding frequencies.”
■ Murray. R. Spiegal
■ “A tabular arrangement of data by class together with the corresponding class
frequencies is called a frequency distribution or frequency table”
Identify the type of distribution
■ No. of accidents on certain days of a month – ■ Discrete
■ No. of shares sold on a day in market ■ Discrete
■ Lengths of 1,000 bolts produced at a factory ■ Continuous
■ No. of books on a library shelf ■ Discrete
■ Speed of an automobile in Kph ■ Continuous
■ No. of individuals in a family ■ Discrete
Advantages of frequency distribution
• Present raw data in an organized and easy-to-read format.
• Data are expressed in a more compact form.
• One can quickly note the pattern of distribution of observations falling in various classes.
• Permits the use of more complex statistical techniques which helps reveal certain hidden
characteristics of the data.
Disadvantages of frequency distribution
• As we group the data, individual observations lose their identity.
• Statistically calculations are based only on the values of the class mark
and not on the values of the observation in that class.
Organizing the data
• To examine large set of numerical data is first to organize and present it in an appropriate tabular
and graphical format.
Raw Data Pertaining to total time hours worked by Machinists
94, 89, 88, 89, 90, 94, 92, 88, 87, 85, 88, 93, 94, 93, 94,
93, 92, 88, 94, 90, 93, 84, 93, 84, 91, 93, 85, 91, 89, 95
• This raw data do not highlight any characteristics/trend, such as highest, lowest, average weekly
hours.
• Unless these data are reorganized, no meaningful inferences can be drawn.
• The raw data can be reorganized in a data array and frequency distribution.
• When a raw data set is arranged in rank order, from smallest to largest observation or vice versa,
the ordered sequence obtained is called an ordered array.
Raw Data Pertaining to total time hours worked by Machinists
■ 1. Provides quick look at the highest and lowest observation in the data within which individual values
vary.
■ 2. Helps in dividing the data into various sections or parts.
■ 3. Helps us to know the degree of concentration around a particular observation.
■ 4. Helps to identify whether any values appear more than once in the array.
Array and Tallies
Number of Total time Hours Tally Number of Workers
(Frequency)
84
85
86
87
88
89
90
91
92
93
94
95
Types of frequency distribution
Exclusive FD
Types
Inclusive FD
Cumulative
FD
Exclusive Method
• When the data are classified in such a way that the upper limit of a class interval is the lower limit of
the succeeding class interval, then it is said to be the exclusive method of classifying data.
Table: Exclusive method of data classification
Dividends Declared in % (Class Interval) Number of companies (Frequencies)
0-10 5
10-20 7
20-30 15
30-40 10
• Such classification ensures continuity of data as the upper limit of one class is the lower limit of
succeeding class.
Dividends Declared in % (Class Interval) Number of companies (Frequencies)
0 but less than 10 5
10 but less than 20 7
20 but less than 30 15
30 but less than 40 10
Inclusive method
• When the data are classified in such a way that both lower and upper limits of a class interval are
included in the interval itself, then it is said to be the inclusive method of classifying data.
Table: Inclusive Method of Data Classification
Number of Accidents (Class Interval) Number of Weeks (Frequencies)
40-44 5
45-59 22
60-74 13
75-89 8
90-104 2
• Certain adjustment in the class interval is needed to obtain the continuity.
Cumulative Frequency Distribution
• Cumulative Frequency series is that series in which the frequencies are
continuously added corresponding to each class interval in the series.
• Distribution which shows the cumulative number of observations below the upper
limit of each frequency distribution.
• A cumulative frequency distribution is of two types:
1. Less than type (Upper limit)
2. More than type (Lower limit)
• Less than cumulative frequency distribution-
- The frequencies of each class interval are added successively from top to bottom.
- Represents the cumulative number of observations less than or equal to the class frequency to which
it relates.
• More than Cumulative Frequency Distribution-
- The frequencies of each class interval are added successively from bottom to top.
- Represents the cumulative number of observations greater than or equal to class frequency to which
the relates.
Practice Questions of Frequency
Distributions
1. Form a frequency distribution from the following data by Inclusive Method, taking 4
as the magnitude of class-intervals:
10, 17, 15, 22, 11, 16, 19, 24, 29, 18, 25, 26, 32, 14, 17, 20, 23, 27, 30, 12, 15, 18, 24, 36, 18,
15, 21, 28, 33, 38, 34, 13, 10, 16, 20, 22, 29, 19, 23, 31
2. Following figures relate to the weekly wages of workers in a factory Wages (in ’00
Rs.) Prepare a frequency table by taking a class interval of 5.
100 100 101 102 106 86 82 87 109 104 75 89 99 96 94 93 92 90 86 78 79 84 83 87 88 89 75
76 76 79 80 81 89 99 104 100 103 104 107 110 110 106 102 107 103 101 101 101 86 94 93
96 97 99 100 102 103 107 107 108 109 94 93 97 98 99 100 97 88 86 84 83 82 80 84 86 88
91 93 95 95 95 97 98 100 105 106 103 85 84 77 78 80 93 96 97 98 98 98 87
3. In a survey, it was found that 64 families bought milk in the following quantities (litres)
in a particular week. Convert the above data into a frequency distribution by ‘Inclusive
Method’ by taking a class interval of 5.
19 16 22 9 22 12 39 19 14 23 6 24 16 18 7 17 20 25 28 18 10 24 20 21 10 7 18 28 24 20 14
23 25 34 22 5 33 23 26 29 13 36 11 26 11 37 30 13 8 15 22 21 32 21 31 17 16 23 12 9 15 27
17 21
4. For the following raw data prepare an exclusive frequency distribution and
classes with the same width 5. Marks in English
12 36 40 16 10 10 19 20 28 30 19 27 15 21 33 45 7 19 20 26 26 37 6 5 20 30 37 17
11 20
5. The following are the weights in kilograms of a group of 55 students.
Prepare a frequency table taking the magnitude of each class-interval as 10 kg.
and the first class-interval as equal to 40 and less than 50.
42 74 40 60 82 115 41 61 75 83 63 53 110 76 84 50 67 78 77 63 65 95 68 69 104 80
79 79 54 73 59 81 100 66 49 77 90 84 76 42 64 69 70 80 72 50 79 52 103 96 51 86
78 94 71
6. Using the following data of hours worked by 50 piece rate workers for a
period of a month in a certain factory. Compute frequency distribution by
taking 4 as the magnitude of class-intervals:
110, 175, 161, 157, 155, 108, 164, 128, 114, 178, 165, 133, 195, 151, 71, 94, 97, 42,
30, 62, 138, 156, 167, 124, 164, 146, 116, 149, 104, 141, 103, 150, 162, 149, 79,
113, 69, 121, 93, 143, 140, 144, 187, 184, 197, 87, 40, 122, 203, 148.
7. Compute Less Than and More than Cumulative Frequency distribution of
marks of 70 students in a test as given
8. Convert the following distribution into ‘more than’ frequency distribution for
discrete and continuous series.
Weekly wages (less than ’00 Rs.) : 20 40 60 80 100
Number of workers : 41 92 156 194 201
9. The credit office of a departmental store gave the following statements for
payment due to 40 customers. Construct a frequency table of the balances due
taking the class intervals as Rs. 50 and under Rs. 200, Rs. 200 and under Rs. 350,
etc. Also find the percentage cumulative frequencies and interpret these values.
Balances due in Rs.
337, 570, 99, 759, 487, 352, 115, 60, 521, 95 563, 399, 625, 215, 360, 178, 827,
301, 501, 199 110, 501, 201, 99, 637, 328, 539, 150, 417, 250 451, 595, 422, 344,
186, 681, 397, 790, 272, 514