Module 1
Introduction to statistics
Introduction
Statistics is a mathematical body of science that pertains to the collection, analysis, interpretation
or explanation, and presentation of data, or as a branch of mathematics. Some consider statistics to
be a distinct mathematical science rather than a branch of mathematics.
Statistics is the study of the collection, analysis, interpretation, presentation, and organization of data.
In other words, it is a mathematical discipline to collect, summarize data. Also, we can say that
statistics is a branch of applied mathematics. However, there are two important and basic ideas
involved in statistics; they are uncertainty and variation. The uncertainty and variation in different
fields can be determined only through statistical analysis. These uncertainties are basically determined
by the probability that plays an important role in statistics.
Statistics Examples
Some of the real-life examples of statistics are:
To find the mean of the marks obtained by each student in the class whose strength is 50. The average
value here is the statistics of the marks obtained.
Suppose you need to find how many members are employed in a city. Since the city is populated with 15
lakh people, hence we will take a survey here for 1000 people (sample). Based on that, we will create the
data, which is the statistic.
Basics of Statistics
The basics of statistics include the measure of central tendency and the measure of dispersion. The
central tendencies are mean, median and mode and dispersions comprise variance and standard
deviation.
Mean is the average of the observations. Median is the central value when observations are arranged in
order. The mode determines the most frequent observations in a data set.
Variation is the measure of spread out of the collection of data. Standard deviation is the measure of
the dispersion of data from the mean. The square of standard deviation is equal to the variance.
Types of Statistics
Basically, there are two types of statistics.
1. Descriptive Statistics
2. Inferential Statistics
In the case of descriptive statistics, the data or collection of data is described in summary. But in the
case of inferential stats, it is used to explain the descriptive one. Both these types have been used on
large scale.
1. Descriptive Statistics
The data is summarised and explained in descriptive statistics. The summarization is done from a
population sample utilising several factors such as mean and standard deviation. Descriptive statistics
is a way of organising, representing, and explaining a set of data using charts, graphs, and summary
measures. Histograms, pie charts, bars, and scatter plots are common ways to summarise data and
present it in tables or graphs. Descriptive statistics are just that: descriptive. They don‟t need to be
normalised beyond the data they collect.
2. Inferential Statistics
We attempt to interpret the meaning of descriptive statistics using inferential statistics. We utilise
inferential statistics to convey the meaning of the collected data after it has been collected, evaluated,
and summarised. The probability principle is used in inferential statistics to determine if patterns found
in a study sample may be extrapolated to the wider population from which the sample was drawn.
Inferential statistics are used to test hypotheses and study correlations between variables, and they can
also be used to predict population sizes. Inferential statistics are used to derive conclusions and
inferences from samples, i.e. to create accurate generalisations.
Methods in Statistics
The methods involve collecting, summarizing, analyzing, and interpreting variable numerical data.
Here some of the methods are provided below.
Data collection
Data summarization
Statistical analysis
What is Data in Statistics?
Data is a collection of facts, such as numbers, words, measurements, observations etc.
Types of Data
Qualitative data- it is descriptive data.
Example- She can run fast, He is thin.
Quantitative data- it is numerical information.
Example- An Octopus is an Eight legged creature.
Types of quantitative data
Discrete data- has a particular fixed value. It can be counted
Continuous data- is not fixed but has a range of data. It can be measured.
Representation of Data
There are different ways to represent data such as through graphs, charts or tables. The general
representation of statistical data are:
Bar Graph
Pie Chart
Line Graph
PictographHistogram
Frequency Distribution
Bar Graph: A Bar Graph represents grouped
data with rectangular bars with lengths
proportional to the values that they represent.
The bars can be plotted vertically or
horizontally.
Pie Chart: A type of graph in which a circle is
divided into Sectors. Each of these sectors
represents aproportion of the whole.
Linegraph: The line chart is represented by a
series of data points connected with a
straight line. The series of data points are
called „markers.‟
Pictograph:
A pictorial symbol for a word or phrase, i.e.
showing data with the help of pictures. Such as
Apple, Banana & Cherry can have different
numbers, and it is just a representation of data.
Histogram:
A diagram is consisting of rectangles. Whose
area is proportional to the frequency of a
variable and whose width is equal to the class
interval.
Frequency Distribution: The frequency of a
data value is often represented by “f.” A
frequency table is constructed by arranging
collected data values in ascending order of
magnitude with their corresponding
frequencies.
Measures of Central Tendency
In Mathematics, statistics are used to describe the central tendencies of the grouped and ungrouped
data. The three measures of central tendency are:
Mean
Median
Mode
All three measures of central tendency are used to find the central value of the set of data.
Importance of Statistics
(vii) Data Analysis: Statistics allows us to organize, interpret, and summarize large amounts of data,
making it easier to identify patterns and trends.
(viii) Informed Decision Making: It aids in making evidence-based decisions by providing insights and
evidence to support or refute hypotheses or claims.
(ix) Risk Assessment: It helps in understanding and quantifying risks, enabling better risk management
strategies in various fields, including finance, healthcare, and engineering.
(x) Generalization: Statistics enables the extrapolation of insights from a sample to a largerpopulation,
providing a way to make broader inferences from limited data.
Limitations of Statistics
(ix) Sampling Errors: If the sample collected is not representative of the population, it can
lead tobiased or inaccurate conclusions.
(x) Causation vs. Correlation: Statistics can show relationships between variables but may
notalways establish a cause-and-effect relationship, leading to misinterpretation.
(xi) Assumptions and Simplifications: Many statistical methods rely on assumptions that
mightnot always hold true in real-world scenarios, affecting the accuracy of results.
(xii) Misuse or Misinterpretation: Improper use or misinterpretation of statistics can lead to
incorrect conclusions or misinform decision-making processes.
(xiii) Data Quality: Statistics heavily relies on the quality of input data. Inaccurate,
incomplete, orbiased data can lead to flawed analysis and results.
Statistical data– Classification
What is Classification of Data?
For performing statistical analysis, various kinds of data are gathered by the investigator or analyst.
The information gathered is usually in raw form which is difficult to analyze. To make the analysis
meaningful and easy, the raw data is converted or classified into different categories based on their
characteristics. This grouping of data into different categories or classes with similar or homogeneous
characteristics is known as the Classification of Data. Each division or class of the gathered data is
known as a Class. The different basis of classification of statistical information are Geographical,
Chronological, Qualitative (Simple and Manifold), and Quantitative or Numerical.
For example, if an investigator wants to determine the poverty level of a state, he/she can do so by
gathering the information of people of that state and then classifying them on the basis of their income,
education, etc.
According to Conner, “Classification is the process of arranging things (either actually or notionally)
in groups or classes according to their resemblances and affinities, and gives expression to the unity
of attributes that may exist amongst a diversity of individuals.”
The main objectives of Classification of Data are as follows:
Explain similarities and differences of data
Simplify and condense data’s mass
Facilitate comparisons
Study the relationship
Prepare data for tabular presentation
Present a mental picture of the data
Basis of Classification of Data
The classification of statistical data is done after considering the scope, nature, and purpose of an
investigation and is generally done on four bases; viz., geographical location, chronology, qualitative
characteristics, and quantitative characteristics.
1. Geographical Classification
The classification of data on the basis of geographical location or region is known
as Geographical or Spatial Classification. For example, presenting the population of different states
of a country is done on the basis of geographical location or region.
2. Chronological Classification
The classification of data with respect to different time periods is known
as Chronological or Temporal Classification. For example, the number of students in a school in
different years can be presented on the basis of a time period.
3. Qualitative Classification
The classification of data on the basis of descriptive or qualitative characteristics like region, caste,
sex, gender, education, etc., is known as Qualitative Classification. A qualitative classification can
not be quantified and can be of two types; viz., Simple Classification and Manifold Classification.
Simple Classification
When based on only one attribute, the given data is classified into two classes, which is known
as Simple Classification. For example, when the population is divided into literate and illiterate, it is
a simple classification.
Manifold Classification
When based on more than one attribute, the given data is classified into different classes, and then sub-
divided into more sub-classes, which is known as Manifold Classification. For example, when the
population is divided into literate and illiterate, then sub-divided into male and female, and further
sub-divided into married and unmarried, it is a manifold classification.
4. Quantitative Classification
The classification of data on the basis of the characteristics, such as age, height, weight, income, etc.,
that can be measured in quantity is known as Quantitative Classification. For example, the weight of
students in a class can be classified as quantitative classification.
Tabulation
What is Tabulation?
Tabulation of data in statistics as well as mathematics is a method of storing classified data in a tabular
form. It may be complex, double, or simple, depending upon the type of categorization.
The purpose of a tabulation chart/data is to display a large volume of complex information in a systematic
fashion that would enable the viewers to draw reasonable outcomes and interpretations from them.
Parts of Table in Tabulation
In order to tabulate data accurately and precisely, one must understand some of the essential parts of a
table which are as follows:
1. Table Number: This is the first section of a table and is presented on top of any table to facilitate
straightforward identification and for further reference.
2. Title of the Table: One of the most related parts of any given table is its title. The title of the table
describes its contents. It is important that the title should be short and crisp and exactly worded to
define the table’s contents efficiently.
3. Column Headings or Captions: Captions are the piece of information on the table which is at the
top of each column that tells the figures under each column.
4. Row Headings: The title of every horizontal row comes under the row heading.
5. Body of a Table: This is the part that includes the numeric information collected from examined
facts. The data in the body is displayed in rows which are read horizontally starting from left to
right and the data in the columns are read vertically from top to bottom.
Types of Tabulation
Tabulation can be classified into the following types:
1. Simple Tabulation or One-way Tabulation
When the data are tabulated to one aspect, it is declared to be a simple tabulation or one-way tabulation.
For example, the tabulation of data on the population of the earth divided by one feature like language is
an example of a simple tabulation.
2. Double Tabulation or Two-way Tabulation
When the given data are tabulated according to two characters at a time, it is stated to be a double
tabulation or a two-way tabulation.
For example, suppose that a table has to show the highest population in various states of India. This can
be achieved by a one-way table. However, if the population has to be analyzed in terms of the total
number of males and females in every state, it will ask for a two-way table.
3. Three-way Tabulation
Similar to the above-mentioned category, three-way charts show information handled from three mutually
dependent and interrelated subjects.
Let us consider the same above example and elaborate on that further with the added category in the table.
Now we need the position of literacy amongst the male and female populations in each state. The
tabulation for such categories has to be placed down in a three-way table.
4. Complex Tabulation
When the data are tabulated according to various characteristics, it is stated to be a complex tabulation.
For example tabulation of data on the population of the planet is divided into three or more characteristics
like religion, language, literacy, gender etc. is an example of a complex tabulation.
Example : Study the below tabular data and answer the questions based on it.
Expenditures of a company (in lakh) per annum over the given years is:
Year Salary(lakhs) Fuel and Bonus(lakhs) Interest on Taxes(lakhs)
Transport(lakhs) Loans(lakhs)
2013 288 98 3.00 23.4 83
2014 342 112 2.52 32.5 108
2015 324 101 3.84 41.6 74
2016 336 133 3.68 36.4 88
2017 420 142 3.96 49.4 98
Calculate the total expenditure of the company over these items during the year 2015 from the table chart
given.
Solution: Total expenditure of the company during 2015 is;
⇒ Rs. (324 + 101 + 3.84 + 41.6 + 74) lakhs
⇒ Rs. 544.44 lakhs
∴ Total expenditure is 544.44 lakhs
Diagrammatic & Graphic representation of data
Diagrammatic and graphic presentation of data means visual representation of the data. It shows a
comparison between two or more sets of data and helps in the presentation of highly complex data in its
simplest form. Diagrams and graphs are clear and easy to read and understand. In the diagrammatic
presentation of data, bar charts, rectangles, sub-divided rectangles, pie charts, or circle diagrams are used.
In the graphic presentation of data, graphs like histograms, frequency polygon, frequency curves,
cumulative frequency polygon, and graphs of time series are used.
Bar Diagram
This is one of the simplest techniques to do the comparison for a given set of data. A bar graph is a graphical
representation of the data in the form of rectangular bars or columns of equal width. It is the simplest one and
easily understandable among the graphs by a group of people.
Browse more Topics under Statistical Description Of Data
Introduction to Statistics
Textual and Tabular Representation of Data
Frequency Distribution
Histogram
Frequency Polygon
Cumulative Frequency Graph or Ogive
Construction of a Bar Diagram
1. Draw two perpendicular lines intersecting each other at a point O. The vertical line is the y-axis and the
horizontal is the x-axis.
2. Choose a suitable scale to determine the height of each bar.
3. On the horizontal line, draw the bars at equal distance with corresponding heights.
4. The space between the bars should be equal.
Properties of a Bar Diagram
Each bar or column in a bar graph is of equal width.
All bars have a common base.
The height of the bar corresponds to the value of the data.
The distance between each bar is the same.
Types of Bar Diagram
A bar graph can be either vertical or horizontal depending upon the choice of the axis as the base. The
horizontal bar diagram is used for qualitative data. The vertical bar diagram is used for the quantitative data
or time series data. Let us take an example of a bar graph showing the comparison of marks of a student in all
subjects out of 100 marks for two tests.
With the bar graph, we can also compare the marks of students in each subject other than the marks of one
student in every subject. Also, we can draw the bar graph for every student in all subjects.
Histogram
We can use another way of diagrammatical representation of data. If we are working with a continuous data
set or grouped dataset, we can use a histogram for the representation of data.
A histogram is similar to a bar graph except for the fact that there is no gap between the rectangular
bars. The rectangular bars show the area proportional to the frequency of a variable and the width of
the bars represents the class width or class interval.
Frequency means the number of times a variable is occurring or is present. It is an area graph. The
heights of the rectangles are proportional to the corresponding frequencies of similar classes.
Construction of Histogram
1. Draw two perpendicular lines intersecting each other at a point O. The vertical line is the y-axis and the
horizontal is the x-axis.
2. Choose a suitable scale for both the axes to determine the height and width of each bar
3. On the horizontal line, draw the bars with corresponding heights
4. There should be no gap between two consecutive bars showing the continuity of the data
5. If the grouped frequencies are not continuous, the first thing to do is to make them continuous
It is done by adding the average of the difference between the lower limit of the class interval and the upper
limit of the preceding class width to the upper limits of all the classes. The same quantity is subtracted from
the lower limits of the classes.
Properties of Histogram
Each bar or column in a bar graph is of equal width and corresponds to the equal class interval
If the classes are of unequal width then the height of the bars will be proportional to the ration of the
frequencies to the width of the classes
All bars have a common base
The height of the bar corresponds to the frequency of the data
Suppose we have a data set showing the marks obtained out of 100 by a group of 35 students in statistics. We
can find the number of students in the various marks category with the help of the histogram.
Line Graph
A line graph is a type of chart or graph which shows information when a series of data is joined by a line. It
shows the changes in the data over a period of time. In a simple line graph, we plot each pair of values of (x,
y). Here, the x-axis denotes the various time point (t), and the y-axis denotes the observation based on the
time.
Properties of a Line Graph
It consists of Vertical and Horizontal scales. These scales may or may not be uniform.
Data point corresponds to the change over a period of time.
The line joining these data points shows the trend of change.
Below is the line graph showing the number of buses passing through a particular street over a period of time:
Solved Examples for diagrammatic Representation of Data
Problem 1: Draw the histogram for the given data.
Marks No. of Students
15 – 18 7
19 – 22 12
23 – 26 56
27 – 30 40
31 – 34 11
35 – 38 54
39 – 42 26
43 – 46 37
47 – 50 7
Total 250
Solution: This grouped frequency distribution is not continuous. We need to convert it into a continuous
distribution with exclusive type classes. This is done by averaging the difference of the lower limit of one
class and the upper limit of the preceding class. Here, d = ½ (19 – 18) = ½ = 0.5. We add 0.5 to all the upper
limits and we subtract 0.5 from all the lower limits.
Marks No. of Students
14.5 – 18.5 7
18.5 – 22.5 12
22.5 – 26.5 56
26.5 – 30.5 40
30.5 – 34.5 11
34.5 – 38.5 54
38.5 – 42.5 26
42.5 – 46.5 37
46.5 – 50.6 7
Total 250
The corresponding histogram is
Problem 2:
Draw a line graph for the production of two types of crops for the given years.
Production in metric tones
Year Crop I Crop II
1968 10 12
1978 12 10
1988 15 21
1998 30 20
2008 18 17
2018 25 25
Solution: The required graph is