Statistics
Statistics may be defined as the collection, presentation, analysis, and
interpretation of numerical data.
Statistics is concerned with scientific methods of collecting, organizing,
summarizing, presenting, and analyzing data, as well as drawing valid
conclusions and making reasonable decisions based on the analysis.
Statistics is concerned with the systematic collection of numerical data and
its interpretation. Statistics helps convert raw numerical data into useful
information.
1. Statistics in Plural Sense
In the plural sense, statistics means numerical facts or data.
These data are collected for a specific purpose and expressed in
numbers.
Examples include marks of students, population, rainfall, income, and
sales.
Example:
The statistics of student attendance are available.
The statistics of India's population are collected every 10 years.
2. Statistics in Singular Sense
In the singular sense, Statistics is the science of collecting,
organizing, presenting, analyzing, and interpreting data.
It helps in making decisions and drawing conclusions from data.
It is a branch of mathematics.
Example:
Statistics is an important subject for research.
Statistics is widely used in business, economics, medicine, and
artificial intelligence.
Process of Statistics:
Example: Performance of 100 Students in an Examination
Instead of looking at all 100 students' marks individually, statistics helps a
teacher to collect the marks.
1. Arrange them in ascending order.
2. Classify them into groups (Classification).
3. Present those using tables or graphs (Tabulation).
4. Calculate average marks (Analysis).
5. Compare the performance of different sections.
6. Conclude whether the overall performance is satisfactory
(Interpretation).
7. Thus, statistics converts raw data into useful information.
Example: Hospital Data
A hospital wants to study the health condition of patients admitted during a
month. Using statistics, the hospital can:
1. Collect the number of patients admitted.
2. Classify patients according to age or disease.
3. Present the data using tables and graphs.
4. Calculate the average number of patients admitted per day.
5. Analyze which disease is the most common.
6. Conclude whether additional doctors or facilities are required.
Example: Weather Department
The weather department records daily temperature and rainfall. Using
statistics, it can:
1. Collect daily weather data.
2. Organize and classify the information.
3. Present the data through charts and graphs.
4. Calculate average temperature and rainfall.
5. Analyze weather patterns and seasonal changes.
6. Forecast future weather conditions.
Steps in Statistical Investigation
1. Identify the Problem
2. Define the Objectives
3. Plan Statistical Study (Population, Sample & Variables)
4. Select Data Source (Primary/Secondary)
5. Choose Data Collection Method
6. Collect Data
7. Organise the Data
8. Classification
9. Tabulation
10. Presentation
11. Statistical Analysis (Mean, Median, Mode, SD, etc.)
12. Interpretation of Results
13. Conclusion
Scope
The scope of statistics refers to the areas where statistical methods are
used to collect, organize, analyze, and interpret data for decision-making.
Today, statistics is used in almost every field.
1. Business and Commerce
Helps in market research.
Forecasts sales and demand.
Assists in pricing and inventory management.
Supports business decision-making.
2. Economics
Measures national income and inflation.
Studies unemployment and poverty.
Helps in economic planning and policy making.
3. Government
Conducts population census.
Plans budgets and welfare schemes.
Analyzes crime, education, and health data.
4. Education
Evaluates student performance.
Conducts examinations and surveys.
Improves teaching methods through data analysis.
5. Medicine and Healthcare
Tests new medicines and vaccines.
Studies diseases and epidemics.
Maintains hospital and patient records.
Statistics is an essential tool for collecting, organizing, presenting,
analyzing, and interpreting data. It plays a vital role in business, science,
engineering, medicine, government, education, agriculture, artificial
intelligence, and many other fields, making it indispensable for informed
decision-making.
Importance of Statistics
Statistics plays a vital role in today's world because almost every decision
is based on data. It is a branch of mathematics that deals with the
collection, organization, presentation, analysis and interpretation of data.
Statistics helps individuals, businesses, researchers and governments
understand facts, solve problems, and make better decisions.
Importance of statistics is defined as the significant role that statistics plays
in collecting, analyzing and interpreting data to support informed decision-
making and problem-solving in various fields.
(A) Helps in Decision- Making:-Statistics provides reliable information
that helps individuals, organizations and governments make informed
decisions based on facts rather than assumptions. Eg: A company studies
customers' demand before launching a new product.
(B) Simplifies Complex Data :- Large amounts of data can be difficult to
understand. Statistics organizes and summarizes the data into tables,
graphs and charts, making it easier to understand. Eg: A school presents
students' examination results in the form of graphs instead of a long list of
marks.
(C) Helps in Planning: - Statistics is useful in planning future activities by
analyzing past and present data. Eg: Government uses population statistics
to plan schools, hospitals and transport facilities.
(D) Supports Scientific Research: - Statistics is an essential tool in
scientific and academic research. Researchers use statistical methods to
collect data, analyses results and test hypotheses. Eg: Medical researchers
use statistics to evaluate whether a new medicine is effective.
(E) Helps in Forecasting :- Statistics helps predict future events by
studying historical data and trends. Eg: Weather departments use statistical
data to forecast rainfall and temperature.
Limitations of Statistics
1. Statistics Cannot Study Qualitative Data Directly
Statistics mainly deals with numerical/quantitative data. It cannot
directly study qualitative characteristics that cannot be expressed in
numbers.
Examples: Honesty, beauty, intelligence, personality, happiness, etc.
These cannot be measured directly using statistics. However, they
can sometimes be converted into numbers through surveys.
Eg: The intelligence of students cannot be assessed directly, but it
can be estimated using their examination marks or IQ scores.
2. Statistics Studies Groups, Not Individuals
Statistics is concerned with large groups for collection of data, called
aggregates. It does not focus on a single individual or a single
observation.
Eg: Statistics can calculate the average marks of 100 students, but it
cannot describe the performance of each student in detail.
So, statistics provides information about the whole group, not
individuals.
3. Statistical Results Are Not Exact
Unlike mathematics, statistical conclusions are not always 100%
accurate. They are based on estimates, averages and probabilities.
Therefore, statistical conclusions are generally true on average and
may not apply to every individual.
Eg: If the average salary of employees in a firm is ₹40,000 per
month, it does not mean every employee earns exactly ₹40,000.
Some may earn more while others may earn less.
4. Statistics Can Be Misused
Statistics can give misleading results if the data is incorrect or if it is
analysed by inexperienced people. Wrong methods, incomplete data
or unrepresentative samples may lead to false conclusions.
Eg: If a company may advertise that “90% of customers are satisfied”
without showing how the sample was selected, this can mislead
people.
So, statistics must be used carefully by knowledgeable and trained
people.
5. Statistics Is Only One Method of Studying a Problem
Statistics alone cannot provide a complete solution to every problem.
Many problems also depend on social, economic, political, ethical
and psychological factors. Therefore, statistical analysis should be
combined with practical knowledge and experience.
Eg: If the crime rate increases in a city, statistics can show the
increase, but it may not fully explain the reasons behind it. Other
factors, such as unemployment, education, poverty and social
conditions, must also be considered while decision-making.
Collection of data
Data collection is the first and one of the most important steps in any
statistical investigation. The quality and accuracy of statistical analysis
depend on the quality of the data collected.
Data collection is the process of gathering, measuring, and recording
information from various sources for the purpose of analysis and decision-
making.
Data collection is the systematic process of gathering information about a
particular subject or problem.
Need for Data Collection :-
1. Provides accurate information for analysis.
2. Helps in solving research problems.
3. Supports planning and decision-making.
4. Forms the basis for statistical studies.
5. Helps identify trends and patterns.
Objectives of Data Collection:-
1. To obtain reliable and accurate information.
2. To study and understand a particular problem.
3. To analyze and interpret data scientifically.
4. To support research and decision-making.
5. To test hypotheses and theories.
6. To predict future trends.
Types of Data:-
1. Primary Data
Primary data is the data collected first-hand by the investigator for a
specific purpose. It is original data and has not been collected before.
Since the investigator collects the information directly, primary data is
generally more reliable and relevant to the study.
Examples
Conducting a survey among students.
Measuring the height of participants.
Recording temperatures during an experiment.
Collecting customer feedback through questionnaires.
Methods of Collecting Primary Data
1. Direct Personal Investigation
The investigator personally contacts respondents and records the
required information.
Advantages - Highly accurate and reliable data, Better understanding
of the problem, Suitable for small-scale investigations.
Disadvantages - Time-consuming, Expensive, Not suitable for large
populations.
Example: A researcher visits farmers to collect information about crop
production.
2. Indirect Oral Investigation
Information is collected from knowledgeable persons instead of
directly from respondents.
Advantages - Useful when respondents are unavailable, Saves time,
Disadvantages - Less reliable, Depends on the knowledge and
honesty of informants.
Example: Collecting information about illegal activities from police
officers.
3. Through Local Correspondents
Information is collected through local agents or correspondents who
regularly send data to a central office.
Advantages - Covers large geographical areas, Economical, Suitable
for continuous data collection.
Disadvantages - Less control over data quality, May contain bias or
errors
Example: Newspapers collect weather information from different
cities through local correspondents.
4. Mailed Questionnaire Method
A list of questions is mailed or sent electronically to respondents, who
fill it out and return it.
Advantages - Low cost, Suitable for large populations, Respondents
get enough time to answer.
Disadvantages - Low response rate, Questions may be
misunderstood, No opportunity to clarify doubts.
Example: Customer satisfaction surveys sent by email.
5. Schedule Method
An enumerator personally visits respondents, asks questions, and fills
in the schedule on their behalf.
Advantages - High response rate, Suitable for illiterate respondents,
More accurate information.
Disadvantages – Expensive, Time-consuming, Requires trained
enumerators.
Example: Population Census conducted by government officials.
6. Observation Method
The investigator collects information by directly observing events or
activities.
Advantages – Real-time data, no dependence on respondents.
Disadvantages – Time-consuming, observer bias may occur.
Example - Observing traffic at a road intersection.
7. Interview Method
The investigator asks questions directly to respondents.
Types – Face-to-face interview, telephone interview, online interview.
Advantages – Detailed information, clarification of doubts.
Disadvantages – Costly, time-consuming.
B. Secondary Data
Secondary data is the data that has already been collected,
processed and published by someone else for another purpose. The
investigator uses this existing data instead of collecting it personally.
Eg: Government census reports, company annual reports, research
journals, books and newspapers.
Sources of Secondary Data :-
(i) Internal Sources – Data available within an organisation.
Eg: Sales records, employee records, financial statements, inventory
records.
(ii) External Sources – Data obtained from outside organisations.
Eg: Government publications, research institutions, international
organisations, books and journals, websites and databases.
Advantages of Secondary Data - Saves time, Less expensive, Easily
available, Covers large populations, Useful for comparative studies.
Limitations of Secondary Data - May be outdated, May not suit the
research objective, May contain bias, Accuracy cannot always be
verified.
Primary Data Secondary Data
1. Collected first-hand by 1. Already collected by
the investigator others
2. Original data 2. Previously published
data
3. More reliable for a 3. Reliability depends on
specific research the source.
4. Expensive and time- 4. Less expensive and
consuming readily available
5. Collected for the current 5. Collected for another
study purpose
Classification of Data
In statistics, the information collected from surveys, experiments,
observations or records is called data. Data which is collected for the first
time is known as raw data or ungrouped data.
Raw data is generally scattered, unorganised and difficult to understand. If
such data is analysed directly, it becomes time-consuming and may not
provide meaningful conclusions.
So, before any statistical analysis is carried out, the collected data must be
organised in a systematic manner. This process of organising data is
called classification. Once the data has been classified, it is presented in
the form of tables, a process known as tabulation.
Classification is the process of arranging raw data into different groups or
classes according to their common characteristics or attributes. The
purpose of classification is to simplify large amounts of data so that it
becomes easy to understand, compare and analyze.
Classification is the process of arranging or grouping data into similar
classes or categories based on common characteristics to facilitate
analysis and interpretation.
For example: Suppose a teacher collects marks of 100 students:
52, 68, 74, 89, 96, 58, 71, 65, 83, 92
Instead of studying each mark individually, the marks can be grouped as
follows:
Marks No. of Students
0–20 2
21–40 8
41–60 28
61–80 42
81–100 20
Now information becomes much easier to understand.
Need for Classification :-
1. It presents data in a simple and easy-to-understand form.
2. It facilitates statistical analysis.
3. It makes data suitable for comparison.
4. It helps identify important relationships, trends and patterns.
5. It helps simplify and summarize large amounts of data.
Types of Classification:-
There are mainly four types of classification:
(A) Qualitative Classification
Qualitative classification is based on attributes or qualities that cannot be
measured numerically.
Eg: Gender, literacy, religion, blood group, etc.
Gender Number of Students
Boys 80
Girls 120
Gender Number of Students
Total 200
Characteristics:
1. Data is classified according to attributes.
2. Attributes may be present or absent.
3. It is generally expressed in terms of categories.
(B) Quantitative Classification
Quantitative classification is based on measurable characteristics.
Data is grouped according to numerical values.
Examples: Age, income, height, weight, marks, salary, etc.
Marks Frequency
0–20 5
21–40 8
41–60 15
61–80 7
Characteristics:
1. Quantitative characteristics can be measured numerically.
2. Variables may be discrete or continuous.
(C) Geographical Classification
Geographical classification means arranging data according to place
or geographical location.
Eg: Population by state, production by country, sales by region, etc.
State Wheat Production (lakh tonnes)
Punjab 80
Haryana 120
Rajasthan 95
Uttar Pradesh 250
Applications:
1. Government reports
2. Agricultural studies
3. Market analysis
4. Regional comparisons
(D) Chronological Classification
Chronological classification means arranging data according to time.
The arrangement may be by years, months, days, etc.
Eg: Sales of a company over different years.
Year Sales (in lakhs)
2022 120
2023 145
2024 170
2025 195
The data is arranged according to time. The arrangement may be
presented from past to present or vice versa.
Tabulation of Data
After data is collected and classified, the next important step in statistical
analysis is tabulation.
Tabulation means presenting classified data in the form of rows and
columns. It helps to present information systematically and makes it easier
to understand, compare and analyse.
The main purpose of tabulation is to summarise large amounts of data into
a simple and meaningful form.
A table helps to:
Present information clearly.
Make data easy to understand.
Facilitate comparison.
Save space and time.
Help in statistical analysis and interpretation.
Parts / Components of a Table
(A) Table Number - The table number is assigned to identify the table.
Eg: Table 3.1
(B) Title - The title indicates the subject matter of the table.
Eg: Distribution of Students by Branch and Gender (Academic Session
2025–26)
(C) Headnote - The headnote is placed below the title and explains
additional information about the data, such as units.
Eg: (Number of students)
(D) Captions - Captions are the headings given to the columns of a table.
(E) Stubs - Stubs are the headings given to the rows of a table.
(F) Body - The body contains the actual numerical data or information
presented in the table.
(G) Totals - Totals are provided at the bottom of rows and/or columns.
They indicate the sum of observations.
(H) Footnote - A footnote appears below the table. It explains
observations, symbols, exceptions or special remarks.
(I) Source Note - The source note indicates where the data has been
obtained.
Eg: Source: Ministry of Education, Government of India (2026).
Model Structure of a Table
Table Number: Title of Table:
Stub Caption Caption
Body
Total
Footnote:
Source note:
Example of a Statistical Table
Table 2.1 : Distribution of Students by Branch and Gender
(Academic Session 2025–26)
Branch Boys Girls Total
CSE 60 40 100
IT 45 35 80
ECE 55 30 85
Mechanical 50 15 65
Total 210 120 330
Footnote: Figures represent enrolled students only.
Source: Admission Office, ABC Engineering College.
Types of Tables
1. Simple Table - One-way table. Presents information about one
characteristic only.
Example:
Branch No. of Students
CSE 60
IT 50
ECE 40
2. Double Table - Two-way table. Presents information according to two
characteristics.
Example:
Branch Boys Girls
IT 45 50
ECE 55 30
3. Triple Table - Three-way table. Presents data according to three
characteristics.
Example:
Branch Boys Girls Total
IT 30 60 90
ECE 40 60 100
4. Manifold Table - A complicated table. Presents data according to more
than three characteristics. It is generally used when several characteristics
need to be analysed together. 4 or more captions in this type.
Depiction of Data
Depiction of Data (also called Presentation of Data) is the process of
presenting collected and classified data in a systematic, meaningful, and
attractive form so that it becomes easy to understand, compare, analyze,
and interpret.
The main objective of depicting data is to convert raw numerical information
into a form that can be easily understood by everyone.
1. Textual Presentation -Textual presentation means presenting data in
the form of words, sentences, or paragraphs. This method is suitable when
the amount of data is small. This information is written in sentence form.
Example - A survey of 100 students showed that:
60 students passed.
40 students failed.
Advantages - Very simple, Easy to prepare, Suitable for small data, No
special statistical knowledge required.
Limitations - Not suitable for large datasets, Difficult to compare values,
Time-consuming to read, Trends are difficult to identify.
2. Tabular Presentation - Tabular presentation means arranging data into
rows and columns. A table summarizes large amounts of data in a compact
form.
Example:-
Subject Students
AI 45
Python 40
Statistics 35
Subject Students
OOP 50
Advantages - Easy comparison. Saves space, Organized presentation.
Suitable for statistical analysis.
Limitations - Less attractive, Difficult for non-technical people, Requires
careful preparation.
3. Diagrammatic Presentation - Diagrammatic presentation represents
data through diagrams instead of words or tables. It is attractive and easily
understood by the general public.
Types of Diagrams
A. Bar Diagram - A bar diagram represents data using rectangular bars of
equal width. The length or height of each bar is proportional to the value
it represents.
B. Pie Diagram (Pie Chart) - A pie diagram is a circular diagram divided
into sectors. Each sector represents the proportion of a category in
relation to the whole.
C. Pictogram - A pictogram represents data using pictures or symbols.
Each symbol stands for a
D. fixed number of units.
E. Cartogram - A cartogram presents statistical data on maps. It is used to
show geographical variations in data such as population, rainfall, or
literacy.
Types of Graphs
A. Histogram - A histogram is a graph used to represent the frequency
distribution of continuous data. It consists of adjacent rectangles whose
heights represent frequencies.
B. Frequency Polygon - A frequency polygon is formed by joining the
midpoints of the tops of the bars of a histogram with straight lines. It
shows the shape of a frequency distribution.
C. Frequency Curve - A frequency curve is a smooth curve drawn through
the points of a frequency polygon. It represents the overall pattern of a
frequency distribution.
D. Scatter Graph (Scatter Diagram) - A scatter graph represents pairs of
values as points on a coordinate plane. It is used to study the
relationship between two variables.
Comparison of Different Methods
Method Best Used For Advantages Limitations
Difficult to
Simple and
Textual Small datasets compare large
descriptive
data
Organized Easy comparison
Tabular Less attractive
numerical data and analysis
General audience Attractive and easy
Diagrammatic Less precise
and presentations to understand
Shows trends, Requires
Statistical analysis
Graphical patterns, and knowledge to
and research
relationships interpret
Advantages of Depiction of Data
Simplifies large volumes of data.
Makes comparison easy.
Saves time.
Highlights trends and relationships.
Improves communication.
Helps in forecasting.
Supports research and decision-making.
Makes reports more attractive.
Limitations of Depiction of Data
May oversimplify complex information.
Incorrect scales can mislead readers.
Some methods require statistical knowledge.
Graphs and diagrams may hide detailed values.
Preparation may take time and skill.