Sources of Collection of Data
1. Primary Source
It is a collection of data from the source of origin. It provides the researcher with first-
hand quantitative and raw information related to the statistical study. In short, the primary
sources of data give the researcher direct access to the subject of [Link]
example,statistical data, works of art, and interview transcripts.
2. Secondary Source
It is a collection of data from some institutions or agencies that have already collected
the data through primary sources. It does not provide the researcher with first-hand
quantitative and raw information related to the study. Hence, the secondary source of data
collection interprets, describes, or synthesizes the primary [Link] example,reviews,
government websites containing surveys or data, academic books, published journals,
articles, etc.
Even though primary sources provide more credibility to the collected data because of the
presence of evidence, but good research will require both primary and secondary sources
of data collection.
Primary and Secondary Data
1. Primary Data
The data collected by the investigator from primary sources for the first time from
scratch is known as primary data. This data is collected directly from the source of
origin. It is real-time data and is always specific to the researcher's needs. The primary
data is available in raw form. The investigator has to spend a long time period in the
collection of primary data and hence is expensive also. However, the accuracy and
reliability of primary data are more than the secondary data. Some examples of sources for
the collection of primary data are observations, surveys, experiments, personal interviews,
questionnaires, etc.
2. Secondary Data
The data already in existence which has been previously collected by someone else for
other purposes is known as secondary [Link] does not include any real-time data as the
research has already been done on that information. However, the cost of collecting
secondary data is less. As the data has already been collected in the past, it can be found in
refined form. The accuracy and reliability of secondary data are relatively less than the
primary data. The chances of finding the exact information or data specific to the
researcher's needs are less. However, the time required to collect secondary data is short
and hence is a quick and easy process. Some examples of sources for the collection of
secondary data are books, journals, internal records, government records, articles, websites,
government publications, etc.
Data Collection Methods - Primary and Secondary Data
Data Collection is the systematic process of gathering, measuring and analyzing
information from various sources to gain an accurate understanding of a specific topic or
problem. It is the first and most fundamental step in research, statistics and data-driven
decision-making, as it provides the relevant information needed to answer research
questions or solve statistical problems. Accurate data collection ensures reliable results,
meaningful insights and informed decisions, while poor or incomplete data can lead to
misleading analysis and incorrect conclusions.
The main objectives of data collection are:
To support decision-making.
To identify trends and patterns.
To measure performance and progress.
To provide evidence for conclusions.
Methods of Collecting Data
Primary Data
Primary data is information collected directly from original sources for a specific research
purpose. It is fresh, relevant and tailored to the study. Advantages of primary data is
High accuracy
More control over data quality
Specific to research objectives
Methods of Collecting Primary Data
There are a number of methods of collecting primary data, Some of the common methods
are as follows:
1. Interviews:Interviews involve direct communication between the investigator and
respondents.
Direct Personal Investigation:The investigator personally collects information from
the source.
Indirect Oral Investigation:Information is collected from third parties who possess
relevant knowledge.
Advantage:Provides real-time, natural data; no reliance on self-reported
information.
Disadvantage:Observer bias; limited to what can be seen; may influence subjects'
behavior.
Suitable Use Case:Behavioral studies, user experience research.
2. Questionnaires:A questionnaire is a structured set of questions prepared to collect
information. The investigator can collect data through the questionnaire in two ways:
Mailing Method:Questionnaires are sent by mail or online.
Enumerator’s Method:The enumerator personally visits respondents and fills the
questionnaire.
Advantage:Can reach a large audience quickly and cost-effectively.
Disadvantage:Responses may be biased or inaccurate; low response rates.
Suitable Use Case:Customer satisfaction surveys, market research.
3. Observations:The observation method involves collecting data by watching and
recording behaviors or events as they naturally occur.
Advantage:Provides real-time, authentic data without reliance on self-reported
information.
Disadvantage:Risk of observer bias and behavior changes.
Suitable Use Case:User behavior studies, classroom analysis, field research.
4. Experiments:The experiment method involves manipulating variables in a controlled
environment to study cause-and-effect relationships.
Advantage:Allows for the establishment of cause-and-effect relationships with high
precision.
Disadvantage:Can be expensive and less realistic.
Suitable Use Case:Drug testing, teaching method evaluation, marketing impact
analysis.
5. Focus Group:A focus group gathers 6–12 participants to discuss a topic under a
moderator’s guidance.
Advantage:Provides diverse and detailed insights.
Disadvantage:Results may not represent the larger population.
Suitable Use Case:Product feedback, brand perception studies, public opinion
research.
6. Local Correspondents:In Local Correspondent method, for the collection of data, the
investigator appoints correspondents or local persons at various places, which are then
furnished by them to the investigator. With the help of correspondents and local persons,
the investigators can cover a wide area.
Secondary Data
Secondary data is collected from information that has already been gathered, processed
and published by others. It is broadly classified into Published Sources and Unpublished
Sources.
Methods of Collecting Secondary Data
Secondary data can be collected through different published and unpublished sources.
Some of them are as follows:
1. Published Sources
Published sources are officially available reports and documents that provide reliable and
structured data for research and analysis.
Government Publications:Central and State Governments publish statistical
reports such as census data, economic surveys and industrial [Link] of
Government publications on Statistics are the Annual Survey of Industries, Statistical
Abstract of India, etc.
Semi-Government Publications:Different Semi-Government bodies also publish
data related to health, education, deaths and births. These kinds of data are also reliable
and used by different informants. Some examples of semi-government bodies are
Metropolitan Councils, Municipalities, etc.
Publications of Trade Associations:Various big trade associations collect and
publish data from their research and statistical divisions of different trading activities
and their [Link] example, data published by Sugar Mills Association regarding
different sugar mills in India.
Journals and Papers:Different newspapers and magazines provide a variety of
statistical data in their writings, which are used by different investigators for their
studies.
International Publications:Different international organizations like IMF, UNO,
ILO, World Bank, etc., publish a variety of statistical information which are used as
secondary data.
Publications of Research Institutions:Research institutions and universities also
publish their research activities and their findings, which are used by different
investigators as secondary data. For example National Council of Applied Economics,
the Indian Statistical Institute, etc.
2. Unpublished Sources
Unpublished sources are another source of collecting secondary data. The data in
unpublished sources is collected by different government organizations and other
organizations. These organizations usually collect data for their self-use and are not
published anywhere. For example, research work done by professors, professionals,
teachers and records maintained by business and private enterprises.
Meaning of Questionnaire
A questionnaire is a research instrument used by any researcher as a tool to collect data or
gather information from any source or subject of his or her interest from the respondents.
It has a specific goal to understand topics from the respondent's point of view. It consists
of a set of written or printed questions with a choice of answers devised for survey or
statistical studies. It is the most popular type of primary data collection, which can be used
to gather both quantitative data ( in the form of numerals) and qualitative data (in the form
of words and figures) or mixed data, which is a continuation of both quantitative and
qualitative data.
Types of Questionnaires
In a broader sense, there are two types of questionnaires:
Structured questionnaire:It is also known as a closed questionnaire where such
questions are asked, which can be answered as yes or no. It includes less number of
researchers and a large number of respondents, and it has definite and concrete
questions. These types of questionnaires are formal and are prepared well in advance.
Unstructured questionnaire:It is based on a more open questionnaire. An open
questionnaire means recording more data, as the respondents can point out what is more
important for them in their own words and methods, as responses can go to any length.
This type of questionnaire is quite flexible and can be applied to several areas of study
as they do not require much planning and time.
Qualities of a Good Questionnaire
1. Limited Number of questions:The number of questions in the questionnaire should be
as limited as possible, and questions should be asked only related to the purpose of the
inquiry.
2. Proper sequence of questions:Questions must be placed in the proper sequence, like
simple and direct questions must be placed at the start of the questionnaire, and hard and
indirect questions must be placed at the last.
3. Simplicity:The language of the questions should be simple and easy to understand, and
the questions should be short. Complex questions must be avoided.
4. Instructions:A good questionnaire must have clear and proper instructions for filling
out the forms.
5. No undesirable questions:Undesirable questions like personal questions, which can
offend the respondents, must be avoided.
6. Non-controversial questions:The questions should be asked in such a way that they
can be answered impartially.
7. Calculations:Questions involving calculations must be avoided, as they can be complex
and time-consuming.
8. Objective-type questions:More focus should be given to objective-type questions,
whereas subjective-type questions should be avoided.
Types of Questions in Questionnaire
Broadly, There are two types of questions, i.e. Closed-ended and Open-ended:
1. Closed-ended Questions: In these kinds of questions, respondents can answer questions
by selecting from a limited number of predefined options already given by the researcher.
In such kinds of questions, the researcher cannot provide an unanticipated answer but
rather chooses from the list of answers already provided. The various types of closed-
ended questions are as follows:
A. Alternate response type:This type of question offers only two options, which can be
either yes or no, fair or unfair, or true or false.
Example:
Question: Have you ever written an article for GFG?
Answer:
1. Yes
2. No
B. Multiple Choice Type of Questions:In such kinds of questions, respondents are
allowed to select one or more options from a list of predefined answers.
Example:
Question: What is your educational qualification?
Answers:
1. Primary
2. Matriculation
3. Graduate
4. Post-Graduate
5. Doctorate
In this question, there are multiple choices for answers. This kind of question
is called a single multiple-choice question.
C. Rating Scale or Ordinal type of Questions:In a rating scale question, the researcher
gives a scale of numbers for the answer to choose from, and the respondent can choose a
number from the given scale that most accurately represents his response.
Example:
Question: pH of water is?
Answer:
1.1-5 (Highly acidic)
2.5-6 (Acidic)
3.6-7 (Moderately Acidic)
4.7-7.1 (Neutral)
Here the options have a scale for answers.
D. Ranking type of questions:These questions ask respondents to order their answers in
order of preference.
Example:
Question: Most peaceful country?
Answers:
Here respondents can rate the countries on a scale of 1 to 5 according to their
preferences.
2. Open-ended Questions:These kinds of questions are explanatory in nature and can go
to any length, as they provide the researcher with rich qualitative data and give an
opportunity to the researcher to gain insight into those fields of study, which are not
covered by the close-ended questions. Such questions have no statistical purpose, as they
let the respondent answer questions of varying lengths, which makes such questions
concluding in nature.
Data Cleaning
Data cleaning is the process of preparing raw data by detecting and correcting errors so it
can be effectively used for analysis. It is a foundational step in data preprocessing that
ensures datasets are suitable for analytical, statistical and machine learning tasks.
Raw data is often noisy, incomplete and inconsistent which can negatively impact
the accuracy of the model.
Clean datasets are also important in EDA (Exploratory Data Analysis), which
enhances the interpretability of data so that the right actions can be taken based on
insights.
Data Cleaning Process
1. Assess Data Quality
The first step in data cleaning is to assess the quality of your data. This involves checking
for:
Missing Values:Identify any blank or null values in the dataset. Missing values can
be due to various reasons such as incomplete data collection, data entry errors or data
loss during transmission.
Incorrect Values:Check for values that are outside the expected range or are
inconsistent with the data type.
Inconsistencies in Data Format:Verify that the data format is consistent
throughout the dataset.
After assessing data quality, several issues can be identified in the dataset:
Rows 1 and 6 are duplicates indicating potential data duplication that may distort
analysis.
Row 7 has a missing value in the "Name" column, which could impact calculations
or summaries.
The "Date" column uses the "YYYY-MM-DD" format, but it is important to
maintain this consistency across all entries.
The score of 100 in row 7 may be an outlier depending on the scoring system, which
could skew statistical analysis.
2. Remove Irrelevant Data
Removing irrelevant or duplicate data ensures the dataset is clean, accurate and
meaningful, preventing skewed analysis and improving overall quality.
Identify duplicate entries using techniques like sorting, grouping or hashing.
Remove duplicate records to ensure each data point is unique and correctly
represented.
Detect redundant observations that do not add new information to the dataset.
Eliminate variables or columns that are irrelevant to the analysis and do not provide
useful insights.
In the deduplicated DataFrame, the duplicate rows 1 and 6 have been removed to ensure
each record is unique
3. Fix Structural Errors
Structural errors occur when data formats, naming conventions or variable types are
inconsistent which can affect analysis accuracy. Correcting these issues ensures uniform
and reliable data representation.
Standardize data formats to maintain consistency in dates, times and other data types
across the dataset.
Correct naming inconsistencies in column names, variable names or labels to ensure
clarity and uniformity.
Ensure consistent data representation such as using the same units for measurements
or the same scales for ratings.
4. Handle Missing Data
Missing data can introduce bias and reduce the reliability of analysis. Properly addressing
missing values helps maintain the integrity of your dataset.
Impute missing values using statistical methods such as mean, median or mode to
fill gaps.
Remove records with missing values when the missing data is extensive or cannot
be accurately imputed.
Apply advanced imputation techniques like regression k-nearest neighbors or
ecision trees to estimate missing values.
The missing value in the 'Name' column (row 7) has been filled with 'Unknown' to indicate
unavailable data, ensuring the dataset remains complete and consistent.
5. Normalize Data
Data normalization organizes the dataset to reduce redundancy and ensure consistency
making it easier to manage and analyze.
Split data into multiple tables, with each table storing specific types of information.
Ensure consistency across the dataset to support efficient querying and accurate
analysis.
6. Identify and Manage Outliers
Outliers are data points that deviate significantly from the rest of the dataset and can affect
analysis accuracy. Properly handling them ensures more reliable insights.
Remove outliers that result from errors or are not representative of the population.
Transform extreme but valid outliers to reduce their impact on the analysis.
Sampling
Sampling is a crucial aspect of research that involves selecting a subset of individuals or
items from a larger population to infer conclusions about the entire population. Two
primary categories of sampling techniques are probability sampling and non-probability
sampling. Understanding the differences, advantages, and applications of each method is
essential for selecting the appropriate sampling strategy for a given research study.
Need of Sampling
To draw conclusions about populations from samples, we must use inferential statistics
which enables us to determine a population’s characteristics by directly observing only a
portion (or sample) of the population. We obtain a sample rather than a complete
enumeration (a census) of the population for many reasons.
Obviously, it is cheaper to observe a part rather than the whole, but we should prepare
ourselves to cope with the dangers of using samples. In this tutorial, we will investigate
various kinds of sampling procedures. Some are better than others but all may yield samples
that are inaccurate and unreliable. We will learn how to minimize these dangers, but some
potential error is the price we must pay for the convenience and savings the samples provide.
Probability Sampling
Probability sampling is a sampling technique where every member of the population has a
known, non-zero chance of being selected. This method ensures that the sample is
representative of the population, which allows for generalization of the results.
Types of Probability Sampling
Simple Random Sampling: Each member of the population has an equal chance of
being selected. This can be achieved using random number generators or drawing lots.
Systematic Sampling: Every nth member of the population is selected, starting
from a random point. For example, selecting every 10th person on a list.
Stratified Sampling: The population is divided into strata (subgroups) based on a
characteristic (e.g., age, gender), and random samples are taken from each stratum.
Cluster Sampling: The population is divided into clusters (e.g., geographic areas),
and entire clusters are randomly selected, then all members of chosen clusters are
surveyed.
Examples and Real-Life Applications of Probability Sampling
Simple Random Sampling: A lottery system to select participants for a survey on
voting behaviour.
Systematic Sampling: Quality control in a factory by testing every 50th product off
the production line.
Stratified Sampling: Conducting a health survey ensuring representation across
different age groups and genders.
Cluster Sampling: Educational research by selecting and surveying all students
from randomly chosen schools.
Non-Probability Sampling
Non-probability sampling does not involve random selection, and not all members of the
population have a known or equal chance of being included in the sample. This can lead to
bias and limits the ability to generalize findings.
Types of Non-Probability Sampling
Convenience Sampling:Samples are taken from a group that is conveniently
accessible to the researcher, such as surveying people in a shopping mall.
Judgmental or Purposive Sampling: The researcher uses their judgment to select
members who are most appropriate for the study.
Snowball Sampling: Existing study subjects recruit future subjects from among
their acquaintances, useful in studying hidden populations.
Quota Sampling: The population is segmented into mutually exclusive subgroups,
just like in stratified sampling. However, the selection within strata is non-random.
Examples and Real-Life Applications of Non-Probability Sampling
Convenience Sampling: Surveying people at a local café about their coffee
preferences.
Judgmental Sampling: Selecting expert doctors for a study on a rare disease.
Snowball Sampling: Researching the social networks of drug users by asking
participants to refer others.
Quota Sampling: Interviewing a fixed number of people from various demographic
groups in a city to understand public opinion.
Difference Between Probability Sampling and Non-Probability Sampling
Feature Probability Sampling Non-Probability Sampling
Definition Every member of the population Not every member has a chance
has a known, non-zero chance of of being selected; based on
being selected. subjective judgment.
Basis of Selection Random selection Non-random selection
Types Simple random sampling, Convenience sampling,
systematic sampling, stratified judgmental sampling, quota
sampling, cluster sampling sampling, snowball sampling
Bias Less bias due to random Higher risk of bias due to
selection subjective judgment
Representativeness More representative of the Less representative of the
population population
Generalizability Results can be generalized to the
Results are less generalizable
entire population
Complexity More complex and time-
Simpler and quicker
consuming
Cost Generally more expensive Generally less expensive
Statistical Analysis Allows for more robust
Limited statistical analysis due
statistical analysis and estimation
to lack of randomness
of sampling error
Application Used in large-scale surveys, Used in exploratory research,
scientific research, and studies pilot studies, and when time or
needing high accuracy resources are limited