Data Collection Methods | Primary and Secondary
Data
Data Collection refers to the systematic process of gathering, measuring, and
analyzing information from various sources to get a complete and accurate picture of
an area of interest. Data collection is a critical step in any research or data-driven
decision-making process, ensuring the accuracy and reliability of the results obtained.
By employing various methods of data collection researchers and organizations can
gather the necessary data to support their objectives effectively.
What is Data Collection?
Data Collection is the process of collecting information from relevant sources to find
a solution to the given statistical inquiry. Collection of Data is the first and foremost
step in a statistical investigation. It's an essential step because it helps us make informed
decisions, spot trends, and measure progress.
Different methods of collecting data include
Interviews
Questionnaires
Observations
Experiments
Published Sources and Unpublished Sources
Here, statistical inquiry means an investigation by any agency on a topic in which the
investigator collects the relevant quantitative information. In simple terms, a statistical
inquiry is a search for truth by using statistical methods of collection, compiling,
analysis, interpretation, etc. The basic problem for any statistical inquiry is the
collection of facts and figures related to this specific phenomenon that is being studied.
Therefore, the basic purpose of data collection is collecting evidence to reach a sound
and clear solution to a problem.
Terms Related to Data Collection
Data: Data is a tool that helps an investigator in understanding the problem
by providing him with the information required. Data can be classified into
two types; viz. Primary Data and Secondary Data.
Investigator: An investigator is a person who conducts the statistical
enquiry.
Enumerators: In order to collect information for statistical enquiry, an
investigator needs the help of some people. These people are known as
enumerators.
Respondents: A respondent is a person from whom the statistical
information required for the enquiry is collected.
Survey: It is a method of collecting information from individuals. The basic
purpose of a survey is to collect data to describe different characteristics
such as usefulness, quality, price, kindness, etc. It involves asking questions
about a product or service from a large number of people
Methods of Collecting Data
There are two different methods of collecting data: Primary Data Collection and
Secondary Data Collection.
Primary Data
Primary data refers to information collected directly from first-hand
sources specifically for a particular research purpose. This type of data is gathered
through various methods, including surveys, interviews, experiments,
observations, and focus groups. One of the main advantages of primary data is that
it provides current, relevant, and specific information tailored to the researcher's
needs, offering a high level of accuracy and control over data quality.
Methods of Collecting Primary Data
There are a number of methods of collecting primary data, some of the common
methods are as follows:
1. Interviews: Collect data through direct, one-on-one conversations with individuals.
The investigator asks questions either directly from the source or from its indirect
links.
1. Direct Personal Investigation: The method of direct personal investigation
involves collecting data personally from the source of origin. In simple
words, the investigator makes direct contact with the person from whom
he/she wants to obtain information.
For example, direct contact with the household women to obtain
information about their daily routine and schedule.
2. Indirect Oral Investigation: In the indirect oral investigation method of
collecting primary data, the investigator does not make direct contact with
the person from whom he/she needs information, instead they collect the
data orally from some other person who has the necessary required
information. For example, collecting data of employees from their
superiors or managers.
Advantage: Provides real-time, natural data; no reliance on self-reported
information.
Disadvantage: Observer bias; limited to what can be seen; may influence
subjects' behavior.
Suitable Use Case: Behavioral studies, user experience research.
2. Questionnaires: Collect data by asking people a set of questions, either online, on
paper, or face-to-face. In this method the investigator prepares a questionnaire to
collect, Information through Questionnaires and Schedules while keeping in mind the
motive of the study. The investigator can collect data through the questionnaire in two
ways:
1. Mailing Method: This method involves mailing the questionnaires to the
informants for the collection of data. The investigator attaches a letter with
the questionnaire in the mail to define the purpose of the study or research.
2. Enumerator’s Method: This method involves the preparation of a
questionnaire according to the purpose of the study or research. However, in
this case, the enumerator reaches out to the informants himself with the
prepared questionnaire.
Advantage: Can reach a large audience quickly and cost-effectively.
Disadvantage: Responses may be biased or inaccurate; low response rates.
Suitable Use Case: Customer satisfaction surveys, market research.
3. Observations: The observation method involves collecting data by watching and
recording behaviors, events, or conditions as they naturally occur. The observer
systematically watches and notes specific aspects of a subject's behavior or the
environment, either covertly or overtly.
Advantage: Provides real-time, authentic data without reliance on self-
reported information.
Disadvantage: Observer bias can influence the results, and the presence of
an observer might alter subjects' behavior.
Suitable Use Case: Studying user interactions with a product in a natural
setting, monitoring wildlife behavior, or assessing classroom dynamics.
4. Experiments: The experiment method involves manipulating one or more
variables to determine their effect on another variable, within a controlled
environment. Researchers create two groups (control and experimental), apply the
treatment or variable to the experimental group, and compare the outcomes between
the groups.
Advantage: Allows for the establishment of cause-and-effect relationships
with high precision.
Disadvantage: Experiments can be artificial, limiting the ability to
generalize findings to real-world settings, and they can be resource-
intensive.
Suitable Use Case: Testing the efficacy of a new drug, assessing the impact
of a new teaching method, or evaluating the effect of a marketing campaign.
5. Focus Group: The focus group method involves gathering a small group of people
to discuss a specific topic or product, facilitated by a moderator. A group of 6-12
participants engages in a guided discussion led by a moderator who asks open-ended
questions to elicit opinions, attitudes, and perceptions.
Advantage: Provides in-depth insights and diverse perspectives through
interactive discussions, revealing the reasoning behind participants' thoughts
and feelings.
Disadvantage: Results can be influenced by dominant participants or
groupthink, and the findings are not easily generalizable due to the small,
non-representative sample size.
Suitable Use Case: Exploring customer attitudes towards a new product,
gathering feedback on a marketing campaign, or understanding public
opinion on social issues.
6. Information from Local Sources or Correspondents: In this method, for the
collection of data, the investigator appoints correspondents or local persons at various
places, which are then furnished by them to the investigator. With the help of
correspondents and local persons, the investigators can cover a wide area.
Secondary Data
Secondary data refers to information that has already been collected, processed,
and published by others. This type of data can be sourced from existing research
papers, government reports, books, statistical databases, and company
records. The advantage of secondary data is that it is readily available and often free
or less expensive to obtain compared to primary data. It saves time and resources since
the data collection phase has already been completed.
Methods of Collecting Secondary Data
Secondary data can be collected through different published and unpublished sources.
Some of them are as follows:
1. Published Sources
Government Publications: Government publishes different documents
which consists of different varieties of information or data published by the
Ministries, Central and State Governments in India as their routine activity.
As the government publishes these Statistics, they are fairly reliable to the
investigator. Examples of Government publications on Statistics are the
Annual Survey of Industries, Statistical Abstract of India, etc.
Semi-Government Publications: Different Semi-Government bodies also
publish data related to health, education, deaths and births. These kinds of
data are also reliable and used by different informants. Some examples of
semi-government bodies are Metropolitan Councils, Municipalities, etc.
Publications of Trade Associations: Various big trade associations collect
and publish data from their research and statistical divisions of different
trading activities and their aspects. For example, data published by Sugar
Mills Association regarding different sugar mills in India.
Journals and Papers: Different newspapers and magazines provide a
variety of statistical data in their writings, which are used by different
investigators for their studies.
International Publications: Different international organizations like
IMF, UNO, ILO, World Bank, etc., publish a variety of statistical
information which are used as secondary data.
Publications of Research Institutions: Research institutions and
universities also publish their research activities and their findings, which
are used by different investigators as secondary data. For example, National
Council of Applied Economics, the Indian Statistical Institute, etc.
2. Unpublished Sources
Unpublished sources are another source of collecting secondary data. The data in
unpublished sources is collected by different government organizations and other
organizations. These organizations usually collect data for their self-use and are not
published anywhere. For example, research work done by professors, professionals,
teachers and records maintained by business and private enterprises.
Example:
The table below shows the production of rice in India.
The above table contains the production of rice in India in different years. It can be
seen that these values vary from one year to another. Therefore, they are known
as variable. A variable is a quantity or attribute, the value of which varies from one
investigation to another. In general, the variables are represented by letters such as X,
Y, or Z. In the above example, years are represented by variable X, and the production
of rice is represented by variable Y. The values of variable X and variable Y
are data from which an investigator and enumerator collect information regarding the
trends of rice production in India.
Data Preprocessing
Data preprocessing is the process of preparing raw data for analysis by cleaning and
transforming it into a usable format. In data mining it refers to preparing raw data for
mining by performing tasks like cleaning, transforming, and organizing it into a format
suitable for mining algorithms.
Goal is to improve the quality of the data.
Helps in handling missing values, removing duplicates, and normalizing
data.
Ensures the accuracy and consistency of the dataset.
Steps in Data Preprocessing
Some key steps in data preprocessing are Data Cleaning, Data Integration, Data
Transformation, and Data Reduction.
1. Data Cleaning: It is the process of identifying and correcting errors or
inconsistencies in the dataset. It involves handling missing values, removing duplicates,
and correcting incorrect or outlier data to ensure the dataset is accurate and reliable.
Clean data is essential for effective analysis, as it improves the quality of results and
enhances the performance of data models.
Missing Values: This occur when data is absent from a dataset. You can
either ignore the rows with missing data or fill the gaps manually, with the
attribute mean, or by using the most probable value. This ensures the dataset
remains accurate and complete for analysis.
Noisy Data: It refers to irrelevant or incorrect data that is difficult for
machines to interpret, often caused by errors in data collection or entry. It
can be handled in several ways:
o Binning Method: The data is sorted into equal segments, and
each segment is smoothed by replacing values with the mean or
boundary values.
o Regression: Data can be smoothed by fitting it to a regression
function, either linear or multiple, to predict values.
o Clustering: This method groups similar data points together,
with outliers either being undetected or falling outside the
clusters. These techniques help remove noise and improve data
quality.
Removing Duplicates: It involves identifying and eliminating repeated data
entries to ensure accuracy and consistency in the dataset. This process
prevents errors and ensures reliable analysis by keeping only unique
records.
2. Data Integration: It involves merging data from various sources into a single,
unified dataset. It can be challenging due to differences in data formats, structures, and
meanings. Techniques like record linkage and data fusion help in combining data
efficiently, ensuring consistency and accuracy.
Record Linkage is the process of identifying and matching records from
different datasets that refer to the same entity, even if they are represented
differently. It helps in combining data from various sources by finding
corresponding records based on common identifiers or attributes.
Data Fusion involves combining data from multiple sources to create a
more comprehensive and accurate dataset. It integrates information that may
be inconsistent or incomplete from different sources, ensuring a unified and
richer dataset for analysis.
3. Data Transformation: It involves converting data into a format suitable for analysis.
Common techniques include normalization, which scales data to a common range;
standardization, which adjusts data to have zero mean and unit variance; and
discretization, which converts continuous data into discrete categories. These
techniques help prepare the data for more accurate analysis.
Data Normalization: The process of scaling data to a common range to
ensure consistency across variables.
Discretization: Converting continuous data into discrete categories for
easier analysis.
Data Aggregation: Combining multiple data points into a summary form,
such as averages or totals, to simplify analysis.
Concept Hierarchy Generation: Organizing data into a hierarchy of
concepts to provide a higher-level view for better understanding and
analysis.
4. Data Reduction: It reduces the dataset's size while maintaining key information.
This can be done through feature selection, which chooses the most relevant features,
and feature extraction, which transforms the data into a lower-dimensional space
while preserving important details. It uses various reduction techniques such as,
Dimensionality Reduction (e.g., Principal Component Analysis): A
technique that reduces the number of variables in a dataset while retaining
its essential information.
Numerosity Reduction: Reducing the number of data points by methods
like sampling to simplify the dataset without losing critical patterns.
Data Compression: Reducing the size of data by encoding it in a more
compact form, making it easier to store and process.
Uses of Data Preprocessing
Data preprocessing is utilized across various fields to ensure that raw data is
transformed into a usable format for analysis and decision-making. Here are some key
areas where data preprocessing is applied:
1. Data Warehousing: In data warehousing, preprocessing is essential for cleaning,
integrating, and structuring data before it is stored in a centralized repository. This
ensures the data is consistent and reliable for future queries and reporting.
2. Data Mining: Data preprocessing in data mining involves cleaning and
transforming raw data to make it suitable for analysis. This step is crucial for
identifying patterns and extracting insights from large datasets.
3. Machine Learning: In machine learning, preprocessing prepares raw data for
model training. This includes handling missing values, normalizing features, encoding
categorical variables, and splitting datasets into training and testing sets to improve
model performance and accuracy.
4. Data Science: Data preprocessing is a fundamental step in data science projects,
ensuring that the data used for analysis or building predictive models is clean,
structured, and relevant. It enhances the overall quality of insights derived from the
data.
5. Web Mining: In web mining, preprocessing helps analyze web usage logs to
extract meaningful user behavior patterns. This can inform marketing strategies and
improve user experience through personalized recommendations.
6. Business Intelligence (BI): Preprocessing supports BI by organizing and cleaning
data to create dashboards and reports that provide actionable insights for decision-
makers.
7. Deep Learning Purpose: Similar to machine learning, deep learning applications
require preprocessing to normalize or enhance features of the input data, optimizing
model training processes.
Advantages of Data Preprocessing
Improved Data Quality: Ensures data is clean, consistent, and reliable for
analysis.
Better Model Performance: Reduces noise and irrelevant data, leading to
more accurate predictions and insights.
Efficient Data Analysis: Streamlines data for faster and easier processing.
Enhanced Decision-Making: Provides clear and well-organized data for
better business decisions.
Disadvantages of Data Preprocessing
Time-Consuming: Requires significant time and effort to clean, transform,
and organize data.
Resource-Intensive: Demands computational power and skilled personnel
for complex preprocessing tasks.
Potential Data Loss: Incorrect handling may result in losing valuable
information.
Complexity: Handling large datasets or diverse formats can be challenging.
Discretization By Histogram Analysis
The histogram is old method used to plot the attributes in a graph. Histo means to plot
and gram means chart. So basically, histogram is a graph of the poles. It is one of the
effective methods to summarize the distribution of a given attribute.
If the attribute is nominal, then a vertical bar is plotted for every known value of the
attribute, in which the height of the bar indicates the count/frequency of that attribute.
Graph is more precisely called as bar chart.
If attribute is numeric, then the range of the values are divided into disjoint but
consecutive partitions. Each such range can be termed as buckets/bins. The range of
every bucket is called width. Each bucket has nearly equal width. For example, for
the price attribute having values 1 to 100, can be divided into bins of 1 to 25, 25 to 50
and so on. for every subrange, a bar is plotted having the height that counts total no
of items in that subrange.
Discretization Technique:
Discretization is one form of data transformation technique. It transforms numeric
values to interval labels of conceptual labels. Ex. age can be transformed to (0-10,11-
20....) or to conceptual labels like youth, adult, senior.
There are different techniques of discretization:
1. Discretization by binning: It is unsupervised method of partitioning the
data based on equal partitions, either by equal width or by equal frequency
2. Discretization by Cluster: clustering can be applied to discretize numeric
attributes. It partitions the values into different clusters or groups by
following top down or bottom-up strategy
3. Discretization By decision tree: it employs top-down splitting strategy. It
is a supervised technique that uses class information.
4. Discretization By correlation analysis: Chi Merge employs a bottom-up
approach by finding the best neighboring intervals and then merging them
to form larger intervals, recursively
5. Discretization by histogram: Histogram analysis is unsupervised learning
because it doesn't use any class information like binning. There are various
partition rules used to define histograms.
Discretization By Histogram:
Histogram analysis is unsupervised learning because it doesn't use any class
information like binning. There are various partition rules used to define histograms.
In equal width histogram, values are partitioned in equal size bins or ranges. in our
earlier example, we have created bin of size 25, which is an equal-width histogram.
In equal frequency histogram, partition is done in such a way that every bucket
contains same number of data tuples.
Histogram algorithm can be applied to every partition recursively to create a concept
hierarchy until the predefined levels are generated. or a minimum interval size is used
to control the recursive procedure. It will specify a minimal width of a partition or
minimum number of values for each partition at every level
Example: The following data shows the price of commonly sold items in sorted
order: 1,1,4,4,4,4,7,7,9,9,9,9,9,11, 13,13,13,17,17,17,17,17,17, 21, 21, 21, 21, 25, 25,
25, 25, 25, 28, 28, 30,30, 30.
Following figure shows histogram for the current data:
Now, we will partition into equal width bins where every bucket has same size width
of 10.
Characteristic:
Histograms are very effective technique of data reduction which can work on sparse
and dense data as well as uniform and highly skewed data. Multidimensional
histograms can be used to capture data up to five attributes and are effective in
determining dependencies between attributes.
Importance of Discretization:
A discretization is important because it is useful:
1. To generate concept hierarchies.
2. Transform numeric data.
3. To ease evaluation and management of data.
4. To minimize data loss.
5. To produce a better result.
6. Generate a more understandable structure viz. decision tree.