0% found this document useful (0 votes)
9 views7 pages

Data Collection Methods and Analysis Techniques

Chapter 2 covers data collection methods, types of data, and algorithms for analysis. It includes multiple-choice questions, fill-in-the-blanks, and short answer questions to assess understanding of qualitative and quantitative data, primary and secondary sources, and the characteristics of big data. The chapter emphasizes the importance of data collection in decision-making and research.

Uploaded by

chitra.merin
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
9 views7 pages

Data Collection Methods and Analysis Techniques

Chapter 2 covers data collection methods, types of data, and algorithms for analysis. It includes multiple-choice questions, fill-in-the-blanks, and short answer questions to assess understanding of qualitative and quantitative data, primary and secondary sources, and the characteristics of big data. The chapter emphasizes the importance of data collection in decision-making and research.

Uploaded by

chitra.merin
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Chapter 2: Arranging and Collecting Data

MCQs

1. What is data collection?


a) The process of deleting unwanted data
b) Gathering and measuring information systematically
c) Randomly selecting data from different sources
d) Modifying data for presentation

Answer: b

2. Which of the following is a primary data source?


a) Census data
b) Social media data tracking
c) Government reports
d) News articles

Answer: a

3. What is an example of a categorical variable?


a) Age
b) Weight
c) Nationality
d) Height

Answer: c

4. A school named ABC has recorded the total marks of every student in the
class. This an example of:

a) Qualitative data
b) Quantitative data
c) Both qualitative and quantitative data
d) None of the above

Answer: b

5. A food delivery app has asked for your feedback on the quality of the food.
You have written two paragraphs to describe the food. This is an example of:
1) Qualitative data
2) Quantitative data
3) Both qualitative and quantitative data
4) None of the above
Answer: a
6. You need to predict what the temperature will be for next Friday. Which
algorithm will you use?
a) Clustering
b) Regression
c) Anomaly detection
d) Binary classification
Answer: b
7. You need to predict if your car tyre will last for the next 1000 km. Which
algorithm will you use?
a) Clustering
b) Regression
c) Anomaly detection
d) Binary classification
Answer: d
8. Which of the following are the benefits of Big data processing?
a) Business can utilize outside intelligence while making decisions
b) Improved customer service
c) Better optimal efficiency
d) All of the above
Answer: d
9. The analysis of large amounts of data to see what patterns or other useful
information can be found is known as
a) Data Analysis
b) Information Analytics
c) Big data Analytics
d) Data Analytics
Answer: c
10. Big data analysis does the following except
a) Collects data
b) Spreads data
c) Organizes data
d) Analyzes data
Answer: b
11. Primary data for the research process be collected through
a) Experiment
b) Survey
c) Both a and b
d) None of the above
Answer: c
12. The advantage of secondary data are low cost, speed, availability, and
flexibility
a) True
b) False
Answer: a
13. The method of getting primary data by watch people is called
a) Survey
b) Informative
c) Observational
d) Experimental
Answer: c
Fill in the Blanks

1. A variable that can take numerical values is called a quantitative variable.


2. A data collection method where observations are made firsthand is called
primary data collection.
3. Big Data refers to large and complex datasets that require advanced processing
techniques.

State Whether True or False

1. The method of gathering data for calculating and analyzing reliable insights is
known as data [Link]
2. Data collection tools have fundamentally changed the way businesses
[Link]
3. Interview is one of the common methods of diagnosing and solving social
[Link]
4. Numerical variables represent types of data which may be divided into
[Link]
5. Ordinal data mixes numerical and categorical data. True
6. Quantitative data comes out as numbers or values that can be [Link]
7. Velocity refers to the size of the data to be gathered, stored and analyzed. False
8. Volume determines the rate at which data is generated. False

Short Answer Questions

1. What are data collection tools?

Data collection tools are methodologies employed to gather data from a targeted
and selected group of people to asses predefined parameters by analyzing the data
and gaining rich insights about the same.

2. What is qualitative data? Give examples.

Qualitative data involves a descriptive judgement using concept words instead of


numbers. Gender, country name, animal species, and emotional state are examples
of qualitative information.

3. Define Big data.

Big data can be defined as the technologies and initiatives that involves data that
is too diverse, fast changing or massive for conventional technologies, skills and
infrastructure to address efficiently.

4. Define quantitative and qualitative data.


o Quantitative data consists of numerical values that can be measured
(e.g., height, weight, temperature).
oQualitative data consists of descriptive information that cannot be
measured numerically (e.g., color, nationality, customer reviews).
5. What are primary and secondary data sources?
o Primary Data: Collected firsthand through surveys, experiments,
and interviews.
o Secondary Data: Previously collected data from books, government
reports, or online sources.
6. What is the importance of data collection?
o Data collection helps in decision-making, trend analysis, and
problem-solving. Businesses rely on customer data to improve
services, and researchers use data to support scientific findings.
7. What is the meaning of velocity in big data?

Velocity determines the rate at which data is generated. It is the measure of how
fast the data is coming in. Facebook has to handle a huge of photographs and other
posts, comments and likes every day. It has to ingest it all, process it, file it and
somehow, later, be able to retrieve it.

8. What happens in personal interviews while collecting data?

Ans. The interviewer asks questions generally in a face to face contact to the
interviewee. These interviews may be organised for employment purpose, say by
UPSC, for achieving a particular social or political purpose.

9. What is scale or rating question in data collection?

Ans. The question displays a scale of answer options from any range (0 to 100, 1 to
10, etc.). The respondent selects the number that most accurately represents their
response.

10. What is data collection software?

Ans. Data collection software is a computerised system for the collection and storage
of qualitative and quantitative data in an electronic form. The benefits of using data
collection software are that they eliminate the use of paper surveys and allow data to
be quickly exported for data analysis and reporting.

Long Answer Questions

1. Explain the different methods of data collection.


o Surveys: Structured questionnaires used to gather public opinion.
o Observations: Recording real-world behaviors (e.g., traffic flow
studies).
oExperiments: Conducting tests under controlled conditions (e.g., drug
trials).
o Web Scraping: Extracting data from websites (e.g., social media
trends).
2. What is Big Data? Discuss its characteristics.
o Big Data refers to large and complex datasets that require advanced
tools for analysis. It has the following characteristics:
 Volume: Huge amount of data.
 Variety: Different types of data (structured, unstructured).
 Velocity: Rapid data generation.
 Veracity: Accuracy and reliability of data.,
3. What is the difference between quantitative and qualitative data?

Ans. Quantitative data means numbers or values that can be measured.


For example, number of times a product has been searched on the World Wide
Web or number of books exported per month. On the other hand qualitative
data involves a descriptive judgment using concept words instead of numbers.
Gender, country name, animal species, and emotional state are examples of
qualitative information.

As compared to quantitative data, the qualitative data is subjective.


Quantitative data helps to understand experiences in depth. The qualitative
method involves elements like feelings, emotions or subjective perception of
the researcher. On the other hand, quantitative method includes a
questionnaire with close-ended questions and using methods or correlation,
regression, mean and mode.

4. Differenciate Univariate and Multivariate data

5. Explain the 5 ways in which you can question your data.

Ans. Based on the type of data, we need to ask five simple questions to the data.
Question 1: Is this A or B?

Some questions can have only two possible answers. To predict this, a family of
algorithms is used, and the mechanism is called Binary Classification or two-
class classification.

Question 2: Is this odd?

Sometimes we find unexpected records in a set of mostly consistent [Link]


are called anomalies and could be a cause of concern. Algorithms used for these
types of questions are called Anomaly Detection Algorithms.

Question 3: How much or how many?

When we need to predict numerical values based on the data. The algorithms
which predict these values are called Regression Algorithms.

Question 4: Can I group the data?

Sometimes data may be separated into distinct groups. This approach is called
Clustering.

Question 5: What should I do now?

Based on trial and error, machines take some actions. These are questions that,
generally, a machine or robot is programmed to do. These types of learning are
called Reinforcement Learning.

Higher Order Thinking Skills (HOTS) Questions


1. If you had to conduct a survey on the effects of online learning on
students’ academic performance, what data collection method would you
use and why?
o Think about primary vs. secondary data, sample size, and data
accuracy.
2. You are given a dataset of students’ scores in Mathematics and Science.
How would you determine if there is a relationship between the two
subjects?
o Discuss methods such as correlation analysis and graphical
visualization.
3. Imagine you are working as a data analyst for an agricultural company.
You need to predict crop production based on rainfall, soil quality, and
temperature data. What kind of data would you collect, and how would
you analyze it?
o Explain the importance of multivariate analysis in prediction models.
4. Why is it important to use a mix of primary and secondary data while
researching a social issue like unemployment?
o Discuss the advantages and limitations of both data types.
5. How can bias in data collection affect the accuracy of research findings?
Provide an example.
o Consider sampling bias, questionnaire design, and data
manipulation.
6. If you were asked to design a questionnaire for collecting data on
students' favorite sports, what factors would you consider to ensure
reliable and unbiased results?
o Discuss question design, response options, and sample selection.
7. A researcher collects data on pollution levels in a city but finds that some
areas have missing data. What strategies can be used to handle missing
values?
o Explain techniques like data imputation, removing incomplete data,
or using alternative data sources.

Common questions

Powered by AI

Quantitative data consists of numerical values that can be measured, such as height, weight, and temperature, allowing for objective comparison and statistical analysis. Qualitative data, however, involves descriptive judgments using concept words like gender, country name, and emotional state, providing deeper insight but lacking numerical precision . The choice between these types can greatly impact research outcomes, as quantitative data often yields clear patterns and correlations, whereas qualitative data can offer rich, subjective context .

Utilizing a combination of primary and secondary data allows researchers to gain a comprehensive view of social issues like unemployment. Primary data, collected specifically for the research, offers current, specific insights but may be costly and time-consuming to collect. Secondary data, however, provides background information quickly and at a lower cost but might not be as specific or up-to-date . By mixing these data types, a study can benefit from both depth and breadth of information, enhancing the reliability and richness of research outcomes .

Designing a reliable questionnaire involves careful consideration of question clarity, avoiding leading or biased language, and ensuring a logical flow that considers the respondent's perspective . Providing comprehensive response options, using clear and concise language, and pre-testing the questionnaire can prevent misunderstandings and bias, enhancing data quality. Sample selection must be representative to ensure the results are generalizable to the intended population, maintaining data integrity .

Clustering algorithms group data into distinct clusters based on similarity, helping identify patterns and trends without prior knowledge of group characteristics, useful in market segmentation or image recognition . In contrast, anomaly detection algorithms identify data points that deviate significantly from the norm, detecting unusual activity that may indicate fraud or malfunction . Understanding these distinctions allows analysts to choose appropriate methods for specific tasks, enhancing the precision of pattern identification and anomaly detection, thus improving decision-making and insights generation .

Data collection underpins business decision-making by providing accurate, up-to-date information that drives strategic initiatives. Collecting reliable data facilitates informed decisions, identifies market trends, and uncovers consumer preferences, which can lead to improved customer services and competitive advantages . It is fundamental because it transforms raw data into actionable insights, shaping decisions that align with business objectives and market demands effectively .

Binary classification algorithms categorize data into two distinct classes, suitable for yes/no or A/B type decisions such as spam detection or patient diagnosis . Regression algorithms, on the other hand, predict continuous numerical values based on data inputs, used for forecasting metrics like sales revenue or temperature changes . The distinction lies in the type of output required—discrete categories for classification versus continuous numerical forecasting for regression—each suited for scenarios requiring specific predictive goals .

Data collection tools streamline data gathering by digitizing and automating processes, allowing for faster, more accurate data collection, which can be analyzed to drive strategic business decisions . These tools can reduce errors, save time, and enhance insights by integrating with analysis software, profoundly transforming business operations . However, limitations include potential over-reliance on digital tools, data security concerns, and the necessity for continuous updates to handle new data types and sources, which can be costly and complex .

Data veracity addresses the trustworthiness and quality of data, crucial for ensuring that decisions are based on accurate and reliable information . High veracity reduces risks associated with erroneous or misleading data, which can negatively impact decision-making, leading to faulty strategies or resource misallocation. Ensuring data veracity strengthens confidence in insights derived from big data, thus optimizing data-driven strategies and outcomes .

Velocity refers to the speed at which data is generated and processed, encompassing the rapid arrival and accessibility of data . This differentiates it from volume, which is about the sheer amount of data, and variety, which pertains to different data types such as structured or unstructured data . Velocity is crucial as it impacts how quickly data can be turned into actionable insights, heavily influencing real-time decision-making processes .

Strategies for handling missing data include data imputation, which involves estimating missing values, deleting rows with missing data, or using more robust models like predictive mean matching to estimate missing values more accurately . Each method impacts data analysis differently; for example, deletion may reduce dataset size, potentially biasing results, while imputation may introduce errors if assumptions about the missing data are incorrect . Choosing the right strategy depends on understanding the data's nature and the extent and pattern of missingness to preserve data integrity and accuracy.

You might also like