0% found this document useful (0 votes)
2 views6 pages

Data Collection

The document discusses the importance of proper data collection and sampling design in research, emphasizing the need for systematic methods to gather accurate information. It outlines the differences between primary and secondary data, various data collection methods, and the significance of determining an appropriate sample size for valid results. Additionally, it highlights the consequences of poorly collected data and provides guidelines for ensuring effective data gathering practices.

Uploaded by

joshua.dragon
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views6 pages

Data Collection

The document discusses the importance of proper data collection and sampling design in research, emphasizing the need for systematic methods to gather accurate information. It outlines the differences between primary and secondary data, various data collection methods, and the significance of determining an appropriate sample size for valid results. Additionally, it highlights the consequences of poorly collected data and provides guidelines for ensuring effective data gathering practices.

Uploaded by

joshua.dragon
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

DATA COLLECTION AND BASIC CONCEPTS IN SAMPLING DESIGN

DATA COLLECTION

Everyone gathers and uses information daily, often in numerical or statistical form. With so much data available (from
media, conversations, and more), it’s important to know how to collect, filter, and interpret it properly—especially in
business, research, and policy-making.

Good data helps us make informed decisions and avoid misleading conclusions. Poorly collected or misused data can
lead to incorrect assumptions and wasted resources. For example, radio or television talk shows regularly ask poll
questions for which respondents must call in or use the Internet to supply their vote. Most likely, the individuals who are
going to call in are those who have a strong opinion about the topic. This group is not likely to be representative of
people in general, so the results of the poll are not meaningful. Whenever we look at data, we should be mindful of
where the data come from.

Even when data tell us that a relation exists, we need to investigate. For example, a study showed that breast-fed
children have higher IQs than those who were not breast-fed. Does this study mean that a mother who breast-feeds her
child will increase the child’s IQ? Not necessarily. It may be that some factor other than breast-feeding contributes to
the IQ of the children. In this case, it turns out that mothers who breastfeed generally have higher IQs than those who
do not. Therefore, it may be genetics that leads to the higher IQ, not breast-feeding.

Data collection is the process of gathering and measuring information on variables of interest, in an established
systematic fashion that enables one to answer stated research questions, test hypotheses, and evaluate outcomes.

Without proper planning for data collection, a number of problems can occur. If the data collection steps and processes
are not properly planned, the research project can ultimately end up with a data set that does not serve the purpose for
which it was intended. For example, if more than one person is involved in the data collection, but data collectors do not
follow consistent data collection practices, they can end up with data with different units, collection processes, and
variable names.

Consequences of Improperly Collected Data:


 Inability to answer research questions accurately.
 Inability to repeat and validate the study.
 Distorted findings resulting in wasted resources.
 Misleading other researchers to pursue fruitless avenues of investigation.
 Compromising decisions for public policy.
 Causing harm to human participants and animal subjects.

Steps in Data Gathering:


1. Set the objectives for collecting data
2. Determine the data needed based on the set objectives.
3. Determine the method to be used in data gathering and define the comprehensive data collection points.
4. Design data gathering forms to be used.
5. Collect data.

SOURCES OF DATA
The statistical data may be classified under two categories, depending on the sources. approaches: Primary Data and
Secondary Data.

Whether conducting research in the social sciences, humanities, arts, or natural sciences, the ability to distinguish
between primary and secondary sources is essential.

1. Primary Sources - Provide a first-hand account of an event or period and are considered to be authoritative. They
represent original thinking, report on discoveries or events, or they can share new information. Often, these sources
are created at the time the events occurred, but they can also include sources that are created later. They are
usually the first formal appearance of original research.
 Primary Data - data documented by the primary source. The data collectors documented the data
themselves.

2. Secondary Data – offer an analysis, interpretation, or a restatement of primary sources and are considered to be
persuasive. They often involve generalization, synthesis, interpretation, commentary or evaluation in an attempt to
convince the reader of the creator`s argument. They often attempt to describe or explain primary sources.
 Secondary Data - are data documented by a secondary source. The data collectors had the data
documented by other sources.

Secondary data are less expensive to collect both in money and time. These data can also be better utilized, and
sometimes the quality of such data may be better because these might have been collected by persons who were
specially trained for that purpose.

The primary data can be collected by the following five methods:

1. Direct Personal Interviews - The researcher has direct contact with the interviewee. The researcher gathers
information by asking questions to the interviewee.

2. Indirect/Questionnaire Method - This method of data collection involves sourcing and accessing existing data that
were originally collected for the purpose of the study.

Key Design Principles of a Good Questionnaire:


 Keep the questionnaire as short as possible.
 Decide on the type of questionnaire (Open Ended or Closed Ended).
 Write the questions properly.
 Order the questions appropriately.
 Avoid questions that prompt or motivate the respondent to say what you would like to hear.
 Write an introductory letter or an introduction.
 Write special instructions for interviewers or respondents.
 Translate the questions if necessary.
 Always test your questions before taking the survey. (Pre-test)

An open-ended question is a type of question that does not include response categories. The respondent is not given
any possible answers to choose from. This type of question is usually appropriate for collecting subjective data. It
permits free responses that should be recorded in the respondent’s own words.

Example:

 Can you describe exactly what the traditional birth attendant did when your labor started?
 What do you think are the reasons for the high dropout rate of village health committee members?

A closed-ended question is a type of question that includes a list of response categories from which the respondent will
select their answer. It is useful if the range of possible responses is known. This type of question is usually appropriate
for collecting objective data.

Example:

 Did you eat any of the following foods yesterday?

Yes No
Fish or meat
Eggs
Milk or cheese
No Take Note!

Question wording and question order have a large effect on the responses obtained.
3. Focus group - a group interview of approximately six to twelve people who share similar characteristics or common
interests. A facilitator guides the group based on a predetermined set of topics.

4. Experiment - a method of collecting data where there is direct human intervention on the conditions that may affect
the values of the variable of interest.

Bear in mind that the experimental method has several limitations that you should be aware of.
 Ethical, moral, and legal Concerns
 Unrealistic Controlled Environments
 Inability to Control for All Variables

5. Observation - a technique that involves systematically selecting, watching and recoding behaviors of people or other
phenomena and aspects of the setting in which they occur, for the purpose of getting (gaining) specified
information. It includes all methods from simple visual observations to the use of high level machines and
measurements, sophisticated equipment or facilities such as:
 Radiographic
 biochemical
 X-ray machines
 Microscope
 Clinical examinations
 Microbiological examinations

It gives relatively more accurate data on behavior and activities but Investigators or observer’s own biases, prejudice,
desires, and etc. and needs more resources and skilled human power during the use of high level machines.

The secondary data can be collected by the following five methods:


1. Published report on newspaper and periodicals.
2. Financial Data reported in annual reports.
3. Records maintained by the institution.
4. Internal reports of the government departments.
5. Information from official publications.

Take Note!
 Always investigate the validity and reliability of the data by examining the collection method employed by
your source.
 Do not use inappropriate data for your research.
 The choice of methods of data collection is largely based on the accuracy of the information they yield.

SAMPLE SIZE

“How many participants should be chosen for a survey”?

One of the most frequent problems in statistical analysis is the determination of the appropriate sample size. One may
ask why sample size is so important. The answer to this is that an appropriate sample size is required for validity. If the
sample size it too small, it will not yield valid results. An appropriate sample size can produce accuracy of results.
Moreover, the results from the small sample size will be questionable. A sample size that is too large will result in
wasting money and time because enough sample will normally give an accurate result.

The sample size is typically denoted by n and it is always a positive integer. No exact sample size can be mentioned here
and it can vary in different research settings. However, all else being equal, large sized sample leads to increased
precision in estimates of various properties of the population.

Take Note!
 Representativeness, not size, is the more important consideration.
 Use no less than 30 subjects if possible. - If you use complex statistics, you may need a minimum of 100 or
more in your sample (varies with method).
Choosing of sample size depends on nonstatistical considerations and statistical considerations.
 Non-statistical considerations – It may include availability of resources, man power, budget, ethics and sampling
frame.

 Statistical considerations – It will include the desired precision of the estimate.

Three criteria need to be specified to determine the appropriate sample size:


1. Level of Precision
- Also called sampling error, the level of precision, is the range in which the true value of the population is
estimated to be.

2. Confidence Interval
- It is statistical measure of the number of times out of 100 that results can be expected to be within a
specified range. For example, a confidence interval of 90% means that results of an action will probably
meet expectations 90% of the time.
- To find the right z – score to use, refer to the table:

Desired Confidence Level Z-score


80% 1.28
85% 1.44
90% 1.65
95% 1.96
99% 2.58
3. Degree of Variability
- Depending upon the target population and attributes under consideration, the degree of variability varies
considerably. The more heterogeneous a population is, the larger the sample size is required to get an
optimum level of precision.

Methods in Determining the Sample Size:

 Estimating the Mean or Average


- The sample size required to estimate the population mean µ with a level of confidence with a
specified margin of error e, given by:

where z = the z-score corresponding to the level of confidence


σ = the population standard deviation
e = the level of precision

Take Note!
When σ is unknown, it is common practice to conduct a preliminary survey to determine s and use it as an estimate of
σ or use results from previous studies to obtain an estimate of σ. When using this approach, the size of the sample
should be at least 30. The formula for the sample standard deviation s is:


2
Σ ( x− x )
s=
n−1

Example:
A soft drink machine is regulated so that the amount of drink dispensed is approximately normally distributed with a
standard deviation equal to 0.5 ounce. Determine the sample size needed if we wish to be 95% confident that our
sample mean will be within 0.03 ounce from the true mean.
Solution:
z = 1.96
σ = 0.5
e = 0.03

Answer: We need a 1068 sample for our study.

 Estimating Proportion (Infinite Population)


- The sample size required to obtain a confidence interval for p with a specified margin of error e is
given by:

where z = the z-score corresponding to the level of confidence


p = population proportion
e = the level of precision

There is a dilemma in this formula:


x
It depends on p= , which we know only after we have taken the sample.
N

There are two ways to solve this dilemma:


1. We could determine a preliminary value for p based on a pilot study or an earlier study.

Example:
If last month 37% of all voters thought that state taxes are too high, then it is likely that the proportion with that
opinion this month will not be dramatically different, and we would use the value 0.37 for p in the formula.

2. To replace p in the formula with 0.5.

When p = 0.5, the maximum value of p ( 1− p )=0.25. This is called the most conservative estimate, since it gives
the largest possible estimate of n.

Where; Confidence level is 95%. The level of precision is 0.05.

Example:
Suppose we are doing a study on the inhabitants of a large town, and want to find out how many households
serve breakfast in the mornings. We don’t have much information on the subject to begin with, so we’re going
to assume that half of the families serve breakfast: this gives us maximum variability. So p = 0.5. We want 99%
confidence and at least 1% precision.

Solution: The z – z-score for confidence level 99% in the z – z-table is 2.58.

n ≥ (2.580.01)2 0.5(1 − 0.5) = 16,641

Answer: We need a 16,641 sample for our study.

 Slovin`s Formula
- a formula that is used to determine what sample size should be chosen to study a given population,
depending on the researcher's error tolerance level.
- Slovin’s formula cannot be used to determine the sample size if the population size is very small.
- As a general rule, it was suggested by Gay (1978) that the sample size should be 20% of then given
population for small population sizes (less than 500), and the

For example, suppose that you want to know about the voting preferences of a given population. It is not
feasible to ask each person about their voting preferences. In such a situation, it is reasonable to select a
random sample from the entire population and interview only those people who are selected. But this
introduces scope for error in the study. We also need to decide how many people should be included in our
sample. Clearly, a larger sample will minimize the error, and a smaller sample greatly increases the error rate.
Generally, a 5% error rate is considered acceptable. Given an acceptable error limit, we can use Slovin’s formula
to decide the size of the randomly selected sample.

The Slovin`s Formula

Where:
N - the size of the population
e - the maximum acceptable error limit
n – sample size

Example 1:
Suppose that a company wants to conduct marketing supply research to know about consumer preferences. The
company estimates that a total of N = 10000 people are regular loyal customers of the company. How many of
these people should be interviewed to understand customer preferences? Take the margin of error to be 5%.
Solution: We want to choose a random sample of size ‘n’ from the entire population of customers.

Solution: Applying Slovin’s formula we get that,

Answer: This means that a total of 385 people should be randomly selected and interviewed to conduct
research on consumer preferences.

You might also like