Midterm Part 01 ─ STAT 2 Methods of Data Collection
Lesson 1: Basic Concepts in Statistics 1. Survey - an investigation of one or more
Statistics characteristics of a population.
✓ Greeks; statistiks i) Census – method of gathering the facts of
✓ The early use of statistics can be traced from the interest or pertinent data on every unit of the
administration of the state regarding the population population.
and property usually for war and finance purposes. ii) Sample Survey – method by which data from a
✓ the science of collection, organization, presentation small but representative cross-section of the
and analysis of data. population are scientifically collected and
✓ one can draw conclusion and make reasonable decision analyzed.
based on the analysis of that data. 2. Observation - makes possible the recording of
behavior but only at the time of occurrence. It is also
Area of Statistics employed when the subjects cannot talk or write.
1. Descriptive statistics - a set of methods involving the 3. Existing records - Data from published materials like
collection, presentation and summarization by means reports, personal files, and historical records will be
of numerical descriptions. utilized.
2. Inferential statistics - a set of methods that allow 4. Simulation - use of a mathematical or physical model
estimation or testing of the characteristics of the to reproduce the conditions of a situation or process. It
population based only from the sample drawn from allows you to study situations that are impractical or
that population. even dangerous to create in real life.
5. Experiment - In performing, a treatment is applied to
Definition of Terms part of a population and responses are observed. Data
✓ Data - used to describe a collection of natural are obtained under controlled conditions.
phenomena descriptors, including the result of
experience, observation or experiment. Classification of Data
✓ Population - entirety of individuals or objects of 1. Primary data – data that are collected directly from
interest. the subjects/objects of the study. These
✓ Parameters - measures of the population. subjects/objects may be people, experimental animals,
✓ Statistics or Estimates - measures of the sample. or the environment.
✓ Variable - characteristics of an individual or object Advantages: Original in character, Accurate, More
that can be measured. information, Confidential information can be collected
tactfully
Scale of Measurement Disadvantages: Expensive, Time-consuming,
1. Nominal Scale - a scale of measurement in which personal prejudice and bias can destroy the purpose,
objects or individuals are assigned into distinct may not be reliable if collected carelessly
categories and have no numerical properties. This is 2. Secondary data – these are previously collected data
the lowest scale. that are found in publications of both government and
2. Ordinal scale - has the property of a nominal scale in non-government institutions, research papers, books,
which categories can be ranked. periodicals, pamphlets, computer files, microfilms or
3. Interval scale - a scale of measurement in which the internet.
objects or individuals have the characteristics of Advantages: Collected quickly and cheaply, when
ordinal scale. The difference between the values is a using official statistics can be more reliable and
constant size. There is no absolute zero. acceptable
4. Ratio scale - a scale of measurement in which objects Disadvantages: Can’t meet specific needs, difficult in
or individuals have all the characteristics of interval assessing the accuracy
scale but it has an absolute zero value.
Sample - representative of a whole
Types of Variables Sampling Technique - a procedure used to determine the
1. Qualitative Variables are variables that can be members of a sample.
classified into categories, according to characteristics Sampling frame - a list, or set of the elements belonging to
or attributes. the population from which the sample will be drawn.
2. Quantitative variables are variables that are
numerical or you can possibly rank them. Determine Sample Size:
i) Discrete variables assume only certain values I. Slovin’s Formula - primarily used in the descriptive
and are countable. studies where the population is known and the margin of error
ii) Continuous variables are variables that can is preidentified.
assume any values between two values. 𝑁
𝑛=
1 − 𝑁𝑒 2
Lesson 2: Data Collection, Organization, and Presentation
Collection of Data II. According to Gay and Mills (2016), the larger the
✓ goal of every statistical study which will be used in population size, the smaller the percentage of the population
making decisions. required to get a representative sample.
✓ Decision made from the results of any statistical study ✓ For N=100 or less, survey the entire population
is only as good as the process used to obtain the ✓ For N=500, 50% should be sampled
information. If the process is flawed, then the resulting ✓ For a population of around 1500, 20% sample size
decision is questionable. ✓ Beyond a certain point (N=5000), sample size of 400
III. According to Fraenkel and Wallen (2011), the guideline A good statistical table has four essential parts:
in selecting a sample size will be as follow: 1. Table heading – includes the table number and table
✓ Descriptive study, N=100 title. The title should briefly explain the contents of the
✓ Correlational studies, N=50 table.
✓ Experimental or Causal-Comparative, N=30 per group 2. Stub – items or classification written on the first
but 15 per group is alright if tightly controlled column and identifies what are written on the rows.
3. Caption or box head – includes the items or
Sampling Techniques classifications written on the first row and identifies
A. Probability or Random Sampling is a procedure wherein what are contained in the columns.
every element of the population is given an equal chance of 4. Body –the main part of the table and it contains the
being selected in the sample. substance or the figures of one’s data.
1. Simple Random Sampling - giving each sampling
unit an equal chance of being included in the sample. In the construction of a table:
i) Fish bowl or Lottery Every table must be self-explanatory.
ii) Table of random numbers The title should be clear and descriptive.
2. Systematic Random Sampling – Samples are selected The title gives information about what, where, how,
by using every kth individual from a population. The and when the data were taken.
first individual selected is a random number between 1
and k. ✓ Graphical presentation, the data are presented in
3. Stratified Random Sampling - Separate the population graphs, charts, or diagrams. Graph is a pictorial
into nonoverlapping groups called strata and then representation of a set of data that shows relationship.
obtaining a simple random sample from each stratum.
i) Proportional Allocation – This process chooses Types of Graphs
sample sizes proportional to the sizes of the 1. Line Graph - shows the relationship between two sets
different subgroups or strata. of quantities. Line graph is similar to the graph drawn
ii) Equal Allocation – This process chooses the in Cartesian plane where the points are plotted using
same number of samples from each group vertical and horizontal axes.
regardless of its size. 2. Bar Graph - consists of vertical or horizontal bars of
4. Cluster Sampling - called area sampling because this equal widths.
is usually applied when the population is large. 3. Pie Chart - appropriate in comparing the parts with
Subjects are selected by dividing the population into the whole.
groups (clusters) and then some of the groups are 4. Pictograph - through the use of pictographs or picture
selected randomly. graphs. In this type of chart, actual pictures or
5. Multi-stage Sampling - The sample is randomly facsimiles of the objects under study are used to
selected through two or more steps or stages. represent values. Each figure is considered a unit
representing a definite number.
B. Nonprobability or Nonrandom Sampling is a procedure 5. Stem-and-leaf diagram - a visual presentation of raw
wherein not all the elements in the population are given a data. In this set up, the numbers (data) are broken into
chance of being included in the sample. tens digit and unit digits. Every row represents the
1. Haphazard or Accidental Sampling. Samples are stem and the numbers on the right are the leaf.
picked as it comes to the researcher. 6. Frequency distribution - arrangement of data that
2. Judgment or Purposive Sampling. The samples are shows the number of times or frequency of occurrence
selected by the researcher subjectively. The researcher of the different values of the variables.
will pick a sample that he/she believes is representative i) Qualitative frequency distributions are usually
to the population of interest. constructed for discrete variables
3. Convenience Sampling. Selecting a sample based on ii) Quantitative frequency distributions are
the convenience of the researcher or using data from usually constructed for continuous variable.
population members that are readily available. 7. Histogram - series of columns or vertical rectangles,
4. Snowball Sampling. A special nonprobability method each having as its base one class interval, and the
used when the desired sample characteristic is rare. frequency or number of cases in that class as its height.
Snowball sampling relies on referrals from initial 8. Frequency polygon - graph of the class mark against
subjects to generate additional subjects. the frequency. The shape of the histogram or the
frequency polygon gives an idea of the shape of the
Data Presentation distribution.
✓ Textual form, the researcher uses the sentences to 9. Cumulative frequency distribution or Ogive - tells
convey the information contained in the data. This is us how many observations lie above or below certain
incorporated with important figures only. Textual form values rather than merely recording the number of
of presentation can be seen on news reports. observations within intervals.
✓ Tabular form, the data are presented in rows and i) The less than ogive is a graph showing the how
columns. This systematic arrangement of data is called many values are below a certain upper class
a statistical table. Through this presentation, data can boundaries.
easily be understood. In addition, you can easily ii) The greater than ogive is a graph showing how
compare and contrast the data. many values are above a certain upper class
boundary.