0% found this document useful (0 votes)
22 views9 pages

Introduction to Statistical Concepts

Uploaded by

Leaneth Escamos
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
22 views9 pages

Introduction to Statistical Concepts

Uploaded by

Leaneth Escamos
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

LESSON 1: INTRODUCTION TO THE STATISTICAL Suppose you wanted to use this scenario as a gauge of the morality of

CONCEPTS students at your school by determining the percent of students who would
Statistics return the money. How might you do this? You could attempt to present
Statistics is the science of collecting, organizing, summarizing, and the scenario to every student at the school, but this would be difficult or
analyzing information to draw conclusions or answer questions. In impossible if the student body is large. A second possibility is to present
addition, statistics is about providing a measure of confidence in any the scenario to 50 students and use the results to make a statement about
conclusions. all the students at the school.

Let’s break this definition into four parts. The first part states that In the PHP100 study presented, the population is all the students at the
statistics involves the collection of information. The second refers to the school. Each student is an individual. The sample is the 50 students
organization and summarization of information. The third states that the selected to participate in the study.
information is analyzed to draw conclusions or answer specific questions.
The fourth part states that results should be reported using some measure Suppose 39 of the 50 students stated that they would return the money to
that represents how convinced we are that our conclusions reflect reality. the owner. We could present this result by saying that the percent of
students in the survey who would return the money to the owner is 78%.
•Statistics is important because it enables people to make decisions based This is an example of a descriptive statistic because it describes the
on empirical evidence. results of the sample without making any general conclusions about the
population. So, 78% is a statistic because it is a numerical summary
• Statistics provides us with tools needed to convert massive data into based on a sample. Descriptive statistics make it easier to get an
pertinent information that can be used in decision making. overview of what the data are telling us.

• Statistics can provide us information that we can use to make sensible If we extend the results of our sample to the population, we are
decisions. performing inferential statistics. The generalization contains uncertainty
because a sample cannot tell us everything about a population. Therefore,
What information is referred to in the definition? inferential statistics includes a level of confidence in the results. So,
rather than saying that 78% of all students would return the money, we
The information referred to the definition is the data. might say that we are 95% confident that between 74% and 82% of all
students would return the money. Notice how this inferential statement
According to the Merriam Webster dictionary, data are “factual includes a level of confidence (measure of reliability) in our results.
information used as a basis for reasoning, discussion, or calculation”.
It also includes a range of values to account for the variability in our
Data can be numerical, as in height, or nonnumerical, as in gender. In results. One goal of inferential statistics is to use statistics to estimate
either case data describe characteristics of an individual. parameters.

Field of Statistics PROCESS OF STATISTICS


A. Mathematical Statistics- The study and development of statistical 1. Identify the research objective.
theory and methods in the abstract. A researcher must determine the question(s) he or she wants answered.
The question(s) must clearly identify the population that is to be studied.
B. Applied Statistics- The application of statistical methods to solve real Identify the research objective.
problems involving randomly generated data and the development of new
statistical methodology motivated by real problems. 2. Collect the information needed to answer the questions.
Conducting research on an entire population is often difficult and
Example branches of Applied Statistics: psychometric, econometrics, and expensive, so we typically look at a sample. This step is vital to the
biostatistics. statistical process, because if the data are not collected correctly, the
conclusions drawn are meaningless. Do not overlook the importance of
Limitation of Statistics appropriate data collection.
● Statistics is not suitable to the study of qualitative phenomenon.
● Statistics does not study individuals. Example: A research objective is presented. For each research objective,
● Statistical laws are not exact. identify the population and sample in the study.
● Statistics table may be misused.
● Statistics is only, one of the methods of studying a problem. Example:
1. The Philippine Mental Health Associations contacts 1,028 teenagers
Definitions of terms: who are 13 to 17 years of age and live in Antipolo City and asked
• Universe is the set of all entities under study. whether or not they had been prescribed medications for any mental
• A Population is the total or entire group of individuals or observations disorders, such as depression or anxiety.
from which information is desired by a researcher. Apart from persons, a
population may consist of mosquitoes, villages, institution, etc. Population: Teenagers 13 to 17 years of age who live in Antipolo City
• An individual is a person or object that is a member of the population Sample: 1,028 teenagers 13 to 17 years of age who live in Antipolo City
being studied.
• A statistic is a numerical summary of a sample 2. A farmer wanted to learn about the weight of his soybean crop. He
• Sample is the subset of the population. randomly sampled 100 plants and weighted the soybeans on each plant.
• Descriptive statistics consist of organizing and summarizing data. Population: Entire soybean crop
Descriptive statistics describe data through numerical summaries, tables, Sample: 100 selected soybean crop
and graphs.
• Inferential statistics uses methods that take a result from a sample, 3. Organize and summarize the information.
extend it to the population, and measure the reliability of the result. Descriptive statistics allow the researcher to obtain an overview of the
• A parameter is a numerical summary of a population. data and can help determine the type of statistical methods the researcher
should use.
EXAMPLE: Consider the scenario.
You are walking down the street and notice that a person walking in front 4. Draw conclusion from the information.
of you drops PHP100. Nobody seems to notice the PHP100 except you. Take Note! If the entire population is studied, then inferential statistics is
Since you could keep the money without anyone knowing, would you not necessary, because descriptive statistics will provide all the
keep the money or return it to the owner? information that we need regarding the population.
LESSON 2: INTRODUCTION TO THE STATISTICAL
CONCEPTS (Cont.) 2. Ordinal Level - This involves data that may be arranged in some
order, but differences between data values either cannot be determined or
DISTINCTION BETWEEN QUALITATIVE AND meaningless. An ordinal scale not only classifies subjects but also ranks
QUANTITATIVE VARIABLES them in terms of the degree to which they possess a characteristic of
interest. In other words, an ordinal scale puts the subjects in order from
Variables are the characteristics of the individuals within the population. highest to lowest, from most to least. Although ordinal scales indicate
that some subjects are higher, or lower than others, they do not indicate
For example, recently my mother and I planted a tomato plant in our how much higher or how much better.
backyard. We collected information about the tomatoes harvested from
the plant. The individuals we studied were the tomatoes. The variable that Example:
interested us was the weight of a tomato. - Food Preferences
- Rank of a Military officer
Variables can be classified into two groups: - Social Economic Class (First, Middle, Lower)
1. Qualitative variables are variable that yields categorical responses. It
is a word or a code that represents a class or category. 3. Interval Level - This is a measurement level not only classifies and
orders the measurements, but it also specifies that the distances between
2. Quantitative variables take on numerical values representing an each interval on the scale are equivalent along the scale from low interval
amount or quantity. to high interval. A value of zero does not mean the absence of the
quantity. Arithmetic operations such as addition and subtraction can be
Example: Determine whether the following variables are qualitative or performed on values of the variable.
quantitative.
Example:
1. Haircolor (Qualitative) - Temperature on Fahrenheit/Celsius Thermometer
2. Temperature (Quantitative) - Trait anxiety (e.g., high anxious vs. low anxious)
3. Number of hamburger sold (Quantitative) - IQ (e.g., high IQ vs. average IQ vs. low IQ)
4. Number of children (Quantitative)
5. Zip code (Qualitative) 4. Ratio Level - A ratio scale represents the highest, most precise, level
of measurement. It has the properties of the interval level of measurement
Quantitative variables may be further classified into: and the ratios of the values of the variable have meaning. A value of zero
1. A discrete variable is a quantitative variable that either a finite means the absence of the quantity. Arithmetic operations such as
number of possible values or a countable number of possible values. If multiplication and division can be performed on the values of the
you count to get the value of a quantitative variable, it is discrete. variable.

2. A continuous variable is a quantitative variable that has an infinite Example:


number of possible values that are not countable. If you measure to get - Height and weight
the value of a quantitative variable, it is continuous. - Time
- Distance and speed
Example: Determine whether the following quantitative variables are
discrete or continuous. LESSON 3: DATA COLLECTION AND BASIC CONCEPTS IN
SAMPLING DESIGN
1. The number of heads obtained after flipping a coin five times.
(Discrete) INTRODUCTION
Everybody collects, interprets and uses information, much of it in
2. The number of cars that arrive at a McDonald’s drive-through between numerical or statistical forms in day-to-day life. It is a common practice
12:00 P.M and 1:00 P.M. (Discrete) that people receive large quantities of information everyday through
conversations, televisions, computers, the radios, newspapers, posters,
3. The distance of a 2005 Toyota Prius can travel in city conditions with a notices and instructions. It is just because there is so much information
full tank of gas. (Continuous) available that people need to be able to absorb, select and reject it. In
everyday life, in business and industry, certain statistical information is
4. Number of words correctly spelled. (Discrete) necessary, and it is independent to know where to find it how to collect it.

5. Time of a runner to finish one lap. (Continuous) The Definition of Data Collection
Data Collection is the process of gathering and measuring information on
variables of interest, in an established systematic fashion that enables one
LEVELS OF MEASUREMENT to answer stated research questions, test hypotheses, and evaluate
It is important to know which type of scale is represented by your data outcomes.
since different statistics are appropriate for different scales of
measurement. A characteristic may be measured using nominal, ordinal, Without proper planning for data collection, a number of problems can
interval and ratio scales. occur. If the data collection’s steps and processes are not properly
planned, the research project can ultimately end up with a data set that
1. Nominal Level - This is the first level of measurement, and it is does not serve the purpose for which it was intended. For example, if
characterized by data that consist of names, labels or categories only. The more than one person is involved in the data collection, but data
data cannot be arranged in ordering scheme. Nominal scales have no collectors do not follow consistent data collection practices, they can end
numerical value. up with data with different units, collection processes, and variable
names.
They are sometimes called categorical scales or categorical data. Such a
scale classifies persons or objects into two or more categories. Whatever CONSEQUENCES FROM IMPROPERLY COLLECTED DATA
the basis for classification, a person can only be in one category, and 1. Inability to answer research questions accurately.
members of a given category have a common set of characteristics. 2. Inability to repeat and validate the study.
3. Distorted findings resulting in wasted resources.
Example: 4. Misleading other researchers to pursue fruitless avenues of
- Method of payment (cash, check, debit card, credit card) investigation.
- Type of school (public vs. private) 5. Compromising decisions for public policy.
- Eye Color (Blue, Green, Brown) 6. Causing harm to human participants and animal subjects.
Take Note!
STEPS IN DATA GATHERING Question wording and question order have a large effect on the responses
1. Set the objectives for collecting data obtained.
2. Determine the data needed based on the set objectives.
3. Determine the method to be used in data gathering and define the Example:
comprehensive data collection points. Two surveys were taken in late 1993/early 1994 about Elvis Presley.
4. Design data gathering forms to be used.
5. Collect data. One survey asked: “In the past few years, there have been a lot of rumors
and stories about whether Elvis Presley is dead. How do you feel about
SOURCES OF DATA this? Do you think there is any possibility that these rumors are true and
Whether conducting research in the social sciences, humanities arts, or that Elvis Presley is still alive, or don’t you think so?”
natural sciences, the ability to distinguish between primary and secondary
sources is essential. Second survey asked: “A recent television show examined various
theories about Elvis Presley’s death. Do you think it is possible that Elvis
Primary Sources - Provide a first-hand account of an event or time is alive or not?”
period and are considered to be authoritative. They represent original
thinking, reports on discoveries or events, or they can share new 8% of the respondents to the first question said it is possible that Elvis is
information. Often these sources are created at the time the events still alive and 16% of respondents to the second question said it is
occurred but they can also include sources that are created later. They are possible that Elvis is still alive.
usually the first formal appearance of original research.
3. Focus Group - is a group interview of approximately six to twelve
The firsthand information obtained by the investigator is more reliable people who share similar characteristics or common interests. A
and accurate since the investigator can extract the correct information by facilitator guides the group based on a predetermined set of topics.
removing doubts, if any, in the minds of the respondents regarding
certain questions. High response rates might be obtained since the 4. Experiment is a method of collecting data where there is direct human
answers to various questions are obtained on the spot. It permits intervention on the conditions that may affect the values of the variable of
explanation of questions concerning difficult subject matter. interest.

Primary Sources of Data can be obtaibed using the following methods: Bear in mind that the experimental method has several limitations that
you should be aware of.
1. Direct personal interviews – The researcher has direct contact with - Ethical, Moral, and Legal Concerns
the interviewee. The researcher gathers information by asking questions - Unrealistic Controlled Environments
to the interviewee. - Inability to Control for All Variables

2. Indirect/Questionnaire Method - This methods of data collection 5. Observation is a technique that involves systematically selecting,
involve sourcing and accessing existing data that were originally watching and recoding behaviors of people or other phenomena and
collected for the purpose of the study. aspects of the setting in which they occur, for the purpose of getting
(gaining) specified information. It includes all methods from simple
visual observations to the use of high-level machines and measurements,
Key Design Principles of a Good Questionnaire sophisticated equipment or facilities.
1. Keep the questionnaire as short as possible.
2. Decide on the type of questionnaire (Open Ended or Closed Ended). SECONDARY SOURCES
3. Write the questions properly.
4. Order the questions appropriately. Secondary Sources - offer an analysis, interpretation or a restatement of
5. Avoid questions that prompt or motivate the respondent to say what primary sources and are persuasive. They often involve generalization,
you would like to hear. synthesis, interpretation, commentary or evaluation to convince the
6. Write an introductory letter or an introduction. reader of the creator's argument. They often attempt to describe or
7. Write special instructions for interviewers or respondents. explain primary sources.
8. Translate the questions if necessary.
9. Always test your questions before taking the survey. (Pre-test) Secondary Sources of Data can be obtained using the following methods:
1. Published report on newspaper and periodicals.
An open-ended question is a type of question that does not include 2. Financial Data reported in annual reports.
response categories. This type of question is usually appropriate for 3. Records maintained by the institution.
collecting subjective data. 4. Internal reports of the government departments.
5. Information from official publications.
A closed-ended question is a type of question that includes a list of
response categories from which the respondent will select his answer. Take Note!
This type of question is usually appropriate for collecting objective data.
• Always investigate the validity and reliability of the data by examining
the collection method employed by your source.

• Do not use inappropriate data for your research.

Secondary data are less expensive to collect both in money and time.
These data can also be better utilized and sometimes the quality of such
data may be better because these might have been collected by persons
who were specially trained for that purpose.

On the other hand, such data must be used with great care, because such
data may also be full of errors because the purpose of the collection of the
data by the primary agency may have been different from the purpose of
the user of these secondary data.
On Designing Questionnaires
Secondly, there may have been bias introduced, the size of the sample
may have been inadequate, or there may have been arithmetic or
definition errors, hence, it is necessary to critically investigate the validity
of the secondary data.

LESSON 4: SAMPLE SIZE


The Need to Determine Appropriate Sample Size
3. Degree of Variability – Depending upon the target population
WHY WE NEED APPROPRIATE SAMPLE SIZE? and attributes under consideration, the degree of variability
“How many participants should be chosen for a survey”? varies considerably. The more heterogeneous a population is,
One of the most frequent problems in statistical analysis is the the larger the sample size is required to get an optimum level of
determination of the appropriate sample size. One may ask why sample precision.
size is so important.
Methods in Determining the Sample Size
The answer to this is that an appropriate sample size is required for SLOVIN'S FORMULA
validity. If the sample size it too small, it will not yield valid results. An Slovin’s Formula is used to calculate the sample size n given the
appropriate sample size can produce accuracy of results. Moreover, the population size and error. It is computed as:.
results from the small sample size will be questionable. A sample size
that is too large will result in wasting money and time because enough
sample will normally give an accurate result. where:
N is the population size
Characteristics of a Good Sample Size e is the level of precision or error estimate
The sample size is typically denoted by n and it is always a positive
integer. No exact sample size can be mentioned here and it can vary in Example. A researcher plans to conduct a survey about food preference
different research settings. However, all else being equal, large sized of BS Stat students. If the population of students is 1000, find the sample
sample leads to increased precision in estimates of various properties size if the error is 5%.
of the population.

Also, take note of the following:


- Representativeness, not size, is the more important consideration.
- Use no less than 30 subjects if possible.
- If you use complex statistics, you may need a minimum of RAOSOFT FORMULA
100 or more in your sample (varies with method) Example. A university has 5,000 students, and we want to survey them at
95% confidence level with a 5% margin of error.
POPULATION
The set of all individuals of interest in a particular study

SAMPLE
A set of individuals from a population, intended to represent the
population itself

Considerations in Determining your Sample Size


Choosing of sample size depends on nonstatistical considerations and
statistical considerations.

• Non-statistical considerations – It may include availability of


resources, man power, budget, ethics and sampling frame.
• Statistical considerations – It will include the desired precision of the
estimate.

3 STATISTICAL CONSIDERATIONS FOR SAMPLE SIZE


1. Level of Precision – Also called sampling error, the level of
precision, is the range in which the true value of the population
is estimated to be.
2. Confidence Interval – It is statistical measure of the number
of times out of 100 that results can be expected to be within a
specified range. For example, a confidence interval of 90%
means that results of an action will probably meet expectations
90% of the time. To find the right z – score to use, refer to the
table:
LESSON 5: BASIC SAMPLING DESIGN TYPES OF PROBABILITY SAMPLING

DESIGN 1. Simple Random Sampling


The goal in sampling is to obtain individuals for a study in - Also known as Fish Bowl Sampling
such a way that accurate information about the population - Most basic method of drawing a probability sample.
can be obtained. - Assigns equal probabilities of selection to each possible sample.
- Results to a simple random sample.
Reason for Sampling
-Important that the individuals included in a sample represent a cross Advantage: It is very simple and easy to use.
section of individuals in the population. Disadvantage: The sample chosen may be distributed over a wide
-If sample is not representative, it is biased. You cannot generalize to the geographic area.
population from your statistical data. When to use: This is preferable to use if the population is not widely
spread geographically. Also, this is more appropriate to use if the
Definition of Terms population is more or less homogenous with respect to the characteristics
• Observation unit - An object on which a measurement is taken. This is of the population.
the basic unit of observation, sometimes called an element. In studying TYPES OF PROBABILITY SAMPLING
human populations, observation units are often individuals.
• Target population - The complete collection of observations we want
to study.
• Sampled population - The collection of all possible observation units
that might have been chosen in a sample; the population from which the
sample was taken.
• Sample - A subset of a population.
• Sampling unit - A unit that can be selected for a sample. We may want
2. Systematic Random Sampling
to study individuals, but do not have a list of all individuals in the target
-It is obtained by selecting every kth individual from the population.
population. Instead, households serve as the sampling units, and the
- The first individual selected corresponds to a random number between 1
observation units are the individuals living in the households.
to k.
• Sampling frame - A list, map, or other specification of sampling units
in the population from which a sample may be selected. For a survey
Obtaining a Systematic Random Sample
using in-person interviews, the sampling frame might be a list of all street
1. Decide on a method of assigning a unique serial number, from 1 to N,
addresses.
to each one of the elements in the population.
• Sampling technique/Sampling Strategies - It is a plan you set forth to
2. Compute for the sampling interval
be sure that the sample you use in your research study represents the
population from which you drew your sample.
• Sampling Bias - This involves problems in your sampling, which
reveals that your sample is not representative of your population. 3. Select a number, from 1 to k, using a randomization mechanism. The
element in the population assigned to this number is the first element of
Some Ways in which Selection Bias Can Occur: the sample. The other elements of the sample are those assigned to the
- Deliberately or purposively selecting a “representative” sample. Mis- numbers and so on until you get a sample of size.
specifying the target population. Failing to include all the target
population in the sampling frame, called under coverage. Including Example:
population units in the sampling frame that are not in the target We want to select a sample of 50 students from 500 students under this
population, called over coverage. method kth item and picked up from the sampling frame.
-Having multiplicity of listings in the sampling frame. Substituting a
convenient member of a population for a designated member who is not
readily available. Failing to obtain responses from all the chosen sample.
(Nonresponse)
- Allowing the sample to consist entirely of volunteers. We start to get a sample starting form i and for every kth unit
subsequently. Suppose the random number i is 5, then we select 15, 25,
Advantage of Sampling Over Complete Enumeration 35, 45, .. .
- Less Labor
- Reduced Cost 2. Systematic Random Sampling
- Greater Speed Advantage: Drawing of the sample is easy. It is easy to administer in the
- Greater Scope field, and the sample is spread evenly over the population.
- Greater Efficiency and Accuracy Disadvantage: May give poor precision when unsuspected periodicity is
- Convenience present in the population.
- Ethical Considerations When to use: This is advisable to us if the ordering of the population is
essentially random and when stratification with numerous data is used.
Two Type of Samples
● Probability Sampling
● Non-Probability Sampling

Probability Sampling Techniques

WHAT IS PROBABILITY SAMPLING?


- Also known as random sampling in general.
-Samples are obtained using some objective chance mechanism, thus 3. Stratified Random Sampling
involving randomization. - It is obtained by separating the population into non-overlapping groups
- They require the use of a complete listing of the elements of the called strata and then obtaining a simple random sample from each
universe called the sampling frame. stratum.
- The probabilities of selection are known. - The individuals within each stratum should be homogeneous (or
- They are generally referred to as random samples. similar) in some way.
-They allow drawing of valid generalizations about the
universe/population.
Example:
A sample of 50 students is to be drawn from a population consisting of When to use: If the population can be grouped into clusters where
500 students belonging to two institutions A and B. The number of individual population elements are known to be different with respect to
students in the institution A is 200 and the institution B is 300. How will the characteristics under study, this preferable to use.
you draw the sample using proportional allocation?

Solution:
There are two strata in this case.

5. Multi - Stage Sampling


- Selection of the sample is done in two or more steps or stages, with
The sample sizes are 20 from A and 30 from B. sampling units varying in each stage.
- The population is first divided into a number of first-stage sampling
Then the units from each institution are to be selected by simple random units from which a sample is drawn. Smaller units, called the secondary
sampling. sampling units, comprising the selected first-stage units then serve as the
sampling units for the next stage. If needed additional stages may be
Advantage: Stratification of respondents is advantageous in terms of added until the units of observation for the survey are clearly identified.
precision of the estimates of the characteristics of the population. The units comprising the samples selected from the previous stage
Sampling designs may vary by stratum to adjust for the differences in the constitute the frame for the stages.
conditions across strata. It is easy to use as a random sampling design.
Disadvantage: Values of the stratification variable may not be easily Obtaining a Multi-Stage Sampling
available for all units in the population especially if the characteristic of 1. Organize the sampling process into stages where the unit of analysis is
interest is homogeneous. It is possible that there are not representative in systematically grouped.
one or two strata. Also, transportation costs can be high if the population 2. Select a sampling technique for each stage.
covers a wide geographic area. 3. Systematically apply the sampling technique to each stage until the
unit of analysis has been selected.
When to use: If the population is such that the distribution of the
characteristics of the respondents under consideration concentrated in Example:
small and spread segment of the population. Thus, this is preferred to use Suppose we wish to study the expenditure patterns of households in
if precise estimates are desired for stratified parts of the population and if NCR. We can select a sample of households for this study using simple
sampling problems differ in the various strata of the population. three-stage sampling.

- First, divide into smaller cities/municipalities and a random sample of


these cities/ municipalities is collected.
- Second, a random sample of smaller areas such as barangays is taken
from within each of the cities/municipalities chosen in the first stage.
- Third, a random sample of even smaller areas such as households is
taken from within each of the areas chosen in the second stage.

Advantage: It is easier to generate adequate sampling frames.


Transportation costs are greatly reduced since there is some form of
clustering among the ultimate or final samples; i.e., they are in the sample
lower-stage units.
Disadvantage: Its complexity in theory may be difficult to apply in the
field. Estimation procedures may be difficult for non-statisticians to
follow.
When to use: If no population list is available and if the population
covers a wide area.

4. Cluster Sampling
- You take the sample from naturally occurring groups in your
population.
- The clusters are constructed such that the sampling units are
heterogeneous within the cluster and homogeneous among the clusters.

Example: A researcher wants to survey academic performance of high Take Note!


school students in MIMAROPA. Used probability sampling if the main objective of the sample survey is
1. He/She can divide the entire population into different clusters. making inferences about the characteristics of the population under study.
2. Then the researcher selects several clusters depending on his
research through simple or systematic random sampling. Non-probability Sampling Techniques
3. Then, from the selected clusters the researcher can either - Samples are obtained haphazardly, selected purposively or are taken as
include all the high school students as subject or he can select volunteers.
several subjects from each cluster through simple or systematic - The probabilities of selection are unknown.
random sampling. - They should not be used for statistical inference.

Advantage: There is no need to come out with a list of units in the TYPES OF NON-PROBABILITY SAMPLING
population; all what is needed is simply a list of the clusters. It is also less 1. Accidental Sampling - There is no system of selection but only those
costly since the elements are physically closer together. whom the researcher or interviewer meets by chance.
Disadvantage: In actual field applications, adjacent households tend to 2. Quota Sampling - There is specified number of persons of certain
have more similar characteristics than households distantly apart. types is included in the sample. The researcher is aware of categories
within the population and draws samples from each category. The size of In the statistics class of 40 students, 3 obtained the perfect score
each categorical sample is proportional to the proportion of the of 50. Sixteen students got a score 40 and above, while only 3 got 19 and
population that belongs in that category. below. Generally, the students performed well in the test with 23 or 70%
3. Convenience Sampling - It is a process of picking out people in the getting a passing score of 38 and above.
most convenient and fastest way to get reactions immediately. This
method can be done by telephone interview to get the immediate Advantages:
reactions of a certain group of sample for a certain issue. ● The data would be more interpreted directly
4. Purposive Sampling - It is based on certain criteria laid down by the ● Can help in emphasizing some important points in data
researcher. People who satisfy the criteria are interviewed. It is used to ● Small sets of data can be easily presented.
determine the target population of those who will be taken for the study.
5. Judgement Sampling - selects sample in accordance with an expert’s Remember!
judgment. -Keep your paragraphs simple and short.
-Always make sure that the readers are provided with additional
WHEN TO USE NON-PROBABILITY SAMPLING? explanations about the relevance of the figures and its implications.
- Only few are willing to be interviewed
- Extreme difficulties in locating or identifying subjects TABULAR PRESENTATION
- Probability sampling is more expensive to implement ● It is a systematic and logical arrangement of data in the form of
- Cannot enumerate the population elements. Rows and Columns with respect to the characteristics of data.
● A table is best suited for representing individual information and
SOURCES OF ERROR represents both quantitative and qualitative information.

SOURCES OF SAMPLING ERRORS Advantages:


Non-sampling Error - Errors that result from the survey process. - Any ✦ More information may be presented.
errors that cannot be attributed to the sample-to-sample variability. ✦ Exact values can be read from a table to retain
Sources of Non-Sampling Error precision.
1. Non-responses ✦ Flexibility is maintained without distortion of data.
2. Interviewer Error ✦ Less work and less cost are required in the preparation.
3. Misrepresented Answers
4. Questionnaire Design ON PREPARING TABLES
5. Wording of Questions The making of a compact table itself is an art. This should contain all the
6. Selection Bias information needed within the smallest possible space. What the purpose
of tabulation is and how the tabulated information is to be used are the
Sampling Error - Error that results from taking one sample instead of main points to be kept in mind while preparing for a statistical table.
examining the whole population.
An ideal table should consist of the following main parts:
- Error that results from using sampling to estimate information regarding A. Title: The title must tell as simply as possible what is in the table. It
a population. should answer the questions:

LESSON 6: PRESENTATION OF DATA ✦ Who? White females with breast cancer, black males
with lung cancer.
INTRODUCTION ✦ What are the data? Counts, percentage distributions,
Data are usually collected in a raw format and thus the inherent rates.
information is difficult to understand. Therefore, raw data need to be ✦ Where are the data from? Example: One hospital, or the
summarized, processed, and analyzed to usefully derive information from entire population covered by your registry.
them. ✦ When? A particular year, time period.

However, no matter how well manipulated, the information derived from B. Boxhead: The boxhead contains the captions or column headings.
the raw data should be presented in an effective format, otherwise, it The heading of each column should contain as few words as
would be a great loss for both authors and readers. Planning how the data possible, yet explain exactly what the data in the columns represent.
will be presentedis essential before appropriately processing raw data. C. Stubs: The row captions are known as the stub. Items in the stub
should be grouped to facilitate interpretation of the data. For
THREE WAYS TO PRESENT DATA example, rows may stand for score of classes and columns for data
Presentation of data refers to an exhibition or putting up data in an related to sex of students. In the process, there will be many rows for
attractive and useful manner such that it can be easily interpreted. scores classes but only two columns for male and female students.
D. Footnotes: Footnotes are given at the foot of the table for
The three main forms of presentation of data are the Textual explanation of any fact or information included in the table which
Presentation, Tabular Presentation and Graphical Presentation. needs some explanation. Thus, they are meant for explaining or
providing further details about the data that have not been covered in
TEXTUAL PRESENTATION OF DATA title, captions and stubs.
● All the data is presented in the form of text, phrases, or paragraphs. E. Sources of Data: We should also mention the source of information
● It involves enumerating characteristics, emphasizing significant from which data are taken. This may preferably include the name of
figures and identifying important features of data. the author, volume, page and the year of publication. This should
● Text is the principal method for explaining findings, outlining also state whether the data contained in the table is of ‘primary or
trends, and providing contextual information secondary’ nature.

Example: PARTS OF THE TABLE


A researcher is asked to present the performance of a section in the
statistics test. The following are the test scores:

Example:
The data presented in textual form would be like this:
cw = is the class width
nc = is the number of classes

Round this value up to a convenient number.

Remember!
Creating the classes for summarizing continuous data is an art form.
There is no such thing as the correct frequency distribution. However,
there can be less desirable frequency distributions. The larger the class
width, the fewer classes a frequency distribution will have.

Construction of Data Tables


FREQUENCY DISTRIBUTION TABLE
The title should be in accordance with the objective of study
A frequency distribution list each category of data and the number of
✦ Comparison
occurrences for each category of data. It may also indicate the proportion
✦ Alternative location of stubs
of samples in percentages
✦ Headings
✦ Footnote
✦ Size of columns
✦ Use of abbreviations
✦ Units (It may also include totals and percentages)

SIMPLE TABLE

COMPOUND TABLE

ON ORGANIZING QUANTITATIVE DATA IN A TABLE


Classes are categories into which data are grouped. When a data set
consists of a large number of different discrete data values or when a data
set consists of continuous data, we create classes by using intervals of
numbers. Make sure that the classes do not overlap. This is necessary to
avoid confusion as to which class a data value belongs. Also, make sure
that the class widths are equal for all classes.

Guidelines for Determining the Lower-Class Limit of the First Class


and Class Width

Choosing the Lower-Class Limit of the First Class:


● Choose the smallest observation in the data set or a convenient
number slightly lower than the smallest observation in the data set.

For example, the smallest observation is 10.2. A convenient lower class


limit of the first class is 10.

Determining the Class Width:


•Decide on the number of classes. Generally, there should be between 5
and 20 classes. The smaller the data set, the fewer classes you should
have.
•Determine the class width by computing:
BAD EXAMPLE OF A TABLE

● Useless Information – Don’t show decimals if they are not


needed.
● Poor Alignment – Make sure alignment makes sense.; Don’t
center numbers, always right justify; try to align decimal
points. Consider the appropriate placement of row titles.
● Difficult to Read – Use commas used when the number
exceeds a thousand.

You might also like