Analytics Module 1
Analytics Module 1
I. INTRODUCTION TO SAMPLING
Universe
In research methodology, the universe is the set of all individuals, objects, or events that are
relevant to a study. The universe is the entire collection of units that are relevant to a study. It can
include people, groups, organizations, or even objects. The universe is important to define because
it helps determine the research question, which specifies what or who is of interest.
Example: In a study of football players in India, the universe would be the set of all football players
in India.
Population
The entire group of people, items, or observations that you want to study. The population is usually
larger than the sample. A population is the entire group that you want to draw conclusions about.
A sample is the specific group that you will collect data from. A population is statistically the
group on which information is being gathered and analyzed. A sample is a representative
selection of the population.
Definition
• A population is any group of individuals that have one or more characteristics in common
and that are in interest of researcher. (Best & Kahn,1995)
• A population refers to any collection of specified group of human beings or of non-human
entities such as objects, educational institutions, time units, geographical areas, prices of
commodities, salary of an individual etc. ( L. Kaul,2010)
Here are some examples of populations and samples:
➢ School population and sample: The population is all students in the school, and the sample
is the students in a specific grade.
➢ Hospital population and sample: The population is all patients in the hospital, and the
sample is the elderly patients.
➢ Employee performance: The population is a team or department, and the sample is a
random selection of employees from that group.
Types of population
When a population contains finite number of units or individuals are called Finite
Population. Example• Five hundred workers in a factory. • Three lakh students in CBSE board. •
Ten thousand books in the library.
The infinite population is also known as an uncountable population in which the counting
of units in the population is not possible. Example: A. The stars in the sky and the four-wheelers
in a town. B. The stars in the sky and the books in a library. C. The car in a town and the books in
a library. D. Number of points in a line and births of insects.
Sample
The term sample refers to a smaller, manageable version of a larger group. It is a subset
containing the characteristics of a larger population. Samples are used in statistical testing when
1
Dr. Akhila Ibrahim K. K.
Essential Statistics for Business Analytics Module 1
population sizes are too large to include all possible members or observations. A sample should
represent the population as a whole and not reflect any bias toward a specific attribute. A sample of
people or things is a number of them chosen out of a larger group and then used in tests or used to
provide information about the whole group.
Creating a sample is an efficient method of conducting research. Researching the whole
population is often impossible, costly, and time-consuming. Hence, examining the sample provides
insights the researcher can apply to the entire population.
For example, if a cell phone manufacturer wants to conduct a feature research study among
students in US Universities. An in-depth research study must be conducted if the researcher is
looking for features that the students use, features they would like to see, and the price they are
willing to pay.
• Sample is a representative part or a single item from a larger whole or group especially
when presented for inspection or shown as evidence of quality
• Sample is a finite part of a statistical population whose properties are studied to gain
information about the whole
Sampling Frame: A complete, accurate and updated list of all units of population required for
selecting a sample from given population is called Sampling Frame.
Parameter and Statistics
Parameter is numeric characteristic of a total or complete population. the value of
parameter is written in Greek letters (N). Example: Population mean and standard deviation
Statistic is a numerical characteristic of a sample drawn from its population. The value of
statistic is expressed in Roman Letters (n). Example: Sample mean and SD.
Sampling
Sampling is a process in statistical analysis in which researchers take a predetermined
number of observations from a larger population. It allows researchers to conduct studies about a
large group by using a small portion of the population. The sampling method depends on the type
of analysis being performed, but it may include simple random sampling or systematic sampling.
Sampling is commonly used in statistics, psychology research, and the financial industry.
It plays a crucial role in the financial sector, particularly in auditing and risk management.
Financial auditors, for instance, rely on audit sampling to evaluate a portion of transactions or
records rather than the entire population, which may be too time-consuming or costly to examine
fully. sampling is a process, which allows us to study a small group of people from the large group
to derive inferences that are likely to be applicable to all the people of the large group. Sampling
is a fundamental technique in research that involves selecting a smaller group, or sample, from
a larger population to represent the whole. Sampling allows researchers to
make generalizations about a population based on the analysis of a well-chosen sample.
Sampling Design
Sampling design is a plan for selecting a sample from a population. The goal of sampling design
is to ensure that the sample can be used to generalize findings to the entire population. Here are
some things to consider when developing a sampling design:
2
Dr. Akhila Ibrahim K. K.
Essential Statistics for Business Analytics Module 1
• Define the universe of study: This is the set of objects being studied, such as a city's
population, a warehouse's workers, or fans of a TV show.
• Consider the sampling unit: This could be geographical, social, or individual.
• Gather the sampling frame: This is the list of names from which the sample will be
drawn.
• Determine sample size: This can be calculated using an equation or a sample size
calculator.
• Factor in budgetary limitations: This can impact the size and type of sample, and may
even lead to using a non-probability sample.
Steps in sampling design
While developing a sampling design, the researcher must pay attention to the following points:
a. Type of universe: The first step in developing a sampling design is to clearly identify the
universe to be studied. The universe can be finite or infinite. In finite universe the number of items
is certain, for example number of workers in a factory, but in case of an infinite universe we do
not know the total number of items, for example listeners of a radio programme.
b. Sampling unit: A decision has to be taken concerning a sampling unit before selecting sample.
Sampling unit may be a geographical one such as state, district, village, etc., or a social unit such
as family, individual, school, etc. The researcher will have to decide one or more of such units that
he has to select for this study.
c. Source list: It is also known as sampling frame from which sample is to be drawn. It contains
the names of all items of a universe. If source list is not available, researcher has to prepare it. The
list should be comprehensive, reliable and appropriate. It is important to be as representative of
the population as possible.
d. Size of sample: This refers to the number of items to be selected from the universe to constitute
a sample. Choosing the number is a major challenge before a researcher. The size of sample should
neither be large, nor too small. The number should be optimum. While deciding the size of sample,
researchers must determine the desired precision as also an acceptable confidence level for the
estimate. The indicators of interest in a research study must be kept in view, while deciding the
size of the sample. Also, the costs factor also kept in mind while choosing the size of sample.
e. Indicators to study: In determining the sample design, one must consider the question of the
specific population parameters which are of interest. For instance, you may be interested in
estimating the proportion of persons with some characteristic in the population, or you may be
interested in knowing some average or the other measure concerning the population. There may
also be important sub-groups in the population about whom you would like to make estimates.
f. Financial constraints: Cost considerations, from practical point of view, have a major impact
upon decisions relating to not only the size of the sample but also to the type of sample.
g. Sampling procedure: Finally, the researcher must decide the type of sample you will use i.e.,
you must decide about the technique to be used in selecting the items for the sample. In fact, this
technique or procedure stands for the sample design itself. There are several sample designs out of
3
Dr. Akhila Ibrahim K. K.
Essential Statistics for Business Analytics Module 1
which you must choose one for the study. You must select the sampling design which fulfils the
sample size, within the cost and has a smaller sampling error.
Characteristics of a good sample design
A good sample design should have the following characteristics:
➢ Sample design must result in a truly representative sample
➢ Sample design must have small sampling error
➢ Sample design must be viable in the context of budget availability
➢ Sample design must be such so that systematic bias can be controlled
➢ Sample should be such that the results of the study can be generalized with a reasonable
level of confidence
Sample size
The sample size is a term used in research for defining the number of subjects included in a
sample. Sample size refers to the number of participants or observations included in a study. This
number is usually represented by n. The ideal sample should be large enough for adequate
representation of their population, needed for generalization and small enough as per availability,
time expense & complexity of data analysis.
According to Best & Kahn (1995) several practical observations about sample size are:
• Larger the sample, smaller the magnitude of sampling error.
• When sample groups are to be divided into their sub-groups for comparison, the
researcher should take large enough sample to get adequate sample size
• In mailed questionnaire studies a larger initial sample should be taken as the percentage
of responses may be as low as 20 to 30 %.
Determinants of optimum sample size
The optimal sample size for a study is determined by several factors, including:
• Population size: The size of the population you are studying
• Confidence level: How sure you want to be that your data is accurate
• Margin of error: How close you want your survey results to be to the actual population
value
• Standard deviation: A measure of how spread out a data set is from its mean
• Precision: How precise you want your estimates to be
• Variability: How different the population is likely to be
• Time and money: How much time and money you have for the study
• Purpose of the study: The purpose of the study
• Incidence rate: For niche samples, the percentage of the target population that meets your
criteria
Sampling techniques
Sampling is basically of two types – probability sampling and non-probability sampling.
1. Probability sampling/ Random sampling: Probability sampling is defined as a sampling
technique in which the researcher chooses samples from a larger population using a method based
on the theory of probability. The researcher sets a few criteria and chooses members of a population
4
Dr. Akhila Ibrahim K. K.
Essential Statistics for Business Analytics Module 1
randomly. This is done so that all the members have an equal opportunity to be a part of the sample
with this selection parameter.
2. Non-probability sampling/ Non-random sampling: It is a sampling technique where the
samples are chosen deliberately and not randomly. Non-probability sampling is defined as a
sampling technique in which the researcher selects samples based on the subjective judgment of
the researcher rather than random selection. It is a less stringent method. This sampling method
depends heavily on the expertise of the researchers. It is carried out by observation, and researchers
use it widely for qualitative research. Non-probability sampling is a sampling method in which not
all members of the population have an equal chance of participating in the study, unlike probability
sampling. Each member of the population has a known chance of being selected. Non-probability
sampling is most useful for exploratory studies like a pilot survey (deploying a survey to a smaller
sample compared to pre-determined sample size). Researchers use this method in studies where it
is impossible to draw random probability sampling due to time or cost considerations. Non-
probability sampling is also known as convenience sampling
Difference between Probability and non-probability sampling
The main difference between probability and non-probability sampling is that probability sampling
uses random selection, while non-probability sampling does not.
Here are some other differences between the two:
• Representativeness
Probability sampling ensures that each unit in a population has an equal chance of being selected,
so the sample is representative of the population. Non-probability sampling does not ensure that
all members of the population have an equal chance of being selected, so the sample may not be
representative.
• Statistical inferences
Probability sampling allows for strong statistical inferences about the whole group. Non-
probability sampling does not allow for strong statistical inferences.
• Data quality
Non-probability sampling may have lower data quality because the selection process is subjective.
• Ease of data collection
Non-probability sampling is easier to use because it's based on convenience or other criteria.
• Suitability
Probability sampling is more suitable for studies that require statistical generalization.
Types of probability sampling
A. Simple random sampling or Unrestricted sampling
It occurs when elements are selected individually and directly from the population. Simple
random sampling as the name suggests, is an entirely random method of selecting the
sample. A simple random sample is a is a randomly selected subset of a population in which
each member of the subset has an equal probability of being chosen to be a part of a sample.
As such, a simple random sample is an unbiased surveying technique. It is one of the best
probability sampling techniques that helps in saving time and resource. It is a reliable
5
Dr. Akhila Ibrahim K. K.
Essential Statistics for Business Analytics Module 1
6
Dr. Akhila Ibrahim K. K.
Essential Statistics for Business Analytics Module 1
7
Dr. Akhila Ibrahim K. K.
Essential Statistics for Business Analytics Module 1
when needed. This method has a predetermined range which is why it is the least time-
consuming.
For example, for evaluating the marks in language subjects of all the students of
standard 6, every 5th student’s mark sheet is selected as a sample. Here, n = 5.
4. Multistage cluster Sampling: Multistage sampling is an extension of cluster sampling
in that, first, clusters are randomly selected and, second, sample units within the
selected clusters are randomly selected. In this design, random selection occurs at both
the cluster or group level and at the sample unit level. Multistage sampling also may
be useful when naturally occurring cluster sizes are rather large, resulting in reduced
precision when compared to the stratified random-sampling approach. In this event,
smaller clusters can be created and sampled. For example, rather than sampling all sixth
graders at a selected school, a secondary cluster, such as classes, could be utilized and
sampled. In this case, only students in sampled classes at sampled schools would be
included. This has the added advantage of reducing the cluster size, thereby enhancing
estimator precision. In multistage sampling, or multistage cluster sampling, you draw
a sample from a population using smaller and smaller groups (units) at each stage. It’s
often used to collect data from a large, geographically spread group of people in
national surveys.
5. Random route sampling: it is a probability sampling technique used in face-to-face
surveys when there is no complete list of households to contact. Interviewers are given
a starting location and instructions on how to walk randomly, such as which direction
to start, which side of the street to walk on, and which crossroads to take. Interviewers
select households based on these instructions until they reach the desired number of
respondents. Random route sampling is used to contact populations that don't have a
register. The goal is to create equal selection probabilities for all households. However,
studies have shown that random route samples can produce biased results. This is
because it's possible that the selected households have unequal probabilities of being
chosen.
Types of non-probability sampling
1. Accidental Sampling: Accidental sampling, also known as grab or opportunity sampling,
is a form of non-probability sampling that involves taking a population sample that is close
at hand, rather than carefully determined and obtained. For instance, a person who is
obtaining opinions for a political poll at a shopping mall by randomly selecting passers-by
is using a form of accidental sampling. Accidental samples are not as experimentally sound
as using random sampling and random assignment.
2. Judgement sampling: Also known as purposive sampling, this method involves selecting
samples based on the researcher's knowledge and credibility. This method is not scientific
and can be influenced by the researcher's preconceived notions. With this method, sampling
is done based on previous ideas of population composition and behavior. An expert with
knowledge of the population decides which units in the population should be sampled. In
8
Dr. Akhila Ibrahim K. K.
Essential Statistics for Business Analytics Module 1
other words, the expert purposely selects what is considered to be a representative sample.
Judgment sampling is subject to the researcher’s biases and is perhaps even more biased
than haphazard sampling. Since any preconceptions the researcher has are reflected in the
sample, large biases can be introduced if these preconceptions are inaccurate. However, it
can be useful in exploratory studies, for example in selecting members for focus groups or
in-depth interviews to test specific aspects of a questionnaire.
3. Convenience sampling: Also known as haphazard, grab, opportunity, or accidental
sampling, this method involves selecting participants based on their availability and
convenience. It's a low-cost way to quickly gather initial insights.
4. Quota sampling: In this method, subjects are selected based on quotas that represent
various demographics of a population. For example, if a company has 1,000 employees,
and 600 drive to work and 400 take the train, you might survey 60 drivers and 40 train-
riders to reflect the proportion seen in the company. This is one of the most common forms of
non-probability sampling. Sampling is done until a specific number of units (quotas) for various
subpopulations have been selected. Quota sampling is somewhat similar to stratified
sampling, which is probability sampling, in that similar units are grouped together.
However, it differs in how the units are selected. In probability sampling, the units are
selected randomly while in quota sampling a non-random method is used—it is usually left
up to the interviewer to decide who is sampled. Contacted units that are unwilling to
participate are simply replaced by units that are, in effect ignoring nonresponse bias.
Market researchers often use quota sampling (particularly for telephone surveys) instead
of stratified sampling to survey individuals with particular socio-economic profiles. This
is because compared with stratified sampling, quota sampling is relatively inexpensive and
easy to administer and has the desirable property of satisfying population proportions.
However, it disguises potentially significant selection bias.
5. Snowball sampling or network sampling: In this method, current study subjects recruit
additional subjects to the study. This method can be used to access a specific, hard-to-find
population. Suppose a researcher wishes to find rare individuals in the population, and
already knows of the existence of some of these individuals and how to contact them. One
approach is to contact those individuals and simply ask them if they know anyone like
themselves, then contact those people, etc. The sample grows like a snowball rolling down
a hill to hopefully include virtually everybody with that characteristic. Snowball sampling
is useful for rare or hard to reach populations such as people with disabilities, homeless
people, drug users, or other persons who may not belong to an organised group or such as
musicians, painters, or poets, not readily identified on a survey list frame. However, some
individuals or subgroups may have no chance of being sampled. In order to be able to
generalize the conclusion to the whole population, some assumptions, which are usually
not met, are required.
9
Dr. Akhila Ibrahim K. K.
Essential Statistics for Business Analytics Module 1
Causes
Some of the most common causes are:
1. Random Variation: It can occur due to chance. Because a sample is only a subset of the
population, it will naturally have some degree of variation due to the randomness of the
selection process. This random variation can cause the sample statistics to differ from the
population parameters.
2. Sampling Bias: Bias can occur if the sample does not represent the population. This bias
can occur due to non-response, self-selection, or convenience sampling.
3. Measurement Error: It can also arise from measurement error, which occurs when the
measurements or observations are not accurate or precise. Measurement error can be due
to the instrument that measures a variable, data collection methods, or the observer making
the measurements.
4. Non-response Bias: Non-response bias can occur when individuals selected for the sample
choose not to participate in the study. If the non-response is related to the variables of
interest, this can lead to biased estimates of population parameters.
10
Dr. Akhila Ibrahim K. K.
Essential Statistics for Business Analytics Module 1
5. Sampling Frame Errors: Sampling frame errors can occur when the list or frame used to
identify the population is incomplete or inaccurate
TYPES
• Biased errors
These errors occur when the sample selection is based on the investigator's personal bias or
prejudice. For example, if an investigator uses non-random sampling instead of simple random
sampling, the results will be biased.
• Unbiased errors
These errors occur due to chance, such as when the investigator has taken care to select the sample
but the sample still differs from the population due to individual differences.
Non-sampling errors
Can occur at any stage of a survey or census, and can be caused by human error. Examples of non-
sampling errors include biased survey questions, data entry errors, and non-responses. Non-
sampling errors can be reduced by using a larger sample size, careful research design, and good
data quality practices.
Both types of errors can affect the validity of research by leading to inaccurate or biased
conclusions. Researchers should try to minimize both types of error as much as possible.
Causes of non-sampling error
Non-sampling errors are errors in data collection that are unrelated to sampling. They can occur at
any stage of a survey or census, including planning, fieldwork, tabulation, and computation. Some
causes of non-sampling errors include:
• Incomplete sampling frame: Some members of the target population are not included in
the sample.
• Nonresponse: Some respondents do not provide data, or they provide inaccurate data.
• Faulty methods: Poor sampling techniques, biased survey questions, or inappropriate data
analysis can lead to non-sampling errors.
• Measurement errors: Errors can occur during measurement or with study tools.
• Data entry errors: Errors can occur during data entry, such as coding errors, editing errors,
or programming errors.
• Questionnaire problems: The content, wording, or layout of a questionnaire can make it
difficult to provide accurate responses.
To reduce non-sampling errors, you can:
• Use a validated questionnaire or measurement tool
• Ensure quality checks for data entry
• Use appropriate data analysis controls
• Test questionnaires on a sample of respondents before finalizing them
Central Limit Theorem
The central limit theorem says that the sampling distribution of the mean will always be normally
distributed, as long as the sample size is large enough. Regardless of whether the population has a
11
Dr. Akhila Ibrahim K. K.
Essential Statistics for Business Analytics Module 1
normal, Poisson, binomial, or any other distribution, the sampling distribution of the mean will be
normal. The central limit theorem (CLT) is a statistical theory that states that the mean of a
sample will be close to the mean of the population if the sample size is large enough. The
theorem also states that the distribution of sample means will approach a normal distribution
as the sample size increases.
Put another way, CLT is a statistical premise that, given a sufficiently large sample size from a
population with a finite level of variance, the mean of all sampled variables from the same
population will be approximately equal to the mean of the whole population. Furthermore, these
samples will approximate a normal distribution, with their variances being approximately equal to
the variance of the population as the sample size gets larger, according to the law of large numbers.
As a general rule, sample sizes of 30 or more are typically deemed sufficient for the CLT to hold,
meaning that the distribution of the sample means is fairly normally distributed. In addition, the
more samples one takes, the more the graphed results should take the shape of a normal
distribution.
12
Dr. Akhila Ibrahim K. K.