Types of error:
1. Sampling error: when a characteristic of a sample differs from that of the whole population. This
error is random and will occur even for samples that are well-chosen to avoid bias.
2. Measurement error: inaccuracies in measurement at the data collection stage. For example, when
we record a person’s height to the nearest centimeter, the recorded height is slightly different from
the person’s exact height.
3. Coverage error: when a sample does not truly reflect the population we are trying to find
information about. For example, suppose we intend to sample a bee population on an island. If we
collect data from 10 bees, we do not have enough data. If we collect data from one beehive only, it
might not be representative of all the bees on the island, as the sample might be a biased sample.
4. Non-response error: when a large number of people selected for a survey choose not to respond
to it. For example, customer satisfaction surveys tend to receive more responses from unsatisfied
customers, and online surveys are less likely to be completed by elderly people who are unfamiliar
with technology.
Sampling methods:
1. Simple random sampling: Each member of the population has the same chance of being selected
in the sample
2. Systematic sampling: Selecting members of the population at regular intervals
3. Convenience sampling: Selecting members because they are easier to select or more likely to
respond
4. Stratified sampling (randomly selected from each strata) or quota sampling (specfically selected
by experimenter): Dividing the population into subgroups then selecting members from the
subgroups (number of members from each strata should be proportional to the fraction of the total
number of members)
Measures of central tendency
● Mode: most frequent number
○ Gives the most usual value
○ Only takes common values into account
○ Not affected by extreme values
● Mean: sum of all values / number of values
○ Written as mu (𝜇) or as x-bar (x̄)
○ Takes all values into account
○ Affected by extreme values
● Median: middle number in ordered list
○ Halfway point of data
○ Only takes middle values into account
○ Not affected by extreme values
Mean - 8.2
Sum of all numbers
Sum of every number squared
Sample standard deviation
Population standard deviation
Number of data
Minimum number
First quartile
Median
Third quartile
Maximum value
Discrete data: concrete, countable numbers
Continuous data: measurable, can be within a range
Frequency: number of occurences of each data value
Relative frequency: frequency divded by the total number of values
Small range = more consistent
Range: the greatest value - the smallest value
Left-shifted data: mean > median
Right-shifted data: median > mean
Standard deviation: average distance between the mean and a value in the data set.
Larger standard deviations indicate greater spread in the data.
More divided data → bigger standard deviation
More concentrated data → smaller standard deviation
Add a number closer to the mean to reduce standard deviation
Approximation = find midpoint of each interval
Upper boundary = Q3 + 1.5 x IQR
Any data larger than the upper boundary is an outlier.
Lower boundary = Q1 - 1.5 x IQR
Any data smaller than the upper boundary is an outlier.
Addition/subtraction Multiplication/division
Mean Yes - adds/subtracts same value Yes - multiplies/divides by same value
Median Yes - adds/subtracts same value Yes - multiplies/divides by same value
Mode Yes - adds/subtracts same value Yes - multiplies/divides by same value
Percentiles Yes - adds/subtracts same value Yes - multiplies/divides by same value
Range No Yes - multiplies/divides by absolute value
IQR No Yes - multiplies/divides by absolute value
Standard deviation No Yes - multiplies/divides by absolute value
Variance No Yes - multiplies/divides by square of value
Skewness No No