This chapter is intended to quickly familiarize the student with the ideas of uncertainty,
randomness, and variation. This is done by introducing the concepts of a random variable
to model random phenomena without the use of formal probability. The immediate
emphasis is on data gathering (sampling techniques), which includes bias, confounding,
blocking, and randomness.
Lecture 1
Any problem that has a fixed result is called deterministic.
What was the average temperature for August 30, 2020?
𝑥𝑥 2
Let 𝑓𝑓(𝑥𝑥) = + 3𝑥𝑥 − 1. 𝐹𝐹𝐹𝐹𝐹𝐹𝐹𝐹 𝑓𝑓(1).
3
Any problem that has an unpredictable result is called random or uncertain. In a random
problem, the wanted information inside a random problem is called a random variable.
The reason the wanted information is called a random variable because of the uncertainty
of obtaining the wanted information.
For example, suppose you toss a coin ten times, and you want to know the number of
heads for this experiment.
The number of heads is the random variable for this case. If we let X be the number of
heads, the possible outcomes are, X = 0, 1, 2, 3, …, 10.
This range from 0 to 10 is called the possible outcomes.
What is the average temperature between 1:30 PM and 2:00 PM for June 17, 2022? The
answer to this question is uncertain; therefore, the random variable is the average
temperature between 1:30PM and 2:00PM on June 17, 2022.
If we let X be the average temperature for this case, X = 88𝑜𝑜 𝐹𝐹, 86𝑜𝑜 𝐹𝐹, 900 𝐹𝐹, 𝑒𝑒𝑒𝑒𝑒𝑒. We
do not know which one will be the average temperature. This is why we call the average
temperature a random variable.
You try these two questions.
Classify each outcome as deterministic or uncertain. If the outcome is uncertain, describe
the random variable of interest and give reasons why it is uncertain.
1. The temperature for this Friday between 1:00PM and 3:00PM.
2. The radius of a circle with a circumference of 12 inches.
What is statistics?
Statistics is the science of collecting, organizing, displaying, summarizing, and
interpreting data in order to answer your statistical questions.
Today, statistics continue to be used to enhance our decisions. Specifically, statistical
techniques are used in two related ways - they assist us in seeing the world more clearly
and in thinking more clearly.
Vocabulary in statistics:
Data = wanted information for your statistical question
The first part of the statistical process for answering a statistical question involves design
– planning an investigative study to obtain relevant data for answering the statistical
question. The design often involves taking a sample from a population. The group to
which we wish to generalize is called the population. A population is finite when the
data could be physically listed. When the data is unlimited, the population is infinite.
If we have data only from a subset of the population, our data are called a sample. In
other words, a portion of the population selected to represent the population is called a
sample.
For example, a paper company is interested in estimating the proportion of trees in a 450-
acre forest with diameters exceeding 2 feet. The company selects 30 plots (100 feet by
100 feet) from the forest and utilizes the information from the plots to help estimate the
proportion for the whole forest. One acre is equal to 43,560 square foot.
The population for this case is the 450-acre forest, and the sample in this situation is the
30 plots (100 feet by 100 feet) from this forest.
You try this one:
Select 50 students currently enrolled in our campus and collect data for these two
variables: X = number of courses enrolled in and Y = total cost of textbooks and supplies
for courses.
What is the population for this case? Finite or Infinite?
What is the sample for this case?
A statistic is the calculated measure of some characteristics of a sample.
A parameter is the calculated measure of some characteristics of a population.
The difference between a population parameter and sample statistic is called the
sampling error.
For you to ponder: A Pharmaceutical firm is performing clinical trials on a new drug that
is intended relieve symptoms for COVID vaccine side effects. Two percent of the 500
clinical trial participants experienced headache and muscle pain as side effects.
1) What is the population being studied?
2) What is the sample being studied?
3) Based on the sample, what percentage of the population do you think would suffer
from headache and muscle pain?
4) Is the two percent a statistic or parameter in this case?
The field of statistics can be subdivided into two areas: descriptive statistics and
inferential statistics or statistical inference.
In other words, the two main branches of statistical methods are inferential and
descriptive. We are going to study both of these statistical methods plus elementary
probability.
Descriptive statistics: present, display, and describe sample data.
In a sample of 200 students in a university, 60 students, or 30%, live in the dormitories.
The 30% is an example of descriptive statistics.
Based on the sample, the school’s paper reported that “30% of all students at the
university lived in the dormitories.” This report is an example of inferential statistics.
Inferential statistics: interpret based on the descriptive statistics, make decisions and draw
conclusions about the population from which the sample was drawn.
You try this:
Which of the following is an example of descriptive statistics and which is an inferential
statistics?
1) A histogram depicting the age distribution of COVID-19 patients.
2) An estimate of the number of Bangkok residents who have contracted COVID-19.
3) A table summarizing the data collected in a sample of new-car buyers for 2020.
4) The proportion of mailed-out questionnaires that were returned.
In an observational study, the wanted results for your study are simply observed – you
don’t get to manipulate the explanatory variable.
Experimental study on the other hand, you can manipulate the explanatory variable or the
treatment. In other words, in an experimental study, treatments are given for the study
purpose, such as treating headache with 250 mg Tylenol pills.
What is the response variable in the study?
The response variable is the wanted result from the research question.
The explanatory variable is the treatment used to get the wanted result for the research
question.
Placebo: any "fake" treatment used on the subject is called placebo.
You want to study the percentage of men who wash their hands after using the bathroom.
This study is called an observational study. Why? No treatment has been imposed on the
subjects.
Suppose we want to know how many minutes one Tylenol tablet with 500 mg temporary
relief minor aches will.
This is an experimental study. Why? A treatment was imposed (giving the pill to the
subjects).
In this situation, our response variable is the number of minutes, and the explanatory
variable is the Tylenol tablet with 500 mg.
Confounding Factor: two variables are confounded when their effects on a response
variable cannot be distinguished from each other.
For example: On your most recent visit to your dentist, you suggest to run a study in
order to establish a cause and effect relationship between eating habits and number of
cavities.
Since your dentist has a large practice and has been in business for a long time, you ask
for permission to search through existing files in order establish a per patient cavity
count.
You follow this up by telephoning the patients, explaining what you are doing, and
asking simple questions about their eating habits. In particular, you are interested in the
number of fruit servings they have per week. While analyzing the data, you notice that if
you divide the patients into two groups (those who eat four or more fruit servings per
week and those who eat one or less fruit servings per week) there are large differences on
the number of cavities.
1) Is this an observational study or an experimental study?
Ans. Observational study. No treatment has been imposed on the subjects.
2) Can it be concluded that eating four or more fruit servings per week prevent cavities?
Ans. Of course not. There are many other factors not accounted for that contribute to
dental hygiene. For example, how about flossing habits?
3) There is serious confounding in this study. Please explain the confounding factor in
this case.
The confounding is due to ignoring other variables that are relevant to maintaining
healthy teeth, such as brushing, flossing, using a mouthwash, etc.
There are two types of data.
Quantitative data versus qualitative data (categorical data or attribute data): Qualitative
data are labels used to identify attributes of elements, and quantitative data are always
numeric.
Discrete data are data that can be considered as individual, separable data, which means
that we are able to count them with numbers – we can have 0, 3, 14, 26, and so on.
Continuous data are data that can be measured.
Quantitative data Qualitative data
Grade point average from school Yes or No
The height of a third grader Eye color
The weight of a baby College or University
The number of pairs of shoes Ethnicity
The number of ounces of water Marital status
Check your understanding of what you have learned so far:
A quality-control inspector selects assembled parts from an assembly line and records the
information concerning each part as: 1: defective or nondefective, 2: the employee
number of the individual who assembled the part, and 3: the weight of the part.
2. What is the population?
The population is all the assembled parts from the assembly line.
3. Is the population finite or infinite? Why?
The population is infinite. Because all assembled parts from the assembly line can’t
be physically listed.
4. What is the sample?
The sample is the selected parts.
5. What are the random variables?
The random variables are defective, nondefective, employee #, and weight of the
part.
6. Classify the three random variables as either qualitative or quantitative.
Qualitative: defective, nodefective, or employee #
Quantitative: weight of the part
Techniques of sampling:
1. Simple Random Sample (SRS)
2. Systematic Sampling
3. Cluster Sampling
4. Stratified Sampling
Now, how do we collect this data?
Sampling plays a key role in data collection. There are various methods of sampling
(your textbook lists some examples). Whatever the method, it is important that the
sample be random and representative of the population. Poor sampling can lead to bias
results.
Simple random sample: A set of data chosen from a population in such a way that each
member of the population has an equal probability of being selected.
Systematic sampling: a random sampling technique in which the researcher selects every
kth item or subject from the population.
Cluster sampling: a type of random sampling where the population is divided into
nonoverlapping areas or clusters and all the elements or subjects are randomly sampled
from the clusters.
Stratified sampling: a stratified sample is obtained by forming strata in the population,
and from each stratum, selecting the sample by using the simple random sample.
Identify the technique of sample that is used.
1. A retailer samples 15 receipts from the past week by numbering all the receipts,
generating 15 random numbers, and sampling the receipts that correspond to these
numbers.
2. A pollster walks around a busy shopping mall, and approaches shoppers passing
by to ask them how they often shop at the mall.
3. Police at a sobriety checkpoint pull over every third car to determine whether the
driver is sober.
4. The superintendent of Santa Clara County wants to test the effectiveness of a new
math program designed to improve quantitative skills among elementary school
children.
There are 75 elementary schools in the county.
The superintendent chooses a simple random sample of 8 schools, and institutes the new
math program in those selected schools. A total of 5200 children attend these 8 schools.
5. A cell phone company wants to draw a sample of 500 customers to gather
opinions about potential new features on upcoming new phone model. The
company draws a random sample of 150 from customers with Blackberry phones,
a random sample of 150 from customers with LG phones, a random sample of
100 from customers with Samsung phones, and a random sample 100 from
customers with Google phones.
Problems for you to ponder:
6. A health magazine presented results of a recent study that analyzed data collected
by the U.S. Census Bureau in 2000. Results reveal that both men and women in
the United States, heart disease remains the number one killer, victimizing
500,000 people annually. Age, obesity, and inactivity all contribute to heart
disease, and all three factors vary considerably from one location to the next. The
highest mortality rates (deaths per 100,000 people) were in New York, Florida,
Oklahoma, and Arkansas, where the lowest were reported in Alaska, Utah,
Colorado and New Mexico.
a) What is the population?
b) What are the random variables?
c. What is the parameter?
d. Classify all the random variables of the study as either categorical or quantitative.
2. Determine the level of measurement (nominal, ordinal, interval, or ratio) for each
of the following situations.
Scales of Measurement
Since statistics deals with data and data are the result of measurement, we need to discuss
one of the most important schemes for classifying a variable involves its scale of
measurement. Researchers generally discuss four different scales of measurement:
nominal, ordinal, ratio, and interval. Before analyzing a data set, it is important to
determine which scales of measurement were used, because certain types of statistical
procedures require certain scales of measurement. For example, a t test is appropriate
when your analysis involves a single criterion variable that is measured on an interval or
ratio scale.
Nominal scales: A nominal scale is a classification system that places people, objects, or
other entities into mutually exclusive categories. The examples for a nominal scale are
blood type, sex, political party, social security numbers, and races.
Ordinal scales: Values on an ordinal scale represent the rank order of subjects with
respect to the variable being assessed. Ordinal data provides information about relative
comparisons, but not the magnitudes of the differences. You cannot use descriptive
statistics on ordinal scale. For example, your grades from History and Statistics courses:
A and B. You cannot take the average of A and B grades.
Interval scales: With an interval scale, equal differences between scale values do have
equal quantitative meaning. For this reason, it can be seen that the interval scale provides
more quantitative information than the ordinal scale. A good example of an interval scale
is the Fahrenheit scale used to measure temperature. With the Fahrenheit scale, the
difference between 60 degrees and 65 degrees is equal to the difference between 70
degrees and 75 degrees: the units of measurement are equal throughout the full range of
the scale. The same thing for years 1980 and 1990 can be arranged in order, and the
difference of 10 years can be found. However, the interval scale also has an important
limitation: it does not have a true zero point. A true zero point means that a value of zero
on the scale represents zero quantity of the scale being measured.
Ratio scales: Ratio scales are similar to interval scales in that equal differences between
scale values have equal quantitative meaning. However, ratio scales also have a true zero
point which gives them an additional property. For example, the system of inches used
with a common ruler is an example of a ratio scale. There is a true zero point with this
system, in that zero inches does in fact indicate a complete absence of length. With this
scale, it is possible to make meaningful statement about ratio. It is appropriate to say that
that an object ten inches long is twice as long as an object five inches long.
Data scales:
Nominal = male, female, color, your name, social security number, religious affiliation,
and more
Ordinal = mild, moderate, severe, your rank, your grade, results of a horse race (arrived
1st, 2nd, etc.), and more
Interval = Fahrenheit, Celsius (not Kelvin system), and score on an IQ test
Ratio = weight, age, number of cell phones, your salary, area, volume, Kelvin, and more
What kind of scale is used in each situation?
a. The temperature (in degrees Fahrenheit) of patients with Covid-19.
b. The age at which the average male marries.
c. Client satisfaction survey responses: poor, average, good, and excellent.
d. The region of the U.S. in which an individual live: North, South, East or West.
e. The number of people with a Type A personality.
f. The local fast-food restaurant offers small, medium, and large soft drinks.
3. A group of 500 registered voters are randomly drawn from voter registration rolls
in the state of California and asked their opinion about Proposition 22, a highly
controversial proposed law allowing firms to hire drivers to be independent
contractors. It is found that 350 of these persons favor Proposition 22. In this
example,
a) what is the sample?
b) What is the population?
c) Is this an example of descriptive or inferential statistics?