Actual work
Statistics is the science of collecting, organizing, summarizing, and
analyzing information to draw conclusions or answer questions. In
addition, statistics is about providing a measure of confidence in any
conclusions.
The process -
1. Identify the research objective
2. Collect data
3. Describe data
4. Perform interference
Descriptive stats consists of organizing and summarizing, and
describing data through numbers, Inferential stats uses methods that
take the result from the sample, extend it to the full population and
measure the reliability of the result.
In a sample of 50 trees, 60% have a disease (Statistic, because its from
the sample.)
The mean weight of peaches at Lucky’s is 5.30OZ (Parameter, no
sample, full population, using numbers and averages.)
In a survey of voters, 77% plan to vote for the incumbent. (Statistics,
survey of voters indicates sample, not full population.)
Two students out of all 29 students in a class have red hair.
(Parameter, all)
Basically if it is a measurement, its quantitative, otherwise, its
qualitative.
The type of calculator you use (Qualitative, Non-measuring)
Political party affiliation (Qualitative, Non-measuring)
The weights of sumo wrestlers (Quantitative, Measuring)
The tuition for your classes (Quantitative, Measuring)
Discrete data usually arises from counting while continuous data
usually arises from measuring. In between things (.5,.2, makes it
continuous)
Number of correct answers on a quiz (Quantitative, Discrete,
counting)
Peoples attitude towards the government (Qualitative)
Percentage of a cars surface which is rusted (Quantitative,
Continuous, measuring a percentage)
Time to learn a song (Quantitative, Continuous, has in between
values)
Which kind of hot dogs have the most calories? State the individual,
population, and sample.
The individual is a hot dog, and the population is ALL hot dogs.
The sample is the specific hot dogs being studied.
Sample - 54 hot dogs on dot plot
Variables - Calories, type of hot dogs
Quantitative, and continuous because calories can be halved.
The world health organization is collecting the density of people per
square kilometer from 56 countries. They want to estimate the mean
density of people per square kilometer for the world.
Individual - A country
Population - All countries / The world
Sample - 56 countries
Variable - Density of people per square kilometer.
Quantitative - Continuous
A news reporter wanted to know what the opinion of her town was,
so she put out a poll on social media. - CONVENIENCE/VOLUNTARY
Every 20th member of the population is selected - SYSTEMATIC
You pull names from a hat - RANDOM
A construction engineer received 1000 concrete blocks, each
weighing approximately 50 pounds. The engineer takes 10 blocks off
the top of the pile to test their strength - CONVENIENCE (taking off
the top)
Obtain a list of patients who had surgery and divide the patients to
type of surgery, draw random. - STRATIFIED (Type of surgery)
State whether the study is observational or experimental
In 1993, Frances Rauscher And His Colleagues Designed A Study To
Test Whether Listening To Mozart Would Help Students Improve
Their Performance On A Spatial Reasoning Task. They Recruited 36
College Students To Participate. The Subjects Were Randomly
Assigned To Three Groups, With 12 Students Per Group. Subjects In
Group 1 Listened To A 10-Minute Selection From A Mozart Piece.
Group 2 Listened To A Relaxation Tape For 10 Minutes. Subjects In
Group 3 Sat In Silence For 10 Minutes. Each Subject Took A Pretest
On Spatial Reasoning Two Days Before And A Post-Test On Spatial
Reasoning Immediately After The 10-Minute Period. The Results
Seemed Surprising: Students Who Listened To Mozart Showed
Significantly Higher Gains In Their Scores On Spatial-Reasoning
Tasks Than Students In The Other Two Groups.
Experimental Study they’re directly changing the results through
playing mozart / silence / relaxation tape
The treatments are the mozart/silence/relaxation tape
The control group is the group listening to no music while studying
The response variable is the gains in scores
The explanatory variable is what they listened to
The research objective is to see how music affects test scores and
studying habits
Some confounding variables could be age, mental health, time of day,
hunger, sleep, music preference
The study is observational since they are not directly interfering with
the results.
Explanatory variable is the type of hot dog
Response variable is the calories
A Medical Researcher Wants To Determine Whether Exercising Can
Lower Blood Pressure. At A Health Fair, He Measures The Blood
Pressure Of 100 Individuals, And Interviews Them About Their
Exercise Habits. He Divides The Individuals Into Two Categories:
Those Whose Typical Level Of Exercise Is Low, And Those Whose
Level Of Exercise Is High.
This is an Observational study because they are just measuring, not
interfering with the results, and not assigning exercising.
A teacher wants to investigate the effect of room temperature on
memory. 100 students at the school volunteered to participate in an
experiment. The students were randomly assigned into two groups.
One group was sent to a room where the temperature was 70 degrees
and the other group was sent to a room that was 80 degrees.
Students were then given 3 minutes to memorize a list of ten
nonsense words. They then listened to recorded music for five
minutes, after which they were quizzed to determine how many of
the ten words they could recall.
This is an experimental study, and utilizes random assignment
through where they are assigned and which groups.
If the study was split between gender/age, it wouldn’t be
randomized, and it might be a biased study. This is called a
confounding variable
In An Experiment, Half Of The Participants Are Randomly Assigned
To Sleep On A New Type Of Mattress, While The Other Half Of The
Participants Sleep On A Standard Mattress. The Next Night, The
Participants Sleep On The Other Type Of Mattress. After The Second
Night Of Sleep, The Participants Are Asked Which Mattress Made
Them Feel More Rested In The Morning.
If A Greater Percent Stated That The New Mattress Made Them Feel
More Rested, Could We Conclude That The New Mattress Caused
Them To Feel More Well Rested?
Yes, because they are randomly assigned which deals with all the
confounding variables.
The following table lists the types of aircraft for the landings that
occurred during a day at a small airport. (“Single” refers to a single
engine and “twin” refers to a twin engine)
Twin 9
Single 10
Helicopter 2
Jet 4
Turboprop 5
Relative frequency = count of category / total count
Twin Relative frequency 9/30 = 0.3 (30%)
In decimal or fraction form, this is called the proportion.
To convert decimals into percentages, x100.
.3 x 100 =30%
10/30 =0.333 (rounded to the 3rd decimal place)
0.333 x 100= 33.3%
2/30 =0.067
0.067 x 100 = 6.7%
5/30 =0.167
0.167 x 100 = 16.7%
4/30 = 0.133
0.133 x 100 = 13.3%
Relative Frequency Distribution
Twin 0.3
Single 0.333
Helicopter 0.067
Jet 0.167
Turboprop 0.133
.3+.333+0.067+.167+.133 = 1 , A whole
Frequency Bar Graph
Keep levels consistent, dont have the graph going 2,5,10, keep even
intervals.
Proportion indicates decimals, so if you have a %, / 100
To answer an or question, add the two values together.
If we surveyed 500 hispanic people, how many would we expect to
have private insurance? (500)(.50) = 250
How many white people could we expect?
500(.75) = 375
Class width > (Max value - starting value or minimum value /
Number of classes)
( 24-14 )/ 5 = 2
A dot plot is a chart where the values are plotted on a line graph,
showing the frequency of each data value.
Qualitative Data Charts - Pie Chart, Bar Graph
Quantitative Data Charts - Dot plots, Histograms
0 2 10 2 5 3 6 2 2 4 1 5 4 15 3
4 5 3 2 3 4 2 6 0 0 0 8 3 3 6
2 4 2 0 4 6 4 7 2 0 1 2 6 5
To find the average, you add up all your values and divide by how
many values you have. (1.4+3.2+4.9+4.4+7.4= 19.24 … 19.23/5 = 3.846)
(2.4+2.4+1.4+1.4+4.8 =12.4 12.4 / 5 = Number of Frequen
2.48)
chocolate cy
chips
14 – 16 7
17 – 19 8
20 – 22 14
23 – 25 3
26 – 28 0
13+6+8+8+15+8+9+12+6+16 =101 / 10 = 10.1 readability level
Median - 6,6,8,8,15,8,9,12,6,16 8+9 = 17 / 2 = 8.5
Mean - All numbers in set / amount of numbers
To calculate the median, first arrange the data set in ascending
order. If there's an odd number of data points, the median is the
middle number. If there's an even number of data points, find
the average of the two middle numbers
Green has a greater dispersion than the blue graph, because it
spreads out further from the middle.
To tell which graph has a smaller or bigger standard deviation,
determine which graph has a standard deviation, look at the
spread or variability of the data. A wider spread of data points
indicates a higher standard deviation, meaning the data are
more dispersed from the mean. Conversely, a narrower spread,
with data points tightly clustered around the mean, signifies a
lower standard deviation. You can also look for the inflection
points on a bell curve, where the distance from the mean to the
inflection point represents one standard deviation.
X = A data value
M = Population mean
= Population standard deviation
Vocab
POPULATION - The entire group of people being studied
SAMPLE - A subset from a population, containing individuals that
are actually being observed.
INDIVIDUAL - A person or object that is a member of the
population.
PARAMETER - A number calculated from the population.
STATISTIC - A number calculated from the sample. Since you can
find samples, it is readily known. It is used to estimate the parameter
value.
VARIABLE - The characteristics of the individuals
QUALITATIVE - Classifies individuals based on an attribute or
characteristics
QUANTITATIVE - These provide numerical measures of
individuals. It can be counted or measured from the individual
DISCRETE data has values that can be listed or countable
CONTINUOUS data can take on any value in some
interval, height, weight, oz, measurements.
SIMPLE RANDOM SAMPLE - Every different possible sample of size
N has the same chance of being selected.
CONVENIENCE SAMPLING - Not drawn from a well defined
random method, results that are easy to get
STRATIFIED SAMPLING - Subdividing the population into at least
two different subgroups that share the same characteristics, then
draw a sample from each subgroup (or sample)
CLUSTER SAMPLING - Divide the population area into sections (or
clusters) then randomly select some of those clusters. Now choose all
members from selected clusters.
SYSTEMATIC SAMPLE - The population items are ordered and then
decided how frequently to sample items.
VOLUNTARY RESPONSE -
OBSERVATIONAL STUDY - The study examines individuals in the
sample but does NOT try to influence the response variable (Does not
establish cause and effect)
EXPERIMENTAL STUDY - A study in which the researchers
CONTROL one variable (Intended to establish cause and effect)
EXPLANATORY VARIABLE - An explanatory variable is what you
manipulate or observe changes in (Meds, Caffeine dose)
RESPONSE VARIABLE - A response variable is what changes as a
result (Pain level, Reaction times)
FACTOR - The variable whos effect on the response variable is to be
assessed by the experimenter
TREATMENT - Any combination of values of the factors (explanatory
variables)
CONTROL GROUP - A group that receives a baseline treatment that
can be used to compare to other treatments
PLACEBO - A fake drug or treatment that just looks like the “real”
treatments in a study.
BLINDING - Nondisclosure of the treatment an experimental unit is
receiving
EXPERIMENTAL UNITS - A person, object, or some other well
defined item upon which a treatment is applied (The individuals) (If
the experimental unit is a person, it is often called a subject.
RANDOMIZED = RANDOM ASSIGNMENT - Subjects are normally
assigned to different groups
CONFOUNDING VARIABLE - A confounding or lurking variable is
another variable, besides the explanatory that affects the response
variable. When it is present, it is difficult for the investigators to
determine whether the differences in the outcome are due to the
treatment or the confounding variable.
FREQUENCY DISTRIBUTION - A mathematical function showing
the number of instances in which a variable takes each of its possible
values.
MEAN - The average of the data
MEDIAN - The middle of your data set
MODE - Most common number in data
STANDARD DEVIATION - A quantity calculated to indicate the
extent of deviation for a group as a whole.
HOMEOWRK
1.1
3 - Parameter, it shows the full population, not just a sample.
5 - Statistic, it only reports the results from the respondents.
43 - A. The research objective is to determine if non-smokers
have a higher IQ than smokers. B. The population being studied
is Israeli military recruits, and the sample is 20211 Israeli military
recruits. C. The average IQ of smokers was 94, while non-
smokers were 101. D. The conclusions of studies
45 - A. The research objective is to figure out the amount of
Americans that believe the government is wasting over 51 cent
for every dollar spent. B. The population is Americans 18 or
older. C. The sample is 1017 people surveyed. D. 35% think that
the government wastes 51 cents of a dollar with a margin of error
of 4% and a 95% level of confidence. E. From this survey, we can
infer that 35% of Americans think that the government wastes 51
cents of each dollar.
1.3
5 - I had The Scarlet Letter, The Jungle, and Crime and
Punishment. To get this random sample, I shuffled the names
onto a website, and took the first 3 names that the wheel chose.
1.4
29 - To get a simple random sample, you could take each class
and submit the entire roster to a random number generator 1-
1280 until you get 128 participants. (10% of 1280 is 128) To get a
stratified sample, you can take each section of the class, and get
4 from each class (128/32). To get a cluster sample, you could take
each class and randomly select 4 clusters. I think a stratified
sample is the best in this situation, because you get accurate,
equal, and full representation of each class.
1.2
9 - I think this study is observational because they are not
directly interfering or altering the results.
11 - I think this study is experimental because they are directly
teaching the math and interfering with the study and results.
19 - A. This is an observational study because it does not do
anything to the individuals participating. This is a cross sectional
study. B. The response variable is the BMI, and the explanatory
variable is the tv in the bedroom. C. Lurking variables might be
the environment of the individuals, what they’re eating, how
much they exercise. D. This means that after they made sure that
economic status can’t be a lurking variable, the answers stayed
the same. This can help us conclude that socioeconomic status
did not affect the studies results. E. We cannot conclude that a tv
in the bedroom causes a high bmi, since it isn’t specified that the
participants were randomly selected.
1.6
9 - A. The response variable is the achievement test. B. There is a
fixed grade level and fixed district. C. The treatments are the
new and old version of teaching the math. D. They assign the
students to 2 groups. E. The control group is the group learning
normal math. F. This kind of study is a True-Experiment
because of its random assignment. G. 500 first grade students. H.
Students may be hungrier in the afternoon, or more tired in the
mornings, which may affect their ability to focus. To fix this, you
could conduct the experiment at the same time (seperated).
2.1
11 - A. The proportion of 18-34 year old respondents that are
most likely to buy when a product is tagged with “Made in
America” is around .40, according to the graph. The proportion
of 35-44 year old respondents more likely to buy is around .6 .
B. The age group that is most likely to buy when made in
America is 55+ (≈.75) C. The age group least likely to buy when
made in America is 18-34. D. The association between higher age
and buying American products is likely correlated with older
ages wanting to buy American products, while younger
generations don’t prioritize where the product comes from.
23- A. B. If I owned the instacart franchise, I would make sure I
had the most drivers on Saturday, to account for more people
ordering on those days.
Monday 5 .12
Tuesday 7 .17
Wednesday 6 .14
Thursday 3 .07
Friday 4 .01
Saturday 9 .21
Sunday 8 .20
2.2
11 - A. 200. B. CW = (160-60) / 11 = 9.091. CW = 10 C. 60 - 69 (2) ,
70 - 79 (3) , 80 - 89 (13) , 90 - 99 (42) , 100 - 109 (58) , 110 - 119 (40) ,
120 - 129 (31) , 130 - 139 (8) , 140 - 149 (2) , 150 - 159 (1). D. The class
with the highest frequency is 100-109. E. The class with the
lowest frequency is 150-159. F. 8+2+1 = 11 / 200 =0.055 x 100 =
5.5%. G. No
31 - A. 0.0 is our lowest value, so we’ll make that our lower class
limit, with a width of 4. (23.6 - 0) / 7 (number of classes) = 3.7. CW
= 4.
0-3 10 .25
4-7 10 .25
8-11 6 .15
12-15 2 .05
15-18 5 .25
19-22 4 .1
22-25 3 .075
B. C. The graph is slightly skewed right.
3.1
17- A. The mean is greater than the median because it is skewed
right. B. The mean and median are equal, since it’s bell shaped.
C. The median is greater than the mean because it is skewed left.
19- A. The mean for the flipped classroom is 77.5, while the
traditional classroom’s mean is 71.81. The median for the flipped
classroom is 76.8, while the traditional classroom’s mean is 70.8.
This reflects that test scores are often higher than the traditional
learning environment. B. If the traditional classroom gets a
typo-ed high score, it changes the mean to 113.21, but the median
stays at a similar number, showcasing how if data is skewed
higher in certain parts, the median is often the better way to
measure data.
3.2
16- I think Column I is histogram A, because it peaks at 53, and
with the median and mean being equal, it would be symmetric,
like the graph. I think Column II is histogram 60, because it
peaks around 60, and with the data, it would be symmetric. I
think Column III is histogram B, and I think column IV is the last
histogram, because the standard deviation is so wide and
inconsistent.
27- I would buy car 2, because on average, its mileage lasts
longer, having a mean of 237.2, while car 1 has a mean of 223.50.
but, its spread is bigger. (21.8 vs 49.06) meaning it can be less
consistent, and show variability. If the goal is to have a higher
mileage and variability, I would choose car 2, but if the goal was
to be consistent, I would pick car 1.
project
My topic choice is how many people feel music affects their mental health.
The question I plan to answer is does music seriously make a change in
people's feelings.
“If you listen to music when you're upset, how does it affect your mood?”
, through a google form that I will advertise using my teachers help, and
flyers. My project will be analyzing if there is a relationship between two
quantitative variables.
The way I'm going to gather my data is by getting two bags of skittles, and
counting the skittles by colour.
My population is going to be all skittles bags, and my sample will be 2
skittle bags.
The variable will be the colour.
This is quantitative and discrete, since we’re counting, but we can't have
half intervals.
My estimate is that there will be more red skittles, since I feel like I have
the most red skittles whenever I eat them.