0% found this document useful (0 votes)
4 views9 pages

Intro to Statistics and Sampling Methods

stat notes

Uploaded by

ralph.coley.25
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views9 pages

Intro to Statistics and Sampling Methods

stat notes

Uploaded by

ralph.coley.25
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Chapter 1 Notes: Intro to Data

Section 1.1: What is Statistics?


Good Question! What do you think statistics is? Or, what “buzz words” come to mind when you think
about stats?

Statistics is the science of how to ___________, ________________, __________________, and


__________________ numerical and non-numerical information.

Statistics takes data from a ___________________ of individuals,


called a ________________and helps us to make general
_____________________________ about ______ individuals in
the overall group, called the _____________________________.

A value that describes a sample is called a ___________________, while values that describe a
population are called _____________________. Statistics may ______________ from sample to sample,
but parameters are _______________.

Common Parameters of Interest:

Corresponding Sample Statistics:

Statisticians collect ______________ to make ______________________ ___________________. Most


commonly, data can be ___________________________, but it doesn’t have to be. Data is _________
piece of information that describes a _______________ to a question of interest. It could be ________
type answers, or any _____________________ response.

What are common places/disciplines where stats are used to make informed decision?
Variable Classification
_______________________ are the focus of a statistical study. These are the aspects being __________

or __________________ from each individual. Variables can be________________ into two categories:

Quantitative Variable: Any _______________________ data for which doing operations such as
addition or subtraction, or that from which you can get a meaningful ___________________.

Examples:

Qualitative Variable: Data that _________________________ into groups by a specific


________________.

Examples:

PRACTICE

How important is music education in school (K–12)? The Harris Poll did an online survey of 2286 adults
(aged 18 and older) within the United States. Among the many questions, the survey asked if the
respondents agreed or disagreed with the statement, “Learning and habits from music education equip
people to be better team players in their careers.” In the most recent survey, 71% of the study
participants agreed with the statement.

1. Identify the individuals of the study and the variable.

2. Do the data comprise a sample? If so, what is the underlying population?

3. Is the variable qualitative or quantitative?

4. Identify a quantitative variable that might be of interest.

5. Is the proportion of respondents in the sample who agree with the statement regarding music
education and effect on careers a statistic or a parameter?
Levels of Measurement
The _________ variable being studied and how that variable is ________________ dictates the type of
___________ done on the data.

Qualitative Data is measured in two ways:

The _______________ level of measurement applies to data that consist of _________, labels, or

categories. There are no implied criteria by which the data can be ___________ from smallest to largest.

The _______________ level of measurement applies to data that can be _____________. However,

differences between data values either cannot be determined or are meaningless.

Quantitative Data is also measured in two ways:

The _____________ level of measurement applies to data that can be arranged in order. In addition,

_______________ between data values are meaningful.

The _____________ level of measurement applies to data that can be arranged in order. In addition,

__________ differences between data values and ___________ of data values are meaningful. Data at

the ratio level have a ________________.

PRACTICE

Identify the level of measurement for each variable below.

1. Taos, Acoma, Zuni, and Cochiti are the names of four Native American pueblos from the
population of names of all Native American pueblos in Arizona and New Mexico.

2. In a high school graduating class of 319 students, Tatum ranked 25th, Nia ranked 19th, Elias
ranked 10th, and Imani ranked 4th, where 1 is the highest rank.

3. Body temperatures (in degrees Celsius) of trout in the Yellowstone River.

4. Length of trout swimming in the Yellowstone River.

5. A senator’s age.

6. The years in which the senator was elected to the Senate are 2000, 2006, and 2012.

7. The senator’s total taxable income last year was $878,314.


Section 1.2: Random Sampling
As we have noted, statistics studies _______________ of observations called a _____________.

How the sample is ___________________ is very __________________.

Those _________________ that make up a _____________ should have the same characteristics as the

_______________________________________. Moreover, samples should be ___________________

of the overall population they come from.

When a sample is _______ ______________________ of the overall _____________________, then

your sampling design will suffer from ________________________________. Any generalizations to the

overall population will likely be ____________________.

Gettysburg Activity
• Select 10 words from the address that you think best represent the words of the Gettysburg
address, record the number of letters in each word, then calculate the average number of
letters of your words. Write the average on your Blue sticky note and stick it in the
appropriate “bin” on the graph on the board.

“Four score and seven years ago our fathers brought forth on this continent, a new nation, conceived in
Liberty, and dedicated to the proposition that all men are created equal.

Now we are engaged in a great civil war, testing whether that nation, or any nation so conceived and so
dedicated, can long endure. We are met on a great battlefield of that war. We have come to dedicate a
portion of that field, as a final resting place for those who here gave their lives that that nation might
live. It is altogether fitting and proper that we should do this.

But, in a larger sense, we can not dedicate—we can not consecrate—we can not hallow—this ground.
The brave men, living and dead, who struggled here, have consecrated it, far above our poor power to
add or detract. The world will little note, nor long remember what we say here, but it can never forget
what they did here. It is for us the living, rather, to be dedicated here to the unfinished work which they
who fought here have thus far so nobly advanced. It is rather for us to be here dedicated to the great
task remaining before us—that from these honored dead we take increased devotion to that cause for
which they gave the last full measure of devotion—that we here highly resolve that these dead shall not
have died in vain—that this nation, under God, shall have a new birth of freedom—and that government
of the people, by the people, for the people, shall not perish from the earth.”

• What would you guess the class average would be? Consider the “balance-point” of the stickys
on the board.

• How do the actual average and class average compare? Why do you think this is?

• What could you have done to get a more representative sample?


Sampling Techniques
The are various ways to incorporate _______________________ when selecting your sample.

Good Sampling Designs:

1. Simple Random Sampling-


A simple random sample of n measurements from a population is a subset of the population
selected in such a manner that ____________ sample of size n from the population has an
__________ chance of being selected. We can do this by numbering all the individuals in the
population and randomly select a set amount of them using a random number generator.

2. Stratified Sampling-
Groups or classes inside a population that share a common characteristic are called________.
In the method of stratified sampling, the population is divided into at least _____ distinct strata.
Then a ___________________________ of a certain size is drawn from each stratum.

3. Cluster Sampling-
In the method of cluster sampling, we begin by dividing the demographic area into sections,
called _________________. Then we _______________ select sections or clusters. _________
member of the cluster is included in the sample.

4. Systematic Sampling-
In this method, it is assumed that the elements of the population are arranged in some natural
sequential order. Then we select a (random) starting point and select every ________ element
for our sample.

5. Multistage Sampling-
Any ___________________ of the above methods.

Bad sampling designs:

6. Convenience Sampling-
No ___________________ is deployed, you simply ask anyone who is available to answer.

7. Voluntary Response Sampling-


In this method, a broad net is cast to get as many responses as possible in a short time (like a
mass email). Subjects are ________ ____________ selected. Only subjects who care to respond
are likely to.

Why should design #6 and #7 never be used?


PRACTICE

An important part of employee compensation is a benefits package, which might include health
insurance, life insurance, child care, vacation days, retirement plan, parental leave, bonuses, etc.
Suppose you want to conduct a survey of benefits packages available in private businesses in Hawaii.
You want a sample size of 100. Some sampling techniques are described below. Categorize each
technique as simple random sample, stratified sample, systematic sample, cluster sample,
or convenience sample.

1. Assign each business in the Island Business Directory a number, and then use a random-number
generator to select the businesses to be included in the sample.

2. Use postal ZIP Codes to divide the state into regions. Pick a random sample of 10 ZIP Code areas and
then include all the businesses in each selected ZIP Code area.

3. Send a team of five research assistants to Bishop Street in downtown Honolulu. Let each assistant select
a block, then they interview an employee from each business found on the block.

4. Use the Island Business Directory. Number all the businesses. Select a starting place at random, and
then use every 50th business listed until you have 100 businesses.

5. Group the businesses according to type: medical, shipping, retail, manufacturing, financial, construction,
restaurant, hotel, tourism, other. Then select a random sample of 10 businesses from each business
type.

Section 1.3: Observational and Experimental Studies


Consider the scenario below:

In a study conducted at Mission Viejo High School, in CA, researchers compared the GPA of students
who took music classes to those students who did not. The average GPA of students taking music classes
is 3.59, and of the non-music students is 2.91. Based on the difference the school is considering making
all student take a music class. Should the school proceed with the new policy?

This scenario is called an _____________________________ _________________. In that the

researchers have not ___________________ a specific ________________________ to the subjects of

the study, but they simply ____________________________ them.

PRO:

CON:
Types of observational Studies:
• Surveys

• Retrospective Studies: Researchers collect data on something that has __________________


occurred.

• Prospective Studies: Researchers identify participants in advance and collect data as events
______________.

o Additional PROs:

Possible to _______________ variables.

Can design the study to your _____________________.

o Additional CONs:
Can be _________________________.
Require very _______________________ samples.
Can be _____________________.

PRACTICE

In early 2007, many dogs and cats died of kidney failure. Should you conduct a survey, retrospective or
prospective study to find out why?

Randomized, Comparative Experiments


One way to get definitive evidence of a ______________________ and ____________________

relationship is by conducting a well-designed _____________________________. Experiments study

the relationship of _______ or __________ variables, called _________________.

In a 2 factor study (either observational or experimental) we have the ___________________________

(or independent) variable and the _______________________________ (or dependent) variable.

Experimenters look at how ______________________ in the explanatory variable affect the response
variable.

With experiments, researchers ___________ specific ______________________________ to individuals,

called the ____________________________________ units, in the study. To get valid results,

experiments must _________________________________________________ participants to each

_______________________________ ______________________. (Random Assignment)


PRACTICE

Identify the following as an observational study or experiment. Then, identify the explanatory and
response variables.

1. A public speaking teacher has developed a new lesson that she believes decreases student
anxiety in public speaking situations more than the old lesson. She designs a study to test if her
new lesson works better than the old lesson. Public speaking students are randomly assigned to
receive either the new or old lesson; their anxiety levels during a variety of public speaking
experiences are measured.

2. A researcher believes that the origin of the beans used to make a cup of coffee affects
hyperactivity. He wants to compare coffee from three different regions: Africa, South America,
and Mexico. He randomly gives a group of individuals one of the three coffees and measures the
hyperactivity they show.

3. A researcher finds 100 women age 30 of which 50 have been smoking a pack of cigarettes a day
for 10 years while the other 50 have been smoke free for 10 years. The researcher measures the
lung capacity for each of the 100 women to determine if there is a significant difference in lung
capacity.

Experimental Design
As when designing surveys, we must take care when designing experiments. Here are the 4 principles of
experimental design:

1. Control
 Make all conditions as ____________________ as possible for all treatment groups.

 Controls allow us to ______________________ the one thing that is being studied.

2. Randomize
 Equalizes the effects of variation that we _________________ control.

 Distributes the uncontrollable factors _________________.

3. Replicate

 Apply each treatment to _______________________ subjects. (Not just 1)

 ______________ the entire experiment on an entirely different population of


experimental units.

4. Block

 Group _____________ individuals together and _______________ within each of these


blocks.

 Blocking helps account for the __________________________ due to the difference


between blocks
Control Groups, Blocking, and Confounding
Control Group: a group involved in an experiment that ____________ __________ receive a

_______________________.

Blind Studies
It is sometimes important, especially when doing studies with _______________ subjects, to disguise
the control group. We do this by not letting the participants know which treatment they are receiving.
This is called a _______________-________________ study. This technique helps eliminate participant
bias and inflated responses.

But it isn’t just the subjects that should be blind! _______________ - _________________ studies make
both the subjects and researchers unaware of the treatment the subjects are receiving. How could this
help with more valid data?

Blocking
The process of blocking is a design that eliminates ___________________________ within each
treatment group. Grouping ________________ individuals together reduces _________________
among the nonsimlar people in unblocked groups.

Example:

Confounding Variables

Example: It is known that throughout the year, murder rates and ice cream sales are highly
positively correlated. That is, as murder rates rise, so does the sale of ice cream. There are three
possible explanations for this correlation:

Possibility #1: Murders cause people to purchase ice cream.

Possibility #2: Purchasing ice cream causes people to murder (or get murdered).

Possibility #3: There is a third variable—a confounding variable—which causes the increase in
both ice cream sales and murder rates.

Confounded Variables: Variables that are ____________being studied by the researcher that could
produce similar effects of the response variable.

Researchers should ______________________ think about any ______________________ that may

mar their experimental results, and introduce ____________________ and

__________________________________ to take are in _____________________ the

___________________ of these variables.

Common questions

Powered by AI

Confounding variables may impact experimental results by introducing an additional influence that can distort the relationship between the independent and dependent variables, leading to incorrect conclusions . For example, in an experiment assessing the correlation between a specific treatment and its effect, a confounding variable might independently affect the outcome, creating false associations . To mitigate these effects, researchers can employ strategies such as random assignment, which helps to evenly distribute potential confounding variables across treatment groups, and the use of control groups to compare and account for potential confounding influences . Additionally, implementing a double-blind design can prevent both participants and researchers from knowing the allocations, thereby reducing bias .

Qualitative variables classify individuals into groups or categories, often described non-numerically, such as gender, ethnicity, or type of car . They are typically measured at the nominal or ordinal level, where nominal data consist of names or categories without a specific order, and ordinal data can be ranked but differences between them are meaningless . Quantitative variables, on the other hand, are numerical and allow for mathematical operations such as addition or subtraction. They are measured at the interval or ratio levels, where interval data can be ordered and meaningful differences can be calculated, and ratio data have a true zero point, allowing both meaningful differences and ratios . These classifications help determine the appropriate statistical methods for analysis.

Observational studies involve observing and measuring specific characteristics without attempting to influence the variables of interest . The main advantage is that they are often easier and less costly to conduct, and they are suitable for ethical and practical reasons when manipulation is impossible. However, they can suffer from confounding variables and do not establish causal relationships . Experimental studies, on the other hand, involve controlling and manipulating the variables to observe effects, allowing researchers to establish cause-and-effect relationships with greater certainty . However, they are often more complex and costly to design and carry out, and they may face ethical constraints .

Sampling error refers to the discrepancy between a statistic obtained from a sample and the corresponding parameter from the entire population. It occurs because the sample is not an exact replica of the population, leading to variations in estimates . This error affects the reliability of statistical inferences as conclusions drawn from the sample may not accurately reflect the population. As a result, ensuring a representative sample and understanding the potential for sampling error are crucial for making valid inferences about the larger population from sample data .

To ensure a sample is representative, researchers can use probability sampling methods, such as simple random, stratified, cluster, or systematic sampling. These methods allow every individual or group in the population a known chance of being selected, reducing selection bias . Stratified sampling ensures representation from each subgroup if the population is heterogeneous, while cluster sampling can be effective while simplifying logistics when geographic diversity is a concern . A representative sample is critical for statistical analysis because it allows for accurate generalizations from the sample to the population, making conclusions drawn valid and reliable .

A well-designed experimental study includes the key components of control, randomization, replication, and blocking. Control is important to ensure that all conditions are as similar as possible for all treatment groups, allowing us to isolate the effect of the one thing being studied . Randomization helps to equalize the effects of variation that cannot be controlled and distributes uncontrollable factors evenly, ensuring that the results are unbiased . Replication involves applying each treatment to multiple subjects or repeating the entire experiment on different populations to confirm the results' reliability . Blocking helps account for differences between groups by grouping similar individuals and randomizing within blocks, which reduces variation due to the differences between blocks . Together, these components ensure the validity and reliability of the experimental results.

Convenience and voluntary response sampling are considered poor designs because they don't ensure a representative sample of the population. Convenience sampling involves selecting individuals who are readily available, which may not reflect the population's diversity . Voluntary response sampling relies on individuals who choose to participate, which often results in a biased sample as those with strong opinions are more likely to respond . These methods affect the validity of survey results by introducing bias and limiting the generalizability of the findings to the broader population, potentially skewing conclusions drawn from the data .

Systematic sampling involves selecting a random starting point within the population and then choosing every nth element, where n is determined by dividing the population size by the desired sample size . To ensure each element has an equal chance, start with a list ordered in a natural sequence and employ a random mechanism (like a random number generator) to select the initial participant . From this point, every nth item is selected, maintaining periodicity to cover the population evenly. This method provides simplicity and can yield representative samples, assuming the list order does not introduce systematic bias .

Stratified sampling is preferred when the population has distinct subgroups or strata that share a similar characteristic but may differ from each other. This method ensures that all subgroups are adequately represented in the sample. A researcher may prefer this method when the goal is to make comparisons between different subgroups with high precision . Stratified sampling is beneficial when the stratum variability is less than the overall population variability, leading to more precise estimates and insights about each subgroup and the population as a whole .

Levels of measurement—nominal, ordinal, interval, and ratio—affect the type of statistical analysis that can be appropriately applied to the data. Nominal level data, which involve categories without a ranking order, can be analyzed using frequencies and modes but not means or medians . Ordinal level data, which involve a rank order, allow for median calculations and non-parametric tests but lack equal intervals between rankings, making means inappropriate . Interval and ratio levels enable a full range of statistical analyses, including calculating means, variances, and applying parametric tests, with the ratio level allowing meaningful zero points for ratios . The appropriate measurement level ensures valid methods and conclusions in analysis.

You might also like