0% found this document useful (0 votes)
22 views80 pages

Geog 246 Notes

The document contains lecture notes for GEOG2460-F23, covering topics in statistics including descriptive statistics, probability, hypothesis testing, and regression analysis. It discusses the importance of sampling methods and types of data, providing examples of observational studies and experimental designs. Additionally, it outlines measures of central tendency such as mean, median, and mode, along with their applications in statistical analysis.

Uploaded by

kinywajunior
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
22 views80 pages

Geog 246 Notes

The document contains lecture notes for GEOG2460-F23, covering topics in statistics including descriptive statistics, probability, hypothesis testing, and regression analysis. It discusses the importance of sampling methods and types of data, providing examples of observational studies and experimental designs. Additionally, it outlines measures of central tendency such as mean, median, and mode, along with their applications in statistical analysis.

Uploaded by

kinywajunior
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

lOMoARcPSD|60615441

GEOG2460-F23 Lecture Notes


Contents
1 Introduction 2
2 Descriptive Statistics 10
3 Probability 30
4 Confidence Intervals for µ 42
5 Hypothesis Testing 55
6 ANOVA (one-way) 71
7 2 Tests 77
8 Regression 86
9 Logistic Regression 109

1
GEOG*2460 F23: Unit 1
Introduction
Descriptive Statistics: The use of numerical or graphical summaries of data to describe a group (of
people, things, etc.).
Inferential Statistics: The use of descriptive statistics of a sample to infer parameters of a population.

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 1. INTRODUCTION 3


Example: Is there a relationship between spanking and IQ?
From “Corporal Punishment by Mothers and Development of Children’s Cognitive Ability: A
Longitudinal Study of Two Nationally Representative Age Cohorts”. Journal of Aggression,
Maltreatment and Trauma. (2009)
Results: spanked children tended to score lower than those that weren’t.
Results from this study were reported as follows:
“Spanking lowers a child’s IQ” (Los Angeles Times)
“Do you spank? Studies indicate it could lower your kid’s IQ” (Houston Chronicle)
“Spanking can lower IQ” (NBC)
“Smacking hits kids’ IQ” (New Scientist)
Example: Is there a relationship between alcohol consumption and lung cancer?
from “Alcohol Consumption and Risk of Lung Cancer: The Framingham Study”, L. Djouss´e et al,
Journal of the National Cancer Institute, 2002.
Example: One-hundred students were surveyed after an exam. Their mark on the exam and their level
of co↵ee consumption were recorded. It was found that students who had lower co↵ee consumption
tended to score higher on the exam.
All three of these studies are examples of an observational study.
Study participants were observed as they are without any control.
Is there a causal relationship between urban vegetation and rates of Coronavirus transmission in the
United States?
from “Urban Vegetation Slows Down the Spread of Coronavirus Disease (COVID-19) in the United
States”, [Link]

The alternative is a randomized experiment in which di↵erent treatments are randomly assigned to
members of the sample. One major benefit of a randomized experiment is that we can try to account for
confounding variables.

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 1. INTRODUCTION 4


Example: Karels and Boonstra (2003) experimentally manipulated bark colour by painting sections of
white birch (a tree with light-coloured bark) with dark brown paint. Karel and Boonstra predicted that
the rate of temperature increase in the bark would be greater in the painted sections.
published as “Reducing Solar Heat Gain during Winter: The Role of White Bark in Northern
Deciduous Trees.”, Arctic Journal of North America, 2003.

Sampling
Sampling is the processes of taking a (usually random) subset of the population in order to estimate the
population’s parameters. It’s much easier to calculate the mean (and other statistics we will talk about
later in the course) on a sample than on a population. The mean of the sample is an example of a
sample statistic, which can be used to estimate the population mean (known as a population
parameter).
In statistical studies, a population refers to the collection of subjects (people, animals, things, events,
etc...) that are being studied. How would you define the population in each of these hypothetical
examples?
A survey on how first year students feel about their residences at the University of Guelph

An epidemiological study on the health of children under the age of 5 living in malaria-a↵ected
regions in West Africa
A study on earthquake magnitudes during the past 100 years along the Pacific Rim
A note on notation: In statistics texts and papers, a variable that is being measured (e.g., height in
cm; # of children per family, etc.) is commonly denoted as x. The population mean is often denoted by
the greek letter mu (µ), and the sample mean is denoted by ¯x. We’ll come across these later in the
course, so it’s important to keep the di↵erence between µ and ¯x in mind.

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 1. INTRODUCTION 5


Types of sampling schemes
Simple random sample. Randomly select a sample of size n from a population of size N. All possible
samples of size n have an equal inclusion probability.
Stratified random sample. First, we stratify an area based on some property (e.g., using a map), and
then draw random samples from each of the strata.
Systematic random sample. E.g., sample every nth individual, but randomize the starting point.
In each of these sampling schemes, we need to have access to each member of the population. In other
words, the inclusion probability for each invidual should never be zero. In some cases, this is simply
not possible, in which case other schemes are sometimes used:
Convenience sample. Only accessible or convenient individuals in a population are selected.
Volunteer sample. All individuals in a sample are “self-selected”.
How representative do you think these types of samples of the population from which they were
sampled?
Example: In World War II, researchers from the Center for Naval Analyses conducted a study of the
damage done to aircraft that had returned from misions. The results are shown in the graphic below.
Researchers concluded that armour should be added to the areas that su↵ered the most damage.

from The Analysis of Biological Data, Whitlock/Schluter, 2nd Ed., p.21 (2015)
Other sources of bias include: wording of questions in a survey inaccurate responses in a survey
missing data and many more...

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: Unit 2


Descriptive Statistics
Measurement Levels (Data Types)
Understanding the type of data you are working with is critical to choosing the techniques and methods you
use to analyze your data.

Categorical/Qualitative Numerical/Quantitative

Nominal Ordinal Interval Ratio


Nominal Data
Nominal data are data whose values have no quantitative meaning, and cannot be ordered in any meaningful
way.
10
Example: Researchers sampled 57 lakes in Sierra Nevada in order to analyze the distribution of chironomids
in lake surface sediment. Data on a number of di↵erent possible factors a↵ecting chironomid habitat were
obtained. One of the factors included in the study was the dominant vegetation type around each lake, which
could be one of: upper subalpine (US), subalpine (S), upper montane (UM), Je↵rey pine woodland (JP), or
pinyon pine-juniper woodland (PP). The data were as follows:
Type Frequency Relative Frequency

US 15 26.3
S 23 40.4
UM 8 14.0
JP 5 8.8
PP 6 10.5
from Elementary Statistics for Geographers, 3rd edition, Burt et al., 2009, p.41.
The data collected in this case are dominant vegetation classes, types or factors. The classes themselves
have no quantitative meaning, and we cannot order them without collecting other types of data (e.g.,
frequency in this case).
Bar charts are useful ways to display nominal data.

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 2. DESCRIPTIVE STATISTICS 7

A Pareto Chart is a type of bar chart that displays the boxes in order of decreasing frequency. Sometimes
an additional line is superimposed on the plot showing the cumulative frequency. These are often used in
quality control (e.g., types of defects present in a product), and can be used to quickly identify the most
common types of phenomena and how prevalent they are. For example, here we can say that subalpine
(S)and upper subalpine (US) are the dominant vegetation types in our dataset, representing over 65% of the
lakes in the study.

Ordinal Data
Ordinal data are similar to nominal data in that their values have no quantitative meaning. The di↵erence
between ordinal and nominal data is that ordinal data do follow a natural order.
Example: Data collected by Statistics Canada in 2011 shows the highest educational level attained by
Canadians. Summarizing slightly gives the following data:
Highest Educational Attainment Relative Frequency

No diploma, certificate or degree 12.7


High school diploma 23.2

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 2. DESCRIPTIVE STATISTICS 8

Trades certificate 12.0


College diploma 26.2
Bachelor’s degree 16.5
Advanced degree 9.4
from Statistics Canada, showing additional levels of division
[Link]

Interval Data
Interval data are quantitative data that follow an arbitrary scale. They lack a meaningful zero-point. We can
quantify the di↵erence between data points, but we cannot compute ratios.
Example: A sample of 39 healthy adults was obtained and their body temperature (in degrees Fahrenheit)
was measured at 8:00am.
The data are as follows:
96.2 96.6 97.0 97.2 97.3 97.4 97.4 97.4 97.4 97.4 97.5 97.6 97.7 97.8 98.0 98.0 98.0 98.0 98.1 98.2 98.2
98.2 98.2 98.4 98.4 98.6
98.7 98.7 98.8 98.8 98.8 98.8 98.9 98.9 99.0 99.0 99.0 99.2 99.4
from Elementary Statistics, 5th edition, Triola, 2014, p.22.
The data obtained from this study is an example of numerical/quantitative data, that is, data that can be
physically measured. Specifically, this is an example of interval data, that is, data for which the di↵erences
between values have meaning, but there is no natural zero value. We can measure di↵erences (in degrees
Fahrenheit) between test subjects, but we cannot make ratio comparisons (“twice as warm”, “half as cold”,
etc.).

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 2. DESCRIPTIVE STATISTICS 9


Ratio Data
Ratio data are identical to interval data except for the fact that ratio data have a meaningful zero-point.
Example: The percent of green light reflected by 1-year-old tree seedlings from several di↵erent species
were measured using a spectrophotometer. The reflectances (in %) are as follows:
8.74 10.22 8.34 8.27 15.38 36.67 10.42 8.37 12.77 12.48 8.45 7.94
37.21 12.12 15.25 12.77 10.62 14.35 11.03 17.24 6.79 7.86 9.33 14.21
44.74 38.71 14.04 12.07 9.31 7.54 8.73 16.22 10.13 13.51 13.33 9.69 Numerical data can be displayed in a
histogram:

Histograms are made by separating data points into bins that represent ranges of values, and then counting
the number of data points within each bin. Visualizing numerical data like this can help us to describe the
distribution of the data.

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 2. DESCRIPTIVE STATISTICS 10


Example: Time between earthquakes. The time (in days) between earthquakes from 1902 to 1977. An
earthquake was included in the list if it had a magnitude of at least 7.5 (on the Richter scale), or over 1000
people were killed.
The data obtained from this study is another example of ratio data, that is, data for which the di↵erences
between values have meaning, and there is a natural zero value.
9 30 33 36 38 40 40 44 46 76 82 83 92 99 121 129 139
145 150 157 194 203 209 220 246 263 280 294 304 319 328 335 365 375
384 402 434 436 454 460 556 562 567 584 599 638 667 695 710 721 735
736 759 780 832 840 887 937 1336 1354 1617 1901

data obtained from “A Handbook of Small Data Sets” by D.J. Hand et al., Chapman & Hall, 1994.
Measures of Central Tendency
Mean
Suppose we measured the height of all the students in a classroom, and we call each of the height
measurements x1, x2, and so on. The mean height (µ) of our “population” is calculated as the sum of all the
measured heights (x1, x2, ...) divided by the population size, N:

(2.1)

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 2. DESCRIPTIVE STATISTICS 11


Median
If we sort a list of numbers from lowest to highest, the median of that list is the number in the exact middle.
This is easy to find if the number of values is odd, e.g.:
1 2 2 4 5 7 9 10 15
If we have an even number of values, the median is the average of the middle 2 values:
1 2 2 4 5 7 9 10 15 17
Example: How does rudeness a↵ect work behaviour? 98 students in a management course were asked to
write down as many uses for a brick in 5 min.
45 of the students were assigned to a “rudeness condition” in which the facilitator berated students: this
group had a mean performance of 8.5 words.
The other 53 students were assigned to a “control group” in which the facilitator made no comments: this
group had a mean performance of 11.7 words.
Example from “Business Statistics”, 2nd Ed., Sharpe et al., Pearson (2014), p.396.
Based o↵ this sample, do you think people are, in general and on average, negatively a↵ected by rudeness?

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 2. DESCRIPTIVE STATISTICS 12

Mode
The mode is the value of the variable that occurs most often within a sample or a population. One advantage
of the mode is that it can be applied to categorical variables.
A bimodal distribution is a dataset that has two modes. For example, below is a histogram of final grades in
a course (these are made up, but based on an experience from my own undergrad studies!):

Example: 50 companies were surveyed to determine what percentage of their revenue was spent on Research
& Development. The data are as follows:
5.2 5.6 5.9 6.0 6.5 6.5 6.5 6.6 6.8 6.9 6.9 6.9 7.1 7.1 7.2 7.2 7.4 7.5 7.5 7.7 7.7 7.8 7.9 8.0 8.0 8.1 8.2 8.2
8.2 8.4 8.5 8.8 9.0 9.2 9.4 9.5 9.5 9.6 9.7 9.9 10.1 10.5

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 2. DESCRIPTIVE STATISTICS 13


10.5 10.6 11.1 11.3 11.7 13.2 13.5 13.5
Example from “Business Statistics”, 2nd Ed., Sharpe et al., Pearson (2014), p42.

Measures of dispersion
Range
Variance and Standard Deviation
Variance tells us something about how “spread out” a distribution is. In calculating the population variance (
2
), we can think of it in several steps:
1. calculate the “distance” between each measurement and the population mean: x µ

2. square that distance: (x µ)2

3. take the sum of all of the squared distances:


(x1 µ)2 + (x2 µ)2 + ... + (xn µ)2

4. divide by the population size:


Note: To calculate the sample variance (s2), the method is almost the same, but we have to use the sample
mean instead of the population mean, and we also subtract 1 from the sample size (n) in the denominator.
Example: Suppose we grow 10 potato plants in an experiment, and measure their stem heights (in cm) 100
days after planting and record the following measurements:
55.2 68.5 75.3 50.5 58.1 64.6 50.1 61.3 60.9 64.7
What is the sample variance?
In summary, the sample variance (s2) is calculated as:

(2.2)
We often take the square root of s2, otherwise known as the sample standard deviation, which is a bit more
interpretable than the variance:

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 2. DESCRIPTIVE STATISTICS 14

(2.3)

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 2. DESCRIPTIVE STATISTICS 15


Z-scores
When we know the mean and standard deviation of a population or a sample, it is often useful to determine
how close an individual observation or value is to the mean, with respect to the standard deviation.
Example: The mean score (in %) on a mid-term exam was 68%, with a standard deviation of 10%. If a
student scores 88% on the exam, how many “standard deviations” is she above the mean score?
This way of viewing an invidual measurement with respect to the mean and standard deviation is called a Z-
score. It is a way to “standardize” or “normalize” a variable. The general equation for a Z-score for a value x
in a population with mean µ and standard deviation is:
x µ
Z= (2.4)
Quantiles and the Inter-Quartile Range (IQR)
Quantiles are a useful way to analyze how data are distributed. To generate quantiles, we first need to sort
our dataset from lowest to highest. So in the case of the spectral reflectance data we looked at earlier, we
have:
6.79 7.54 7.86 7.94 8.27 8.34 8.37 8.45 8.73 8.74 9.31 9.33 9.69
10.13 10.22 10.42 10.62 11.03 12.07 12.12 12.48 12.77 12.77 13.33
13.51 14.04 14.21 14.35 15.25 15.38 16.22 17.24 36.67 37.21 38.71 44.74
Dividing these by increments of 25%, we generate quartiles, which is just a special type of quantile:
0% 25% 50% 75% 100%
6.7900 8.7375 11.5500 14.2450 44.7400
If we divide the data into 100 categories, we get percentiles, which are indicated above each quartile above.
A percentile literally means “This percent of the data are below this value”. E.g., you would interpret the
first number to mean “0% of the data are below 6.79”. Which summary statistic is this equal to?
The inter-quartile range (IQR) is defined as the distance between the 1st and 3rd quartiles, also known as the
25th and 75th percentiles:
IQR = P75% P25% (2.5)
Another useful way to plot numerical data that shows the IQR very clearly is to use a boxplot:

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 2. DESCRIPTIVE STATISTICS 16

All of the values inside the box fall within the IQR. The dark horizontal line in the middle is the median.
There are a few stray points at the top – these are called outliers.

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 2. DESCRIPTIVE STATISTICS 17


Central Tendency in Spatial Data: Points
Spatial data are so called because they have a spatial dimension. This is usually represented by some kind of
X and Y coordinate (e.g., longitude and latitude), which allows us to map the data.
Example: Fake town populations
X Y Population
3.3 4.3 34000
1.1 3.4 6500
5.5 1.2 8000
3.7 2.4 5000
1.1 1.1 1500
from Statistical Methods for Geography, 2nd Ed. Rogerson, 2006, p.45
Y

X
You can calculate the mean centre, by averaging the X coordinates and averaging the Y coordinates.
Is there a way to incorporate the attribute associated with each spatial location when computing the centre
mean? For example, the population of each city?
X Y Population weight
3.3 4.3 34000
1.1 3.4 6500
5.5 1.2 8000
3.7 2.4 5000
1.1 1.1 1500
By adjusting the size of the circle proportionally with the weight, we can visualize locations with larger data
values.

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 2. DESCRIPTIVE STATISTICS 18


Y

X
You can calculate the weighted mean centre, by using the value of the variable to weight each point
(basically, give points with larger values more importance in the formula).
Standard Distance

Example: Location of crimes reported in the city of Davis, California:


Data from [Link]

Areal Data
Another example of spatial data is areal data, in which a measurement is associated with a specific region.
Areal data are graphically represented in GIS as polygons.
Example: 2019 Canadian Federal Election

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 2. DESCRIPTIVE STATISTICS 19


Figures taken from: [Link]

Canadian
federal election

Figuretakenfrom:[Link] of the 2019 Canadian federal election

Descriptive Statistics of Areal Data: Location Quotient


A common statistic used in economic geography for describing areal data is the location quotient (LQ).

Example: A country is divided into four regions called “D”, “E”, “F” and “G”.
The employment in various sectors for each region is shown below:

D E F G Nation
Manufacturing 5 70 15 10 100
Services 40 220 60 80 400
Other 105 210 125 60 500

Total 150 500 200 150 1000


From Burt et al., Elementary Statistics for Geographers, 2009, p. 124

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 2. DESCRIPTIVE STATISTICS 20


What is the location quotient (LQ) for the Manufacturing sector in Region D? This will tell us something
about the relative share of the manufacturing sector in Region D compared to other regions.

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: Unit 3


Probability
The probability of an event is the chance or likelihood of the event occuring.
Probabilities range from 0 (impossible) to 1 (certain), and are sometimes expressed as percentages (0 to 100%).
Simplest definition is that the probability is an event is # of ways an event can happen# of total possible
outcomes .
Example: Rolling a fair six-sided die.
There are six possible outcomes {1,2,3,4,5,6}.
What is the probability of rolling a 1?
What is the probability of rolling an even number?
What is the probability of rolling a number greater than a 1?
Example: Rolling two six-sided dice simultaneously.
There are 36 possible outcomes
{1 1,1 2,1 3,1 4,1 5,1 6,2 1,2 2,2 3,...}.
What is the probability that you roll two 1’s?
What is the probability that at least one of the two die is a 1?
What is the probability of rolling a sum of 7?
30
Example: Plea bargaining. Data from a sample of 1028 subjects is included below.
plea guilty not guilty

sentenced to prison 392 58 not sentenced to prison 564 14


from Elementary Statistics, 5th edition, Triola, 2015, p.209
If you randomly choose one subject, what is the probability that...
...they had not been not sentenced to prison?
...they had been sentenced to prison?
...they had been sentenced to prison after entering a guilty plea?
...they had been sentenced to prison after entering a not-guilty plea?
Example: Yahtzee! What is the probability of...?

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 3. PROBABILITY 22

obtained from
[Link]
Example: You are dealt a hand of five cards from a fifty-two card deck.
What is the probability of...?
No value Pair 2 Pairs 3 of a kind Straight Flush Full house 4 of a kind
Probability 0.501 0.423 0.0475 0.0211 0.00392 0.00198 0.00144 0.000240

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 3. PROBABILITY 23


Example: Roll a single, fair, six-sided die. Let’s call the event that we get an odd number “A”. Let’s call the
event in which the outcome is less than or equal to 3 “B”.
What is P(A)?
What is P(B)?
If we know that A has occurred, what is P(B)?
This is referred to as a conditional probability. We could rephrase the above question to ask:
“What is the probability of A occurring conditional upon B?”
“Given B, what is the probability of A”
“What is P(A|B)?”
Sample Space: The collection, or “set”, of all the possible outcomes.
Conditional probability: The probability that A occurs given B occurred is denoted as: P(A|B)
Complement: The complement of an event A includes all events except for A: A¯
Union: For two events, A and B, the probability that either event occurs is denoted as: P(A [ B)
Intersect: For two events, A and B, the probability that both event occurs is denoted as: P(A \ B)
Example: A standard deck of cards consists of 13 “numbers” (2-10, Jack, Queen, King and Ace) and four suits
(hearts, diamonds, clubs and spades). If a card is drawn from a deck at random, what is the probability that the
card is a club or a King? What is the probability that the card is a King of clubs?
Example: The time (in days) between earthquakes from 1902 to 1977. An earthquake was included in the list if
it had a magnitude of at least 7.5 (on the Richter scale), or over 1000 people were killed.
data obtained from “A Handbook of Small Data Sets” by D.J. Hand et al., Chapman & Hall, 1994.
9 30 33 36 38 40 40 44 46 76 82 83 92 99 121
129 139 145 150 157 194 203 209 220 246 263 280 294 304 319
328 335 365 375 384 402 434 436 454 460 556 562 567 584 599
638 667 695 710 721 735 736 759 780 832 840 887 937 1336 1354
1617 1901

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 3. PROBABILITY 24


We can use histograms to determine the observed probabilities of a given variable (time between earthquakes in
this case) across a range of possible values. Often it is helpful to represent these as a continuous function (a
smooth curve) rather than discrete probabilities (bins or “steps”). This continuous function is called a
probability distribution function (or “pdf” for short). When a pdf is graphically displayed, it is sometimes
called a density curve.

There are many “flavours” of probability distribution functions. The time between earthquakes model is an
example of an exponential pdf, which can be very useful to model the time between discrete events that occur
at a constant rate.
Example: Calls received in a call centre during the afternoon (ie. 1pm to 4pm) happen at a constant rate:

What is the probability that the next phone call will happen:
within the next five minutes? within the next ten minutes?
between five and ten minutes from now? more than five minutes from now?

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 3. PROBABILITY 25


Example: Personal income levels in Ontario can be modelled using a log-normal
distribution

What is the probability that a randomly selected Ontario resident will have a personal income: less than
100,000? less than 150,000? greater than 50,000? between 100,000 and 150,000?
One of the most commonly encountered probability distributions is the normal distribution, which can be used
to model many real-world observations.
Normal distributions always have their characteristic “bell-curve” shape, but can di↵er based on where the peak
is located and how spread out the distribution is. As we saw in the previous Unit, these two characteristics can
be described by two parameters of the distribution: the mean (µ) and the standard deviation ( ).

When µ and are known (or can be estimated), we can follow these guidelines for normal distributions:
1. ⇡ 68% of the observations lie between µ and µ +

2. ⇡ 95% of the observations lie between µ 2 and µ + 2

3. ⇡ 99.7% of the observations lie between µ 3 and µ + 3

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 3. PROBABILITY 26

Example: It is known that the gestation period in cats is approximately normally distributed with a mean of 62
days and a standard deviation of 1.5 days.

What is the probability that a cat selected at random would have a gestation period of:
less than 59 days? greater than 2 s.d. from the mean? greater than 63.5 days?

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 3. PROBABILITY 27


R has a very handy function for finding the probability that a randomly sampled value will be less than a given
threshold X (if the mean and standard deviation are known, that is): pnorm(X, mean, sd)
Using the previous example, what is the probability that a cat selected at random would have a gestation period
of:
less than 61 days? greater than 64 days? between 60.5 and 63.5 days?
Example: If we know that the mean weight of male African Elephants is 11,000 pounds with a standard
deviation of 1600 pounds, would you say a 12,000 pound elephant is particularly heavy?

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: Unit 4


Confidence Intervals for µ
We ended the last unit by talking about a particular measured value (ie. the weight of a single
elephant). In particular, we were interested in whether this individual was particularly large compared
to the population distribution. We answered this question by estimating the probability that we would
have randomly selected an elephant from the population with a higher weight. If that probability was
very small, we might have concluded that the 12,000 pound elephant was indeed particularly large. If
not, we might have concluded that this elephant is fairly common among the population.
Example:
10 elephants were randomly sampled from a population. The mean weight of the sample was 12,000
pounds, while the mean weight for the population was 11,000 pounds. Is our sample particularly heavy
or light, on average?
42

Recall: inferential statistics involves the estimation of population parameters using sample statistics.

If we calculate the sample statistic X from a sample of size n, how do we know we are anywhere close
to the actual population mean (µ)?

What is the probability that X is close to µ?

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 4. CONFIDENCE INTERVALS FOR µ 29


The mean heights of Canadian men was taken to be µ=70 inches with a standard deviation of =2.9
inches. A random sample of 5 male students in GEOG*2460 was taken in 2019. Let’s see how they
each compared to the national distribution:

What is the sample mean?


It was easy to see whether individual men are tall or short because we can place them on a distribution.
In order to decide if our sample mean X is particularly tall or short, we need a distribution of all the
possible sample means.
Just like the population distribution shows all of the possible individual measured values, the sampling
distribution (of the sample mean) shows all

possible values of the sample mean X and how frequently they occur (or, their probabilities).
Imagine you carry out an experiment where you sample from a population, take

the sample mean (X), and repeat this many times over. After plotting a histogram of all the sample
means you obtained, you can find the sampling distribution of the sample mean.
[Link]
If the original population has any distribution and a mean µ and a standard deviation , then the
Central Limit Theorem tells us a few things about the

sampling distribution of X:
It will be more normal than the population distribution

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 4. CONFIDENCE INTERVALS FOR µ 30


It has the same mean as the population distribution (µx¯ = µ)
It has a smaller standard devation ( x¯ = pn) x¯ is also known as the Standard error of the sample
mean
Is our sample mean of Canadian men particularly tall or short?
Here’s what we know so far:

population mean, µ = 70 standard error of the sample mean;

Example: A sample of 9 female students in GEOG2460 was taken, and their mean height was
calculated to be 67.1 inches. What is the average height of 18-25 year old Canadian women, if we
know that the standard deviation is 2.7 inches?

Sampling
Population Sample

Descriptive
Statistics

Parameter Statistic
Inferential
Statistics

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 4. CONFIDENCE INTERVALS FOR µ 31


How can we use this single sample mean X to estimate µ?
Remember the sampling distribution (of the sample mean) from earlier. It stated that:

X is normally distributed with mean µ and standard deviation pn


Using some algebra, we could show that this is the same thing as saying the following:

1. X µ is normally distributed with mean 0 and standard deviation pn.

2. is normally distributed with mean 0 and standard deviation 1.


Recall from Unit 2: If you have a variable X with a mean of µ and a s.d. of , the Z-score of any
individual can be calculated as:
X µ
Z=

In the same way, if we have a sample with a particular sample mean X, and we know that the mean of
all the possible sample means is µ and the s.d. of all the

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 4. CONFIDENCE INTERVALS FOR µ 32

I want a “common” range of Z values.

There is a 90% probability that 1.64 < Z < 1.64.

There is a 90% probability that 1.


There is a 90% probability that 1.64pn < X µ < 1.64pn.

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 4. CONFIDENCE INTERVALS FOR µ 33

There is a 90% probability that X 1.64pn < µ < X + 1.64pn.


To use this formula to estimate µ, we need to know the value of .
In the case of women’s heights, we know = 2.7 inches.
Use this to find a “90% Confidence Interval” for µ.
Example: We want to estimate the true mean calorie content µ in a specific brand of frozen dinners.
This brand of frozen dinners has a true standard deviation of = 15 calories.
In the original study, the researchers purchased a random sample of 10 frozen dinners and found the
calorie content in each meal:
227 249 252 225 254 239 224 226 267 244

Calculate the sample mean, X

Use the X to find an 80% Confidence Interval for µ.

✓X Zpn,X + Zpn◆
What can we do to make the confidence Interval for µ narrower?
✓X Zpn,X + Zpn◆
Example: Nicotine withdrawal and time perception
A study conducted at Pennsylvania State University in 2003 investigated whether or not a person’s
ability to concentrate was impaired during nicotine withdrawal. A sample of 20 smokers engaged in a
24-hour withdrawal period and then were asked to estimate how much time had passed during a 45-
second period:
39 45 47 50 52 53 55 56 56 57 59 64 65 67 67 69 70 70 72 73
from Statistics: The Exploration and Analysis of Data, Peck, 2012, pp. 485.
Calculate a 95% Confidence Interval for the population mean time elapsed.
Confidence Interval for µ (when is not known).

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 4. CONFIDENCE INTERVALS FOR µ 34

Example: Nicotine withdrawal and time perception (continued)

Assumptions:
In order to use the T distribution described above, you need
1. a simple random sample
You want a sample that is representative of the population and, as such, you want to randomize your
sample as much as possible. If the sample is not randomized, you have to be very, very careful about
using the sample to infer information about the population.
2. from a normally distributed population
How do we check this and what are we looking for?

Practice Problem: Exposure to Cosmic Radiation in Flight Crews


A study to consider the e↵ect of exposure to cosmic radiation due to increasing altitude was conducted.
A random sample of 100 flight crew members was chosen. The sample had a mean cosmic radiation
dose of 219 mrems and standard deviation of 35 mrem.

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 4. CONFIDENCE INTERVALS FOR µ 35


data from Statistics: The Exploration and Analysis of Data, Peck, 2012, pp. 432.
Calculate a 99% Confidence Interval for the true mean radiation dosage.

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 4. CONFIDENCE INTERVALS FOR µ 36


Practice Problem: Phosphorous and algal blooms
Scientists have determined that if the concentration of phosphorous in a lake exceeds 0.025 mg/L, the
lake is likely to switch from a clear lake to a turbid (or cloudy) lake, which could lead to algal blooms
and dead zones for fish.
An environmental consultant collected 25 water samples from a local lake. The sample mean was 0.028
mg/L and the sample standard deviation was 0.01 mg/L.
Find a 95% Confidence Interval for the true mean phosphorous concentration.

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: Unit 5


Hypothesis Testing
Example: Educational testing in the UK
It has always been assumed that university students, on average, have a level 14 vocabulary. To test this
assumption, a sample of 54 university students agreed to take the WAIS-R subtest. The sample had a
mean score of 13.2 with a standard deviation of 2.21. from “A Handbook of Small Data Sets” by D.J.
Hand et al., Chapman & Hall, 1994.
This group of students is clearly, on average, performing below level 14.
Or are they? Is this group of students performing enough below level 14 for us to conclude that all
students are, on average, performing below level 14?
Step 1: Construct hypotheses for your question of interest.
Step 2: Assume that the null hypothesis is true and calculate a test statistic.
Recall: the T-statistic is similar to the Z-statistic, except we can use s instead of

55
Step 3: Use the sampling distribution (of the test statistic) to determine how extreme your test statistic
value is compared to all possible values of the test statistic.

Step 4: Make a formal conclusion based on your p-value.

p-value conclusion

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 5. HYPOTHESIS TESTING 38


p  0.01 very strong evidence against H0

0.01 < p  0.05 strong evidence against H0

0.05 < p  0.1 suggestive, but inconclusive evidence against H0

0.1 < p little evidence against H0

p-value conclusion
p  0.01 very strong evidence in support of HA

0.01 < p  0.05 strong evidence in support of HA

0.05 < p  0.1 suggestive, but inconclusive evidence in support of HA

0.1 < p little evidence in support of HA

Smaller p-values result from more negative test statistics.

How do the values of (X,s,n) a↵ect the value of T?


Example: Emissions from eco-engines
A car manufacturer has developed a new engine that they feel is environmentally friendly. In order to
legally claim on their advertising that their engine is an “eco-engine”, it must be shown to emit, on
average, less than 20 parts per million (ppm) of carbon. A sample of 10 engines is randomly selected and
the sample has a mean of 17.57 ppm of carbon with a standard deviation of 2.95.
data from Business Statistics, Sharpe/De Veaux/Velleman/Wright, Pearson, 2014.
Test whether or not the true mean engine emission is less than 20 ppm.

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 5. HYPOTHESIS TESTING 39


Example: Commuting Times
A commuter train claims that the time on the train between a specific station and the downtown terminal
is 23 minutes. A random sample of 30 trains in a single month were chosen and the mean travel time was
found to be 25 minutes with a standard deviation of 6 minutes.
in “Elementary Statistics for Geographers” (3rd ed.) by Burt et al. (page 340).
Is there any evidence that the true mean travel time is longer than advertised?

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 5. HYPOTHESIS TESTING 40


Example: Overgrazing
Overgrazing is a serious problem leading to desertification in many semi-arid environments. Researchers
in Jordan identified 14 sites where grazing could occur. At each of these sites they installed large
herbivore exclosures. After one year, they measured the net primary productivity (NPP) within the
exclosures and immediately outside the exclosures. data obtained from Dr. Ze’ev Gedalof (U Guelph)
Site Outside(O) Within (W)
A 0.7 1.8

B 1.1 0.8

C 1 0.5

D 0.9 0.7

E 0.4 1.4

F 0.6 1.7

G 0.7 0.9

H 1.3 2.6

I 0.5 1.9

J 0.8 2.1

K 0.9 0.0

L 0.4 1.9

M 0.6 1.2

N 0.8 0.2
Do the data provide any evidence to suggest that exclosures are associated, on average, with increased
NPP?
Example: What is a normal human body temperature?
You are taught in school that normal human body temperature is 98.6 degrees Fahrenheit. Is this claim
accurate?
in “The Analysis of Biological Data” (2nd ed.) by Whitlock/Schluter. (page 310).
Researchers took a sample of 25 healthy adults and found a sample mean of 98.524 degrees with a
standard deviation of 0.678 degrees.

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 5. HYPOTHESIS TESTING 41


A di↵erent study of a sample of 130 individuals found a mean of 98.25 degrees with a standard deviation
of 0.733 degrees.

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 5. HYPOTHESIS TESTING 42


All of these examples so far compared a single sample to a previously “known” mean (µ). Can we use
similar procedures for hypothesis testing to compare 2 samples with each other?
Example: In 1899, Herman Bumpus collected samples of house sparrows that either survived or died as
the result of a severe winter storm. He took many measurements of physical characteristics of the two
samples, but one measurement in particular was taken as evidence of natural selection in sparrows. The
humerus bone was measured, and the results (in inches) were as follows:
Sparrows that survived:
0.687, 0.703, 0.709, 0.715, 0.721, 0.723, 0.723, 0.726, 0.728, 0.728, 0.729, 0.730, 0.730, 0.733, 0.733,
0.735, 0.736, 0.739, 0.741, 0.741,
0.741, 0.743, 0.749, 0.751, 0.752, 0.752, 0.755, 0.756, 0.766, 0.767,
0.769, 0.770, 0.780
Sparrows that perished:
0.659, 0.689, 0.702, 0.703, 0.709, 0.713, 0.720, 0.720, 0.726, 0.726,
0.729, 0.731, 0.736, 0.737, 0.738, 0.738, 0.739, 0.743, 0.744, 0.745, 0.752, 0.752, 0.754, 0.765
data from Statistical Sleuth, Ramsay/Schafer, 2013
Population 1 (all birds that survived): µ1 = mean humerus length in pop. 1.
Population 2 (all birds that perished): µ2 = mean humerus length in pop. 2.
Is there any reason to believe that sparrows that survived, on average, have longer humerus lengths than
birds that perished?

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 5. HYPOTHESIS TESTING 43


SampleData:
Sample1( n 1 =35 survived): X 1 =0 . 738,s 1 =0 . 0198
Sample2( n 2 =24 perished): X 2 =0 . 728,s 2 =0 . 0235

Assumptions: We will be using T distributions again, so the same assumptions must be satisfied as in the
one-sample case, but now they must be satisfied for both populations. The two samples must also be
independent.

The test statistic T given below comes from a T distribution with degrees of freedom also given below
that we can use to obtain a p-value.

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 5. HYPOTHESIS TESTING 44

, with degrees of freedom .


The output from R is as follows
Welch Two Sample t-test
data: survived and perished t = 1.7207, df = 43.824, p-value = 0.04618
alternative hypothesis: true difference in means is greater than 0 95 percent confidence interval:
0.0002363793 Inf sample estimates:
mean of x mean of y
0.7380000 0.7279167
Alternate Formula (you might see in another course) for when it is reasonable to conclude that that 1

= 2.

, with degrees of freedom df = n1 + n2 2 Example: The Stare


as a Stimulus to Flight in Human Subjects:
Researchers either stared or did not stare at the drivers of automobiles stopped at a campus stop sign. The
researchers timed how long it took each driver to proceed from the stop sign to a mark on the other side
of the intersection.
from Mind on Statistics, Heckard, 2012, pp. 435-436 - originally Ellsworth et al., 1972 Do people being
stared at, on average, drive faster (in an attempt to flee)?
Population 1 (all people being stared at): µ1 = mean crossing time in pop. 1.
Population 2 (all people not being stared at): µ2 = mean crossing time in pop. 2.

Sample 1 (stare): Statistics: X1 = 5.59,s1 = 0.82.


Data: (n1 = 13): 4.5, 4.5, 4.8, 4.9, 5.0, 5.6, 5.7, 5.8, 5.8, 6.1, 6.3, 6.5, 7.2

Sample 2 (no-stare): Statistics: X2 = 6.63,s1 = 1.36.


Data: (n2 = 14): 4.7, 4.7, 5.2, 5.5, 5.7, 6.0, 6.5, 6.9, 7.1, 7.5, 7.8, 8.1, 8.3, 8.8 Assumptions:

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 5. HYPOTHESIS TESTING 45

Results
Welch Two Sample t-test
data: stare and nostare
t = -2.4122, df = 21.596, p-value = 0.01241
alternative hypothesis: true difference in means is less than 0 95 percent confidence interval:
-Inf -0.297971 sample estimates:
mean of x mean of y
5.592308 6.628571
Example: Researchers were interested in testing the di↵erences in recreation behaviour between urban
and suburban regions of a city - specifically, they focused on swimming frequencies. Is there a di↵erence
between the true mean swimming frequency of urban and suburban residents?
Population 1 (all urban residents): µ1 = mean swimming frequency in pop. 1.
Population 2 (all suburban residents): µ2 = mean swimming frequency in pop. 2.

Sample 1 (urban): Statistics: X1 = 48.63,s1 = 19.88.


Data: (n1 = 8): 20, 32, 38, 42, 50, 57, 70, 80

Sample 2 (suburban): Statistics: X2 = 63.63,s2 = 12.66. Data: (n2 = 14): 39, 58, 58, 62, 66, 73, 73, 80
data from Statistical Methods for Geography, Rogerson, 2006, p106.
Assumptions:

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 5. HYPOTHESIS TESTING 46


Results
Welch Two Sample t-test
data: urban and suburban
t = -1.8003, df = 11.876, p-value = 0.09725
alternative hypothesis: true difference in means is not equal to 0 95 percent confidence interval:
-33.175105 3.175105 sample estimates:
mean of x mean of y
48.625 63.625
Example: E↵ects of trickle irrigation of maize yields
A country is implementing a number of innovations to increase the e ciency of food production and
thereby reduce its reliance on food imports. You have been asked to determine if maize production under
a new trickle irrigation system is as high as the productivity under the old spray irrigation system. The
new system uses less water because of lower evaporation losses. A decision has been made to implement
this new system if it can be shown that maize yields are not lower than those under the old system. An
experimental design is developed using 25 fields for each system and the maize yields are measured in
each field.
Adapted from: Burt et al., Elementary Statistics for Geographers
> [Link](trickle, spray, alternative = "less") Welch Two Sample t-test
data: trickle and spray t = -0.60721, df = 32.254, p-value = 0.274
alternative hypothesis: true difference in means is less than 0 95 percent confidence interval:
-Inf 188.6074 sample estimates:
mean of x mean of y
985.0168 1090.4461
Suppose the decision to adopt the new trickle irrigation system would be made only if it were proven to
be better than the old spray system, in terms of maize yields?
Welch Two Sample t-test
data: trickle and spray t = -0.60721, df = 32.254, p-value = 0.726 alternative hypothesis: true difference in
means is greater than 0 95 percent confidence interval:
-399.4658 Inf sample estimates:
mean of x mean of y
985.0168 1090.4461
Example: E↵ect of tra c calming schemes

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 5. HYPOTHESIS TESTING 47


City planning o cials are interested in measuring the success of tra c calming measures on tra c
volumes. They choose a random sample of 20 locations and measured tra c volume before and after the
measures take e↵ect.
Location Before After Location Before After
1 1005 509 11 559 666
2 646 476 12 892 758
3 583 888 13 794 610
4 1064 433 14 979 582
5 817 993 15 757 722
6 703 665 16 1079 822
7 788 531 17 854 519
8 1025 540 18 902 494
9 546 964 19 767 402
10 527 867 20 953 728
data from Elementary Statistics for Geographers, Burt et al, 2007, page 364.
Is there any reason to believe that the measures are e↵ective?
Paired t-test
data: before and after t = 2.3168, df = 19, p-value = 0.01592
alternative hypothesis: true difference in means is greater than 0 95 percent confidence interval:
38.94663 Inf sample estimates:
mean of the differences 153.55

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: Unit 6


ANOVA (one-way)
So far, we have discussed how to test hypotheses using one sample (is there evidence in support of an
assumed population mean?) and two samples (is there evidence that the means of two populations are
equal?). How do we approach this problem if we want to compare three or more samples?
Example: Width of the joint of the first tarsus for genus Chaetocnema
from “A Handbook of Small Data Sets” by D.J. Hand et al., Chapman & Hall, 1994.
The length (in microns) of the width of the joint of the first tarsus is measured for three di↵erent
species in the genus Chaetocnema.
Is there a di↵erence between the mean tarsus width of the three species?
Population 1 (all of species 1): µ1 = mean tarsus width in pop. 1.
Population 2 (all of species 2): µ2 = mean tarsus width in pop. 2.
Population 3 (all of species 3): µ3 = mean tarsus width in pop. 3.
To determine if there is statistical evidence against the null hypothesis, we use a procedure called
ANOVA (ANalaysis Of Variance).
71
Assumptions: We will be using a di↵erent distribution than we used with 1- and 2-sample hypothesis
tests. The same assumptions must be satisfied, but now they must be satisfied for all populations. All
samples must also be independent and their population standard deviations must also be equal.

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 6. ANOVA (ONE-WAY) 49

The data are as follows:


Species 1: 160 163 171 173 174 185 186 188 191 200
Species 2: 184 186 199 201 208 211 211 217 223 242
Species 3: 122 125 130 131 132 135 138 146 151 158
To analyze the signal versus the noise in these three samples, we will use the sum of squares of the
residuals (recall variance in Unit 02). The first values we need to calculate are the “sum of squares
between groups” (SSB) and the “sum of squares within groups” (SSW):
For ANOVA, there are two di↵erent degrees of freedom that we are concerned with:
df1 = t 1; df2 = t(n 1)
where t is the number of “treatments” (or groups, or samples) and n is the size of each group.
Finally, we calculate the F-statistic, which is defined as the ratio between the signal and the noise.

That’s alot of calculations for such small samples! Software like R makes this easy. The output of an
ANOVA for the tarsus dataset looks like this:
Df Sum Sq Mean Sq F value Pr(>F) species 2 25780 12890 64.88 4.87e-11 ***
Residuals 27 5364 199

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 6. ANOVA (ONE-WAY) 50


---
Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 Example: Researchers are interested in the
relationship between crime rates and city size. They selected a random sample of 8 “small cities”, 8
“medium cities” and 8 “large cities” and recorded the number of assaults per month per 1000
population. The sample data are as follows.
Small cities: 21 27 31 58 63 64 74 93 Medium cities: 28 71 72 74 76 86 91 118 Large cities: 42 73 92
100 108 118 146 166
data from Elementary Statistics for Geographers, Burt et al, 2007, page 459.
Assumptions: qq-plot for each sample, box plot of all three samples

Df Sum Sq Mean Sq F value Pr(>F)


size 2 10753 5376 5.743 0.0102 *
Residuals 21 19659 936
---
Signif. codes: 0 *** 0.001 ** 0.01 * 0.05 . 0.1 1
If we decide to use a p-value threshold of 5% to reject H0, then our p-value of 0.01 allows us to reject
the null hypothesis in favour of the alternative hypothesis that at least one of the city size categories
has a statistically di↵erent mean number of assaults per 1000 people per month. But how do we know
which mean(s) is/are di↵erent from the others? One option is to test each pair using a two-sample t-test
we learned in the previous unit:
(1) small vs. medium (2) medium vs. large (3) small vs. large
But when we perform multiple t-tests one after another like this, we can run into a serious problem! If
we reject the null hypothesis with a p-value of 0.05, there is still a 5% chance that the null hypothesis
was actually true. This is called a Type I error. Maybe 5% is an acceptable risk to take, but what is the
probability of at least one Type I error if we perform 3 t-tests in a row?
For our 3-group example, it would be 1 0.95 ⇥ 0.95 ⇥ 0.95 = 0.14. This gets worse as we analyze
more groups (and therefore more pairs):
P(at least one Type I error) = 1 (0.95)#pairs

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 6. ANOVA (ONE-WAY) 51


Pairwise comparisons using t tests with pooled SD data: crime and size
large medium
medium 0.2260 small 0.0084 0.4366
P value adjustment method: bonferroni

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: Unit 7

2
Tests

Recall our data types (“measurement levels”) from Unit 02:

Categorical/Qualitative Numerical/Quantitative

Nominal Ordinal Interval Ratio


When we carried out one-sample hypothesis tests in Unit 5, we were testing assumptions about
distributions of certain populations by testing some assumed population mean (µ) using a sample with
mean ¯x and standard deviation s.
How can we test assumptions about distributions of categorical data?
77
Example: Suppose it is commonly believed that around a quarter of Americans believe that winning the
lottery is the most practical way to gain wealth over a lifetime. A survey conducted in 2006 asked
Americans what they believe to be the most practical way to accumulate 200000 in net wealth in their
lifetimes. A sample of 1000 people was randomly chosen and 210 of them indicated that playing the
lottery would be the most practical way to obtain financial wealth.
Is there any statistical reason to believe that the percentage of the American population that believes
that the lottery is their most practical option di↵ers from 25%?
Data drawn from Mind on Statistics, Heckard, 2012, pp. 435-436

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 7. 2


TESTS 53
To test beliefs or assumptions about how categorical data are distributed, we can perform a 2 Test. The 2
statistic is defined as:

where O is the observed value (e.g., frequency of a given category) and E is the expected value (e.g.,
2
calculated using an assumed probability). follows a distribution that looks similar to the F-
distribution, but it is controlled by a single parameter - the degrees of freedom (df).

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 7. 2


TESTS 54
Example: In Pokemon Go, one of the goals is to collect wild Pokemon. You run around until a wild
Pokemon appears. In each portion of the game, the likelihood of obtaining specific species of Pokemon
is di↵erent. There are a number of websites that publish the theoretical probabilities of obtaining
certain species. In one such published website in 2017, one of the routes had the following theoretical
probabilities: Sandshrew (40%), Trapinch (20%), Cacnea (20%), and Baltoy (20%).
Prof. Skelton (who taught this course in 2017) asked a former student of his to go to that route and
capture 50 Pokemon. The student reported back that he captured 15 Sandshrew, 12 Trapinch, 15
Cacnea and 8 Baltoy. Does this sample provide any evidence that the published theoretical probabilities
are incorrect?

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 7. 2


TESTS 55
Example: Does fire a↵ect deer habitat?
After a fire, four distinct areas of a forest were identified. The inner burn consisted of 520 acres, the
inner edge consisted of 210 acres, the outer edge consisted of 240 acres and the unburned area
consisted of 2030 acres.
A sample of 75 deer was obtained and it was found that 2 deer were in the inner burn, 12 deer in the
inner edge, 18 deer in the outer edge and 43 deer in the unburned area. Do these sample data provide
evidence that fire a↵ects deer habitat?
Data drawn from Statistics for the Life Sciences, 4th ed., Samuels, 2012, pp. 349

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 7. 2


TESTS 56
Example: A superintendent of a school district in the U.S. wants to know whether students from the
various high schools in her district tend to perform equally well on their SAT’s. More specifically, she
is interested in knowing whether the percentage of students that perform above the national median is
approximately the same from school to school. She randomly sampled 100 students from five high
schools in the district and asked them to take the SAT. The number of students in each school who
scored above the national median score were as follows: School 1 - 42 students, School 2 - 45 students,
School 3 51 students, School 4 - 47 students, and School 5 - 60 students.
Is there any reason to believe that these five schools are not performing at the same level, with respect
to SAT scores?
Adapted from Chapman McGrew and Monroe, An Introduction to Statistical Problem Solving in
Geography, 2nd ed., p. 158

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 7. 2


TESTS 57
Example: An instructor of a course wanted to determine if sketching during exams had an influence on
grades. For each student in the class, the instructor recorded whether or not the student had provided a
sketch while answering a specific question, and also recorded whether or not their answer to that
question was correct. The results were as follows:
sketch no sketch

correct 53 28 incorrect 32 63
Is there any statistical reason to believe that sketching is associated with higher grades?

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 7. 2


TESTS 58
Example: While conducting a study on housing prices in di↵erent neighbourhoods, researchers
obtained a random sample of 51 houses in one of three di↵erent neighbourhoods and whether or not the
house sold for more or less than the average price. The data are as follows:
neighbourhood
#1 #2 #3

above average price 10 6 7 below average price 6 15 7


from “Statistical Methods for Geography”, Rogerson, 2007, p231.
Test whether or not there is a tendency for above or below average sale prices to cluster in certain
neighbourhoods.

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 7. 2


TESTS 59
Example: Researchers were interested in testing whether or not there is an association between soil
type and crop type. A random sample of 83 fields was obtained. The data are as follows.
Crop
Corn Oats Hay

Elevasil 10 8 3 Doreton 7 10 6
Lamoile 4 5 12 Valton 4 6 8
from “Elementary Statistics for Geographers” by Burt et al., 2009, p.406.

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: Unit 8


Regression
Example: Fifteen students in an undergraduate course were randomly selected and the number of hours
of sleep the night before the midterm and the midterm mark were recorded for each student.

Simple Linear Regression (SLR)


“Regression”: describe the distribution of Y as a function of values of X.
“Linear”: the function above is a linear combination of explanatory variables.
“Simple”: one explanatory variable.
86

Population Sample

Parameter Statistic
Residual: vertical distance from a data point to a straight line.

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 8. REGRESSION 61

Computers will test di↵erent slopes and y-intercepts until they find the line that results in the minimum
value of the sum of the squares of the residuals.
This line is called the “best fit line”.
There is a lot of information given by R when you conduct a SLR!
Coefficients:
Estimate Std. Error t value Pr(>|t|) (Intercept) 39.0474 3.9652 9.848 2.15e-07 hours 5.7483
0.6707 8.571 1.04e-06

Interpolation:
Extrapolation:
As with most of the hypothesis tests we’ve covered so far, linear regression also depends on a number
of assumptions about the data:
The observations must be independent.
In addition, ✏, the random error term that represents how Y varies above and below the best-fit line,
1. is normally distributed (normality),

2. has a mean of 0 (linearity),

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 8. REGRESSION 62


3. has an equal variance at each value of X (homoscedasticity).

from [Link]
How do we check these assumptions? 1. Normality (qq-plot of the residuals)

2. Linearity (residual plot: plot residuals vs y-values on line of best fit)

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 8. REGRESSION 63

3. Homoscedasticity (residual plot: plot residuals vs y-values on line of best fit)

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 8. REGRESSION 64

Population Sample

Parameter Statistic
Is the linear relationship between Y and X significant?

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 8. REGRESSION 65


Example: The average annual mortality rate (per 100 000 men) and the calcium concentration (in parts
per million) in the drinking water were measured for 61 large towns in England and Wales. How are
calcium concentration and mortality among men related?
data from “A Handbook of Small Data Sets” by D.J. Hand et al., Chapman & Hall, 1994.
Model:
µ(Y |X) = 0 + 1X
Assumptions:

Results:
Coefficients:
Estimate Std. Error t value Pr(>|t|) (Intercept) 1676.3556 29.2981 57.217 < 2e-16 hardness -
3.2261 0.4847 -6.656 1.03e-08

Example: Researchers were interested in quantifying the relationship between average precipitation and
elevation across southern Scotland. A random sample of 20 sites was obtained and the elevation and
average precipitation recorded.
data from “Linear Regression in Geography” in CONCEPTS AND TECHNIQUES IN MODERN
GEOGRAPHY No 15, Ferguson, p.7
Model:
µ(Y |X) = 0 + 1X
Assumptions:

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 8. REGRESSION 66

Results:
Coefficients:
Estimate Std. Error t value Pr(>|t|) (Intercept) 895.3223 149.7609 5.978 1.18e-05 elevation
2.3774 0.4438 5.357 4.32e-05

Pearson Correlation Coe cient (⇢ or r) is a measure of the strength and direction of the linear
correlation between two variables.

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 8. REGRESSION 67

graphics from [Link]/wiki/Pearson_correlation_coefficient


Examples r = 0.91 (sleep v grade) r = 0.65 (mortality v calcium)
r = 0.78 (elevation v rainfall)
Autocorrelation is related to the degree to which individual observations of a variable correlate with
each other.
Temporal autocorrelation: Repeated observations of a variable (measured at a specific location) tend to
be more highly correlated with each other the closer in time they are made.
E.g. Precipitation
Spatial autocorrelation: Observations of a variable at multiple locations in space tend to be more highly
correlated the closer they are to each other.

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 8. REGRESSION 68


Tobler’s 1st Law of Geography

How might temporal or spatial autocorrelation impact our regression analysis?


(Think about our assumptions!)
Coe cient of Determination: R2 (or sometimes r2) is the proportion of variation existing in Y that can
be explained by X.
Examples
R2 = 0.8219 (sleep v grade)
Multiple R-squared: 0.8219, Adjusted R-squared: 0.8082
R2 = 0.4288 (mortality v calcium)
Multiple R-squared: 0.4288, Adjusted R-squared: 0.4191
R2 = 0.6145 (elevation v rainfall)
Multiple R-squared: 0.6145, Adjusted R-squared: 0.5931
Example: City transportation planners are interested in estimating the total number of trips made by
households each day. A city is divided into 12 tra c zones where houses within in each zone share
similar characteristics. Planners measured the average number of trips per household per day (the “trip
generation rate”) and the average household income ($000).
from “Elementary Statistics for Geographers” by Burt et al., 2009, p.468.
Model:
µ(Y |X) = 0 + 1X
Assumptions: (checked and satisfied)
Results:
Coefficients:
Estimate Std. Error t value Pr(>|t|) (Intercept) 0.80307 0.76632 1.048 0.319 income
0.11175 0.01555 7.185 2.98e-05
Multiple R-squared: 0.8377, Adjusted R-squared: 0.8215

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 8. REGRESSION 69


Planners measured the average number of trips per household per day (the “trip generation rate”) and
the average household income ($000). City transportation planners also measured the average
household size for each of the 12 zones.
Model:
µ(Y |X1,X2) = 0 + 1X1 + 2X2
Assumptions: (checked and satisfied)
Results:
Coefficients:
Estimate Std. Error t value Pr(>|t|) (Intercept) -3.008099 0.523261 -5.749 0.000277 income
0.116734 0.005492 21.255 5.31e-09 size 1.178818 0.138811 8.492 1.37e-
05
Multiple R-squared: 0.982, Adjusted R-squared: 0.978

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 8. REGRESSION 70


Example: Oenologists are interesting in quantifying a relationship that uses a variety of factors to help
predict the quality of a wine. A random sample of 38 Pinot Noir wines is obtained and a number of
quantities are recorded for each (clarity, aroma, body, flavor, oakiness), as well as a quality score. data
from Introduction to Linear Regression Analysis, Montgomery et al, 2012, p564 Model:
µ(Y |X1,X2,X3,X4,X5) = 0 + 1X1 + 2X2 + 3X3 +4X4 + 5X5
Assumptions: (checked and satisfied)
Results:
Coefficients:
Estimate Std. Error t value Pr(>|t|) (Intercept) 3.9969 2.2318 1.791
0.082775 data$Clarity2.3395 1.7348 1.349 0.186958 data$Aroma 0.4826
0.2724 1.771 0.086058 data$Body 0.2732 0.3326 0.821 0.417503
data$Flavor 1.1683 0.3045 3.837 0.000552 data$Oakiness -0.6840
0.2712 -2.522 0.016833

Multiple R-squared: 0.7206, Adjusted R-squared: 0.6769


Model:
µ(Y |X1,X2,X3) = 0 + 1X1 + 2X2 + 3X3
Assumptions: (checked and satisfied)
Results:
Coefficients:
Estimate Std. Error t value Pr(>|t|)
(Intercept) 6.4672 1.3328 4.852 2.67e-05
data$Aroma 0.5801 0.2622 2.213 0.033740
data$Flavor 1.1997 0.2749 4.364 0.000113
data$Oakiness -0.6023 0.2644 -2.278 0.029127
Multiple R-squared: 0.7038, Adjusted R-squared: 0.6776

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 8. REGRESSION 71


Example. Researchers conducted a study on the Meadowfoam plant, which is valued for its seed oil.
The Meadowfoam seed oil is highly stable and non-greasy and is frequently used in cosmetics and hair-
care products.
data from Statistical Sleuth Third Edition, Ramsey/Schafer 2013, pg. 238
How are the number of flowers related to light timing: either at photoperiodic floral induction (PFI), or
24 days before PFI?

62.3 55.3 49.6 39.4 31.3 36.8


at PFI
77.4 54.2 61.9 45.7 44.9 41.9

77.8 69.1 57.0 62.9 60.3 52.6


before PFI
75.6 78.0 71.1 52.2 45.6 44.4

How are the number of flowers related to light intensity (µmol/m2/sec)?


Intensity
150 300 450 600 750 900

62.3 55.3 49.6 39.4 31.3 36.8


77.4 54.2 61.9 45.7 44.9 41.9

77.8 69.1 57.0 62.9 60.3 52.6


75.6 78.0 71.1 52.2 45.6 44.4

Model:
µ(Y |X1) = 0 + 1X1
Assumptions: (checked and satisfied)
Results:
Coefficients:
Estimate Std. Error t value Pr(>|t|) (Intercept) 77.385000 4.161186 18.597 6.06e-15 intensity -
0.040471 0.007123 -5.682 1.03e-05
Multiple R-squared: 0.5947, Adjusted R-squared: 0.5763

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 8. REGRESSION 72

How are the number of flowers related to both light timing and light intensity?
150 300 450 600 750 900
62.3 55.3 49.6 39.4 31.3 36.8
at PFI
77.4 54.2 61.9 45.7 44.9 41.9

77.8 69.1 57.0 62.9 60.3 52.6


before PFI
75.6 78.0 71.1 52.2 45.6 44.4

Model:
µ(Y |X1,X2) = 0 + 1X1 + 2X2
Assumptions: (checked and satisfied)
Results:
Coefficients:
Estimate Std. Error t value Pr(>|t|) (Intercept) 71.305833 3.273772 21.781 6.77e-16 intensity -
0.040471 0.005132 -7.886 1.04e-07 timing 12.158333 2.629557 4.624 0.000146
Multiple R-squared: 0.7992, Adjusted R-squared: 0.78

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 8. REGRESSION 73

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 8. REGRESSION 74


Model:
µ(Y |X1,X2) = 0 + 1X1 + 2X2 + 3X1X2
Assumptions: (checked and satisfied)
Results:
Coefficients:
Estimate Std. Error t value Pr(>|t|)
(Intercept) 71.623333 4.343305 16.491 4.14e-13
intensity -0.041076 0.007435 -5.525 2.08e-05
timing 11.523333 6.142360 1.876 0.0753
intensity:timing 0.001210 0.010515 0.115 0.9096
Multiple R-squared: 0.7993, Adjusted R-squared: 0.7692

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: Unit 9


Logistic Regression
Example: Park attendance and distance
A park obtains a random sample of 22 individuals and records their distance from the park (in km) and
whether or not they visited the park.
from Statistical Methods for Geography, 2nd ed., Rogerson, 2006, p150
Can we write an equation that helps us predict if a person will visit a park given their distance from the
park?
109
Probability versus Odds
Example: A fair 6-sided die is rolled. What is the probability of rolling a 2?
What are the odds of rolling a 2?

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 9. LOGISTIC REGRESSION 76


Results of a Logistic Regression
Call:
glm(formula = visit ~ distance, family = binomial(link = "logit"), data = parks)
Deviance Residuals:
Min 1Q Median 3Q Max
-2.05681 -0.73259 -0.03399 0.81980 1.63793
Coefficients:
Estimate Std. Error z value Pr(>|z|) (Intercept) 3.1967 1.4849 2.153 0.0313 * distance -
0.6050 0.2592 -2.334 0.0196 *
---
Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1
(Dispersion parameter for binomial family taken to be 1)
Null deviance: 30.498 on 21 degrees of freedom
Residual deviance: 22.373 on 20 degrees of freedom AIC: 26.373
Number of Fisher Scoring iterations: 4

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 9. LOGISTIC REGRESSION 77


Example: Football kickers
Researchers picked a random week of NFL football, and then a random sample of 20 field goal
attempts and recorded the distance of the kick and whether or not the player was successful.
from Statistical Methods for Geography, 2nd ed., Rogerson, 2006, p151
Coefficients:
Estimate Std. Error z value Pr(>|z|) (Intercept) 4.94417 2.48519 1.989 0.0467 * distance
-0.11269 0.06381 -1.766 0.0774 .
---
Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 9. LOGISTIC REGRESSION 78

Example: Toxicity of potassium cynate on trout eggs. Researchers placed potassium cynate in six
di↵erent concentrations into vials containing trout eggs. The number of surviving eggs and deceased
eggs were then counted after 19 days had elapsed.
Concentration (mg/L) Deceased Alive

90 77 925
180 47 846
360 66 800
720 143 451
1440 220 889
2880 696 293
Handbook of Small Data Sets, Hand et al, 1994.
Coefficients:

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 9. LOGISTIC REGRESSION 79


Estimate Std. Error z value Pr(>|z|)
(Intercept) 2.788e+00 6.590e-02 42.31 <2e-16 *** concentration -1.231e-03 3.619e-05 -34.02 <2e-16
***

Example: Toxicity of potassium cynate on trout eggs (continued).


Researchers placed potassium cynate in six di↵erent concentrations into vials containing trout eggs.
The potassium cynate was either applied immediately, or after several hours in order to let the eggs
water harden. The number of surviving eggs and deceased eggs were then counted after 19 days had
elapsed.
Concentration Water-hardened Immediate
(mg/L) Deceased Total Deceased Total

90 37 438 40 564
180 27 404 20 489
360 37 440 29 426

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 9. LOGISTIC REGRESSION 80

720 21 288 122 306


1440 185 567 35 542
2880 316 496 380 493

Downloaded by Kinywa Junior (kinywajunior@[Link])


lOMoARcPSD|60615441

GEOG*2460 F23: UNIT 9. LOGISTIC REGRESSION 81


Coefficients:
Estimate Std. Error z value Pr(>|z|)
(Intercept) 2.899e+00 9.401e-02 30.835 <2e-16 ***
concentration -1.303e-03 5.187e-05 -25.113 <2e-16 ***
hardened -2.242e-01 1.319e-01 -1.700 0.0892 .
concentration:hardened 1.440e-04 7.247e-05 1.986 0.0470 *

Downloaded by Kinywa Junior (kinywajunior@[Link])

Common questions

Powered by AI

Quartiles divide a dataset into four equal parts, each representing a quarter of the data. The inter-quartile range (IQR), which is the difference between the 75th percentile (Q3) and the 25th percentile (Q1), measures the spread of the middle 50% of the data. It is significant as it helps identify the variability of the dataset and potential outliers, being less influenced by extreme values than full range measures .

In linear regression, the statistical significance of coefficients is determined through the t-value and p-value. The t-value measures the size of the difference relative to the variation in the sample data; a higher absolute t-value indicates a more significant difference. The p-value, derived from the t-distribution, indicates the probability of observing such an extreme value by random chance. A p-value below a standard significance level (usually 0.05) suggests the coefficient is significantly different from zero, implying it has meaningful influence in predicting the dependent variable .

Homoscedasticity assumes that the variance of the residuals is constant across all levels of the independent variable(s). When this assumption is met, it ensures that the model's predictions are equally reliable across all values of the independent variables. If homoscedasticity is violated, it could lead to inefficient estimations and unreliable hypothesis tests, potentially skewing the results and leading to incorrect conclusions .

The Central Limit Theorem indicates that the sampling distribution of the sample mean will be approximately normally distributed, regardless of the distribution of the population. This theorem allows us to understand that the mean of all possible sample means equals the population mean, and the standard deviation of the sample means is the standard error. This enables the use of normal probability theory to estimate the proximity of the sample mean to the population mean .

Logistic regression is used to model the probability of a binary or categorical outcome based on one or more predictor variables. It uses a logit link function to convert the linear predictor into a probability. The odds ratio is crucial as it represents the change in odds for a one-unit increase in the predictor variable, providing insight into the strength and direction of the association between predictors and outcomes .

Temporal autocorrelation refers to the correlation of a variable with itself over successive time intervals. It can lead to underestimation of variances of coefficients, inflating type I errors. This makes standard hypothesis tests invalid as they assume independence of observations. Recognizing and adjusting for autocorrelation using methods like differencing or specific models is crucial for valid time series analysis .

A confidence interval gives a range of values that is likely to contain the population parameter with a certain level of confidence, usually 95% or 90%. It is constructed using the sample mean and the standard error of the mean. This interval provides an estimate of the parameter's credibility, indicating where the true population parameter might fall within a given probability range .

To calculate sample variance, we first determine the distance between each measurement and the sample mean, square each distance, and then sum all the squared distances. This sum is divided by the sample size minus one (n-1) to account for the degrees of freedom, providing an unbiased estimate of the population variance from a sample .

The coefficient of determination, or R-squared, quantifies the proportion of variation in the dependent variable that is explained by the independent variable(s) in the regression model. A higher R-squared indicates a better fit of the model to the data. However, it does not imply causation and can be misleading if used alone, as it doesn't factor in overfitting or the number of predictors .

The width of a confidence interval is inversely related to the sample size; larger samples tend to produce narrower intervals due to the reduced standard error. To narrow a confidence interval, one can increase the sample size, which decreases the standard error, or decrease the confidence level, which reduces the range but also the certainty. It's essential to balance these factors to maintain statistical power while achieving precision .

You might also like