0% found this document useful (0 votes)
5 views38 pages

Basic Statistics

The document is an introductory guide to basic statistics, covering key concepts such as types of statistics, levels of measurement, and central tendencies. It explains the differences between population and sample statistics, as well as the various types of data and their measurements. Additionally, it includes practical assignments and examples to reinforce understanding of statistical methods.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views38 pages

Basic Statistics

The document is an introductory guide to basic statistics, covering key concepts such as types of statistics, levels of measurement, and central tendencies. It explains the differences between population and sample statistics, as well as the various types of data and their measurements. Additionally, it includes practical assignments and examples to reinforce understanding of statistical methods.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

🗣 Learning happens between “I don’t know” and “I figured it out.

BASIC
STATISTICS
UNIVERSITY +

Swit Poizon
🗣 If you can teach it, you’ve learned it.

i
Swit Poizon
🗣 If you can teach it, you’ve learned it.

Table of Contents
BASIC STATISTICS .................................................................................................................... 1

What is statistics?.......................................................................................................................... 1

Types of statistics....................................................................................................................... 2

Levels of Measurement ................................................................................................................... 3

Assignment ..................................................................................................................................... 4

CENTRAL TENDENCIES ............................................................................................................ 4

MEASURES OF DISPENSION .................................................................................................. 8

CORRELATION & REGRESSION ......................................................................................... 16

Types of correlation .................................................................................................................. 16

Introduction to Regression ......................................................................................................... 17

ANOVA ........................................................................................................................................ 21

QUESTIONS ................................................................................................................................ 25

ii
Swit Poizon
🗣 Learning happens between “I don’t know” and “I figured it out.”

BASIC STATISTICS

What is statistics?
Statistics is the study of how to collect, organize, analyze and interpret numerical information
and data. It can also be defined as the science concerned with developing and studying methods
for collecting, analyzing interpreting and presenting data. Statistics is both science of uncertainly
and the technology of extracting information from data. It help in making decisions. Is the body
of knowledge use in making sense of data?

Individuals and variables

Individuals are people or objects included in a study.

 5 individual could be 5 people.


 5 record, or 5 reports.

A Variables is a characteristic of the individuals to be measured or observed.

 The age of the individual person.


 The time an individual record was entered.

Population, Parameter and Sample Statistics.

Population: A population is a group of people or object with a common theme. When every
member of that group is considered it is a population.

Sample: A sample is a small portion of the population. It can be a representative sample. But it
can also be a biased sample.

Population data: in population data, data from every individual in the population is available.

𝐸𝑛𝑡𝑖𝑟𝑒 𝑝𝑜𝑝𝑢𝑙𝑎𝑡𝑖𝑜𝑛 = 𝑐𝑒𝑛𝑠𝑢𝑠

Sample data: in sample data, data is only available from some of the individuals in the population.
Very commonly used in research.

Parameter: A parameter is a measure that describes the entire population.

1
Swit Poizon
🗣 If you can teach it, you’ve learned it.
Statistic: A statistic is a measure that describes only a sample of the population.

Types of statistics
1. Descriptive Statistics: This involves methods of organizing, picturing and summarizing
information from samples and population. This uses data to provide descriptions of the
population either through numerical calculation, graphs or tables.
2. Inferential Statistics: This involves methods of using information from a sample to draw
conclusions regarding the population. This uses data to make inferences and predictions
about a population based on the sample of data taken from the population.

What is data? Data is a raw fact. It is not meaningful so we cannot base on it to make inform
decisions.

What is information? Information is a processed data. It is meaningful and can be used to make
decisions.

Types of data

There are two types of data, there are: Quantitative (Continuous) data and Qualitative
(Categorical) data.

1. Quantitative Data: Is a numerical measurement of something (continuous). Example;


Temperature, time, year, height, weight etc. The quantitative data is divided in two that is
interval and ratio.
 Interval data: is a numeric data where the distance between values is meaningful and
consistent, but there is no true zero point (zero does not mean “nothing”). Example:
temperature in Celsius or Fahrenheit, IQ Scores, Calendar years etc.
 Ratio Data: is a numeric data with a meaningful zero point, allowing for comparisons
of both intervals and ratios. (e.g., “twice as much”). Example: Height, Weight, income,
etc.

2
Swit Poizon
🗣 If you can teach it, you’ve learned it.
2. Qualitative data: Refers to a “quality” or categorical characteristic of something (categorical).
Example: sex/gender, race, level of education, etc. The qualitative data is also divided into two
that is Nominal data and Ordinal data.
 Nominal Data: Data that represents categories or labels with no inherent order or
ranking. Examples; Gender, blood types, eye color, brands, shapes, sizes etc.
 Ordinal Data: Data that represents categories with meaningful order or ranking, but
the different between the categories are not consistent or measurable. Examples:
education level, race rankings, position etc.

Levels of Measurement
Types of scales
Before conducting statistical analysis, it’s important to understand how a variable is
measured. Measurement types fall into four key scale categories: nominal, ordinal,
interval, and ratio.
 Nominal scales categorize data without implying order (e.g., gender, religion, favorite
color). They're the simplest form of measurement.
 Ordinal scales introduce order (e.g., satisfaction ratings like “very satisfied” to “very
dissatisfied”), but the intervals between levels are not necessarily equal, so you can't
assume consistent differences.
 Interval scales have equal spacing between values (e.g., temperature in Fahrenheit),
but lack a true zero. This means ratios (e.g., "twice as hot") are not meaningful.
 Ratio scales have all the properties of interval scales but also include a true zero point
(e.g., money, temperature in Kelvin), making ratios meaningful (e.g., $50 is twice as
much as $25).

What level of measurement is used for psychological variables?

In psychological measurement, rating scales (such as 1 to 7 for pain or attitudes) are typically
ordinal, since we can't be sure that equal numeric differences reflect equal psychological changes.

3
Swit Poizon
🗣 If you can teach it, you’ve learned it.

Assignment
1. Define the types of categorical data and give two examples each.
2. What is the differences between:
a) Primary and secondary data.
b) Discrete data and continuous data.
c) Explain the sources of data and mention two methods of collection under each.

CENTRAL TENDENCIES
1. MODE
The mode is the most commonly occurring number or the number that repeats itself most
in a set of data. The mode is not calculated value.
Example: 4, 2, 3, 4, 4, 6, 4
In the above example 4 is the mode. This is because it repeats itself the most.
Trial
The age of 10 KG pupils are recorded as 4, 4, 5, 5, 6, 5, 6, 6, 6, 6, and 6. Find the modal
age.
Ans: the modal age is 6 because it appears most.

4
Swit Poizon
🗣 If you can teach it, you’ve learned it.

2. MEDIAN:
The median of a set of numbers is the middle value or the arithmetic mean of two middle
values when they are ordered in ascending or descending order of magnitude.

FINDING MEDIAN OF A RAW DATA.

METHOD I:

(a) Re-arrange the numbers in either ascending or descending order


(b) Count to locate the middle value as the median.
(c) If the numbers, a and b, appears as the middle values, find the sum of the two and
divide by 2 to obtain the median.
𝑎+𝑏
That is; 𝑚𝑒𝑑𝑖𝑎𝑛 = 2

METHOD II:

(a) Arrange the numbers in order of magnitude and count to ensure that they are up to
“N” number.
(b) If N is odd, then the median is the middle term.
1
That is; (𝑁 + 1)𝑡ℎ term or position. Note that this will only help you find the
2

position of the meddle value(s).


(c) If N is even, then the median is the arithmetic mean of the two middle terms.

5
Swit Poizon
🗣 If you can teach it, you’ve learned it.
1 1
That is; (𝑁)𝑡ℎ and (𝑁 + 1)𝑡ℎ term or positions.
2 2

Example 1: find the median of the following numbers.

47, 30, 56, 31, 55, 43 and 44

Ans:
Method I: Re-arrange = 30 31 43 44 47 55 56

The median number is 44

Method II: Re-arrange = 30 31 43 44 47 55 56

Total observation is 7

N = 7 and N is odd

1
Median = (𝑁 + 1)𝑡ℎ term or position
2

1
=> (7 + 1)𝑡ℎ
2

1
=> (8)𝑡ℎ
2

=> 4th position

Therefore, 4th position is 44

Example 2: find the median of the following numbers: 4, 5, 4, 6, 7, 5, 7, 4, 6, 5

Ans:

Method I:

Re-arrange 4 4 4 5 5 5 6 6 7 7

The middle numbers are 5 and 5

6
Swit Poizon
🗣 If you can teach it, you’ve learned it.
5+5
Median = =5
2

Median is 5

Method II

The total number of entries, N = 10 (even)

1 1
Median = (𝑁)𝑡ℎ and (7 + 1)𝑡ℎ term
2 2

1 1
(10)𝑡ℎ And (10) + 1 𝑡ℎ term
2 2

5th and 6th position

5+5
=5
2

Median = 5

MEAN ( 𝒙 )

The mean is the ratio of the sum of all the values in a raw data to the number of entries. Note that
mean is also known as Average. The mean is denoted by 𝒙. Thus, given the set of values:
𝑥1 , 𝑥2 , 𝑥3 , 𝑥4 , 𝑥5 , … … … . . 𝑥𝑛
𝑥1 , 𝑥2 , 𝑥3 , 𝑥4 , 𝑥5 ,………..𝑥𝑛
Mean ( 𝑥 ) = 𝑛

∑𝑥
(𝑥)= , (𝑛 𝑖𝑠 𝑡ℎ𝑒 𝑝𝑜𝑠𝑖𝑡𝑖𝑜𝑛 𝑜𝑓 𝑡ℎ𝑒 𝑙𝑎𝑠𝑡 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑡ℎ𝑒 𝑒𝑛𝑡𝑟𝑖𝑒𝑠)
𝑛

Examples 1: Find the mode, median and mean of the following numbers: 4, 24, 10, 17, 19, 21,
and 10

Ans:

Mode = 10

Median = 4, 10, 10, 17, 19, 21, 24

7
Swit Poizon
🗣 If you can teach it, you’ve learned it.
Therefore, 17 is the median

∑𝑥 4+10+10+17+19+21+24 95
Mean ( 𝑥 ) = = = = 15
𝑛 7 7

Assignment 2

1. Find the mean and median of the following:


(a) 26, 23, 27, 28, 29, 22, 23, 27, 20, 20, 24, 25
(b) 170cm, 171cm, 173cm, 174cm, 177cm, 177cm, 184cm, and 186cm
(c) 18, 17, 19, 22, 17, 18, 18, 21, 19, 18
2. The data below shows the number of siblings of class 2 pupil at Koforidua PPS. Find the
means and median.
13, 13, 13, 15, 17, 18, 18, 18, 19, 19,
19, 20, 20, 22, 22, 22, 23, 24, 25, 25,
26, 26, 27, 27, 27, 27, 28, 29, 29, 30,

Frequency and frequency diagrams

Frequency: Is the number of times an event scores in a given data.

Frequency diagrams for ungrouped Data

It is the table that shows the event and the number of times each event occurs. It is usually
divided into 4 sections as shown below.

8
Swit Poizon
🗣 If you can teach it, you’ve learned it.
Range (x) Tally Frequency (f) fx

∑𝑓 ∑ 𝑓𝑥

Worked Examples:

1. The ages of 20 school children were recorded as follows,


13 9 15 17 15 11 13 17 9 9
9 11 9 11 15 11 15 11 11 11
Solution

Age (x) Tally Frequency (f) fx


9 //// 5 45
11 //// // 7 77
13 // 2 26
15 //// 4 60
17 // 2 34

∑ 𝑓 = 20 ∑ 𝑓𝑥 = 242

∑ 𝑓𝑥 242
Mean 𝑥 = ∑𝑓
= = 12.1
20

How to construct frequency distribution table


These are the scores of 40 students in a 50-item Math exam.
13, 13, 13, 15, 17, 18, 18, 18, 19, 19,
19, 20, 20, 22, 22, 22, 23, 24, 25, 25,
26, 26, 27, 27, 27, 27, 28, 29, 29, 30,
32, 33, 35, 36, 37, 39, 42, 43, 45, 48

9
Swit Poizon
🗣 Learning happens between “I don’t know” and “I figured it out.”

Sturges Formula 𝐾 = 6 𝑐𝑙𝑎𝑠𝑠 𝑖𝑛𝑡𝑒𝑟𝑣𝑎𝑙 𝑅𝑎𝑛𝑔𝑒


𝐶𝑙𝑎𝑠𝑠 𝑠𝑖𝑧𝑒 (𝑐) =
𝑘
𝐾 = 1 + 3.322 log 𝑁 𝑅𝑎𝑛𝑔𝑒 (𝑅) = 𝑀𝐴𝑋 − 𝑀𝐼𝑁
35
𝑐=
𝐾 = 1 + 3.322 log 40 𝑅 = 48 − 13 6
𝑅 = 35 𝑐 = 5.83
𝐾 = 6.32
𝑐=6

Coefficient of Variation (CV)


 Measures of Dispersion – Coefficient of Variation
 Coefficient of variation (CV) measures the spread of a set of data as a proportion
of its mean.
• It is the ratio of the sample standard deviation to the sample mean
• It is sometimes expressed as a percentage
• There is an equivalent definition for the coefficient of variation of a
population
• A standard application of the Coefficient of Variation (CV) is to
characterize the variability of geographic variables over space or time
• Coefficient of Variation (CV) is particularly applied to characterize the
interannual variability of climate variables (e.g., temperature or
precipitation) or biophysical variables (leaf area index (LAI), biomass,
etc)
Coefficient of Variation (CV)
• It is a dimensionless number that can be used to compare the amount of
variance between populations with different means

1
Swit Poizon
🗣 If you can teach it, you’ve learned it.

Quartiles, interquartile range and semi-interquartile range

In statistics, a quartile is one of three values that divide a dataset into four equal parts, with each
part representing 25% of the data.

Box plots, also known as whisker plots, are used for summarizing data distribution and detecting
outliers. They display the median, quartiles and potential outliers in a dataset.

Example;

2
Swit Poizon
🗣 If you can teach it, you’ve learned it.
First quartile or lower quartile (Q1): The set of data points between the minimum
value and the first quartile.
Second quartile or middle quartile (Q2): The set of data points between the lower
quartile and the median.
Third quartile or Upper quartile (Q3): The set of data between the median and the
upper quartile.
Fourth quartile: The set of data points between the upper quartile and the maximum
value of the data set.
The above diagram is a boxplot representation.

𝑸𝒖𝒂𝒓𝒕𝒊𝒍𝒆𝒔 = 𝑁𝑡ℎ 𝑝𝑜𝑠𝑖𝑡𝑖𝑜𝑛 + 𝑑𝑒𝑐𝑖𝑚𝑎𝑙((𝑁 + 1)𝑡ℎ − 𝑁𝑡ℎ)

𝑄1 = 5𝑡ℎ 𝑝𝑜𝑠𝑖𝑡𝑖𝑜𝑛 + 0.25(6𝑡ℎ − 5𝑡ℎ)

3
Swit Poizon
🗣 If you can teach it, you’ve learned it.

4
Swit Poizon
🗣 If you can teach it, you’ve learned it.

Grouped frequency Table

Class mark (x) Frequency (f)


25 – 29 3
30 – 34 4
35 – 39 7
40 – 44 6
45 – 49 2

Class limits
25 - 29, 30 – 34………..
Lower class limit → 𝟐𝟓 − 𝟐𝟗 ← Upper class limit

𝑙𝑜𝑤𝑒𝑟 𝑐𝑙𝑎𝑠𝑠 𝑙𝑖𝑚𝑖𝑡 + 𝑢𝑝𝑝𝑒𝑟 𝑐𝑙𝑎𝑠𝑠 𝑙𝑖𝑚𝑖𝑡


Class midpoint = 2
25+29
Midpoint =
2
Class boundaries: to find class boundaries one need to find half of difference between
lower class limit 30 of the second class and the upper class limit 29 of the first class.
1
That is, 2 (30 − 29) = 0.5

Now subtract that answer you have gotten 0.5 from each of the lower class limit and also
add that same 0.5 to each of the upper class limit.
We will get 24.5 – 29.5, 29.5 – 34.5…………
Class size: class size is the difference between the upper boundary and lower boundary
29.5 – 24.5 = 5. It means that 5 is your class size.

5
Swit Poizon
🗣 If you can teach it, you’ve learned it.
Examples:
The data below shows the ages of children at Sowutoum ACA.
Find the mean age of the children.

Ages frequency
1–5 3
6 – 10 7
11 – 15 9
16 - 20 2

Solution

Ages frequency midpoint fx CF Boundaries Class Width/size (c)


1–5 3 3 9 3 0.5 – 5.5 5
6 – 10 7 8 56 10 5.5 – 10.5 5
11 – 15 9 13 117 19 10.5 – 15.5 5
16 - 20 2 18 36 21 15.5 – 20.5 5
∑f=21 ∑fx=218

∑𝑓𝑥 218
Mean (𝑥) = = = 10.38095238
∑𝑓 21

𝑥 ≈10.3810 (4. dp)

Quartiles
1
𝑁−𝐶𝐹𝑏𝑄1
First quartile Q1 of a grouped data is 𝑄1 = 𝑙 + (4 ).𝐶
𝑓𝑄1

1
𝑁−𝐶𝐹𝑏𝑓𝑄2
Second Quartile Q2 of grouped data is 𝑄2 = 𝑙 + (2 ).𝐶
𝑓𝑄2

1
𝑁−𝐶𝐹𝑏𝑓𝑄2
Third Quartile Q3 of a grouped data is 𝑄3 = 𝑙 + (2 ).𝐶
𝑓𝑄3

6
Swit Poizon
🗣 If you can teach it, you’ve learned it.
Where 𝒍 is the lower boundary of that class.

𝑵 is the total frequency

𝑪𝑭𝒃𝑸𝟏 is the preceeding commulative frequency

𝑪 is the class size or width

𝑓𝑄1 is the frequency of that class

𝑵𝒐𝒕𝒆 𝑡ℎ𝑎𝑡 𝑏𝑒𝑓𝑜𝑟𝑒 𝑒𝑣𝑒𝑟𝑦𝑡ℎ𝑖𝑛𝑔 ℎ𝑒𝑟𝑒 𝑦𝑜𝑢 ℎ𝑎𝑣𝑒 𝑡𝑜 𝑓𝑖𝑟𝑠𝑡 𝑓𝑖𝑛𝑑 𝑡ℎ𝑒 𝑝𝑜𝑠𝑖𝑡𝑖𝑜𝑛 𝑜𝑓 𝑡ℎ𝑒 𝑡ℎ𝑎𝑡 𝑞𝑢𝑎𝑟𝑡𝑖𝑙𝑒 𝑐𝑙𝑎𝑠𝑠

MODE OF A GROUPED DATA

∆1
𝑀𝑜𝑑𝑒 = 𝑙 + ( ).𝐶
∆2 + ∆1

Where 𝒍 is the lower boundary of that class.

𝑪 is the class size or width

∆1 is the difference between frequencies of modal and pre − modal class

∆2 is the difference between frequencies of modal and post − modal class

Example

From the following frequency distribution, find the standard deviation, variance, mean and mode
using the formula for grouped data.

Class interval frequency


10 – 20 9
20 – 30 18
30 – 40 31
40 – 50 17
50 – 60 16
60 – 70 9

7
Swit Poizon
🗣 If you can teach it, you’ve learned it.

MEASURES OF DISPENSION

Measures of Dispersion

Measures of dispersion, also called measures of variability, "describe the extent to which the
values of a variable are different." (Wallace & Van Fleet, 2012, p. 293) The most common
measures of dispersion are range, variance, standard deviation, and the coefficient of variation.

Types of Measures of Dispersion

Before diving into the types of measures of dispersion, it's essential to understand why they
matter. Simply knowing the average (mean) of a data set is not enough—two different data sets
can have the same average but completely different spreads. Measures of dispersion help
quantify this spread, providing a clearer picture of the data’s variability.

8
Swit Poizon
🗣 If you can teach it, you’ve learned it.
For example, a class where all students score between 85 and 90 has less variation than a class
where scores range from 50 to 100, even if both classes have the same average score. Dispersion
measures allow us to compare different data sets and understand their consistency.

Measures of dispersion can be categorized into two broad types:

1. Absolute Measures of Dispersion


2. Relative Measures of Dispersion

Absolute Measures of Dispersion

Absolute measures of dispersion express variations in a data set with respect to the deviations
from the central value. These measures retain the same unit as the data set. The common absolute
measures of dispersion include:

Range: The difference between the maximum and minimum values in a data set.

Variance: The average squared deviation from the mean, indicating the spread of data.

Standard Deviation: The square root of variance, measuring how data deviates from the mean.

Mean Deviation: The average of absolute deviations from a central value (mean, median, or
mode).

Quartile Deviation: Half the difference between the third quartile (Q3) and the first quartile
(Q1).

Relative Measures of Dispersion

Relative measures of dispersion are dimensionless and expressed as ratios or percentages,


making them useful for comparing data sets with different units. Common relative measures
include:

Coefficient of Range: Ratio of the difference between the highest and lowest values to their
sum.

Coefficient of Variation: Ratio of standard deviation to the mean, expressed as a percentage.

9
Swit Poizon
🗣 If you can teach it, you’ve learned it.
Coefficient of Mean Deviation: Ratio of mean deviation to the central value from which it is
calculated.

Coefficient of Quartile Deviation: Ratio of the difference between Q3 and Q1 to their sum

Range

Range is a very simple statistic, as it is merely the difference between the largest and smallest
values. It's not very useful from a statistical standpoint; because it relies only on the outermost
values of a dataset, two datasets with the same range could still have drastic differences in their
overall distribution.

Variance

Variance can refer to two different things: population variance (sometimes called parametric
variance) and sample variance. Population variance, typically represented by a lower-case sigma
squared (𝜎 2 ), can only be truly calculated if you have observations for every member of your
population, which is almost never the case. It's used more often by statisticians and using it
should never be your first instinct. Sample variance, typically represented 𝑏𝑦 𝑆 2 , is what you
should almost always use instead. Its formula can be expressed as:

Where x is each value of the variable, 𝑥̅ is the mean for the variable, and n is the total number of
observations. Like the equation for the mean, it looks more complicated than it is. Because it's a
squared measurement, it's rarely reported and is instead more useful for statisticians. Standard
deviation is calculated from the variance, and is generally a more understandable and useful
measurement.

Standard Deviation

Like variance, standard deviation can theoretically be calculated for both a population and a
sample. The population standard deviation (σ) is rarely used; sample standard deviation (s)

10
Swit Poizon
🗣 If you can teach it, you’ve learned it.
should be your first choice. Easy to calculate, it is simply the square root of the variance. It tends
to be more understandable than variance because it expresses the dispersion of a dataset in that
dataset's original units. It can be written as:

Coefficient of Variation

If you need to compare the variation for different measurement variables, the coefficient of
variation is another useful measure of dispersion. It expresses the amount of variation as a
proportion of the total. It is calculated by dividing the standard deviation by the mean:

Example:

11
Swit Poizon
🗣 If you can teach it, you’ve learned it.

Example:

The table below shows the weights(kg) of members in a sport club. Calculate the mean,
median and mode of the distribution.

Masses 40 - 49 50 - 59 60 - 69 70 - 79 80 - 89 90 -99
Frequency 6 8 12 14 7 3

Example:

The table below shows the height of teak trees harvested by a farmer.

Height (m) 3 4 5 6 7 8
No of trees 4 6 4 5 6 2
a. Find the median height
b. Calculate, correct to one decimal place, the
i. Mean
ii. Standard deviation

Finding the sample mean, variance and standard deviation from a frequency distribution
table (Group data)

Example:

The table below shows the scores (out of 60) of the applicants in an aptitude test. Calculate
the mean, median and mode of the distribution.

Scores 51 - 60 41 - 50 31 - 40 21 - 30 11 - 20 1 -10
f 4 7 10 8 6 5

12
Swit Poizon
🗣 If you can teach it, you’ve learned it.
∑𝑓𝑥
Sample Mean = where f = frequency, x = class marks or scores and ∑f or n = number
∑𝑓

sample.

∑𝒇(𝒙−𝒙) ∑𝒇(𝒙−𝒙)
Sample Variance = Sample Standard Deviation = √
∑𝒇−𝟏 ∑𝒇−𝟏

Example:

The scores in Business Math of randomly selected grade 11 GAS students are shown below.
Solve for the sample mean, sample variance and sample standard deviation

Scores Frequency
60 – 64 1
65 – 69 2
70 – 74 3
75 – 79 4
80 – 84 6
85 – 89 8
90 – 94 5
95 – 99 7

Low SD

𝒔𝟐 when the s is small, most data points are close to the average value (mean)

𝒔𝟐 𝒘𝒉𝒆𝒏 𝒊𝒔 𝒍𝒂𝒓𝒈𝒆, 𝒅𝒂𝒕𝒂 𝒑𝒐𝒊𝒏𝒕𝒔 𝒂𝒓𝒆 𝒔𝒑𝒓𝒆𝒂𝒅 𝒇𝒖𝒓𝒕𝒉𝒆𝒓 𝒂𝒘𝒂𝒚 𝒇𝒓𝒐𝒎 𝒕𝒉𝒆 𝒎𝒆𝒂𝒏, 𝒊𝒏𝒅𝒊𝒄𝒂𝒕𝒊𝒏𝒈 𝒎𝒐𝒓𝒆

𝒗𝒂𝒓𝒊𝒂𝒃𝒊𝒍𝒊𝒕𝒚 𝒊𝒏 𝒕𝒉𝒆 𝒅𝒂𝒕𝒂.

𝒊𝒇 𝒔 𝒊𝒔 𝒛𝒆𝒓𝒐, 𝒊𝒕 𝒎𝒆𝒂𝒏𝒔 𝒂𝒍𝒍 𝒕𝒉𝒆 𝒅𝒂𝒕𝒂 𝒑𝒐𝒊𝒏𝒕𝒔𝒔 𝒂𝒓𝒆 𝒆𝒙𝒂𝒄𝒕𝒍𝒚 𝒕𝒉𝒆 𝒔𝒂𝒎𝒆 𝒂𝒔 𝒕𝒉𝒆 𝒎𝒆𝒂𝒏

13
Swit Poizon
🗣 If you can teach it, you’ve learned it.
GRAPHS

1. Bar Graph:

A bar chart makes numerical information


easy to see by showing it in a pictorial form.

The width of the bar has no significance.


The height of each bar represents the
quantity. Example;

2. Pie diagram or chart:


The information is displayed using sectors of a circle.

3. Histogram:

A histogram displays the frequency of either continuous or


grouped discrete data in the form of bars.

The bars are joined together.

The bars can be of varying width.

The frequency of the data is represented by the area of the


bar and not the height.

[When class intervals are different it is the area of the bar which represents the frequency not the
height]. Instead of frequency being plotted on the vertical axis, frequency density is plotted.

𝑓𝑟𝑒𝑞𝑢𝑛𝑐𝑦
Frequency density = 𝑐𝑙𝑎𝑠𝑠 𝑤𝑖𝑑𝑡ℎ

14
Swit Poizon
🗣 If you can teach it, you’ve learned it.

4. Frequency Polygon

A frequency polygon is a type of line graph where the


frequencies of classes are plotted against their
midpoints. This graphical representation closely
resembles a histogram and is typically used for
comparing data sets or showing cumulative frequency
distributions. It uses a line graph to represent quantitative
data.

Frequency polygons are one of the great methods to represent statistical data so that it can be
read easily. In statistics, we deal with lots of data, and reading it quickly is necessary for solving
statistical problems effectively.

Frequency polygons help us to achieve the same result. In this article, we will learn about
frequency polygons, their formula, examples, and others in detail.

15
Swit Poizon
🗣 If you can teach it, you’ve learned it.

CORRELATION & REGRESSION


Introduction to correlation

Correlation measures the strength and direction of the relationship between two variable. It’s
fundamental concept in statistics and machine learning, particularly useful in exploratory data
analysis and feature selection.

Types of correlation
There are three main types of correlation: positive, negative and no correlation.

 Positive correlation means as one variable increases, the other tends to increase.
 Negative correlation means as one variable increases, the other tends to decrease.
 No correlation means theirs is no relationship between the variables.

Positive Relationship Negative Relationship No Relationship

NOTE:

𝑟 = ± 0 − 0.19 𝑣𝑒𝑟𝑦 𝑙𝑜𝑤 𝑜𝑟 𝑝𝑟𝑜𝑏𝑖𝑏𝑙𝑦 𝑚𝑒𝑎𝑠𝑢𝑟𝑎𝑏𝑙𝑒 𝑐𝑜𝑟𝑟𝑒𝑙𝑎𝑡𝑖𝑜𝑛

𝑟 = ± 0.2 − 0.39 𝑎 𝑙𝑜𝑤 𝑐𝑜𝑟𝑟𝑒𝑙𝑎𝑡𝑖𝑜𝑛 𝑡ℎ𝑎𝑡 𝑤𝑎𝑟𝑟𝑎𝑛𝑡 𝑓𝑢𝑟𝑡ℎ𝑒𝑟 𝑖𝑛𝑣𝑒𝑠𝑡𝑖𝑔𝑎𝑡𝑖𝑜𝑛

𝑟 = ± 0.4 − 0.59 𝑎 𝑟𝑒𝑎𝑠𝑜𝑛𝑎𝑏𝑙𝑒 𝑐𝑜𝑟𝑟𝑒𝑙𝑎𝑡𝑖𝑜𝑛


𝑟 = ± 0.6 − 0.79 𝑎 ℎ𝑖𝑔ℎ 𝑐𝑜𝑟𝑟𝑒𝑙𝑎𝑡𝑖𝑜𝑛

𝑟 = ± 0.8 − 1 𝑣𝑒𝑟𝑦 ℎ𝑖𝑔ℎ 𝑐𝑜𝑟𝑟𝑒𝑙𝑎𝑡𝑖𝑜𝑛

If it is exactly ± 𝟏 then we say it is a perfect


positive or negative correlation.

16
Swit Poizon
🗣 If you can teach it, you’ve learned it.

Introduction to Regression
Regression analysis is a statistical method used to model the relationship between a dependent
variable and one or more independent variables. It’s widely used in predictive modeling and
machine learning.

Simple linear Regression

Simple linear regression models relationship between two variables using a linear equation. It’s
the simplest form of regression and serves as a foundation for more complex regression
techniques.

Multiple linear regression

Multiple linear regression extends simple line regression to include multiple independent
variables. It’s useful when trying to predict a dependent variable based on multiple factors.

Regression Analysis

Regression analysis is a widely used technique which is useful for evaluating multiple
independent variables. As a result, it is particularly useful for assess and adjusting for confounding.
It can also be used to assess the presence of effect modification.

Regression line – is a straight line that describes how a response variable y changes as an
explanatory variable x changes. We often use a regression line to predict the value of y for a given
value of x. Regression, unlike correlation, requires that we have an explanatory variable and a
response variable.

Remember y = mx + b? Now we just call it something slightly different

17
Swit Poizon
🗣 If you can teach it, you’ve learned it.
Some properties of the regression

 The line always passes through the point ( 𝑥, 𝑦) called the centroid.

𝑵 𝑵

̂ 𝒊 = ∑ 𝒚𝒊
∑𝒚
𝒊=𝟏 𝒊=𝟏

̂𝒊 ) = ∑ = (∑ ∑𝒊 ) = 𝟎
∑(𝒚𝒊 − 𝒚
𝒊=𝟏

𝒙 𝒚 ̂
𝒚 ∑𝒊 ̂ = 𝟒. 𝟕 + 𝟎. 𝟗𝒙 ( use this to find your 𝑦̂ )
𝒚
1 7 5.6 1.5
2 4 6.5 -2.5
3 9 7.4 1.6
4 7 8.3 -1.3
5 10 9.2 0.8
15 37 37 0

QUESTION:

An agricultural research organization tested a particular chemical fertilizer to try to find out
whether an increase in the amount of fertilizer would lead to a corresponding increase in the food
supply. Comment on it.

Fertilizer 2 1 3 2 4 5 3
Bag of rice 4 3 4 3 6 5 5
a) Find Karl Pearson’s Moment product correlation coefficient.
b) Fit the simple regression line.

18
Swit Poizon
🗣 If you can teach it, you’ve learned it.

SOLUTION

Y1 and Y
7
𝑟 = 0.8112630728
6
y = 0.6711x + 2.3684
𝑦̂ = 0.6711𝑥 + 2.3684 5
X Y ̂
𝒀
4
2 4 3.7106
3
1 3 3.0395
3 4 4.3817 2

2 3 3.7106 1

4 6 5.0528
0
5 5 5.7239 0 1 2 3 4 5 6
X
3 5 4.3817 Y Y1 Linear (Y1)

QUESTION

The table below shows the number of math classes missed during a school year for nine
students, and their final exams scores.

Number of classes missed (x) 2 10 3 22 15 2 20 18 9


Final Exam Score (y) 99 72 90 35 60 80 40 43 75

a. Write the linear regression equation for this data set. Round all values to the nearest
hundredth.
b. State the correlation coefficient for your linear regression. Round your answer to the
nearest hundredth
c. State what the correlation coefficient indicates about the linear fit of the data.

19
Swit Poizon
🗣 If you can teach it, you’ve learned it.

20
Swit Poizon
🗣 If you can teach it, you’ve learned it.
ANOVA
What is ANOVA?

ANOVA is a statistical method used to test the variance or differences between the means of
three or more groups.

ANOVA stands for Analysis of Variance

ANOVA can blend in linear and nonlinear models. It used to determine whether or not there is a
statistically significant differences between means of three or more independent groups. The two
most common types of ANOVA are one way ANOVA and two-way ANOVA.

In ANOVA, there are two components of effect such as experimental effect & error
(experimental & independent differences).

ANOVA involves categorical independent variables

ANOVA neglects the influence of covariates.

ANOVA TABLE

Source of Source of Degree of Mean sum of


𝑭 = 𝒓𝒂𝒕𝒊𝒐
variance squares freedom squares

𝑆𝑆𝑅
Regression 𝑆𝑆𝑅 1 𝑀𝑆𝑅 =
1

𝑴𝑺𝑹
𝑆𝑆𝐸 𝑭=
Error 𝑆𝑆𝐸 𝑁−2 𝑀𝑆𝐸 = 𝑴𝑺𝑬
𝑁−2

Total 𝑆𝑆𝑇 𝑁−1

21
Swit Poizon
🗣 If you can teach it, you’ve learned it.

Hypothesis Test

𝐻0 : 𝑏 = 0 (𝑛𝑜 𝑎𝑠𝑠𝑜𝑐𝑖𝑎𝑡𝑖𝑜𝑛 𝑏𝑒𝑡𝑤𝑒𝑒𝑛 𝑣𝑎𝑟𝑖𝑎𝑏𝑙𝑒𝑠)

𝐻1 : 𝑏 ≠ 0 (There is an association between variables)

𝑴𝑺𝑹
Test statistic, 𝑭 = 𝑴𝑺𝑬 ≈ 𝑭∝ (𝟏, 𝑵 − 𝟏)

Decision Rule: We reject 𝐻0 𝑖𝑓 𝐹 ≥ 𝑭∝ (𝟏, 𝑵 − 𝟐)

∝ 𝑖𝑠 𝑙𝑒𝑣𝑒𝑙 𝑜𝑓 𝑠𝑖𝑔𝑛𝑖𝑓𝑖𝑐𝑎𝑛𝑡 (𝑒. 𝑔. 5% = 0.05)

The F-table will be used to determine the F value.

If you reject, you say there is a relationship and then look at the means to determine which mean
is higher or lower than the others. If you accept, you say there is no relationship, and state what
the mean is about for all of the groups.

𝑵
𝟐 𝑤𝑒 𝑟𝑒𝑗𝑒𝑐𝑡 𝐻0 , 𝑖𝑓 𝐹 ≥ 𝐹(1, 𝑁 − 2), 𝑏 ≠ 0
𝑺𝑺𝒙𝒙 = ∑ 𝒙𝟐 𝒊 − 𝑵𝒙
𝒊=𝟏 𝑊𝑒 𝑎𝑐𝑐𝑒𝑝𝑡 𝐻0 , 𝑖𝑓 𝐹 < 𝐹(1, 𝑁 − 2), 𝑏 = 0

𝑵
𝟐
𝑺𝑺𝒚𝒚 = ∑ 𝒚𝟐 𝒊 − 𝑵𝒚
𝒊=𝟏

𝑵𝑶𝑻𝑬: 𝑺𝑺𝒚𝒚 = 𝑺𝑺𝑻

𝑺𝑺𝑹 = 𝒃𝟐 𝑺𝑺𝒙𝒙

𝑺𝑺𝑬 = 𝑺𝑺𝑻 − 𝑺𝑺𝑹

𝑺𝑺𝑻 = 𝑺𝑺𝑬 + 𝑺𝑺𝑹

[Link]

22
Swit Poizon
🗣 If you can teach it, you’ve learned it.

Solution

𝑆𝑆𝑦𝑦 = (295) − 5(7.4)2 𝑆𝑆𝑥𝑥 = 55 − 5(3)2 𝑆𝑆𝑅 = 𝑏 2 𝑆𝑆𝑥𝑥


𝑆𝑆𝑦𝑦 = 295 − 273.8 𝑆𝑆𝑥𝑥 = 55 − 5(9) 𝑆𝑆𝑅 = (0.9)2 (10)

𝑆𝑆𝑦𝑦 = 21.2 𝑆𝑆𝑥𝑥 = 55 − 45 𝑺𝑺𝑹 = 𝟖. 𝟏

𝑺𝑺𝒚𝒚 = 𝑺𝑺𝑻 = 𝟐𝟏. 𝟐 𝑺𝑺𝒙𝒙 = 𝟏𝟎

𝑆𝑆𝑅 8.1
𝑁−2=5−2=𝟑 𝑀𝑆𝑅 = =
𝑆𝑆𝐸 = 𝑆𝑆𝑇 − 𝑆𝑆𝑅 1 1
𝑆𝑆𝐸 = 21.2 − 8.1 𝑁−1=5−1=𝟒 𝑴𝑺𝑹 = 𝟖. 𝟏

𝑺𝑺𝑬 = 𝟏𝟑. 𝟏 𝑆𝑆𝐸 13.1


𝑀𝑆𝐸 = =
𝑁−2 3
𝑴𝑺𝑬 = 𝟒. 𝟒

23
Swit Poizon
🗣 If you can teach it, you’ve learned it.
ANOVA TABLE

Source of Source of Degree of Mean sum of


𝑭 = 𝒓𝒂𝒕𝒊𝒐
variance squares freedom squares

Regression 8.1 1 𝑀𝑆𝑅 = 8.1

𝐹 = 1.84090909
Error 13.1 3 𝑀𝑆𝐸 = 4.4

Total 21.2 4 -

∝ −𝑙𝑒𝑣𝑒𝑙 𝑜𝑓 𝑠𝑖𝑔𝑛𝑖𝑓𝑖𝑐𝑎𝑛𝑡 (5% = 0.05)

𝑭𝟎.𝟎𝟓 = (𝟏, 𝟑) = 𝟏𝟎. 𝟏𝟑 > 𝟏. 𝟖𝟒𝟎𝟗𝟎𝟗

𝑐𝑜𝑛𝑐𝑙𝑢𝑠𝑖𝑜𝑛: 𝑊𝑒 𝑎𝑐𝑐𝑒𝑝𝑡 𝐻0 𝑠𝑖𝑛𝑐𝑒 1.840909 < 10.13

(𝑖. 𝑒 𝑏 = 0, ℎ𝑒𝑛𝑐𝑒 𝑡ℎ𝑒𝑟𝑒 𝑖𝑠 𝑛𝑜 𝑎𝑠𝑠𝑜𝑐𝑖𝑎𝑡𝑖𝑜𝑛 𝑏𝑒𝑡𝑤𝑒𝑒𝑛 𝑎𝑔𝑒 𝑎𝑛𝑑 𝑡ℎ𝑒 𝑤𝑒𝑖𝑔ℎ𝑡 𝑜𝑓 𝑎 𝑐ℎ𝑖𝑙𝑑. )

24
Swit Poizon
🗣 If you can teach it, you’ve learned it.

QUESTIONS
1. Describe a practical problem or investigation in your field of interest where a regression
model would be appropriate.
2. The following data shows the advertising expenditure (X, in hundreds of cedis) and
corresponding sales (Y, in thousands of cedis) for a company over time:

x 1 2 3 4 5
y 1 1 2 2 4

i) Use the data to estimate sales as a linear function of advertising expenditure.


ii) Test the usefulness of your model at the 5% significance level.
iii) Estimate the sales when the advertising expenditure is 450 cedis.

3. (a) Define the terms "sample" and "population" in statistics.


(b) The data below represent the score on a test of current events knowledge of HR
students in a tertiary institutions in Ghana.

Score Frequency

1 11

3 35

5 21

6 3

7 5

8 3

9 15

25
Swit Poizon
🗣 If you can teach it, you’ve learned it.

i) Represent the data on a box plot


ii) Calculate the standard deviation and interpret the result

Question 2: Levels of Measurement


a) Define the four levels of measurement and classify which level suits the examples
below
i. Temperatures inside 10 pizza ovens
ii. Ranking of golfers in a tournament.
iii. Categories of magazines in a physician’s office (sports, women’s, health, men’s,
news)
iv. Number of amps delivered by battery chargers

b) A sample of IT students in a tertiary institutions in Ghana in asked hours spent in


watching television within a week.

Hours Frequency

6–10 3

11–15 8

16–20 4

21–25 5

i. Compute the mean, mode, and median


ii. Determine the proportion of students watched TV for more than 15 hours in a
week.

26
Swit Poizon

You might also like