0% found this document useful (0 votes)
14 views41 pages

Stats For Health Science

The document provides an overview of statistics, focusing on data collection, types of data, and methods for presenting data. It categorizes statistics into descriptive and inferential types, detailing various data types (NOIR) and presentation methods such as tables, charts, and graphs. Additionally, it discusses measures of central tendency, including mean, median, and frequency distribution.

Uploaded by

yumitanar
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
14 views41 pages

Stats For Health Science

The document provides an overview of statistics, focusing on data collection, types of data, and methods for presenting data. It categorizes statistics into descriptive and inferential types, detailing various data types (NOIR) and presentation methods such as tables, charts, and graphs. Additionally, it discusses measures of central tendency, including mean, median, and frequency distribution.

Uploaded by

yumitanar
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

SCIENCE LECTURE NOTE

1 STATISTICS 1.3 Data collection


1.3.1 Definiton
1.1 Definition
Data are individual pieces of factious (actually
Statistics refers to the study of collecting, occurring) information recorded and used for
analyzing, interpreting, presenting and the purpose of analysis. It is a read information
organizing data in a particular manner in order to from which statistics are created.
give information.
1.3.2 Types of Data
1.2 Types of Statistics Base on their Mathematical properties data are
Statistics have majorly categorized into two divided into four groups, i.e NOIR.
types: 1. Norminand data are alphabet or numerical
in name only, they have no order or
• Descriptive Statistics
direction e.g Gender, Marital status.
• Inferential Statistics 2. Ordinal Data, it place events in rank or
order and allow for setting up inequalities
1.2.1 Descriptive Statistics Descriptive Statistics is e.g undergraduate (U), graduate (G), post
a way to organize, represent and describe a graduate (P), Doctorate ( D ).
collection of data using tables, graphs and ∴U < G< P < D
summary measures. For example, the collection
of students that applied for nursing at Maryam 1
Abacha American University. 3. Interval (score/mark) Data: In addition to
Description Statistics are also categorized into ranking and inequalities allows for forming
four different categories: differences and no absolute zero e.g
Temperature, standardized scores.
• Measure of central tendency (mean,
median and mode) 4. Ratio Data: Allow for forming quotient in
addition to inequalities and differences e.g
• Measure of dispersion (Range, Variance,
Height, weight, age etc.
Standard deviation, Mean deviation,
Quartile deviation and coefficients of Discrete and continuous Data (birth numeri-
dispersion) cal)

• Measure of position • Continuous data can take any numerical


• Measure of frequency values within the range and have infinite
number of possible values. e.g Height of
1.2.2 Inferential Statistics This type of Statistics Nigerian boys 6 − 25
is used to interpret the meaning of Descriptive • Discrete data can take only certain values by
Statistics. There are many types of statistical a finite jump. e.g Number of students in a
tests used for this type, some of them are: class.
• z-test
• t-test 1.4 Methods for Data collection
• f-test
• Anova
2 PRESENTATION OF DATA
Data collection is defined as a method of 3. Interview
collecting, analyzing data for the purpose of
validation and research using some techniques. 4. Observation
5. Forms

6. Online tracking

7. Social median monitoring

8. Document ( Review )

2 Presentation of data
When data are collected, they may be
presented in uncorrected form in which they
are gathered.

e.g 2.1 Tabular form


The record below shows the scores obtained by
some students in a mathematics test:

12 15 20 25 30 20 30 15 12 12

15 12 25 12 20 12 20 12 15 25

12 15 12 15 12 15 12 20 20 30

30 25 15 20 30 15 25 30 12 15

12 15 12 15 12 30 20 12 15 20 from
the data, it will observed that:

• The data is discrete because it is obtained by


• Primary source is collected from the first counting.
hand experience and it is not used with past.
• The least score in the data is 12.
• secondary source is the data that has been
used in the past. e.g (Document review). • The highest score is 30.
• data can be qualitative (contextual in The data may be presented by the use of a table
nature) or quantitative (Numerical in nature) called frequency table.
2 PRESENTATION OF DATA

2.1.1 Frequency table The frequency table


1. Surveys
shows the number of times a particular event
2. Transactional occurs in a given data or information.

2
Class intervals Tally Frequency
18 − 20 llllllll l 11
21 − 23 2 llll
PRESENTATION 4OF DATA
2 24 − 26 llll lll 8
2.1 Tabular form 27 − 29 llll ll 7
30 − 32 llll 5
33 − 35 llll 4
Example
36 − 38 llll 4
Construct a frequency table from the above data. 39 − 41 llll ll 7
Range = highest − least = 40 − 18 = 22 and
Solution

Scores Tally Frequency


12 llllllllllll l 16
Therefore in order to get 8 equal intervals, each
15 llllllll lll 13 grouping must contain 3 units.
20 llll llll 9 2.1.3 Cumulative Frequency
25 llll 5
30 llll ll 7 Cumulative frequency is the cumulative total of
the frequency in each class giving the total
NOTE: The above data is un-grouped data. This
frequency below the class boundary.
is because each entry is recorded on its own.

Example
2.1.2 Grouped Frequency table
Use the above data and construct cumulative
When the number of entries in a given data is frequency table.
large, the construction of a frequency
distribution may be difficult, hence, there will be
Solution
need to group the data.
Class intervals Tally f Cf
Example 18 − 20 llllllll l 11 11
21 − 23 llll 4 15
The following is the shoe sizes of students of
24 − 26 llll lll 8 23
MAAUN.
27 − 29 llll ll 7 30
18 34 20 39 36 32 29 19 25 38 30 − 32 llll 5 35
33 − 35 llll 4 39
25 22 25 28 37 40 3927 23 19 36 − 38 llll 4 43
39 − 41 llll ll 7 50
30 19 21 25 40 27 18 32 19 40
where F is frequency and Cf is cumulative
26 18 39 31 19 29 40 30 21 26 frequency.

40 20 37 33 26 35 35 29 25 18 Exercises
use frequency table and arrange the data into 8 1. The following are the mark of 50 students of
equal intervals. MAAUN in a mathematics examination:

Solution 65 70 60 47 51 55 59 63 68 63

The highest size = 40 47 53 72 53 67 62 64 70 57 56

The least size = 18 73 56 48 51 58 63 65 62 49 64

3
2 PRESENTATION OF DATA
53 59 63 50 48 72 67 56 61 64
60 52 49 62 71 58 53 69 63 59
Prepare the data, using 7 equal class
intervals.
(a) A grouped frequency table

4
2 PRESENTATION OF DATA
2.2 Graphical Presentation of data Example 1
The table show the number of admission list of
(b) A cumulative frequency table MAAUN over a five year period.
2.2.2 Pie chart

PIE CHART

.
2. In a mock examination for the final tear Year 2018 2019 2020 2021 2022
Chemistry class, the following were Stds 4000 1500 3000 2500 5000
obtained by 50 students: The Pie chart is also known as the divided circle.
The circle is divided into sectors. The angle of
71 63 70 45 59 82 61 79 37 89
each sector is proportional to the frequency the
33 56 39 42 64 73 59 67 62 60 item that is being represented. The process is by
finding the sum of all frequencies and divide it
46 36 61 87 91 67 54 72 39 43 by 360°to get the angle sector.

Example 1
using the class interval 31 − 40, 41 − 50, 51
The following is the scores of Nursing students in
− 60 construct
STA1302 test:
(a) A frequency table for the data.
(b) A cumulative frequency table for the GRADE NUMBER OF STUDENT
data. A 12
B 18
C 26
2.2 Graphical Presentation
D 16
of data
Represent the data in pie chart.
2.2.1 Bar chart
Bar chart is the most popular types of diagram Solution
used to show statistical information. Bar chart
Firstly we have to construct a table of sector of
consists of rectangular bars of equal width and
angle
each bar is separated by equal gaps.

5
2 PRESENTATION OF DATA
Solution

GRADE STD ANGLE A table is first to prepare to calculate sectorial


angle for each department.
A 12

B 18

C 26

D 16

TOTAL 72 =360o
2.2 Graphical Presentation of data

Example 2

Below is the breakdown of students to be admitted into MAAUN in the year 2023.
Departments Number of students
Mathematics 140
MBBS 600
Nursing 250
Dentistry 120
Chemistry 230
Physics 260
Computer 200

Draw a pie chart for the data

6
2 PRESENTATION OF DATA
2.2.3 Histogram while the height of each rectangle is
proportional to the corresponding class
The Histogram is a bar chart whose
frequency.
rectangle have no spaces between any pair
of adjacent bars. The bases of the
rectangles of a Histogram lie in the x−axis, Example 1

DEPT STD ANGLEUse the following table and draw a

Mathematics 140 Class Frequency


0-2 6
3-5 12
MBBS 600
6 -8 3
9 and above 3
Nursing 250 Histogram.
the
Solution 3 Dentistry 120 MEASURE OF CENTRAL TENDENCY

Chemistry 230

Physics 260

Computer 200

1800 =360o

7
2 PRESENTATION OF DATA
Example 2 Example
Draw a Histogram using the data below: In an experiment in which 12 dice were
CLASS FREQUENCY thrown 4000 times to determine the number
50 - 55 2 of times that six occur is in the table below:
55 - 60 3
60 - 65 6 No of sizes Frequency
65 - 70 5 0 428
70 - 75 4 1 1126
75 - 80 1 2 1163
3 797
4 284
Solution 5 112
6 65
7 14
8 10
9 1
10 0
11 0
12 0
Total 4000

Solution

2.2.4 Frequency polygon


If the mid-points of the tops of the bars are
joined, then the resulting shape is called
frequency polygon. e.g

3 MEASURE OF CEN
TRAL TENDENCY
A measure of central tendency is a single value
that attempts to describe a set of data by iden
tifying the central position within that set of
data.

8
3.1 Mean of ungrouped data 3 MEASURE OF CENTRAL TENDENCY

3.1 Mean of ungrouped data


3.1.1 Mean 3.1.2 Mean of grouped data
The mean is the average or a calculated central When data are grouped in classes with
value of a set of numbers and is used to measure corresponding frequencies, such collections are
central tendency of the data. called a grouped data.

Example 1 Example
The following lists the ages of a group of 7 50 students sat for an examination, where the
people. 45,39,53,45,43,48,50. Calculate the maximum mark was 100, and the following
mean age of the group. result were obtained:

Solution Mark Frequency


10 - 19 1
20 - 29 7
30 - 39 8
40 - 49 15
50 - 59 10 60 -
69 5
70 - 79 3
80 - 89 1
Calculate the arithmetic mean

Example 2
Solution
The following scores out of 10 points were
recorded for a pupil in twenty test: Mark Midpoint (x) f fx
8,9,9,2,4,6,1,1,2,4,3,7,9,8,1,3,1,2,5,9 10 - 19 14.5 1 14.5
20 - 29 24.5 7 171.5
(a) Construct a frequency table for the marks. 30 - 39 34.5 8 276.0 40 - 49 44.5 15 667.5
50 - 59 54.5 10 545.0 60 - 69 64.5 5 322.5
(b) Calculate the arithmetic mean.
70 - 79 74.5 3 223.5
80 - 89 84.5 1 84.5
Solution
50 2305.0
x f fx
1 4 4 Pfx 2305
2 3 6 x= =
3 2 6
Pf 50
4 2 8
( a) 5 1 5
6 1 6
3.1.3 Median
7 1 7
8 2 16 Median is one of the three measures of central
9 4 36 tendency which identify the central position of
P f = 20 P fx = 94 the given data. For calculation of median, the
( b)

9
3.1 Mean of ungrouped data 3 MEASURE OF CENTRAL TENDENCY

data has to be arranged in ascending or If n is odd, then


descending order.

Example 1
If n is even, then
The ages of the students in nursing class has
been listed as 42,40,50,65,35,58,32. Find the
median of the given set of numbers.
Solution
Example 1
By rearranging numbers we have the followings
The table below shows the distribution of marks
scored by some students in mathematics test:
data = 65,58,50,42,40,35,32
Marks 22 24 36 42 45 48 56 60
The middle number is 42. Now,
median = 42 Frequency 11 2 7 13 10 3 9 5
Find the median.
or by using the formula when ’n’ is odd
Solution

Marks 22 24 36 42 45 48 56 60
Freq 11 2 7 13 10 3 9 5
C.f 11 13 20 33 43 46 55 60
Since the total number of students is even, then

Example 2
The following lists the ages of a group of 8
people. 45,39,53,45,43,48,50,47. Calculate the
median age of the group.
The 30 th mark is 42 and 31st is also 42
Now,
Solution
If ’n’ is even, then the average of two middle
numbers is the median
Example 2

53,50,48,47,45,45,43,39 Calculate the median age from the following


table:
45 and 47 are the middle numbers Ages 10 12 13 14 16 17 18 19
Freq 7 15 11 7 12 9 4 6
Solution

Ages 10 12 13 14 16 17 18 19
3.1.4 Median from frequency distribution
=

10
3.1 Mean of ungrouped data 3 MEASURE OF CENTRAL TENDENCY

Freq 7 15 11 7 12 9 4 6 Prepare a cumulative frequency table


C.f 7 22 33 40 52 61 65 71
Class boundaries f Cf
Since the members are odd, then
109.5 − 118.5 9 9
118.5 − 127.5 3 12
127.5 − 136.6 4 16
136.5 − 145.5 5 21
145.5 − 154.5 2 23
But 36th falls under 40
154.5 − 163.5 5 28
Now, Median=14 years 163.5 − 172.5 12 40 n = 40
3.1.5 Median from grouped data The formula for Next is to find the median class
calculating the median of a grouped frequency
distribution is given as follows:
But 20th falls in the 21 cumulative frequency

Now,
Where: L1 = 136.5 n
• L1 = Lower class boundary of the median = 40
class.
Fb = 16 f
• n = Total frequency. =5
• Fb = Cumulative frequency before the
median class.
• f = Frequency of the median class.
• c = Size of the median class.

Example
Calculate the median weight from the following
table:

Weight Frequency
110 - 118 9
119 - 127 3
128 - 136 4
137 - 145 5
146 - 154 2
3.1.6 Mode
155 - 163 5 164 - 172
12 The mode of a given data is the item which occur
most often in the distribution it is the value with
the highest number of frequency.
Solution

11
3.1 Mean of ungrouped data 3 MEASURE OF CENTRAL TENDENCY

3.1.7 Mode from ungrouped data 70 − 79 30


80 − 89 17 90 −
Example 99 4
The record of the marks scored by a number of 100 − 109 16
students in an oral test in English is as follows:
10,10,5,9,15,10,20,10,9,5,9,10,25,9,5,25. Solution
Find the model mark.
Class Boundaries Frequency
39.5 − 49.5 9
Solution
49.5 − 59.5 2
Mark 5 9 10 15 20 25 59.5 − 69.5 22 69.5 −
Frequency 3 4 5 1 1 2 79.5 30
From the above table, the highest frequency is 5 79.5 − 89.5 17
and this correspond to mark 10 89.5 − 99.5 4
99.5 − 109.5 16
∴ The mode is 10.
3.1.8 Mode from grouped data The formula for L1 = 69.5 fx = 30
− 22 = 8 fy = 30 − 17
calculating mode from group data was as
= 13 c = 79.5 − 69.5
follows: = 10
fx
Mode L1 + c

Where, fx + fy
• L1 = Lower boundary of the class with the
highest frequency. 8 + 13
80
• fx = Difference between the frequencies of = 69.5 + = 69.5 + 3.8 = 73,3kg
the modal class and the class before it.

• fy = Difference between the frequencies of 21


the modal class and the class after it.
3.1.9 Quartiles
• c = Interval of the modal class.
Quartiles are the values which divide the whole
distribution into four equal parts, namely
Example
• First quartile Q1.
The frequency distribution of the weights of 100
participants in a women conference held at • Second quartile Q2.
MAAUN is as shown below: • Third quartile Q3.
Weights Frequency
40 − 49 9
50 − 59 2
60 − 69 22

12
3.1 Mean of ungrouped data 3 MEASURE OF CENTRAL TENDENCY

• Fourth quartile Q4.

The second quartile invariably is the median


while the fourth quartile is the whole
distribution. The formulas for calculating
quartiles from ungrouped data are as follows:

• (lower quartile)

• ( median )

• (upper

quartile)

• Interquatile = Q3 − Q1
Below are the formulas for calculating quartiles
from grouped data:

13
3.1 Mean of ungrouped data 3 MEASURE OF CENTRAL TENDENCY
Interquartile = Q3 − Q1 = 18 − 5 = 13

Example 2
• Q2 = Median The data below shows the age distribution of
students in Nursing class

Age Frequency
20 − 29 12
Where, 30 − 39 13
• L1 is the lower class boundary of the 40 − 49 39
corresponding quartile 50 − 59 28
• n is the total frequency. 60 − 69 11 70 −
79 7 80 − 89
• Fb Cumulative frequency before the 10
corresponding quartile.
Calculate The first and third quartile for the
• c is the class size of the corresponding distribution.
quartile.
Solution
Example 1
The first step is to prepare a table with the
Find the upper quartile, lower quartile, median respective class limits and their cumulative
and interquartile for the following set of frequency.
numbers: 27,19,5,7,6„9,15,12,18,2,1.
Age Frequency Cum. Freq.
Solution 19.5 − 12 12
29.5
Firstly we put the numbers in order 29.5 − 13 25
39.5
1,2,5,6,7,9,12,15,18,19,27
39.5 − 39 64
Now n = 11 49.5
49.5 − 28 92
3(n + 1)th 59.5
Upper quartile = Q3 = term 59.5 − 11 103
4 69.5
69.5 − 7 110
3(11 + 1)th 36th
79.5
= term = term
4 4 79.5 − 10 120
89.5
= 9th term = 18 Now, n = 120

n + 1th
Median = term 2
But 30th lies in the 3rd cumulative frequency
4.2 deviation

14
4 MEASURE OF DISPERSION

th th
11 + 1 12
= term = term
2 2
= 6th term = 9

Find the range of the data


6,6,7,9,11,13,16,21,and 32.
For Q3 we have
Solution
The maximum is 32 The
minimum is 6

The range = 32 − 6 = 26

4.1.1 Range from grouped data For a grouped


data the range follows in a similar manner with
4 MEASURE OF DISPERSION ungrouped data.

Measure of dispersion describe the spread of Example


the data. they include the range, Interquartile
range, semi interquartile, standard deviation Find the range in the table below:
and variance.
Marks 6- 11- 16- 21- 26-
10 15 20 25 30
4.1 The Range Freq. 3 5 2 6 4
Range is the simplest and most straight forward Solution
measure of dispersion. It is the difference
The maximum mark is 30 The
between maximum and minimum value in the
minimum mark is 6
given data.

Range = 30 − 6 = 24
Example

15
4 MEASURE OF DISPERSION
4.2 Mean deviation The first step is to find the arithmetic mean

The mean deviation is the arithmetic mean of all


absolute deviations from the mean and it can be
calculated by the following: Now,

Where x = mean and N =total number of the


item.

4.2.1 Mean deviation from ungrouped data


Example
4.2.2 Mean deviation with frequencies
Find the mean deviation of the following heights
Example
of plants in an Agricultural laboratory, 2 ,5,6,7,7,
9 Find the mean deviation for the data in the
following table:
Solution
Marks 10 12 13 14 15 16
4.3 Interquartile range and semi interquartile Freq. 9 2 3 6 3 2 range

Solution Example 1
Let D = |x − x |
x f fx D fD The scores below are for a group of 24 students in a statistics
10 9 90 2.6 23.4 test:
12 2 24 0.6 1.2
13 3 39 0.4 1.2 22,10,10,12,12,12,14,14,15,15,17,17,
14 6 84 1.4 8.4
15,14,14,12,12,22,22,26,22,10,19,20. Calculate for the
15 3 45 2.4 7.2
16 2 32 3.2 6.8 distribution.
= 25 = 314 = 48. 2 (a) The range
314 (b) Interquartile range
x= = 12. 6 (c) Semi interquartile
25
P f |x − x | 48. 2
Solution

M.D = = = 1.9 Firstly, we have to re-arrange the scores ac-


N 25 cording to size

4.2.3 Mean deviation of a grouped data 10,10,10,12,12,12,12,12,14,14,14,14,

Example 15,15,15,17,17,19,20,22,22,22,22,26
From the table below, calculate the mean (a) The maximum score is 26 The
deviation of the ages. minimum score is 10
Range = 26 − 10 = 16

16
Ages 2- 6- 10- 14- 18- 4 MEASURE OF DISPERSION
5 9 13 17 21
Freq. 4 3 6 2 5 (b)
Solution 1st = 10,10,10,12,12,12

Let D = |x − x | 2nd = 12, 12, 14, 14, 14, 14


Ages Mid pt(x) f fx D fD 3rd = 15, 15, 15, 17, 17, 19
2-5 3.5 4 14 8.2 32.8
4th = 20, 22, 22, 22, 22, 26
6-9 7.5 3 22.5 4.2 12.6
10-13 11.5 6 69 0.2 1.2 Q 1 = 12, Q 2 = 14, Q 3 = 19 and Q 4 = 26
14-17 15.5 2 31 3.8 7.6 Interquartile range = Q 3 − Q 1
18-21 19.5 5 97.5 7.8 39
= 19 − 12 = 7
=20 =234 =93.2
( c)
P fx 234
x= = = 11. 7 Q3 − Q1 7
P f 20 Semi interquartile = = = 3. 5
P f |x − x | 93. 2 2 2
M.D = = = 4. 7
N 20 Example 2

4.3 Interquartile range and semi The table below shows the distribution of scores by science
students in a MTH1301 test.
interquartile range Marks 20- 30- 40- 50- 60-
29 39 49 59 69
The interquartile range is the difference between the
upper and lower quartiles. Freq. 10 7 3 12 6

Interquartile range = Q3 − Q1 Estimate from the scores The semi


interquartile range is half of the in- (a) The Range
terquartile range. (b) Interquartile range

(c) semi interquartile

17
4 MEASURE OF DISPERSION
4.4 Variance and Standard deviation

Solution Scores
Frequncy C.f 4.4 Variance and Standard deviation
19.5-29.5 10 10 Variance is the arithmetic mean of the squares
29.5-39.5 7 17 of the deviation of the observations from the
39.5-49.5 3 20 true mean, it is also called the mean squared
49.5-59.5 12 32 deviation’
59.5-69.5 6 38 The standard deviation is the square root of the
variance. The standard deviation is also referred
(a) The maximum score is 69 The to as "root mean squared deviation".
minimum score is 20 Range = 69 −
The followings are the formulas for calculating
20 = 49 variance and standard deviation


(b) Interquartile range = Q3 − Q!

which lies under 4th row If the observations has frequencies, then the
formulas are:

L1 = 49.5, n = 38, Fb = 20, f = 12, c = 10 •

4.4.1 Variance and Standard deviation from


ungrouped data

Example
For the following distribution, calculate the
lies in first row variance and standard deviation. 2 ,5,6,7,7, 9.

L1 = 19.5, n = 38, Fb = 0, f = 10, c = 10


Solution

Interquartile = 56.6 = 29 = 27.6

18
P|x − x|2 = (−4) 2 +(−1)2 +02 +12 +12 +32 = 16 + 4.4.3 Variance and Standard deviation for group
data

1 + 0 + 1 + 1 + 9 = 28 Example

The data below is that of shoe sizes of 100 men for


the NYSC.

Size Frequency
50-54 7
4.4.2 Variance and Standard deviation for 55-59 15
frequency data 60-64 44
65-69 13
Example 70-74 9
75-79 12
The table below shows the age distribution of 50
5 MOMENTS
pupils in a secondary school.
Ages 10 12 13 14 15 16
No. of pupils 18 4 6 12 6 4
Solution
Solution
Size x f fx A A2 fA2
50-54 52 7 364 11.9 141.61 991.27 55-59 57 15
Put A = |x − x| 855 6.9 47.61 714.15
60-64 62 44 2728 1.9 3.61 158.84 65-69 67 13 871
x f fx A A2 fA2 3.1 9.61 124.93
10 18 180 2.6 6.76 121.68 70-74 72 9 648 8.1 65.61
12 4 48 0.6 0.36 1.44 590.49
13 6 78 0.4 0.16 0.96 75-79 77 12 924 13.1 171.61
14 12 168 1.4 1.96 23.52 15 6 90 2.4 5.76 2059.32
34.56 100 6390 4939
16 4 64 3.4 11.56 46.24
50 628 228.4

5 MOMENTS
√ Moment is a quantity that expresses the
average or expected value of the first, second,
S.D = 4.6 = 2.1 third or fourth power of the deviation of each
component of a frequency distribution from a
given value, typically mean or zero. The first
moment is the mean, the second moment is the

19
variance, the third moment is skewness and the Grouped data
fourth is kurtosis.
• For Population:
When deviations are computed from
arithmetic mean, then such moments are called Ungrouped data
moments about mean (mean moment) or
sometimes called central moments, denoted by Grouped data
mr and given as follows:
when the the deviation of the values are
computed from origin or zero, then such
• For Sample: moments are called the moments about origin,
Ungrouped data denoted by and are given by:

5.2 moments ungrouped data

• For Sample:Example 2
Ungrouped data Find the first, second, third and fourth mo-
ments about the mean for the set of 2 ,3,7,8, 10
Grouped data
Solution
• For population:

Ungrouped data

Grouped data

5.1 Relation between moments


• For sample: = =0
m1 = 0 m2 = Note: m1 is always equals to zero
m0 − (m0 ) 2
2 1

m3 = m03 − 3m01m02 + 2(m01)3


m4 = m04 − 4m01m03 + 6(m01)2m02 − 3(m01)4 • For Population:

µ1 = 0

20
µ2 = µ02 − (µ01)2
µ3 = µ03 − 3µ01µ02 + 2(µ01)3 µ = Note: m2 is the Variance S2
µ − 4µ0 µ0 + 6(µ0 )µ0 − 3(µ0 )4
0
44 1 3
1 2 1

5.2 Computation of moments from ungrouped data

Example 1
Find the
ments of the set of 2 ,3,7,8, 10 first,
Note: m3 is always skewness.
second,
third and fourth mo-

Solution

Note: m4 is always kurtosis.

Or in tabular form,

Then,

21
Computation of from 5 MOMENTS
5.3 Moments grouped
data
x D D2 D3 D4
2 -4 16 -64 256 3 =3 9 -27
81
7 1 1 1 1
8 2 4 8 16
10 4 16 64 256 0 46 -18
610

PD 0 m 1 = = =
0N 5

PD2 46
m2 = = = 9.2

N 5
PD 3 −18
m3 = = = −3.6

N 5
PD 4 610
m4 = = = 122

N 5 Example 4
Using example 2 and example 3, prove the
Example 3
followings:
Find the first, second, third and fourth moments (a) m2 = m02 − (m01)2
about the origin 4 for the set of (b) m3 = m03 − 3m01m02 + 2(m01)3
2,3,7,8,10 (c) m4 = m04 − 4m01m03 + 6(m10)2m02 −
3(m01)4
Solution
Solution

m1 = 0, m2 = 9.2, m3 = −3.6, m4 = 122 m01 = 2,


m02 = 13.2, m03 = 59.6, m04 = 330

(a)
m2 = m02 − (m01)2

9.2 = 13.2 − 22

9.2 = 13.2 − 4

22
9.2 = 9.2 122 = 330−4(2)(59.6)+6(22)(13.2)−3(24)

LHS = RHS 122 = 330 − 476.8 + 316.8 − 48

(b) 122 = 122


m3 = m03 − 3m01m02 + 2(m01)3
LHS = RHS
−3.6 = 59.6 − 3(2 × 13.2) + 2(2)3

−3.6 = 59.6 − 79.2 + 16 5.3 Computation of


Moments from grouped
−3.6 = −3.6
data
LHS = RHS
Example 1
(c)
From the following data, find the first four
moments about the mean for the marks of
MTH1301.
5.3 Moments grouped data
Marks 61 64 67 70 73
Freq. 5 18 42 27 8
Solution Solution

Assume A = 67, then D = x − A Assume A = 10, then, D = x − A


x D f fD fD 2 fD 3 fD 4 x f D fD fD2 fD3 fD4
61 -2 5 -10 20 -40 80 5 1 -5 -5 25 -125 625 64 -1 18 -18 18 -18 18 6 2 -4 -8 32 -128 512
67 0 42 0 0 0 0 7 5 -3 -15 45 -135 405 70 1 27 27 27 27 27 8 10 -2 -20 20 -20 20
73 2 8 16 32 64 128 10 51 0 0 0 0 0
=100 =15 =97 =33 =253 11 22 1 22 22 22 22
12 11 2 22 44 88
176 PfD 15
0 13 5 3 15 45 135 405 m1 = (c) Pf = 3 100 = 0.45 14 3 4 12 48 192 768

where c is class the interval of x. 15 1 5 5 25 125 625 =131 0 =8 =346 =74 3718
for c = 1 we have
0 1 PfD 8
m1 = (c) Pf = 131 = 0.06

0 2 PfD2 346 m2 = (c) Pf =


131 = 2.64

23
Now, m03 = (c)3 PPfDf 3 = 13174 = 0.56

0 4 PfD4 3718 m4 = (c) Pf =


131 = 28.38

m1 = m01 − m01 = 0

m2 = m02 − (m01)2 = 2.64 − 0.062 = 2.64 m 3 =

m3 = m03 − 3m02m01 + 2(m01)3 m03 − 3m02m01 + 2(m01) 3

= 8.91 − 3(8.73)(0.45) + 2(0.45)3 = 0.56 − 3(2.64)(0.06) + 2(0.06)3


= 8.91 − 11.7855 + 0.18225
= 0.08
= −2.6933
m4 = m04 − 4m03m01 + 6m02(m1)2 − 3(m01)4
m4
= =
m04 28.38−4(0.56)(0.06)+6(2.64)(0.06)2−3(0.06)4

4m03m01 + 6m02(m01)2 − 3(m01)4
= 204.93−4(8.91)(0.45)+6(8.73)(0.45)2−3(0.45)4

= 204.93 − 16.038 + 10.60695 − 0.12301875


= 199.3759

Example 5
Compute the first four moments and measure of Skewness and Kurtosis for the
following data:

X 5 6 7 8 9 10 11 12 13 14 15
F 1 2 5 10 20 51 22 11 5 3 1 6.2 Kurtosis 6 SKEWNESS AND
KURTOSIS
x frequency
0 1
1 5
2 10
3 6
4 3

24
6 SKEWNESS AND KUR-
TOSIS

6.1 Skewness
Is the degree of symmetry or departure from symmetry of a distribution,
skewness=P (x − x3)3 where n = Pf Solution n × S

Coefficient of skewness

6.2 Kurtosis
Is the degree of peakedness of a distribution, 2 3 4x f fx (D) f(D
f(D f(D usually taken relative to a normal distribution.
0 1 0 -2.2 4.84 -10.648 23.43

23
kurtosis1 5 5 -1.2 7.2 -8.64 10.37 moment coefficient of kurtosis
106 2018 -0.20.8 3.840.4 3.072-0.08 0.02 2.46 where 4 3 12 1.8
9.72 17.496 31.49 m = moment and S = standard deviatiion
TOTAL 25 55 26 1.2 67.77
When a distribution goes to the high peak, then is called leptokurtic

When a distribution goes to the flat-topped,

25
then is called platykurtic kurtosis

When a distribution is not very peaked or very flat-topped,


then is called mesokurtic

Example 2
Calculate population skewness and population
Kurtosis from the following grouped data class
frequency
2-4 3
Example 1 4 - 6 4

Calculate population Skewness, population 6 - 8 2 Kurtosis from the following data: 8 -


10 1

26
7.3 Index Number 8 ERRORS
Solution Is a comparison of two quantities which can have
different units. e.g 5 miles per hour or N34 per
square foot.

Examples
class x f (D)fx f(D2 f(D3
A mountain climber is 3200 meters from the
70.28
2-4 3 3 9 -2.2 14.52 - 1 day
31.944 0.01 He climbs 50 meters per hour for 8 hours
peak.
per day. How many meters
4-6 5 4 20 -0.2 0.16 -0.032 6-8 7 2 14 21.00 = 400days will it be before he
reaches the peak?
208.51 day
1.8 6.48 11.664
8-10 9 1 9 3.8 14.44 or 400 meters
SolutionPer day
299.8 to find
Thethe number
per dayof days it will
TOTAL 54.872
rate
10 52 35.6 34.56 take him to reach the top,
divide 3200 by the daily rate
rPf(x − x)2
3200 meters
f(D4 8 meters
population S.D = = 8 days
400 meters
= 50 × per day
n

r35.6 √
= = 3.56 = 1.8868
10

kurtosis

7.3 Index Number


In statistics, an index number is a measure
7 RATE, RATIOS AND INDEX designed to show changes in a variable or a
group or related variables with respect to time,
NUMBERS company performance, prices, productivity or
other characteristics.
7.1 Ratio
Is a comparison of two numbers.A ratio can be
8 ERRORS
written using colon, 3 : 5, or as a fraction 3.
Statistical errors refer to the difference between
5 or 3 to 5, or 3 out of 5. the collected data and actual value of facts. Such
errors may be arise due to the followings:
7.2 Rate • Selection of working samples.

27
DISTRIBUTION
• Incorrect information given by the Exercise
respondents.
1. The theoretical value using physics formula
• Collection of data by estimates. is 0.64 seconds, but Nasiru measures 0.62
seconds as an approximate value. find
• Personal prejudice of investigation.

(a) Absolute error


8.1 Measures of statistical errors • (b) Relative error
Absolute error: An Absolute error is the (c) Percentage error
difference between the true value and its
estimated or approximated value. 2. The mark of a mathematics student in a test
was 91, and it has been recorded as
• Relative error: Relative error is defined as
81. Find
the ratio of the absolute error to the actual
value.
• Uncertainty error: (a) Absolute error of the mark.
(b) Relative error of the mark.
Note: Accuracy, Uncertainty, Variation, (c) Percentage error of the mark.
Deviation all are other names of error. 9 BINOMIAL

Example 1 9 BINOMIAL DISTRIBUTION


The length of a rope 7m was measured as 6.5m.
A Binomial distribution is a probability of success
find
or failure outcome in an experiment or survey
(a) Absolute error
that is repeated multiple times. Binomial is also
(b) Relative error
a type of distribution that has two possible
(c) Percentage error
outcomes, for example a coin has only two
possible outcomes: heads or tails and taking a
Solution test could have two possible outcomes: pass or
fail.
(a)
Note: in probability we have
A.E = |measured value − Actual value| p+q=1 =⇒ p=1−q =⇒ q =
1−p
|6.5 − 7| = 0.5
(b) The Binomial distribution formula is:

pr(x) = nCxpxqn−x
where,
• pr(x) = Binomial probability
(c) • x = total number of success (pass or fail,
P.E = Relative error × 100% heads or tail. e.t.c)
= 0.071 × 100% = 7.1% • p = probability of success of an individual
trial.

28
DISTRIBUTION
• q = probability of failure of an individual
trial.
• n = number of trials.
• C = combinatoris
Note: Binomial distribution formula can also
be written as:

Example 3
The probability of MAAUN students graduating is
Example 1 0.4. Determine the probability that out of five
students
Find the probability of getting 4 heads in 6 tosses
of a fair coin. a) None of the students will graduate.

b) 1 student will graduate.


Solution
c) At least 1 student will graduate.

d) All students will graduate.

Solution
a)
p = 0.4 and q = 1 − 0.4 = 0.6

pr(none will graduate) = 5C0p0q5−0

= 1(0.4)0 × (0.6)5

Example 2 = 0.08
9 BINOMIAL
The overall percentage of failures in a certain
examination is 20. If six candidates appear in the
examination, what is the probability that at least b)
five pass the examination?

Solution
c) pr(at least 1 will graduate)
probability of failure
= 1 − pr(none will graduate)

= 1 − 0.08 = 0.92
probability of atleast five pass = pr(5 or 6) d)

pr(all will graduate) = 5C5p5q5−5

29
DISTRIBUTION
Example 1
If the probability of a defective bolt is 0.1, find (a)
Exercises the mean (b) the variance (c) the standard
deviation for the distribution bolts in a total of
1. Find the probability of getting 3 heads in 5 400.
tosses of a fair coin.

Solution
n = 400, p = 0.1 and q = 1−p = 1−0.1 = 0.9
2. What is the probability of guessing correctly ( a )
at least 6 out of a ’true or false’ test?
mean = µ = np = 400 × 0.1 = 40
(b)
3. If on an average one ship in every ten is
variance = σ2 = npq = 400 × 0,1 × 0.9 = 36
wrecked, find the probability that out of 5
ships expected to arrive, 4 at least will arrive ( c )
safely. √ √ standard
deviation = σ = npq = 36 = 6

Example 2
4. Find the probability that in a family of 5 A die is tossed thrice. A success is getting 1 or 6
children there will be on a toss. Find the mean, variance and standard
deviation of the number of successes.
a) At least 1 girl.
b) At least 1 boy and 1 girl.
Solution
c) At most 2 girls.

Assume the

9.1 Mean, Variance and Standard


deviation
the formula of mean of Binomial distribution is
µ = np the formula of 10 NORMAL
variance of Binomial distribution is
σ2 = npq
Exercises
the formula of standard deviation of Binomial
distribution is 1. Find the expected number of males in a
√σ= random sample of 100 families. find also its
npq standard deviation.

30
DISTRIBUTION

Note:

Expected number of families = Mean


Note: f(x) is called probability density function if
Answer µ = 50 and σ = 5
• f(x) ≥ 0 for every value of x.
2. Find the Binomial distribution whose mean

is 5 and variance is

The normal distribution is given by the equation

3. What is the standard deviation of binomial


distribution with n = 18 and p =
0.4?
where
Answer = 2.08
• µ = mean
4. What is the variance of binomial distribution
with n = 25 and p = 0.35? Round your • σ = standard deviation
answer to 2 decimal places.
• π = 3.14159...
Answer = 5.69 • e = 2.71828...
on subtitution
5. Find the mean, variance and standard
deviation for each of the followings:

a) n = 100, p = 0.75
b) n = 300, p = 0.3 ,
c) n = 20, p = 0.5 standard deviation
now, (2) is called standard form of Normal
d) n = 10, p = 0.8 distribution.

10 NORMAL DISTRIBUTION Example 1


A function f(x) is define as follows:
Is a continuous distribution. It is derived as the
limiting form of the Binomial distribution for
large values of n and p and q are not very small.

show that it is a probability density function. 10


then f(x) is defined as the distribution function.
NORMAL
Let f(x) be a continuous function, then

31
DISTRIBUTION
Solution (b)

If f(x) is a probability density function then

(i) f(x) ≥ 0

(iii) f(x) > 0 for 2 ≤ x ≤ 4

∴ The given function is a probability density


function.

Example 2
The diameter of an electric cable is assumed to
be continuous random variate with probability
density function

f(x) = 6x(1 − x), 0≤x≤1

(a) Show that f(x) is a probability density


function.

(b) Find mean and variance of the function.

Solution
(a) If f(x) is a probability density function then

(i) f(x) ≥ 0

Z1 Z1
= 6x(1 − x)dx = (6x − 6x2)dx
0 0

= (3x2 − 2x3)|10 = 3 − 2 = 1

(iii) f(x) > 0 for 0 ≤ x ≤ 1

32
10.1 Normal curve 10 NORMAL DISTRIBUTION

10.1 Normal curve

The total area under this curve is 1. The area


under the curve is divided into two equal parts
by z = 0. Left hand side area and right hand side

33
10.1 Normal curve 10 NORMAL DISTRIBUTION

area to z = 0 is 0.5. The area between the


ordinate z = 0 and any other ordinate can be
noted from the following table:
.

34
10.1 Normal curve 10 NORMAL DISTRIBUTION

35
10.1 Normal curve 10 NORMAL DISTRIBUTION
Example 1
On the final examination of mathematics, the
mean was 72, and the standard deviation was
15. Determine the standard scores of students
receiving grades.
(a) 60 (b)
93
(c) 72 Area between z = 0 and z = 1.2 is 0.3849

Solution
(a)

(b)
Area between z = − 0. 68 and z = 0 is 0. 2518

(c)

Example 2
Find the area under the normal curve in each of Required area
the following cases:
= (Area between z = 0 and z = 2.21)+
(a) z = 0 and z = 1.2
(Area between z = 0 and z = −0.46)
(b) z = −0.68 and z = 0 = (Area between z = 0 and z = 2.21)+

(c) z = −0.46 and z = 2.21 (d) z = 0.81 and z = (Area between z = 0 and z = 0.46)
1.94
= 0.4865 + 0.1772 = 0.6637
(e) to the left of z = −0.6

(f) to the right of z = −1.28

Solution
Required area

= (Area between z = 0 and z = 1.94)−

(Area between z = 0 and z = 0.81)

36
10.1 Normal curve 10 NORMAL DISTRIBUTION
= when x = 64 then
0.4738
− z = ± 1. 16
0.2910
=
0.1828

Area between 0 and 45 − µ = 0. 50 − 0. 31


σ
Since the area is greater than 0.5, then Area
= 0. 19

Required area

= 0.5 − (Area between z = 0 and z = −0.6)

= 0.5 − 0.2257 = 0.2743

between 0 and z
Required area = 0.8621 − 0.5 = 0.3621
= (Area between z = 0 and z = −1.28) + 0.5
=⇒ z = 1.09
= 0.3997 + 0.5 = 0.8997
Example 4
Example 3 Students of a MAAUN were given an aptitude
Find the value of z in each of the cases test. Their marks were found to be normally
(a) Area between 0 and z is 0.3770 distributed with mean 60 and standard
deviation 5. What percentage of students scored
(b) Area to the left of z is 0.8621 Solution
more than 60 marks?
Solution

37
10.1 Normal curve 10 NORMAL DISTRIBUTION

if x > 60 then, z > 0


Area lying to the right of z = 0 is 0.5
∴ the percentage of students getting more than
60 marks is 50 %

Example 5
In a normal distribution, 31% of the items are
under 45 and 8% are over 64. Find the mean and
standard deviation of the distribution.

Solution when x = 45 Then

from the table when area is 0.19 then z = 0.496

Area between

from table when area is 0.42 then z = 1.405

solving (1) and (2) we get µ = 50 and σ = 10

38
11 t-DISTRIBUTION
The t-distribution is a type of normal distribution
that is used for smaller sample sizes. Normally
distributed data from a bell shape when plotted 11 T-DISTRIBUTION
on a graph, with more observations near the
mean and fewer observations in the tails.
Example 2
The t-distribution known as (students
MAAUN claims that its staffs at the analyst level
distribution) is used to test the significance of (i)
earn an average of ₦500 per hour. A sample of
The mean of a small sample.
30 staffs at the analyst level is selected, and
(ii) The difference between the means of two
their average earnings per hour were ₦450 ,
small samples or to compare two small
with a standard deviation of ₦30. calculate
samples.
tdistribution value.
(iii) The correlation coefficient.

Let x1,x2,x3,...,xn, be the members of random Solution


sample drawn from a normal population with
mean µ. If x be the mean of the sample
then,

where:

• x is the sample mean.

• µ is the population mean.


Example 3
• S is the standard deviation. Mathematics department conduct a test to 50
• n is the size of the given sample. randomly students. And the result they found
from that was, the average score was 120 with a
variance of 121. Assume that the t-distribution
Example 1 is 2.407. what is the population mean for the
Consider the following variables: test?

• population mean = 310


Solution t = 2.407, x = 120, S = 11, n = 50, µ
• standard deviation = 50 =?
• size of the sample = 16

• sample mean = 290

calculate the t-distribution.

Solution

39
µ = 120 − 3.74 = 116.24 A machine which produces mica insulating
washers for use in electric device to turn out
washers having a thickness of 10 mm. A sample
Example 4 of 10 washers has an average thickness 9.52 mm
with a standard deviation of 0.6 mm. Find out t.
12.1 Properties of chi-square 12 CHI-SQUARE DISTRIBUTION

Solution Solution

Chi-square is a measure of actual divergence of the observed and expected frequencies. If


is the observed frequency and fe the expected

frequency of a class interval, then χ2 is defined Female No stop = (3 − 3.36)2 = 0.0386

as3.36
χ2 = 0.0092+0.00089+0.0356+0.01+0.00097+0.0386

. = 0.09526

12.1 Properties of chi-square


• The mean of the distribution is equal to the
number of degree of freedom. i.e µ = v

• The variance is equal to two times the


number of degree of freedom. i.e σ2 = 2 × v.

• when the degrees of freedom are greater


than or equal to 2, then the maximum value
for Y occurs when χ2 = v − 2.

• As the degrees of freedom increase, the χ2


curve approaches a normal distribution.

40
Example 1
Calculate the chi-square value for the following data:

Male Female
Full stop 6(observed) 6( observed )
6.24(expected) 5.76( expected
)
Rolling stop 16(observed) 15( observed )
16.12(expected) 14.88( expected
)
No stop 4(observed) 3( observed )
3.64(expected) 3.36( expected
)

41

You might also like