0% found this document useful (0 votes)
15 views12 pages

Probability and Statistics Concepts Explained

Mitwpu ps imp unit 1-2
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
15 views12 pages

Probability and Statistics Concepts Explained

Mitwpu ps imp unit 1-2
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Probability and Statistics

Unit 1: Diagram
1) Classification :
[Link] and it’s types
i. Create a Frequency Table using class interval of width 10 starting from
150-159. The Heights (in cm) of 15 students
[ 150,155,160,166,170,172,133,175,177,180,181,182,185]
ii. Create a Frequency distribution Table using class interval of size 10
starting from 30-40 .The marked scored by 20 students in a test is
[ 33,45,39,41,47,30,49 ,50,55,60, 66 ,70,75,81,85,89 ,90,92,95,99 ]
2) Diagram
a. Histogram
i. Draw histogram :
Class-interval Frequency
0-5 2
5-10 5
10-15 8
15-20 4

b. Frequency Curve & Frequency Polygon


i. An axis , consider mid-point on the class & for Y-axis draw the frequency
u 1−u 2
curve & Frequency polygon : mid− point = 2
Class-interval: Class-interval Frequency
0-10 4
10-20 6
20-30 10
30-40 5
40-50 3

c. Bar graph
i. Draw bar graph :
Subject Students
Math 40
Science 35
History 25
Geography 30
English 45
Probability and Statistics

d. Double Bar graph


i. Draw Bar-graph:
Year Product A Product B
2022 150 100
2023 200 180
2024 250 240

e. % Bar Diagram
i. Draw %-Bar diagram Category 2023 2024
Food 20000 25000
Travel 10000 15000
Utilities 5000 7000

f. Sub-divided Bar graph


i. Draw sub-divided bar graph
Category 2023 2024
Food 20000 25000
Travel 10000 15000
Utilities 5000 7000

3) Measures of central tendency :


a. Definition with its type
4) Mean
a. Definition
b. Formula : grouped: mean ¿ Ungrouped: mean ¿
c. Question :
i. Find the arithmetic mean of numbers : 5,8,10,12,15
ii. Find the arithmetic mean of
Values (x) Frequency ( f ) Values (x) Frequency ( f )
4 10 2 3
5 8 4 5
9 10 6 4
2 9 8 2
Probability and Statistics

iii. Calculate the mean from the following frequency distribution :


Class interval Frequency Class interval Frequency
0-10 4 0-10 4
10-20 6 10-20 6
20-30 10 20-30 10
30-40 5 30-40 5
40-50 5 40-50 5

5) Median :
a. Definition
(n+1)
b. Formula : If n is odd midpoint can be would easily by median=
2

If even then remove mean of 2 midpoint :


median=
n
2 ()
thobs +
n
2 ()
th obs

2
c. Question :
i. 6,8,10,12,15,18,20
ii. 12,18,12,14,18,20,14,22,14
iii. 5,9,13,15,18,20
iv. 6,12,13,20,24,35
d. Median for Discrete frequency distribution :
Step 1: Arrange the data in order if required .
Step 2: Prepare a column of less than cumulative frequency .
Step 3: Locate the obs. At which nth is getting crossed for the first time in
the column of CF (Cumulative frequency ).
Step 4: Corresponding frequency is median
i. Question :

Values (x) Frequency (f)


Values (x) Frequency (f)
5 2
10 2
10 4
20 3
15 6
30 5
20 25
40 4
25 3
50 1
Probability and Statistics

e. Median for continuous frequency distribution ( Class interval) :

(
n
)
Formula : median=l+ 2 −cf ∗h

Class interval Frequency


0-10 5
10-20 8
20-30 15
30-40 12
40-50 10

6) Measures of Dispersion :
a. Definition of Measures of Dispersion .
b. Range :
Range is the difference between the largest & Smallest values in a
dataset . Formula : Range=maximum values−minimum values
c. Coefficient of range :
Coefficient of range is a relative measures of dispersion . It helps compare the
variability between different data sets .
maximum−minimum H −L
Formiula :Coefficient of range= maximum+ minimum = H + L
d. Question :
i. Calculate the range and coefficient of range for the given data :
5,9,12,15,25
ii. 4,6,8,9,20,50,55
iii. Class interval
Class interval Frequency
0-10 5
10-20 8
20-30 12
30-40 6
40-50 4

A. Standard deviation (SD) σ :


:- Standard deviation is a measure of the deviation or spread of a dataset . It
tells how much the values deviate from the mean .
Probability and Statistics


2
Formula : 1) Ungroup data sd ( σ )= ε ( xi−x )
n
Where ,
Xi = individual value
x = Mean , n = Number of observation


2
After simplifying : σ = ε x −¿ ¿
n
2) Class interval and frequency : σ =√ εf ¿ ¿ ¿
Where , x = Midpoint of class interval .
f = Frequencies
x=¿Mean
For the Frequency distribution :
σ=

Question :
εf2
εf
∗¿

iv. Find the SD , CV


 5,10,15,20,25
 13,18,19,11,9,12
 Find SD , CV
Class interval Frequency (f)
10-20 3
20-30 5
30-40 7
40-50 5

7. Moments :
a. Definition :
Moments are statistical measures that describe the shape of distribution .
There are 2 types of Moments :
A) Raw Moments
B) Central Moments
μr =ε ¿ ¿

μr =ε ¿ ¿

r =1,2,3,4
Probability and Statistics
st
1) Calculate 1 four central moments for the following data
i. 2,4,6,8
ii. 1,1,2,3,5
2) X = 1,2,3,4
F = 3,5,4,2
Class interval Frequency
0-10 5
10-20 9
20-30 4
30-40 2

8. Skewness :
Definition : Skewness measures the asymmetric distribution .
μ3
Formula : Skewness :r 1=
¿¿¿
Interpretation :
r 1=0→ Perfectly symmetric distribution
r 1 >0 → positively skewed
r 1 <0 → negatively skewed

Question :
Q1) X = 3,7,8,9
F = 2,1,3,2
Q2)
Class interval Frequency
0-5 4
5-10 6
10-15 10
15-20 7
20-25 3

9. Kurtosis ( β 2 )
a. Kurtosis measures the peakedness or flatness of a distribution .
μ4
Formula : kurtosis ( β 2 )=
¿¿¿
Types :
 Mesokurtic ( β 2=3 ) → Normal distribution shape
Probability and Statistics
 Leptokurtic ( β 2 >3 ) → More peaked , heavely tails
 Platykurtic ( β 2 <3 ) → Flatter peak , light tails )
Q1) Calculate classification of kurtosis for the following
X f data :
2 1
1) 4 3 3,4,2,7
2) 6 5
8 2
10 1

[Link]
a. Definition :
Quartiles are 3 observation Q1 , Q2 And Q3 that divide , entire sorted data
into 4 equal parts , similarly percentiles are into 99 observation which will
divide the sorted data into 100 equal parts .
Q1=¿

Q1) Calculate the quartiles for the following :


6,8,10,12,14,16,18
Q2)
X F
10 2
20 3
30 5
40 4
50 1

Q3) Class interval Frequency


0-10 5
10-20 8
20-30 15
30-40 12
40-50 10
Probability and Statistics

b. Calculate P38 & P67


[Link] :

i. 5,8,9,12,18,20,22
b. Calculate P29 & P83
ii. X = 16,18,20,22,24
F = 12,20,38,20,10
c. Find P42 and P64
Iii Clas interval Frequency
10-20 7
20-30 12
30-40 21
40-50 6
50-60 4
Probability and Statistics

Unit 2: Correlation Analysis and Regression


1) Correlation And Regression
a. Definition of simple correlation & it’s type with scattered diagram .
b. Definition of Coefficient correlation and it’s type
c. Karl’s Perarson’s correlation Coefficient ( ‘r’ )
i. Definition :
Karl pearson’s correlation coefficient (r)is a numerical value that
measures the strength & direction of the linear relationship
between two quantitative variables.
nεxy−εx. εy
Formula : r =
√¿ ¿ ¿

Interpretation :
r =+ 1→ perfect positive relation

r =−1→ Perfect negative relation


r =0 → No linear relation

Question :
i. Find the karl pearson’s correlation coefficient (r)
Q1. Q2. Q3. Q4.
X y X y X y X y
10 20 40 21 9 8 40 21
12 25 43 29 8 6 43 29
14 28 48 25 4 9 48 25
16 30 50 35 10 15 50 35
18 31 54 40 12 18 54 40
20 33 58 42 58 42
Probability and Statistics

d. Spearman’s Rank correlation coefficient ( R)


i. Definition :
Spearman’s rank correlation coefficient (R) is a measure of the
strength and direction of the relationship between two variables.
2
6ε d
Rank ( R )=1− → for no repetation
n ( n2 −1 )

Rank ( R )=1−6−¿ ¿ -> For rank repetition

Q1. Q2. Q3.


Mark 1 Mark 2 X Y X Y
90 85 87 18 87 18
67 87 76 16 76 16
84 54 65 19 65 19
79 89 59 12 76 16
75 68 43 15 43 16
29 8 29 8

Q4.
X Y
45 35
70 80
65 70
40 40
80 90
40 45
50 60
70 80
85 90
60 50
Probability and Statistics
e. Regression Analysis :
The technique of predicting the given value of other variable when 2
variable are co-relate is called regression analysis.

Equation of Yon X :
Y −Y =byx ( X− X )
Equation of X on Y :
X −X =byx ( Y −Y )

When bxy and byx are regression coefficient calculated as follow :


σ y cov (x , y) nεxy−εx . εy
byx=r . = =
σx σx
2 √¿¿¿

σ x cov ( x , y) nεxy−εx . εy
b xy=r . = =
σy σy
2 √¿¿¿

Q1. Find the regression equation &estimate Y when X=2.5 & X when Y=5.2
X Y
2 3
6 5
4 6
8 9
10 12
Q2. Find the regression equation &estimate Y when X=22 & X when Y=20
X Y
35 23
25 27
29 26
31 21
27 24
24 30
33 30
Probability and Statistics

You might also like