0% found this document useful (0 votes)
8 views47 pages

Data Description Methods in Statistics

Probability and Statistics course content

Uploaded by

sonaak292
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
8 views47 pages

Data Description Methods in Statistics

Probability and Statistics course content

Uploaded by

sonaak292
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd

Week 2

DISC: 203 – PROBABILITY &


STATISTICS
Methods for Describing Sets of Data

1 Lecturer: Muhammad Asim


METHODS FOR DESCRIBING SETS OF
DATA

 Describing qualitative data


 Describing quantitative data

 Numerical measures to describe data

 Detecting outliers

2
DESCRIBING QUALITATIVE DATA

 Pie Diagram
 Bar Chart

3
EXAMPLE
Construct a pie diagram and a bar chart for highest
degrees attained by the CEOs.
CEO Highest Degree Company
A Masters AZ
B None BZ
C Doctorate CZ
D None DZ
E Bachelors EZ
F Masters FZ
G Bachelors GZ
H Bachelors HZ
I Doctorate IZ
J Bachelors JZ
K Doctorate KZ
L Bachelors LZ
M Masters MZ
N Masters NZ
O None OZ
P Doctorate PZ
Q Doctorate QZ
R Doctorate RZ
S Bachelors SZ
T Masters TZ
U Bachelors UZ
V Bachelors VZ 4
W None WZ
X Bachelors XZ
Y Bachelors YZ
FREQUENCY CHART FOR EDUCATION
ATTAINED BY EXECUTIVES IN THE
SAMPLE

Class Class Relative


Class Frequency Frequency
None 4 0.16
Bachelors 10 0.40
Masters 5 0.20
Doctorate 6 0.24
Total 25 1.00

5
PIE DIAGRAM

6
BAR CHART – CLASS FREQUENCY

7
BAR CHART– CLASS RELATIVE
FREQUENCY

8
DESCRIBING QUANTITATIVE DATA

 Histogram
 Scatterplot

 Comparing central tendency and variability

9
HISTOGRAM

A histogram is a graphical depiction of the


class frequency or class relative frequency
table.

10
EXAMPLE
 A large company is interested in determining
the length of service of its employees.
Twenty five employees are randomly chosen,
and the length of service (in years) is
recorded for each. The data are as follows:
3.1, 1.8, 6.4, 10.2, 11.2, 15.6, 11.6, 6.8, 1.5,
2.9, 3.4,7.2, 0.5, 7.7, 8.4, 0.7, 3.9, 8.2, 8.0,
5.5, 10.3, 12.1, 3.9, 0.9, 4.3
Construct a frequency histogram for the data.

11
HISTOGRAM
 A Class Frequency table for length of service of employees in
the company is:
0.5 0.7 0.9 1.5 1.8 2.9 3.1 3.4
3.9 3.9 4.3 5.5 6.4 6.8 7.2 7.7
8.0 8.2
8.4 10.2 10.3 11.2 11.6 12.1 15.6
Class Class Frequency
0.0- 5.0 11
5.0-10.0 8
10.0-15.0 5
15.0-20.0 1

12
CLASS FREQUENCY HISTOGRAM

13
EFFECT OF SIZE OF DATA SET ON
HISTOGRAM

14
HISTOGRAM

 While histograms provide good visual


representation of data sets (particularly of large
data sets)–they do not let us identify individual
measurements.

15
SCATTER PLOT

16
METHODS FOR DESCRIBING SETS OF
DATA

 Describing qualitative data


 Describing quantitative data

 Numerical measures to describe data

 Detecting outliers

17
NUMERICAL MEASURES
1. Central 2. Variability 3. Using mean and
Tendency  Range standard deviation to
 Mean  Variance
describe data
 Chebyshev’s Rule
 Median  Standard
 The empirical rule
 Mode deviation
 Skewness

4. Relative
standing
 Percentiles and
quartiles
 Z-score
18
COMPARING CENTRAL TENDENCY
VERSUS VARIABILITY
 Consider a histogram with a trend line:

Distribution
Cente
r

19

Spread/ Variability
EXERCISE

 Which car would you buy?

20
MEASURES OF CENTRAL TENDENCY

Mean: The sample mean is defined as:


n

x i
Mean  x  i 1
n
Symbols for sample and population mean:

x Sample Mean
 Population Mean

 is often unknown so x is often used as an


approximater or estimator for 
21
EXAMPLE: SUMMATION NOTATION
 Suppose we have some sample data values
x1=2, x2=3, x3=6, x4=4, x5=6,
• And we want to sum (2+3+6+4+6) or in other words x1+
x2+ x3+ x4 + x5
• We may rewrite this as
5

Where x
i 1
i

x
Or we may rewrite as
i 1
i  x1  x2  x3  x4  x5

2  3  6  4  6 21
5 22
x i 1
i 21
MEASURES OF CENTRAL TENDENCY

Median: Middle number when the measurement in


sample are arranged in ascending (or descending)
order.

Calculating a sample median m

 Arrange n measurements from smallest to largest.


 If n is odd, m is the middle number

 If n is even, m is the mean of middle two number

Example: Calculate m for n={2,4,5,5,6,7,20} and 23

n={4,5,5,6,7,20}
MEASURES OF CENTRAL TENDENCY

 Mode: Most frequently occurring measurement in the


data set
 Modal class: identifies the area in which the data are

most concentrated.
 Example: Find mode in previous example

 Finding mode in quantitative (continuous) data Find

Modal Class Take mid point of modal class as mode

24
n={2,4,5,5,6,7,20}

25
MEASURES OF VARIABILITY
 Range : Largest measurement in data set –
smallest measurement in data set
 Average Absolute Deviation: Average

distance of each measurement from the


mean
 Sample Variance = s2 : for a sample of n

measurements is equal to the sum of


squared distances from the mean divided by
n
(n-1)

2
 (x  x)
i
2

Sample Variance s  i 1
n 1 26
MEASURES OF VARIABILITY
 Sample Standard Deviation = s : is defined
as a positive square root of the sample
variance
s  s2
 Note that for population N

 i
( x   ) 2

Variance  2  i 1
N
Standard Deviation    2
27
MEASURES OF VARIABILITY
 Example
 Sample 1={1,3,5,7,9}
 Sample 2={3,4,5,6,7}
 Standard deviation is a measure of variability. So while
comparing two sample with equal means, but one sample
having higher standard deviation would imply the higher
variability in the sample with higher sample standard
deviation

28
CHEBYSHEV’S RULE

Applies to any set of data regardless of the


shape of the distribution
 No useful information is provided within 1

standard deviation (μ – σ, μ + σ) of the mean.


 At least ¾ of the measurements will fall

within 2 standard deviations of the mean (i.e.


μ – 2σ, μ + 2σ).
 At least 8/9 of the measurements will fall

within 3 standard deviations of the mean (i.e.


μ – 3σ, μ + 3σ).
 Generally, for any number k > 1, at least 1-

1/k2 of the measurements shall be within k 29


deviations of the mean (i.e. μ – kσ, μ + kσ).
EXERCISE

 What percentage of values will approximately


fall between 10 and 26 for a data set with
mean of 18 and standard deviation of 2.

30
EXERCISE

 What percentage of values will approximately


fall between 10 and 26 for a data set with
mean of 18 and standard deviation of 2.
 Answer:

 k = 4. So, at least 93.75% (1-1/16) of

the measurements will fall within the


interval [10, 26].

31
THE EMPIRICAL RULE

Applies to any data with symmetric or mount shape


distribution
 Approximately 68% of the measurements will fall

within 1 standard deviation (μ – σ, μ + σ) of the


mean.
 Approximately 95% of the measurements will fall

within 2 standard deviation (μ – 2σ, μ + 2σ).


 Approximately 99.7% of the measurements will

fall within 3 standard deviation (μ – 3σ, μ + 3σ).

32
EMPIRICAL RULE

33
EXAMPLE
 A manufacturer of automobile batteries claims that the
average length of life for its grade A battery is 60
months. However, the guarantee on this brand is for
just 36 months. Suppose the standard deviation of the
life length is known to be 10 months, and the frequency
distribution of the life-length data is known to be
mound-shaped.
(a) Approximately what percentage of the manufacturer’s
grade A batteries will last more than 50 months,
assuming that the manufacturer’s claim is true?
(b) Approximately what percentage of the manufacturer’s
batteries will last less than 40 months, assuming
manufacturer’s claim is true?
(c) Suppose your battery lasts 37 months. What could 34
you infer about the manufacturer’s claim?
 A manufacturer of automobile batteries claims that the average
length of life for its grade A battery is 60 months. However, the
guarantee on this brand is for just 36 months. Suppose the
standard deviation of the life length is known to be 10 months, and
the frequency distribution of the life-length data is known to be
mound-shaped.
(a) Approximately what percentage of the manufacturer’s grade A
batteries will last more than 50 months, assuming that the
manufacturer’s claim is true? 84%
(b) Approximately what percentage of the manufacturer’s batteries
will last less than 40 months, assuming manufacturer’s claim is
true? 2.5%
(c) Suppose your battery lasts 37 months. What could you infer
about the manufacturer’s claim?

35
MEASURES OF RELATIVE STANDING
zscore: a relative standing measure that
uses mean and standard deviation
x μ x x
Population z score  , Sample z score 
 s
 Suppose a collected salary data with n = 200

ranges from $18,000 to $30,000 with mean


of $24,000 and standard deviation of $2,000.
Joe’s salary is $22,000. What is his z score?
 z = (22,000-24,000)/2,000 = -1.0

 Which tells us that Joe’s salary is 1.0 sample

standard deviation below average salary 36


EMPIRICAL RULE IN TERMS OF Z-SCORE
For mound shaped distributions,
 Approximately 68% of the measurements shall have z

scores between -1 and + 1


 Approximately 95% of the measurements shall have z

scores between -2 and + 2


 Approximately 99.7% of the measurements shall have z

scores between -3 and + 3

34% 34%

2.5% 13.5 13.5 2.5%


% %

37
MEASURES OF RELATIVE STANDING

We may also be interested in relative quantitative


location of a particular data set. (example SAT
scores)
 Percentile Rank: For any n measurements

(arranged descending or ascending), pth percentile


of a number is such that p% of the measurements
are same or fall below pth percentile and (100 – p)%
fall above it.

 Example: (1,4,5,6,7,11,15,20,22,30}
 pth percentile for 22= 90%
38
QUARTILES
 Based on quartiles of data set.
 Quartiles are the values that partition the data set

into four groups, each containing 25% of the


measurements.
 Q is the 25th percentile of the data set. Middle
L
quartile m
is the median and
QU is the 75th
percentile
 Interquartile range

(IQR) = QU-QL
39

QL m QU
METHODS FOR DESCRIBING SETS OF
DATA

 Describing qualitative data


 Describing quantitative data

 Numerical measures to describe data

 Detecting outliers

40
BOX PLOT

Box Plot: useful for comparing different samples and for


detecting outliers.
 Box with hinges at Q and Q ;
L U
 A line inside the box for Median ;
 Whiskers (one from Q to the smallest measurement
L
between the inner fences and the other from QU to
the largest measurement between the inner fences);
 Inner and Outer Fences. Neither visible. Inner at Q -
L
1.5(IQR) and QU + 1.5(IQR) ; and outer at QL - 3(IQR)
and QU + 3(IQR)
 mild outliers (*) : values within and including the
inner fences and outer fences;
 extreme outliers (o): values outside the outer fences.
EXAMPLE – BOX PLOT
Consider the 20 customer satisfaction
ratings:
1, 3, 5, 5, 7, 8, 8, 8, 8, 8, 8, 9, 9, 9, 9, 9, 10,
10, 10, 10
Construct a box plot for this data, identifying
mild and extreme outliers.

42
BOX PLOT

43
FOR NEXT WEEK

 Download R at CRAN (Comprehensive R


Archive Network) website:
[Link]

Then download RStudio:


[Link]

44
RSTUDIO’S LOOK

45
R RESOURCES

 R Manual:
[Link]
/[Link]

 RStudio's own introductory primers:


[Link]

46
TEXT REFERENCE
 Chapter 2

47

You might also like