0% found this document useful (0 votes)
4 views19 pages

Standard Deviation and Variability Explained

Chapter 4 covers descriptive statistics, focusing on numerical descriptions, measures of center, variability, and standardized data. It discusses various statistical measures such as mean, median, mode, variance, and standard deviation, along with concepts like percentiles, quartiles, and box plots. Additionally, it addresses correlation, covariance, and the characteristics of skewness and kurtosis.

Uploaded by

24007614
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views19 pages

Standard Deviation and Variability Explained

Chapter 4 covers descriptive statistics, focusing on numerical descriptions, measures of center, variability, and standardized data. It discusses various statistical measures such as mean, median, mode, variance, and standard deviation, along with concepts like percentiles, quartiles, and box plots. Additionally, it addresses correlation, covariance, and the characteristics of skewness and kurtosis.

Uploaded by

24007614
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

CHAPTER 4: Descriptive Statistics

CHAPTER 4: Descriptive Statistics.............................................................................. 1


I. 4.1 Numerical Description................................................................................. 3
1. Center:........................................................................................................... 3
2. Variability = dispersion.................................................................................. 3
3. Shape = distribution (phân bố)......................................................................3
II. 4.2 Measures of Center..................................................................................... 3
............................................................................................................................. 3
1. Mean: a familiar measure of center...............................................................3
2. Median: middle value..................................................................................... 3
3. Mode.............................................................................................................. 4
4. Shape............................................................................................................. 4
5. Geometric Mean: multiplicative average (giá trị trung bình nhân)................5
6. Midrange........................................................................................................ 6
7. Trimmed Mean............................................................................................ 6
8. Overall.............................................................................................................. 8
1. Variation......................................................................................................... 8
is the “spread” of data points about the center of the distribution in a sample.
Consider the following measures of variability.....................................................8
2. Range............................................................................................................. 9
3. Variance and Standard Deviation...................................................................9
4. standard deviation..................................................................................... 9
4.1. Two-sum formula.................................................................................. 10
4.2. Characteristics of the standard deviation............................................10
5. Coefficient of Variation..............................................................................11
6. Mean Absolute Deviation........................................................................11
II. 4.4 Standardized Data.................................................................................... 12
1. Chebyshev's Theorem.................................................................................. 12
1.1. Definition............................................................................................. 12
1.2. Formula................................................................................................ 13
1.3. Example............................................................................................... 13
1.4. Characteristics..................................................................................... 13
2. Empirical...................................................................................................... 13
2.1. Definition:............................................................................................ 13
2.2. Definition:............................................................................................ 14
3. Transform a data set into standardized values............................................14
4. Estimating data (to calculate the standard deviation by only knowing the
range)................................................................................................................ 15
III. Percentiles, Quartiles, and Box Plots............................................................15
1. Percentile..................................................................................................... 15
1.1. Percentiles: 100 groups........................................................................16
1.2. Deciles: 10 groups............................................................................... 16
1.3. Quintiles: 5 groups............................................................................... 16
2. Quartiles: 4 groups...................................................................................... 16
2.1. Median: Q2............................................................................................... 16
2.3. Method of medians:.............................................................................17
3. Box plots...................................................................................................... 19
3.3. Fences and Unusual Data Values.........................................................19
3.4. Midhinge.............................................................................................. 20
IV. 4.6 Correlation and Covariance....................................................................20
V. 4.7 Grouped Data............................................................................................ 20
VI. 4.8 Skewness and Kurtosis...........................................................................20
I. 4.1 Numerical Description
1. Center:
2. Variability = dispersion
3. Shape = distribution (phân bố)
II. 4.2 Measures of Center

1. Mean: a familiar measure of center

2. Median: middle value


• If n is odd, the median is the middle observation in the ordered data set.
• If n is even, the median is the average of the middle two observations in the ordered data
set.
3. Mode
• The most frequently occurring data value.
• May have multiple modes or no mode.
• The mode is most useful for discrete or categorical data with only a few distinct data
values. For continuous data or data with a wide range, the mode is rarely useful.
4. Shape
• Compare mean and median or look at the histogram to determine degree of
skewness.
5. Geometric Mean: multiplicative average (giá trị trung bình nhân)
a. The growth rates
6. Midrange
• The midrange is the point halfway between the lowest and highest values of
X. Easy to use but sensitive to extreme data values

7. Trimmed Mean

 The trimmed mean is calculated like any other mean, except that the highest and lowest
k percent of the observations in the sorted data array are removed. The trimmed mean
mitigates the effects of extreme high values on either end.

 For example, for the n = 33 P/E ratios, we want a 5 percent trimmed mean (i.e., k = .05).
To determine how many observations to trim, multiply k by n, which is 0.05 x 33 = 1.65 or 2
observations.
So, we would remove the two smallest and two largest observations before averaging the
remaining values.
Cách tính trim mean khi loại bỏ 20%:
1. Bước 1: Sắp xếp tập dữ liệu (nếu chưa sắp xếp):
{1,2,3,4,5,100}
2. Bước 2: Xác định số lượng phần tử cần loại bỏ:
o Tổng số phần tử của tập dữ liệu là 6.
o 20% của 6 là: 0.2×6=1,2 và làm tròn xuống là 1 phần tử.
o Điều này có nghĩa là sẽ loại bỏ 1 phần tử nhỏ nhất và 1 phần tử lớn nhất.
3. Bước 3: Loại bỏ 1 giá trị nhỏ nhất và 1 giá trị lớn nhất:
o Loại bỏ giá trị nhỏ nhất: 1
o Loại bỏ giá trị lớn nhất: 100
Các giá trị còn lại là:
{2,3,4,5}
4. Bước 4: Tính trung bình của các giá trị còn lại:

Kết luận:
 Trim mean khi loại bỏ 20% (tức là 1 giá trị ở mỗi đầu) của tập dữ liệu {1, 2, 3, 4, 5, 100}
là 3.5.
8. Overall

III. Measure of variability


1. Variation
is the “spread” of data points about the center of the distribution in a sample. Consider the
following measures of variability
 Variance is a statistical measure that indicates how much individual data
points in a dataset differ from the mean (average) of that dataset. It
quantifies the spread or dispersion of the data. A higher variance means the
data points are more spread out, while a lower variance indicates they are
closer to the mean. (Variance show how much data differ from the MEAN)

2. Range
3. Variance and Standard Deviation

4. standard deviation
 In describing variability, we most often use the standard deviation
 The standard deviation is a single number that helps us understand how
individual values in a data set vary from the mean. Because the square root
has been taken, its units of measurement are the same as X (e.g., dollars,
kilograms, miles). To find the standard deviation of a population we use:
4.1. Two-sum formula

[Link] of the standard deviation


a. The standard deviation is nonnegative
b. When every observation = mean, then the standard deviation is zero (i.e., there is no
variation).
For example, if every student received the same score on an exam, the numerators
of formulas would be zero because every student would be at the mean. At the
other extreme, the greatest dispersion would be if the data were concentrated at x
min and xmax (e.g., if half the class scored 0 and the other half scored 100).

c. can be compared only for data sets measured in the same units
example, prices of hotel rooms in Tokyo (yen) cannot be compared with prices of
hotel rooms in Paris (euros)

d. Also, standard deviations should not be compared if the means differ substantially
5. Coefficient of Variation
6. Mean Absolute Deviation
 This statistic reveals the average distance from the center
 The MAD is appealing because of its simple, concrete interpretation

Overall:
II. 4.4 Standardized Data
1. Chebyshev's Theorem
1.1. Definition
 for any data set, no matter how it is distributed, the percentage of
observations that lie within k standard deviations of the mean
 giúp mô tả tỷ lệ phần trăm dữ liệu nằm trong một khoảng cách xác định tính
từ giá trị trung bình

[Link]

[Link]

[Link]
 Don’t need to visual the shape (không cần giả định về hình dạng của phân
phối).
 Nó giúp cung cấp một mức độ tin cậy về tỷ lệ phần trăm dữ liệu nằm trong
các khoảng cách khác nhau so với giá trị trung bình
2. Empirical
2.1. Definition:
 it says that for data from a normal distribution, we expect the interval μ 6 kσ
to contain a known percentage of the data:
 Normal distribution: is symmetric and is also known as the bell-shaped curve.

[Link]:

3. Transform a data set into standardized values.


 A standardized variable (Z) redefines each observation in terms of the
number of standard deviations from the mean.
 vị trí tương đối của từng giá trị trong một tập dữ liệu so với trung bình và giúp
so sánh dữ liệu từ các nguồn khác nhau dễ dàng hơn
4. Estimating data (to calculate the standard deviation by only knowing the range)
[Link], Quartiles, and Box Plots
1. Percentile
• Percentiles are data that have been divided into 100 groups.

[Link]: 100 groups


[Link]: 10 groups
[Link]: 5 groups
2. Quartiles: 4 groups
are scale points that divide the sorted data into four groups of approximately equal
size.

2.1. Median: Q2
2.2. Q1, Q3: dispersion
 the interquartile range Q3 – Q1 measures the degree of spread in the middle
50 percent of data values.
[Link] of medians:
Step 1: Sort the observations.
Step 2: Find the median Q2.
Step 3: Find the median of the data values that lie below Q2.
Step 4: Find the median of the data values that lie above Q2.
3. Box plots

[Link] and Unusual Data Values

Inner fences Outer fences:

Lower fence Q1 – 1.5 (Q3 – Q1) Q1 – 3.0 (Q3 – Q1)

Upper fence Q3 + 1.5 (Q3 – Q1) Q3 + 3.0 (Q3 – Q1)


[Link]

IV. 4.6 Correlation and Covariance


V. 4.7 Grouped Data
VI. 4.8 Skewness and Kurtosis

You might also like