0% found this document useful (0 votes)
3 views6 pages

Descriptive Statistics and Data Visualization

The document covers various topics in descriptive statistics, including definitions of population and sample, methods for visualizing data (like frequency tables and graphs), and calculations of percentiles, quartiles, and correlation coefficients. It also discusses specific data sets, such as oil reserves and student grades, providing examples of how to represent and analyze the data. Additionally, it explores concepts like the Lorenz curve and Gini index related to income distribution.

Uploaded by

220120025
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views6 pages

Descriptive Statistics and Data Visualization

The document covers various topics in descriptive statistics, including definitions of population and sample, methods for visualizing data (like frequency tables and graphs), and calculations of percentiles, quartiles, and correlation coefficients. It also discusses specific data sets, such as oil reserves and student grades, providing examples of how to represent and analyze the data. Additionally, it explores concepts like the Lorenz curve and Gini index related to income distribution.

Uploaded by

220120025
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Descriptive Statistics

1. Describe what is the meaning of population and sample with examples.

2. Given a data set describe frequency, frequency table, relative frequency,


line graph, bar graph, relative frequency line graph, relative frequency
bar graph, frequency polygon, pie chart , histogram (relative frequency
histogram, frequency histogram), cumulative frequency plot and cumu-
lative relative frequency plot with examples (for each item).

3. The following are the estimated oil reserves, in billions of barrels, for
four regions in the Western Hemisphere.
United States 38.7, South America 22.6, Canada 8.8, Mexico 60.0.
(i) At what angle will the two lines defining the sector (inside the pie
chart) for the United States meet?
(ii) Find the same (angle) for the other three countries.

4. Let frequency (relative frequency) table is given for data set .


(i) Then how to represent this data in a pie chart?
(ii) If a data value had relative frequency r, at what angle would the
lines defining its sector meet?

5. Let the following is the data of CGPA of students in a batch


{8.88, 8.90, 8.93, 8.90, 8.93, 8.96, 8.88, 8.94, 8.96, 8.88, 8.94, 8.99, 8.98}.
Now for the above data set describe (or draw) all the followings :
frequency table, relative frequency table, line graph, bar graph, relative
frequency line graph, relative frequency bar graph, frequency polygon, pie
chart , histogram (relative frequency histogram, frequency histogram),
cumulative frequency and cumulative relative frequency .

6. Let the following data is the set of marks of 20 students in a exam.


{91, 35.5, 40.2, 45, 47, 80, 40.2, 50, 55.5, 60.3, 55, 60, 68, 70, 75.8, 68, 60, 80, 85, 68}
(i) Find the values of 50.3 percentile and 60 percentile for the above
data set.
(ii) What are the percentile for the data value 60.3, 68 and 77 w.r.t
the above data set?, (in other words find the percentiles of the three
students who got 60.3, 68 and 77 marks in the exam.)

7. Let the following data is the set of per annum income (in lakh) of a
group of 15 working people.

1
{6, 8.2, 9.7, 6.3, 11.2, 10.5, 6.5, 8.2, 10.58, 9.75, 12, 10.5, 11.9, 12.5, 17}
(i) Represent the above data in stem and leaf plot (table), (write the
mathematical formula to write data value using stem and leaf for this
stem and leaf plot).
(i-a) Also represent this same data in frequency table as well.
(ii) Find the values of 1st, 2nd and third quartile of the above data
set.
(ii-a) Describe the box plot of a data set.
(ii-b) Now find the box plot for the above data set.
(iii) Find the values of 55 percentile, 72.7 percentile, 60 percentile and
80 percentile for the above data set.
(iii) Suppose three persons have income 8.2 lakh per annum, 12.5 lakh
per annum and 14.3 lakh per annum. Now find the percentile for each
of this three persons w.r.t the above data set.

8. Describe the sample mean, sample median, sample mode with example.

9. Find the formula of the sample mean of a data set that is presented
in a frequency table listing the k distinct values x1 , x2 , · · · , xk having
corresponding frequencies n1 , n2 , · · · , nk .

10. Find the formula of the sample mean of a data set that is presented
in a frequency table listing the k distinct values x1 , x2 , · · · , xk having
corresponding relative frequencies f1 , f2 , · · · , fk .

11. The following data represent the lifetimes (in hours) of a sample of 15
transistors:
{112, 121, 126.7, 108.5, 141, 104, 136, 134, 108.5, 126.7, 112, 108, 112, 126.7, 121}
(i) Write the above data set in stem and leaf plot (table) and also in
frequency table. (write the mathematical formula to write data value
using stem and leaf for this stem and leaf plot)
(ii) Determine the values of the sample mean, median, and mode.

12. Let {x1 , x2 , · · · , xn , · · · , x2n−1 , x2n } = {xi }2n


i=1 is a data set of size 2n.
the values of the data set are may not be in increasing or decreasing
order. Let A ⊂ {1, 2, · · · , n, · · · , 2n} and all the elements of A are
distinct and the total number of elements of A is n (cardinality of the
set A is n). Denote Ac = {1, 2, · · · , n, · · · , 2n}\A. Let x̄A is the sample
mean of the data set {xi }i∈A and x̄Ac is the sample mean of the data

2
set {xi }i∈Ac . Find the sample mean x̄ of the original data set {xi }2n
i=1 ,
in terms of x̄A and x̄A .
c

13. Let {x1 , x2 , · · · · · · , x198 } is a data set consisting of 198 values and xj ≤
xj+1 ∀ j. Let the sample mean of the initial 99 values {x1 , x2 , · · · , x99 } 
is 100 whereas the sample mean of the final 99 values {x99 , x100 , · · · , x198 }
is equal to 120.
(i) what is the value of the sample mean for the entire data set
{x1 , x2 , · · · · · · , x198 }.
(ii) Prove that the value of sample median lies in the open interval
(100, 120).

14. Describe the sample variance and sample standard deviation with ex-
amples.

15. Let the following data set represent the temperature (in Centigrade)
for 12 cities on a given day.
{9.2, 14.1, 9.8, 12.4, 16.0, 12.6, 9.2, 18.9, 14.1, 14.5, 20.4, 16.9}
(i) Write the above data set in stem and leaf plot (table) and also in
frequency table. (write the mathematical formula to write data value
using stem and leaf for this stem and leaf plot)
(ii) Find the values of sample mean, sample median, sample variance
, sample standard deviation.
(iii) Find the values of 25 percentile and 60 percentile.
(iv) Suppose two cities A and B has temperature (in Centigrade) 12.6
and 20 find the percentile for both the city A and B w.r.t the above
data.

16. X = {x1 , x2 , · · · , xn } and Y = {y1 , y2 , · · · , yn }, where yi = c + dxi ∀ i,


where c, d ∈ R.
(i) Find the sample mean and sample variance for both the data set X
and Y . Deduce the relation (mathematical formula ) between sample
mean of X and sample mean of Y .
(ii) Also find the relation (mathematical formula ) between sample
variance of X and sample variance of Y .

17. Describe the scatter diagram of data sets that consist of pairs of values
with example.

3
18. Describe the sample correlation coefficient of data set that consist of
pairs of values with example.

19. When we can say that the sample data pairs are positively correlated
or negative correlated ?

20. Let X = {xi }10 i=1 = {x1 , x2 , x3 , x4 , x5 , x6 , x7 , x8 , x9 , x10 } =


{25, 36, 40.1, 26.2, 28.5, 42.3, 30, 31.5, 37.6, 35} represent the temperature
(in Centigrade)of Dharwad of for a given 10 days in the summer season
and Y = {yi }10 i=1 = {y1 , y2 , y3 , y4 , y5 , y6 , y7 , y8 , y9 , y10 } =
{55, 95, 110, 62, 70, 120, 78, 82, 100, 90} represent the unit of consumed
electricity by a house (in Dharwad) on this same 10 days.
(i) Now find the value of the sample correlation coefficient of data set
that consist of pairs of values (x1 , y1 ), (x2 , y2 ), · · · , (x9 , y9 ), (x10 , y10 ) .
(ii) Draw
 the scatter diagram for the data set that consist of pairs of
values (x1 , y1 ), (x2 , y2 ), · · · , (x9 , y9 ), (x10 , y10 ) .

21. Let X = {x1 , x2 , x3 , x4 , x5 } = {4, 10, 8, 6, 2} represent the average hours


(per day) spent by five students individually watching entertainment
on the internet. Y = {y1 , y2 , y3 , y4 , y5 } = {8, 2, 4, 6, 10} represent the
grades of those five students in a semester.
(i) Now find the value of the sample correlation coefficient of data set
that consist of pairs of values {(xi , yi )}5i=1 .
(ii) Draw the scatter diagram for the data set that consist of pairs of
values {(xi , yi )}5i=1 .

22. When we can say that the sample data pairs are positively correlated
or negative correlated ?

23. 
Let r be the sample correlation coefficient of the a data pairs given by
(xi , yi ) : i = 1, 2, · · · , n . Then prove that −1 ≤ r ≤ 1.

24. Let yi = a + bxi , i = 1, 2, · · · , n, a ∈ R, b > 0 then prove that


r = 1, here r is the sample correlation coefficient of the data a pairs
(xi , yi ) : i = 1, 2, · · · , n . Now do the same calculation when b < 0
and a ∈ R.
k
X
1
25. Let {x1 , x2 , · · · , xn } is a data set of size n. Let x̄k = k
xi be the
i=1
sample mean of the first k data values, here 1 ≤ k ≤ n. For k ≥ 2 we

4
define the sample variance of the first k data values as
k
1 X
s2k = (xi − x̄k )2 , here 2 ≤ k ≤ n and define s21 = 0.
(k − 1) i=1

(i) Prove that


1 
x̄k+1 = x̄k + xk+1 − x̄k .
k+1
(ii) Using above now prove that
 
2 1 2 2
sk+1 = 1 − sk + (k + 1) x̄k+1 − x̄k .
k

26. Describe what is normal data set, normal histograms and approximately
normal with examples.

27. Explain the Lorenz curve (L) and Gini index (G) of income (or wealth)
of working members of a population.

28. Let the incomes (in INR Lakh) of a group of five people is given by the
set {9, 7, 22, 5, 17}. Draw the Lorenz curve and also find the value of
the Gini index for this income data.

29. Let there are n numbers of working people and there income (non-
negative)is given by the set {x1 , x2 , · · · , xn }, xj ≤ xj+1 , j = 1, 2, · · · , n.
Let L nj is the value of the Lorenz curve at the point nj for j =
1, 2, · · · , n, in other words
 
j x1 + x2 + · · · + xj
L = , j = 1, 2, · · · , n.
n x1 + x2 + · · · + xn−1 + xn

(i) Proved (mathematically) that


x1 + x2 + · · · + xj x1 + x2 + · · · + xj+1
≤ , j = 1, 2, · · · , n.
j j+1

(ii) Proved (mathematically) that


 
j j
L ≤ , j = 1, 2, · · · , n.
n n

5
(iii) Proved (mathematically) that the Gini index G can be given by
1 s1 + s2 + sn−1
G=1− −2 ,
n nsn
in the above we denote sj = x1 + x2 + · · · + xj for j = 1, 2, · · · , n.

You might also like