Descriptive analytics
1. Frequency Distributions
2. Graphs
3. Measures of Central Tendency, Variability, Degrees of Freedom,
IQR
4. Normal Distribution
5. Scatterplot; Correlation
6. Least Square Solutions
7. Regression
SARANKUMAR K - CIT CHENNAI 1
Frequency distribution
Detect data patterns and make sense out of data
Collection of observations produced sorting observations into
classes showing the frequency of occurrence in each class
SARANKUMAR K - CIT CHENNAI 2
Frequency distribution
Frequency Distribution For Quantitative Data:
Ungrouped more than 100 values only can be partially
displayed
observations are sorted into classes of single values
Top class includes largest observation and Bottom Class
includes smallest observation
Much more informative when it is less than 20 values
SARANKUMAR K - CIT CHENNAI 3
Frequency distribution
Grouped class intervals with 10 values
Range = difference between the largest and smallest observations
Class interval = range / desired number of classes round off to
nearest convenient interval
SARANKUMAR K - CIT CHENNAI 5
Frequency distribution
Gap between classes always equal one unit of measurement
Never bigger than one unit of measurement
Guidelines one, and only one, class; not overlap mention zero
frequencies also upper boundary and lower boundary equal
intervals
SARANKUMAR K - CIT CHENNAI 7
Frequency distribution
Real Limits:
Located at the midpoint of the gap between adjacent tabled
boundaries
Actual Width = ½ of 1 unit of measurement below the lower tabled
boundary – ½ of 1 unit of measurement above the upper tabled
boundary
SARANKUMAR K - CIT CHENNAI 8
Frequency distribution
Cumulative Frequency showing total no. of observations in each
class from lowered ranked classes
Cumulative Relative Frequency frequency of each class by total
frequency
SARANKUMAR K - CIT CHENNAI 9
Frequency distribution
Cumulative Percentage cumulative frequency of each class by Total
cumulative frequency x 100%
SARANKUMAR K - CIT CHENNAI 10
graphs
Histogram bar-type graph
common boundaries between adjacent bars emphasizes the
continuity of data with continuous variables
SARANKUMAR K - CIT CHENNAI 12
graphs
Frequency Polygon line-type graph
Also emphasizes the continuity of continuous variables
SARANKUMAR K - CIT CHENNAI 14
graphs
Stem and Leaf Display
Sorting on the basis of leading and trailing digits
SARANKUMAR K - CIT CHENNAI 17
graphs
SARANKUMAR K - CIT CHENNAI 19
graphs
Bar graph bar-type
Gaps between two adjacent bars emphasize the discontinuous
nature of the data
SARANKUMAR K - CIT CHENNAI 21
Central tendency and
variability
Measure of central tendency central value of given distributions
Describing Variability
How far apart scores are from the mean
How far apart scores are from each other
Measures of variability measures the spread of observations
Range difference b/w max and min score
Variance
Standard Deviation
Degrees of freedom number of values free to vary one or
more mathematical restrictions in sample to get population
characteristic
SARANKUMAR K - CIT CHENNAI 24
variability
Positively skew mean value > median value
Negatively skew mean value < median value
SARANKUMAR K - CIT CHENNAI 27
Inter-quartile range (iqr)
Measures Range covered by Middle 50%
Steps to find IQR:
Arrange in Ascending Order
Find Median
Find Median for lower 50% of data mark it as Q1
Find Median for upper 50% of data mark it as Q3
Find IQR Q3 – Q1
SARANKUMAR K - CIT CHENNAI 28
Normal distribution
Between Mean and z score B, B’
Below or above z score C, C’
Between x and y (not included mean) values are subtracted
Between x and y (included mean) values are added
More than p points above or below mean C, C’
Within q points above or below mean B, B’
Finding one score 20% - 0.2000; 5% - 0.0500 (add two zeros at
end) write approx. z score upper, lower(C, C’)
SARANKUMAR K - CIT CHENNAI 29
Scatter plots
Cluster of dots represent all pairs of scores
SARANKUMAR K - CIT CHENNAI 30
Correlation coefficient
describe relationship between pair of variables by no. b/w -1 and
+1 Pearson Correlation coefficient (r)
0.2 – very low; 0.4 – low; 0.6 – moderate; 0.8 – high; 1 – very high
SARANKUMAR K - CIT CHENNAI 31
regression
Finding relationship between two continuous variables one is
independent variable and other is dependent variable
SARANKUMAR K - CIT CHENNAI 37
Self evaluation time !
SARANKUMAR K - CIT CHENNAI 38
SARANKUMAR K - CIT CHENNAI 39
SARANKUMAR K - CIT CHENNAI 40
SARANKUMAR K - CIT CHENNAI 41