0% found this document useful (0 votes)
12 views81 pages

Data Representation and Analysis Techniques

The document outlines research methodology, focusing on data representation, population and sample, and types of data. It explains statistical inference, the distinction between known statistics and unknown parameters, and various data visualization techniques such as bar charts, histograms, and pie charts. Additionally, it discusses the importance of data collection and interpretation to reveal meaningful information and patterns.

Uploaded by

Zuha Irshad18
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
12 views81 pages

Data Representation and Analysis Techniques

The document outlines research methodology, focusing on data representation, population and sample, and types of data. It explains statistical inference, the distinction between known statistics and unknown parameters, and various data visualization techniques such as bar charts, histograms, and pie charts. Additionally, it discusses the importance of data collection and interpretation to reveal meaningful information and patterns.

Uploaded by

Zuha Irshad18
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

RESEARCH

METHODOLOGY

DATA REPRESENTATION

1
POPULATION & SAMPLE

Statistical Inference

Sample: Subset of Population: Collection of


collection of all possible all possible entities
entities (observation units) (observation units)
Statistical Inference is
the science of drawing
Data on sample is what is Data on the whole
inferences/ conclusions
available. population is usually not
about a population of
available.
interest.
KNOWN
Statistics are used to UNKNOWN
describe samples. Parameters are used to
These can vary across describe populations.
samples. These are constants for a
population.

2
VARIABLES & OBSERVATIONS
VARIABLES

O Entity Height Weight Age Gender


B
S (inches) (pounds) (years) (Category)
E
R Person 1 67 170 33 Male
V Person 2 61 120 38 Female
A
T Person 3 72 220 62 Male
I * * * * *
O * * * * *
N
S

3
TYPES OF DATA: CATEGORICAL AND
NUMERICAL

Categorical Numerical

4
TYPES OF DATA: CROSS-SECTIONAL AND
TIME ORDERED

Period Plant 1 Plant 2 Plant 3 Plant 4

Jan Cross Sectional Data


Feb

Mar

Apr Time
Ordered
May
Data
Jun

July

Questions
Length of plant?

5
Results of a student survey
ID V1 V2 V3 V4 V5 V6 V7 V8 V9 V10

1 1 2 3 1 1 3 2 2 2 2

2 3 1 3 3 1 2 1 5 2 1

3 2 3 3 1 2 2 2 1 2 2

4 1 2 1 3 2 1 3 4 3 2

5 1 1 1 1 1 3 3 4 2 2

6 1 2 5 5 2 2 2 4 2 2

7 1 2 5 2 1 3 3 4 2 2

8 1 1 5 1 1 3 3 1 2 2

9 3 2 2 3 2 2 1 5 2 1

10 2 2 3 5 1 2 1 1 2 1

11 1 1 1 1 2 2 2 4 2 2

12 1 1 1 1 1 3 2 1 2 2

One can’t just stare at this and grasp what the data is “saying.”
6
The numbers don’t “speak for themselves
DATA COLLECTION

Data Needs to be “Boiled Down” to


Reveal Meaningful Information,
Patterns, and Relationships

7
DATA REPRESENTATION
 Tables
 Bar Charts
 Frequency bar charts
 Histograms
 Pie Charts

 Scatter Plots/Line Graphs

 Area maps

8
PROPORTION AND PERCENTAGE

• Proportions and percentages are also called 9


relative frequencies.
BAR CHART
 Bar height represents counts or percentages
 Easier to compare categories with bar

 Called Pareto Charts when ordered from tallest


to shortest

10
BAR CHART
 The highway speed limits on roads within three
states

State Urban Rural


Florida 65mi/h 70 mi/h

Texas 70 mi/h 70 mi/h

Vermont 55mi/h 65 mi/h

11
BAR CHART
 Choose a scale and interval for the vertical axis.

80

State Urban Rural 60

Florida 65mi/h 70 mi/h


40
Texas 70 mi/h 70 mi/h
Vermont 55mi/h 65 mi/h 20

12
BAR CHART
 Draw a pair of bars for each state’s data. Use
different colors to show urban and rural
 If the document is to be printed, use textures
inseated of colors
80

60
State Urban Rural
Florida 65mi/h 70 mi/h
40
Texas 70 mi/h 70 mi/h
Vermont 55mi/h 65 mi/h 20

13
0
Florida Texas Vermont
BAR CHART
 Label the axes and give the graph a title.
 Make a key to show what each bar represents

Speed Limit on Interstate Roads


80
Speed Limit (mi/h)

60 Urban
Rural
40

20

14
0
Florida Texas Vermont
BAR CHART

Comparing Investors

Savings

CD

Bonds

Stocks

0 10 20 30 40 50 60

Investor C Investor B Investor A 15


16
FREQUENCY BAR CHART
 A Frequency Table showing a classification of the
AGE of attendees at an event.

17
FREQUENCY BAR CHART
 A graphical display of distribution of frequencies

Histogram

7 6
6 5
Frequency

5 4
4 3
3 2
2
1 0 0
0
5 15 25 36 45 55 More
18
FREQUENCY BAR CHART
 Sort Raw Data in Ascending Order:
12, 13, 17, 21, 24, 24, 26, 27, 27, 30, 32, 35, 37,
38, 41, 43, 44, 46, 53, 58
 Find Range: 58 - 12 = 46

 Select Number of Classes: 5

 Compute Class Interval (width): 10


(range/classes = 46/5 then round up)
 Determine Class Boundaries (limits): 10, 20, 30,
40, 50
 Compute Class Midpoints: 15, 25, 35, 45, 55

 Count Observations & Assign to Classes 19


FREQUENCY BAR CHART
 Example 2 Sodium in
Cereals

20
FREQUENCY BAR CHART

21
Sodium in Cereals
STACKED BAR CHART
 A way to compress and merge bar graphs is to
“stack” all the bars of an ordinary bar graph on
top of one another to form a single bar
6000
9000

5000 8000
7000
4000 6000
5000
3000 NUST UET
4000
UET NUST
2000 3000
2000
1000 1000
0
0
Students Teachers Admin
Students Teachers Admin

22
STACKED BAR CHART
 We can then combine such stacked bars to “tell
the story” of the change over time

23
HISTOGRAMS

 Histogram of Percent of Population 65+

 This histogram is logically equivalent to a 24


frequency bar chart
HISTOGRAM VS FREQUENCY BAR CHAT
 The preceding histogram is essentially no different from
a frequency bar chart because all class intervals all have
the same width (in this case,1 percentage point wide).

 Otherwise (i.e., if the class intervals are not all of equal


width), a bar chart and a histogram of the same data
may look quite different
 The bar chart presents a misleading picture of the data,
 The histogram presents a more accurate picture.

 The histogram, unlike the bar chart takes account of the


interval property of the variable.

 This can be illustrated by focusing on data for which


unequal class intervals were created. 25
HISTOGRAM VS FREQUENCY BAR CHART

Frequency Percent
Less than $15,000 145 13.7%
$15,000 to $25,000 121 11.4%
$25,000 to $35,000 102 9.7%
$35,000 to $50,000 154 14.6%
$50,000 to $80,000 246 23.3%
$80,000 to $120,000 167 15.8%
More than $120,000 120 11.4%
Total 1055 100%

Income data
26
FREQUENCY BAR CHART (INCOME DATA)
 The bar chart appears to display a distribution of income that
is approximately “uniform” – that is, all bars are
approximately the same height, except for a distinctive peak
 The impression the bar graph conveys to the eye is that there
are more well-off than not-so-well-off people.
 However, this impression is quite misleading, as you can begin
to under-stand when you look more closely at the income class
intervals and notice that they are not of equal width.

27
HISTOGRAM (INCOME DATA)

28
HISTOGRAM (INCOME DATA)
 The fundamental difference between a bar graph and a
histogram:
 In a bar graph, frequency is represented by the height of the
bars (all of which have the same width);
 In a histogram, frequency is represented by the area of the
“bars” (which may have different widths, reflecting the
different “widths” of the class intervals).

 With equal class intervals, the area of a bar depends


only on its height, so Histogram ≈ Frequency Bar Chart

 But with unequal class intervals, the area of a bar


depends on both its height and its width, so Histogram ≠
Frequency Bar Chart 29
CONSTRUCTING THE HISTOGRAM

 To draw this a histogram, first draw a horizontal


line, representing the possible values of the
variable.
 Place markers at equal intervals to mark equal
increments in the value of the variable.
 In the figure above, markers are placed at $0K,
$20K, $40K, etc., up to $260K

30
CONSTRUCTING THE HISTOGRAM

 Next put other [red] markers along the scale at


the points that separate the class intervals we
are using
 On this case at $0, $15K, $25K, $35K, $50K, $80K,
and $120K.
 Set an upper bound more or less arbitrarily
 $250K in this case

31
CONSTRUCTING THE HISTOGRAM
 Next place a rectangle (or a “bar”) on each class
interval, so that the area [not height] of each
rectangle is proportional to the frequency
associated with that class interval.
 How tall should each rectangle be?
 The width of each rectangle is the width of the class
interval, and:
 Area = Height × Width so Height = Area / Width
 Since Area here represents Frequency, we have the
formula:
 Height = Frequency / Width

32
CONSTRUCTING THE HISTOGRAM
 Now we can calculate the relative heights of all the bars/rectangles.

Class Interval Width Freq. Freq/Width Height


0-15 15 13.7 13.7 / 15 = 0.913
15-25 10 11.4 11.4 / 10 = 1.140
25-35 10 9.7 9.7 / 10 = 0.970
35-50 15 14.6 14.6 / 15 = 0.973
50-80 30 23.3 23.3 / 30 = 0.777
80-120 40 15.8 15.8 / 40 = 0.395
120-250 130 11.4 11.4/130 = 0.088

 Now we can draw the appropriate scale on the vertical axis.


 The tallest rectangle has a (relative) height of about 1.14, so the axis should
extend a bit higher than this.
 Having constructed the bars/rectangles, we should remove the vertical axis and
scale.
 Otherwise, readers are likely to (mis)interpret it as representing frequency, like
the vertical axis in a bar graph.
33
HISTOGRAM VS FREQUENCY BAR CHART

34
CONTINUOUS DENSITIES
 The INCOME histogram was based on a small number of
(rather wide) class intervals and a modest number of
observations

 Suppose we have INCOME data that is recorded very


precisely, e.g., to the near dollar or even cent.

 Suppose also we have a huge --- approaching infinite ---


number of observations

 We could then refine INCOME into narrower and narrow


class intervals, redrawing the histogram accordingly

 If we pushed this process to the limit, we would end up


with what would be an essentially continuous (and
probably fairly smooth) density curve 35
Approaching a continuous density curve

36
Approaching a continuous density curve

37
Approaching a continuous density curve

38
Approaching a continuous density curve

39
Approaching a continuous density curve

40
We Approach a Continuous Density [Normal] Curve

41
INTERPRETING HISTOGRAMS

Left and right sides


are mirror images

42
INTERPRETING HISTOGRAMS

43
OUTLIERS

 An outlier falls far from the rest of the data

44
PIE CHARTS
 A “pie” is “divided up,” into parts
 They are especially appropriate for displaying
“shares”:
 Electoral votes, or seats in a legislature, budget
break down among different spending categories

45
PIE CHARTS
Property 14%

Investment Category Amount Percentage


Savings 15%

Stocks 46.5 42.27


Stocks 42%
Bonds 32 29.09

Property 15.5 14.09


Bonds 29%
Savings 16 14.55

Total 110 100

46
 But pie charts are not helpful if there are many
categories – Just show the quantities in tabular
form

47
IS PICTURE WORTH A THOUSAND WORDS?
 Not Always: Sixty eight senators voted for the bill
and thirty two voted against.

48
SCATTER PLOT/ LINE GRAPH
Sales Vs. Years Experience
Shows relationship between two variables.
Can one be used to predict the other? 70
60

Sales (in Ks)


50
40
30
20
Sales for the last 120 months 10
0
1200 0 10 20 30
1000 Experience (in Years)
Sales (in Ks)

800
NSA
600
400 SA

200
Regression Analysis are used to predict
0 one variable’s value based on the other.
1 16 31 46 61 76 91 106 Correlation analyses is used to measure
Time (month #) the strength of linear relationship among
two variables.

49
AREA MAPS
 A graph used to plot variables by geographic
locations
 HIV Prevalence in Adults in Africa, end 2003

50

Source: UNAIDS, 2003


TIPS FOR DISPLAYING DATA
 3D may not be a good idea because the data may
appear distorted, can be misinterpreted, or may be
misleading.

51
TIPS FOR DISPLAYING DATA
 If 2 or more lines are plotted on a graph, a key or legend is
necessary. A different color or symbol should be used for each line.
 The color of the background of the graph, and the lines on the
graph should be clearly distinguishable from each other.
 The color of lines on a multi-line graph should be distinguishable
from each other.

52
TIPS FOR DISPLAYING DATA
 If the document is to be printed, color should not
be the only distinguishing attribute of different
lines/bars

6000

5000

4000

3000 NUST
UET
2000

1000

0
Students Teachers Admin

53
TIPS FOR DISPLAYING DATA
 Keep graphs simple -- make the data do the talking.
Don’t "liven" up you chart with extra colors, 3D, or
pictures.

 Use meaningful titles and labels - let the audience


think about what the data means, not what the
data is or could be.

54
SOME STATISTICAL
MEASURES

55
RANGE

Range = max − min


• The range is strongly affected by outliers.

56
MEAN
 The mean is the sum of
the observations
divided by the number
of observations
 It is the center of mass

57
MEAN & RANGE
Example 1. Two dice were thrown 10 times and their
scores were added together and recorded. Find the mean
and range for this data.
7, 5, 2, 7, 6, 12, 10, 4, 8, 9

Mean = 7 + 5 + 2 + 7 + 6 + 12 + 10 + 4 + 8 + 9
10
= 70 = 7
10 58

Range = 12 – 2 = 10
MEDIAN
The median is the middle value of a set of data once
the data has been ordered.

59
MEDIAN
The median is the middle value of a set of data once
the data has been ordered.

60
MEDIAN
Order Data • Midpoint of the observations
1 78 Order Data when ordered from least to
2 91 1 78 greatest
3 94 2 91 Order observations
4 98 3 94 If the number of observations
5 99 4 98 is:
6 101 5 99 Odd, the median is the middle
7 103 6 101 observation
8 105 7 103 Even, the median is the average
9 114 8 105 of the two middle observations
9 114
61
10 121
MEAN & MEDIAN
 Mean and median of a symmetric distribution are close
 Mean is often preferred
 In a skewed distribution, the mean is farther out in the
skewed tail than is the median
 Median is preferred because it is better representative of a
typical observation

62
RESISTANT MEASURE
 A measure is resistant
if extreme
observations (outliers)
have little, if any,
influence on its value
 Median is resistant to
outliers
 Mean is not resistant to
outliers

63
MODE

 Value that occurs most often


 Highest bar in the histogram
 Mode is most often used with categorical data

64
STANDARD DEVIATION
 Each data value has an associated deviation from

the mean,
 The sum of the deviations is always zero

 A deviation is positive if it falls above the mean

and negative if it falls below the mean

65
STANDARD DEVIATION
• Standard deviation gives a measure of
variation by summarizing the deviations of
each observation from the mean and
calculating an adjusted average of these
deviations:
• Find mean
• Find each deviation
• Square deviations
Sum squared

deviations
• Divide sum by n-1 66
• Take square root
PERCENTILE
 The pth percentile is a value such that p percent
of the observations fall below or at that value

67
QUARTILES
• Splits the data into four
parts
Arrange data in order
The median is the second
quartile, Q2
Q1 is the median of the lower
half of the observations
Q3 is the median of the upper
half of the observations

68
QUARTILES

Quartiles divide a ranked


data set into four equal parts:
Q1= first quartile = 2.2
•25% of the data at or below
Q1 and 75% above
•50% of the obs are above
the median and 50% are M = median = 3.4

below
•75% of the data at or below
Q3 and 25% above Q3= third quartile = 4.35

69
INTER QUARTILE RANGE
 The inter quartile range is the distance between
the third and first quartile, giving spread of
middle 50% of the data: IQR = Q3 − Q1

70
THE 5 NUMBER SUMMARY

• The five-number
summary of a dataset
consists of:
Minimum value
First Quartile
Median
Third Quartile
Maximum value

71
Finding the median, quartiles and inter-quartile range.
Example 1: Find the median and quartiles for the data below.
12, 6, 4, 9, 8, 4, 9, 8, 5, 9, 8, 10
Order the data
Q1 Q2 Q3

4, 4, 5, 6, 8, 8, 8, 9, 9, 9, 10, 12

Lower Upper
Median
Quartile Quartile
= 8
= 5½ = 9

Inter-Quartile Range = 9 - 5½ = 3½ 72
Finding the median, quartiles and inter-quartile range.

Example 2: Find the median and quartiles for the data below.
6, 3, 9, 8, 4, 10, 8, 4, 15, 8, 10
Order the data
Q1 Q2 Q3

3, 4, 4, 6, 8, 8, 8, 9, 10, 10, 15,

Lower Upper
Quartile Median Quartile
= 4 = 8 = 10

Inter-Quartile Range = 10 - 4 = 6 73
Discuss the calculations below.

Battery Life: The life of 12 batteries recorded in hours is:


2, 5, 6, 6, 7, 8, 8, 8, 9, 9, 10, 15
Mean = 93/12 = 7.75 hours and the range = 15 – 2 = 13 hours.

2, 5, 6, 6, 7, 8, 8, 8, 9, 9, 10, 15
Median = 8 hours and the inter-quartile range = 9 – 6 = 3 hours.

74
Box and Whisker Diagrams.

Box plots are useful for comparing two or more sets of data

Anatomy of a Box and Whisker Diagram.


Lowest Lower Upper Highest
Value Quartile Median Quartile Value
Whisker Whisker
Box

4 5 6 7 8 9 10 11 12

75
Box Plots
Box Plot

Lowest value Median Highest value


Lower Quartile Upper Quartile

0 5 10 15 20 25

76
Drawing a Box Plot.
Example 1: Draw a Box plot for the data below

Q1 Q2 Q3

4, 4, 5, 6, 8, 8, 8, 9, 9, 9, 10, 12

Lower Upper
Median
Quartile Quartile
= 8
= 5½ = 9

4 5 6 7 8 9 10 11 12 77
Drawing a Box Plot.
Example 2: Draw a Box plot for the data below

Q1 Q2 Q3

3, 4, 4, 6, 8, 8, 8, 9, 10, 10, 15,

Lower Upper
Quartile Median Quartile
= 4 = 8 = 10

3 4 5 6 7 8 9 10 11 12 13 14 15 78
Drawing a Box Plot.
Question: Stuart recorded the heights in cm of boys in his
class as shown below. Draw a box plot for this data.
QL Q2 Qu

137, 148, 155, 158, 165, 166, 166, 171, 171, 173, 175, 180, 184, 186, 186

Lower Upper
Quartile Median Quartile
= 158 = 171 = 180

130 140 150 160 170 180 cm 190


79
Box Plots
Lower Quartile = 52kg Lower Quartile = 66kg
Median = 78kg
Highest = 112 Median = 62kg
Highest = 96
Lowest = 34 Lowest = 40 Upper Quartile = 89kg
Upper Quartile = 70kg
GIRLS

BOYS

20 40 60 80 100 120 m kg
18
IQRs 23
80

MEDIANS 78 and 62
Comparison statements
• Boys have a larger MEDIAN so
– on average, they are heavier than girls
– the average boy is 16kg heavier than the
average girl
• Boys have larger Interquartile Range so
– boys’ weights vary more than girls’ weights
• Why do we NOT use the Range?
– One extreme value can distort the comparison
81

You might also like