0 ratings 0% found this document useful (0 votes) 5 views 15 pages Module 1
The document outlines the foundational concepts of data analysis, focusing on the distinction between variables and cases, as well as the importance of organizing data into a data matrix. It discusses different types of variables, methods for visualizing data, and key statistical measures such as mean, median, and standard deviation. Additionally, it emphasizes the significance of understanding variability and the use of tools like box plots and z-scores for effective data interpretation.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content,
claim it here .
Available Formats
Download as PDF or read online on Scribd
Go to previous items Go to next items
Every Analysis Starts with Two Questions
At the heart of any study are two simple concepts: the subjects we are interested in,
and the characteristics we measure about them.
VARIABLES —The
"What" We Measure
Age
CASES —The "Who" or What"
Cases are the individuals, objects, 1
or entities we are studying, Har Color
They can be players, teams,
companies, or counties.
Variable vs. Constant
{A variable must have variation. if we study only Spanish teams, the "City" they are based in is a variable (Barcelona, Madrid,
Velencia), but the "Country" is @ constant (Spain forall teams}. Variation is what we analyze.The Data Matrix: Bringing Order to Information
Raw data is chaotic. The data matrix is the standard format for organizing information, forming
the core element of every statistical study. I's a simple grid with a clear purpose.
i {ny Each column represents a specie
{la variable being measured
ui wy ee
‘Age Weight Goals Scored. Team Hair Color
Player 22 755 4 Eagles Brown
a
a ravers) e re Each eell contains
Exch row eoresents the observation fora
ainglecase nour >| Player? 25. 803 = 12 ————_—_—“vencare on 3
ramp, each row en varabie
ounmietootal ——piayer2g 29 z ries a
player.
Real-world data is
Player 56 aa 824 7 Bears Black etter income We
‘missing data'in our
Player 12 24 790 a Lions Blonde Sana
Player 220 a B18 9 Tigers Red
Player 400 2Not All Data Is Created Equal: Classifying
Categorical Variables
The type of a variable dictates the questions we can ask and the tools we can use.
Understanding the ‘level of measurement’ is the most important rule of the game. The
first family of variables deals with categories or labels.
oO
@@ 2|'B
Nominal: Categories without order. Ordinal: Categories with a meaningful order.
The values are distinct categories with no intrinsic The categories have a natural ranking or order, but
ranking or order. itis not possible to argue that one the differences between the ranks are not necessarily
category is better, worse, more, or less than another. equal or measurable.
Examples: Examples:
+ Hair Color (Blonde, Brown, Black) + Competition Rank (1st, 2nd, 3rd)
+ Nationality (Spanish, French, Mexican) + Survey response on a player's skill (Poor, Average,
+ Team Membership (Real Madrid, Barcelona) Good, Excellent)The World of Numbers: Classifying Quantitative Variables
The second family of variables involves numerical values where mathematical operations
make sense. These are further divided by whether they have a “true zero.”
ey
Interval: Ordered values with equal Ratio: All the features of Interval, plus a
intervals. true zero.
The data has a meaningful order, and the differences This level has order, equal intervals, and a meaningful
between values are equal and consistent. However, zero point, which represents the complete absence of
there is no meaningful "true zero” point. the variable.
Example: Examples:
Player's Age. The difference between a 16 and 18- Goals Scored (0 goals means no goals were scored).
year-old is the same as between a 12 and 14-year-old. Player Height (0 cm means no height).
But an age of 0 doesn't mean “no age exists.”
Discrete vs. Continuous
Discrete variables are separate numbers (e.g,, 1, 2, 3 goals—not 1.21 goals). Continuous variables can take any value
within an interval (e.g,, height can be 170cm, 170.1cm, or 170.1246cm).From Raw Data to Visual Story:
Summarizing with Tables and Charts
A data matrix is organized but not insightful. To see the patterns and tell a story, we must
summarize and visualize our data.
Start:RawData _—_—Step 1: Frequency Table Step 2: Visualization
Hai calor ote 200
Block Frequoney of Player Hat Color (n= 400) a
= oer nnn nacnag coma een a
onde Bonde —7S~=~*«OKSSCSCOR sock
Bock —— > bom ease sex He
Brown Back 00K 2K S
omer Ober 90 75K‘ <
ra a Bonde Brown Bick Other
Good for showing parts Clearer for comparing
of a whole. categories.Choosing the Right Chart: The Limits of the Pie
While pie charts can be effective for a few categories, they become messy and
uninformative as data complexity increases. A bar graph offers greater clarity and flexibility.
‘Too Many Categories
A
sox EX
ay
olan
‘Srguay Argentina
Uruguay
zx
Netnernds y France
‘Beigum
Goatia
Germany
haly
Portugal
‘This is aesthetically colorful bt analytically useless. Its impossible
[Link] counties or discern patterns
A Clearer Picture
Span
Be
sgeetina
‘cormany
tay
Peruse
Netbeans
eum
° 2» “0 o 0 100
Number of ayers
‘This graph contans the same amount of information, but itis far more
clgestble, We can easly rank countries and compare their frequencies.
For categorical data with more than 4-5 categories, always prefer a bar graph over a pie chart.Seeing the Shape of Your Data: The Histogram
For quantitative data, we use a histogram to see the distribution of values. Itis similar to a bar graph, but the bars touch to
represent a continuous, underlying scale,
Distribution of Body Weight in Spanish Football Players
20 enceaninan sede
1
Er
5
@ 3 7 75
Body Wich ea)
We crete intervals (2g, 725kg 1075kg) and count the payers in each. We can see ata glance that most players weigh around 75kg.
Common Distribution Shapes
4. Symmetic (ell- Shaped)
20 ro
2. Skewed tothe Right [Link]
Has ane peak (riod! ands symmetric.
Example: Player income. Moat players eam Np
‘mount, but afew spettars eam mush me,
reatng along waht a
sare: Ag of people in a stadium canna shir 3
‘ees metch Youll se a pool enlaren 6-8 years
‘0 a anole or parots (30-40 esr ld),What i:
a ‘Typical’ Value? Finding the Center
‘After visualizing a distribution, our first task is to locate its center of gravity. We have three primary tools for this,
known as the "Three M's" of central tendency.
1. The Mode: The Most
Frequent Value
The value or category that appears most
often in the dataset,
Best for Categorical (Nominal) data.
Piayer Continents
other
rica
south
‘America
Europe
The mode is ‘Europe,’ as it's the most
common category. Note: The mode is the
category ‘Europe’, not the value 70%
2. The Median: The Middle
Value
The value that sits in the exact middle of
the data when its ordered from smallest
to largest. 50% of observations are below
it, and 50% are above.
Best for: Quantitative data, especially
when skewed or with outliers.
6, 7, 7, 8 8, 8, 9
Middle value
6,7,7,8,8,8,9.
The median is 8,
3. The Mean: The Arithmetic
Average
The sum of all values divided by the
‘number of observations, Itis the "balance
point” of the data,
Best for: Quantitative, symmetric data
without significant outliers.
6+7+7+8+8+8+9
7
(6+7+7+8+8+8+9) /7 = 76.
The mean is 78,The Outlier Test: When the Mean Can Mislead
‘The mean is sensitive to extreme values (outliers), while the median is not. This difference is critical for accurately
describing a dataset.
PART 1: THE SITUATION BEFORE
‘Six quests and a bartender ae ina football club cantina. Their annual
incomes are al around €35,000.
Median Mean
0k
20k
30K
Annual income
Cre
‘The mean and median ere very close, both providing e good
‘description ofthe typical income,
PART 2: THE OUTLIER ARRIVES
Famous footballer Franco Galon, who earns €70 milion a year, walks in.
Median income: €38,000 ‘Mean income: 8,000,000
<0 clom «20m = €30m —_€40m
‘annual income
som 60m €70m
Median income: €36,000 (verely chenges)
Mean Income: >€8,000,000 (drastically changes)
Franco Galon isan outlier. His income has @ disproportional effect on
‘the mean, making ta misieading measure of the ‘center'in this case.
WHICH MEASURE SHOULD YOU USE?
@ SyoW eH |, bgp Use nemMode, ) it uanthatve mh
Categorical? ‘extreme outliers or skew?
Ist Quantitative ana
[| usetne Median, (2) county eymmetic?
> Gf Use the MeanHow Spread Out is the Data? Measuring Variability
Knowing the center is only half the story. Two datasets can have the same center but look
completely different. We need to measure their variability, or dispersion
A Tale of Two Teams
Team 1 (Low Variability) ‘Team 2 (High Variability)
Tightly custered © Wily cpersod
sound themean eee @ 2055 the scale
— eco e | ee e e
eee e
e
° 5 10 15 20 2 oo 5 0 15 2» 8 20
Mean Percentage of Body Covered wth Tatoos Mean
Both teams have the same mean tattoo density (5%), but the players in Team 2 are much more different from each other. Team 2 has higher vaibity.
Measures of Spread
‘The Range The Interquartile Range (IQR)
The diference between the highest and lowest value The range ofthe mide 0% of 102-5
the data Calculated asthe Sr
‘Team 1 Range = 19.3 - 10.8 = 8.5 ‘Quartile (Q3) minus the 1st
Tam 2 Range = 277 -0= 277 Quartie (a). —= ~
o} ina) 08
Easy to calculate, but only uses the two most extreme values and canbe _—_‘TheIQRisnot affected by outers because i ignores the lovest 25% and
misleading. highest 25% ofthe data,The Gold Standard of Spread: Variance and Standard Deviation
While the IQR is robust, the most widely used measures of variability are the variance and standard deviation
because they take into account every single value in the dataset.
Step 1: Calculate Deviations from Step 2: Square the Step 3:Sum the Squares ‘Step 4: Take the Square Root for
the Mean’ tattoo data (Mean = 15) Deviations and Calculate Variance ‘Standard Deviation
Rep ee [Perei | asm or squore = 03974
Standard Deviation (
° 0- 152-15, (-15)¢= 225
87 87-15=-63 (63 = 3969 (sumot squares)/n-t)= | {(@ Key Concopt: The standard deviations |
° » 0-15=-20 » (152-1439 = 639,74 10 = B97 Fonts tase tt becth le erone
5 RE CaREaaaT y units ands igh interpretable
(A Problem: The units are now)
5 (157 = 15.23 percent squared which
ishard to interpee. ‘Comparison
° 0-15=0 (457 = 225
Team2 Team
a7 0-15 = 63 (63) = 30.69 ‘Standard Deviation Standard Deviation
a 8.0 2.52
Insight: The sum of | | insight: Squaring makes -
these devations s | al values postive so they ‘The larger standard deviation confirms the
alvaye 200, ont cance ou ‘greater variabty we saw in Team 2 data,The Box Plot: A Five-Number Story in One Graphic
Is there a way to visualize the center, the spread of the middle 50%, the full range, and potential outliers all at
once? Yes: the box plot.
Anatomy of a Box Plot Comparative Power
‘An observation is an
‘outlier if ities more than
15 xR below QI or
15 x IQR above 03.
Range of data,
excluding outers. \
03 (Third Quartile) Fa
The middle 50%
of data (the IQR) <— Median (22)
© <— tiers
Maximum
01 Fst Quarti) Team 1 Box Plat ‘Team 2 80x Plot
‘Side-by-side, the box plots immediately and powerfully
‘show that while the madians are the same, the
variability in Team 2s far greater than in Team 1.
Range of data,
excluding outers.
—— MinimumZ-Scores: The Great Equalizer
‘An observation of 19. Is meaningless on ts own that high? Low? Common? AZ-score provides
nivel Sort 6 revexrecshga valle bme of how any lander vitions Kista the ten
Observation — Mean
Standard Deviation
Team’ (Low Variability) ‘Team? (High Variability)
Mean=15, SD=252 Applying the Formula Mean=15, SD=80
193-15 How common is a tattoo density of 19.3% 193-15
2-57 = 1172 in our two teams? =a = 0.54
252 30
This value is 1.72 standard deviations above the mean for Team 1 ‘This same vale is only 0.54 staeard deviations above the mean for Team 2.
Conclusion: A density of 19:2 is much mere common and less exceptional in the high-variabilty Team 2
Rule of Thumb for interpretation
. __ Key insight: Z-seaes beyond #2 are
Ce /— ewe; beyond #3 ae raeThe Full Analysis: A Chemistry Grade Case Study
Let's put it all together. We will now perform a complete descriptive analysis on the average
chemistry grades for eight high schools.
1. The Data & Visualization 2. Measures of Center 3. Measures of Spread
‘Avg. Grade Mean Median Range
eave 686 7.25 81-41=40
‘School 3 At e eo ane QR
Sms > Observation: Te outierpuls | | 7,65 - 6.45 = 1.2
sons a2 ‘he mean eiow he medion,
School? 78
StolB 77 og 12 9 4 6 6 7 8 9
4. The Box Plot Summary
en es od Max (63) Calculation
Medion (75) Z = (4.1 - 6.86) / 1.27 = -2.17
2116.45) ‘Min (6.2) ‘Conclusion
“The average grade for School 3 is 2.17 standard
BEE ale ©). deviations below the mean, confirming itis an
Saas exceptional value worthy of investigation.Your Toolkit for Describing Data
You now possess the foundational toolkit for exploring and describing any dataset. You have
learned to move from raw data to organized, visualized, and summarized insights.
pea Summarize Key Features
Visualize
Define Data Organize Data mein se ie
inl J)
Pen RUDE eeu Tent
distribution, its center, AND its spread. The right tools depend on.
CIN eeu RU ccc coag ee
These descriptive tools are the essential first step before making predictions, testing hypotheses, or
drawing broader conclusions from data—the world of inferential statistics.