0% found this document useful (0 votes)
5 views15 pages

Module 1

The document outlines the foundational concepts of data analysis, focusing on the distinction between variables and cases, as well as the importance of organizing data into a data matrix. It discusses different types of variables, methods for visualizing data, and key statistical measures such as mean, median, and standard deviation. Additionally, it emphasizes the significance of understanding variability and the use of tools like box plots and z-scores for effective data interpretation.

Uploaded by

Mai Anh Nguyen
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF or read online on Scribd
0% found this document useful (0 votes)
5 views15 pages

Module 1

The document outlines the foundational concepts of data analysis, focusing on the distinction between variables and cases, as well as the importance of organizing data into a data matrix. It discusses different types of variables, methods for visualizing data, and key statistical measures such as mean, median, and standard deviation. Additionally, it emphasizes the significance of understanding variability and the use of tools like box plots and z-scores for effective data interpretation.

Uploaded by

Mai Anh Nguyen
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF or read online on Scribd
Every Analysis Starts with Two Questions At the heart of any study are two simple concepts: the subjects we are interested in, and the characteristics we measure about them. VARIABLES —The "What" We Measure Age CASES —The "Who" or What" Cases are the individuals, objects, 1 or entities we are studying, Har Color They can be players, teams, companies, or counties. Variable vs. Constant {A variable must have variation. if we study only Spanish teams, the "City" they are based in is a variable (Barcelona, Madrid, Velencia), but the "Country" is @ constant (Spain forall teams}. Variation is what we analyze. The Data Matrix: Bringing Order to Information Raw data is chaotic. The data matrix is the standard format for organizing information, forming the core element of every statistical study. I's a simple grid with a clear purpose. i {ny Each column represents a specie {la variable being measured ui wy ee ‘Age Weight Goals Scored. Team Hair Color Player 22 755 4 Eagles Brown a a ravers) e re Each eell contains Exch row eoresents the observation fora ainglecase nour >| Player? 25. 803 = 12 ————_—_—“vencare on 3 ramp, each row en varabie ounmietootal ——piayer2g 29 z ries a player. Real-world data is Player 56 aa 824 7 Bears Black etter income We ‘missing data'in our Player 12 24 790 a Lions Blonde Sana Player 220 a B18 9 Tigers Red Player 400 2 Not All Data Is Created Equal: Classifying Categorical Variables The type of a variable dictates the questions we can ask and the tools we can use. Understanding the ‘level of measurement’ is the most important rule of the game. The first family of variables deals with categories or labels. oO @@ 2|'B Nominal: Categories without order. Ordinal: Categories with a meaningful order. The values are distinct categories with no intrinsic The categories have a natural ranking or order, but ranking or order. itis not possible to argue that one the differences between the ranks are not necessarily category is better, worse, more, or less than another. equal or measurable. Examples: Examples: + Hair Color (Blonde, Brown, Black) + Competition Rank (1st, 2nd, 3rd) + Nationality (Spanish, French, Mexican) + Survey response on a player's skill (Poor, Average, + Team Membership (Real Madrid, Barcelona) Good, Excellent) The World of Numbers: Classifying Quantitative Variables The second family of variables involves numerical values where mathematical operations make sense. These are further divided by whether they have a “true zero.” ey Interval: Ordered values with equal Ratio: All the features of Interval, plus a intervals. true zero. The data has a meaningful order, and the differences This level has order, equal intervals, and a meaningful between values are equal and consistent. However, zero point, which represents the complete absence of there is no meaningful "true zero” point. the variable. Example: Examples: Player's Age. The difference between a 16 and 18- Goals Scored (0 goals means no goals were scored). year-old is the same as between a 12 and 14-year-old. Player Height (0 cm means no height). But an age of 0 doesn't mean “no age exists.” Discrete vs. Continuous Discrete variables are separate numbers (e.g,, 1, 2, 3 goals—not 1.21 goals). Continuous variables can take any value within an interval (e.g,, height can be 170cm, 170.1cm, or 170.1246cm). From Raw Data to Visual Story: Summarizing with Tables and Charts A data matrix is organized but not insightful. To see the patterns and tell a story, we must summarize and visualize our data. Start:RawData _—_—Step 1: Frequency Table Step 2: Visualization Hai calor ote 200 Block Frequoney of Player Hat Color (n= 400) a = oer nnn nacnag coma een a onde Bonde —7S~=~*«OKSSCSCOR sock Bock —— > bom ease sex He Brown Back 00K 2K S omer Ober 90 75K‘ < ra a Bonde Brown Bick Other Good for showing parts Clearer for comparing of a whole. categories. Choosing the Right Chart: The Limits of the Pie While pie charts can be effective for a few categories, they become messy and uninformative as data complexity increases. A bar graph offers greater clarity and flexibility. ‘Too Many Categories A sox EX ay olan ‘Srguay Argentina Uruguay zx Netnernds y France ‘Beigum Goatia Germany haly Portugal ‘This is aesthetically colorful bt analytically useless. Its impossible [Link] counties or discern patterns A Clearer Picture Span Be sgeetina ‘cormany tay Peruse Netbeans eum ° 2» “0 o 0 100 Number of ayers ‘This graph contans the same amount of information, but itis far more clgestble, We can easly rank countries and compare their frequencies. For categorical data with more than 4-5 categories, always prefer a bar graph over a pie chart. Seeing the Shape of Your Data: The Histogram For quantitative data, we use a histogram to see the distribution of values. Itis similar to a bar graph, but the bars touch to represent a continuous, underlying scale, Distribution of Body Weight in Spanish Football Players 20 enceaninan sede 1 Er 5 @ 3 7 75 Body Wich ea) We crete intervals (2g, 725kg 1075kg) and count the payers in each. We can see ata glance that most players weigh around 75kg. Common Distribution Shapes 4. Symmetic (ell- Shaped) 20 ro 2. Skewed tothe Right [Link] Has ane peak (riod! ands symmetric. Example: Player income. Moat players eam Np ‘mount, but afew spettars eam mush me, reatng along waht a sare: Ag of people in a stadium canna shir 3 ‘ees metch Youll se a pool enlaren 6-8 years ‘0 a anole or parots (30-40 esr ld), What i: a ‘Typical’ Value? Finding the Center ‘After visualizing a distribution, our first task is to locate its center of gravity. We have three primary tools for this, known as the "Three M's" of central tendency. 1. The Mode: The Most Frequent Value The value or category that appears most often in the dataset, Best for Categorical (Nominal) data. Piayer Continents other rica south ‘America Europe The mode is ‘Europe,’ as it's the most common category. Note: The mode is the category ‘Europe’, not the value 70% 2. The Median: The Middle Value The value that sits in the exact middle of the data when its ordered from smallest to largest. 50% of observations are below it, and 50% are above. Best for: Quantitative data, especially when skewed or with outliers. 6, 7, 7, 8 8, 8, 9 Middle value 6,7,7,8,8,8,9. The median is 8, 3. The Mean: The Arithmetic Average The sum of all values divided by the ‘number of observations, Itis the "balance point” of the data, Best for: Quantitative, symmetric data without significant outliers. 6+7+7+8+8+8+9 7 (6+7+7+8+8+8+9) /7 = 76. The mean is 78, The Outlier Test: When the Mean Can Mislead ‘The mean is sensitive to extreme values (outliers), while the median is not. This difference is critical for accurately describing a dataset. PART 1: THE SITUATION BEFORE ‘Six quests and a bartender ae ina football club cantina. Their annual incomes are al around €35,000. Median Mean 0k 20k 30K Annual income Cre ‘The mean and median ere very close, both providing e good ‘description ofthe typical income, PART 2: THE OUTLIER ARRIVES Famous footballer Franco Galon, who earns €70 milion a year, walks in. Median income: €38,000 ‘Mean income: 8,000,000 <0 clom «20m = €30m —_€40m ‘annual income som 60m €70m Median income: €36,000 (verely chenges) Mean Income: >€8,000,000 (drastically changes) Franco Galon isan outlier. His income has @ disproportional effect on ‘the mean, making ta misieading measure of the ‘center'in this case. WHICH MEASURE SHOULD YOU USE? @ SyoW eH |, bgp Use nemMode, ) it uanthatve mh Categorical? ‘extreme outliers or skew? Ist Quantitative ana [| usetne Median, (2) county eymmetic? > Gf Use the Mean How Spread Out is the Data? Measuring Variability Knowing the center is only half the story. Two datasets can have the same center but look completely different. We need to measure their variability, or dispersion A Tale of Two Teams Team 1 (Low Variability) ‘Team 2 (High Variability) Tightly custered © Wily cpersod sound themean eee @ 2055 the scale — eco e | ee e e eee e e ° 5 10 15 20 2 oo 5 0 15 2» 8 20 Mean Percentage of Body Covered wth Tatoos Mean Both teams have the same mean tattoo density (5%), but the players in Team 2 are much more different from each other. Team 2 has higher vaibity. Measures of Spread ‘The Range The Interquartile Range (IQR) The diference between the highest and lowest value The range ofthe mide 0% of 102-5 the data Calculated asthe Sr ‘Team 1 Range = 19.3 - 10.8 = 8.5 ‘Quartile (Q3) minus the 1st Tam 2 Range = 277 -0= 277 Quartie (a). —= ~ o} ina) 08 Easy to calculate, but only uses the two most extreme values and canbe _—_‘TheIQRisnot affected by outers because i ignores the lovest 25% and misleading. highest 25% ofthe data, The Gold Standard of Spread: Variance and Standard Deviation While the IQR is robust, the most widely used measures of variability are the variance and standard deviation because they take into account every single value in the dataset. Step 1: Calculate Deviations from Step 2: Square the Step 3:Sum the Squares ‘Step 4: Take the Square Root for the Mean’ tattoo data (Mean = 15) Deviations and Calculate Variance ‘Standard Deviation Rep ee [Perei | asm or squore = 03974 Standard Deviation ( ° 0- 152-15, (-15)¢= 225 87 87-15=-63 (63 = 3969 (sumot squares)/n-t)= | {(@ Key Concopt: The standard deviations | ° » 0-15=-20 » (152-1439 = 639,74 10 = B97 Fonts tase tt becth le erone 5 RE CaREaaaT y units ands igh interpretable (A Problem: The units are now) 5 (157 = 15.23 percent squared which ishard to interpee. ‘Comparison ° 0-15=0 (457 = 225 Team2 Team a7 0-15 = 63 (63) = 30.69 ‘Standard Deviation Standard Deviation a 8.0 2.52 Insight: The sum of | | insight: Squaring makes - these devations s | al values postive so they ‘The larger standard deviation confirms the alvaye 200, ont cance ou ‘greater variabty we saw in Team 2 data, The Box Plot: A Five-Number Story in One Graphic Is there a way to visualize the center, the spread of the middle 50%, the full range, and potential outliers all at once? Yes: the box plot. Anatomy of a Box Plot Comparative Power ‘An observation is an ‘outlier if ities more than 15 xR below QI or 15 x IQR above 03. Range of data, excluding outers. \ 03 (Third Quartile) Fa The middle 50% of data (the IQR) <— Median (22) © <— tiers Maximum 01 Fst Quarti) Team 1 Box Plat ‘Team 2 80x Plot ‘Side-by-side, the box plots immediately and powerfully ‘show that while the madians are the same, the variability in Team 2s far greater than in Team 1. Range of data, excluding outers. —— Minimum Z-Scores: The Great Equalizer ‘An observation of 19. Is meaningless on ts own that high? Low? Common? AZ-score provides nivel Sort 6 revexrecshga valle bme of how any lander vitions Kista the ten Observation — Mean Standard Deviation Team’ (Low Variability) ‘Team? (High Variability) Mean=15, SD=252 Applying the Formula Mean=15, SD=80 193-15 How common is a tattoo density of 19.3% 193-15 2-57 = 1172 in our two teams? =a = 0.54 252 30 This value is 1.72 standard deviations above the mean for Team 1 ‘This same vale is only 0.54 staeard deviations above the mean for Team 2. Conclusion: A density of 19:2 is much mere common and less exceptional in the high-variabilty Team 2 Rule of Thumb for interpretation . __ Key insight: Z-seaes beyond #2 are Ce /— ewe; beyond #3 ae rae The Full Analysis: A Chemistry Grade Case Study Let's put it all together. We will now perform a complete descriptive analysis on the average chemistry grades for eight high schools. 1. The Data & Visualization 2. Measures of Center 3. Measures of Spread ‘Avg. Grade Mean Median Range eave 686 7.25 81-41=40 ‘School 3 At e eo ane QR Sms > Observation: Te outierpuls | | 7,65 - 6.45 = 1.2 sons a2 ‘he mean eiow he medion, School? 78 StolB 77 og 12 9 4 6 6 7 8 9 4. The Box Plot Summary en es od Max (63) Calculation Medion (75) Z = (4.1 - 6.86) / 1.27 = -2.17 2116.45) ‘Min (6.2) ‘Conclusion “The average grade for School 3 is 2.17 standard BEE ale ©). deviations below the mean, confirming itis an Saas exceptional value worthy of investigation. Your Toolkit for Describing Data You now possess the foundational toolkit for exploring and describing any dataset. You have learned to move from raw data to organized, visualized, and summarized insights. pea Summarize Key Features Visualize Define Data Organize Data mein se ie inl J) Pen RUDE eeu Tent distribution, its center, AND its spread. The right tools depend on. CIN eeu RU ccc coag ee These descriptive tools are the essential first step before making predictions, testing hypotheses, or drawing broader conclusions from data—the world of inferential statistics.

You might also like