0% found this document useful (0 votes)
4 views15 pages

FSP Notes

Module 01 introduces statistics as the art of learning from data, covering concepts such as data collection, descriptive and inferential statistics, and measures of central tendency (mean, median, mode). It emphasizes the importance of understanding randomness, sampling, and the implications of different statistical measures in decision-making. The module also explores the significance of probability, Chebyshev's inequality, and the normal distribution in real-world scenarios.

Uploaded by

homepicked2007
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF or read online on Scribd
0% found this document useful (0 votes)
4 views15 pages

FSP Notes

Module 01 introduces statistics as the art of learning from data, covering concepts such as data collection, descriptive and inferential statistics, and measures of central tendency (mean, median, mode). It emphasizes the importance of understanding randomness, sampling, and the implications of different statistical measures in decision-making. The module also explores the significance of probability, Chebyshev's inequality, and the normal distribution in real-world scenarios.

Uploaded by

homepicked2007
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF or read online on Scribd
Module 01: Introduction to Statistics Statistics —> Art of learning from data Collection > Description > Analysis > Inference Data ie ee Already available data Design an appropriate 2 experiment fo collect data % Inferential statistics —> drawing inferences 2 Probobility =» Role of chance Tobe able to draw logical conclusions from data, we. usually make some assumptions about the chance of obtaining the different data values. Model —> randomness rence requires some understanding of the theory of probabi Population: Total collection of ‘Sample: Subgroup of the population elements. use for examining the fees a ‘* In practice, a given sample generally cannot be assumed to be representative of a' population urless that sample has been chosen in'a random manner. ive Statistics > Measures of Central Tendency Measures of central tendency are statistical tools to summarize a dataset by identifying the central value within it. mean median mode Mean represents the average value of all the items, i.e. sum of all individual items / total number of items. Dataset: 5, 8, 10, 10, 12, 12, 15, 15, 15, 18, 20, 20, 25, 30, 35, 40, 50, 60, 100. Mean = Sum. N 5+8+l0+10+I2+12... 20 = 5i2 20. 25.6 GB Median > 1) Arrange the data in ascending order: 5, 8, 10, 10, 12, 12, 15, 15, 15, 18, 20, 20, 25, 30, 35, 40, 50, 60, 100. 2) Find the middle values (N/2 th and N/2 +1) for an even dataset. 5, 8, 10, 10, 12, 12, 15, 15, 15, 18, 20, 20, 25, 30, 35, 40, 50, 60, 100. 3) Median = (15 + 18) / 2 = 16.5 GB The median reflects the middle range of data usage (representative value). Mode > Most frequent value 5, 8, 10, 10, 12, 12, 15, 15, 15, 18, 20, 20, 25, 30, 35, 40, 50, 60, 100. Mode = 12 and 15 (each occurs 3 times). {inferences} 5, 8, 10, 10, 12, 12, 15, 15, 15, 18, 20, 20, 25, 30, 35, 40, 50, 60, 100. Mean (25.6 GB) Median (16.5 GB) Mode (12 GB, 15 GB) Provides the overall average, Reflects the middle range of data. ‘Shows the most common data incorporating all values in the) Outliers (Ike 100 G8) do not heavily skew it.) consumption values "mast dataset. Prone to outliers. Better represents the majority of students. Not prone to outliers Which of these measures appear to be better central measure for this particular dataset? popular chelces) Significance & Users > Where is the mean useful? Scenario 1: Evaluating overall | | Scenario 2: Comparing usage | Scenario 3: Budgeting usage. Telecom companies groups. If comparing data resources. A university might use the mean to usage between urban and —_ planning to provide mobile determine the average data | | rural students, the mean | data subsidies could use the demand across their customer | given gives a quick overview | mean to estimate the total base to plan infrastructure | of which group uses more | | amount of data required for upgrades or pricing strategies. data on average. all students. Where can the mean be misleading? ~ Individual-level insights: The mean (25.6 GB) suggests that most students use around 26 GB of data. Use the median (16.5 GB) to show typical usage. - Decision-making for popular plans: If a telecom company relies on the mean to design data plans, they might set a high base plan that does not match most students’ needs. Use the mode (12 GB, 15 GB) to identify the most popular plan instead. Descriptive Statistics -» Step 1: Organizing the data -> Example! Trends in streaming platform preferences among students, You surveyed 60 students fo find their favorite platform, The options inchde Netflix, Amazon Prime, Disney: Hotstar, YouTube, ZeeS and SonyLiv. [Streaming | Netflix: 12 votes Platform louver ef ee neg ‘Amazon Prime: 10 votes Netflix 2 Disney+ Hotstar: 15 votes Frequency Table -> I Ameren Prine: YouTube: 12 votes Disneys Hotstor Zeeb:7 votes yecTube SonyLiv:4 votes ae | Step 2: Visualizing the data > |_Sonyliv _ 207 Line Graph 15+ Tol 5 Netflix “Amazon Disney YouTube Zee —Sonyliv Prime” Hotstar Bor Graph Relative Frequency Table > ‘Streaming Relative Platform Frequency li 20% Senco eDeurecTan iene a ae Disney+ Hotstar 25% YouTube 20% Zee5 1.67% SonyLiv Frequency Polygon Consider a data set consisting of n values. If f is the frequency of a particular value, then the ratio f/n is called its relative frequency. Histogram Disneys Hetstar siping or 5% (Gas ies Nefflie 20% % Standard deviation & probability > The notation (P(X'S1G)) represents the probability that the random variable X takes a value greater than a. - X is a random variable, meaning it represents different possible outcomes of a rocess (e.g., dice roll, test scores, AI response time). - P(X>a) means we are calculating the probability that X will be greater than a specific value a. ~ It Is equivalent to asking: “What is the chance that X will be more than a?” % Chebyshev's Tdentity -> Chebyshev's inequality is a fundamental theorem in probability. Provides a bound on how much of the data. dévites from the mean. Chebysher's Ineqiality states that for any lataset: > mean (average value ot the dataset) ‘o > standard deviation (spread of data) K © number of stondard do The probability that @ value is at least k standard deviations away from the mean is at most 1/k?, fs > A random variable (e.g., test scores, stock prices etc.) viations away from the mean ‘Why should you know about Chebyshe ? YA) TE does ‘hot assume normality, making it broadly applicable. 2) Tt gives a worst-case bound; meaning it guarantees that extreme values do not occur too often (helps in setting safety bounds). V4) It is widely used in anomaly detection, AI model evaluation, and risk assessment: Example 1-> Fraud detection in AI Systems A company tracks daily transactions of an e-commerce AI system. mean transaction Value is: Fraudulent, transactions usually have = 20 dollars extremely high or low values. AL fraud detection Standard deviation o = 50 dollars. ) models éan use this to set risk thresholds. E.g., How often will a transaction be more than 3 standard deviations away from the mean? P(IX-20] 2 3(50)) < fe = 2 = 1% At most 11.11% of transactions will be below ~130 dollars or above 170 dollars. Scenario: AI-Powered delivery system An Al-driven delivery system “predicts estimated delivery times for a food delivery app. After analyzing data from thousands of orders, you find: The mean delivery time (jz) is 30 minutes. The standard deviation () is 5 minutes. Your company wants to guarantee customers that deliveries will be within 20 minutes to 40 minutes most of the time. Use Chebyshev's inequality to determine the maximum proportion of deliveries that might fall oudside this range. Delivery Scenario Continued : a ae or ee eae ‘Answer: At most 25% of deliveries might take less than 20 minutes or more than 40 minutes. Scenario: The “Ver Late” Di The mean delivery time (i) is 30 minutes rab te ireted lar (0) is 5 minutes. The considers a ‘ ” if it takes more than 3 standard Te mgr edn ay “yt Use shev's to find out: At most, what deliveries wil falecno ten se nantes Ot oentege of SS Using Chebysher’s inequality: P(|X- 30|2 3(5)) < 4. = $= 0.001 Answer: At most 1.11% of deliveries might take more than 45 minutes. Normal Distribution CData Sef) -> Pl ee ea uw scianRn cE eee "Asi eda aie di natant ni lbetherspoadtencnge, Xi ste es hl flows normal dation sows peak al he macn and con tf the lah fle thin oye andand deviation (0) of he mean. SO he fg add in ‘# Normal Distribution in Real-World Scenarios -» Example of normally distributed data: Natural phenomenon human height (cm) Measurement errors Social and economic phenomenon Biological processes Industrial control processes + Financial data (in some cases) * Canal Lit Theorem “Ihe cetrl lini theorem states that, given auf lrg sunle sz, the ditibtion ofthe tangle ean of a random varable wil approach a nermal detibuton, regardless of the erga distebuton of the poplin This happens othe sample sz m increases, 4 Paired Dataset > A paired dataset comessts of two related variables collected for each individual ‘or observation. Paired datasets are best visualized using seatter plots. Friend | Time on social media (hours) [Sleep duration (hour) i (ean 8 z — 3 is 3 4 6 6 5 5. —6 5 Beros 7 ? 33 5 3 8 3 9 45 10) 0 + + Scatter Plot > Coe ee ee ee ete POEL tise the he valatss: Covrelation # causation How can comelation, be quantified? cowelation coefficient (r) -> : The sample eomcation coefficient menses the strength and diction of pe 2 DY-D {xentee oat eal VEC-3! DYi-w? (Zand J = means of x and y ranges fom 1101 aia creat pa ee Quantitative Qualitative Numerical. Measurable quantities Categorical. Non-numerical data that can be expressed numerically. that represents categories or labels. (Discrete) (Continuous) (Nominal) (Ordinal) Countble date with Can foe any vole Categories without « Categories with « specie valet siti «hen rage rail order Bee ember of students, Exo Heit weigh ioe eal Beate, eel ee neem * Concept of Percentile > ‘A percentile is a statistical measure that indicates the relative standing of a value within 4 datoge. 11 represents the percentage of datapoints thal fll below a given value. e p-" percentile of a dataset is the value below which p% of the data falls. Examples > = Ifa student scores in the 80" percentile on an exam, it means they performed better than Sos of'shuserts but worse thar 20% net = In healthcare, if a baby's weight is in the 90" percentile, it means the baby is heavier than 90% of ‘other babies of the same age. * Quartiles -> The p* percentile of a dataset is the value below which p% of the data falls. The 25 percentile (Q)) is the vatue below which 25% of the data lies. The 50" percentile (Q2 or median) is the value below which 50% of the data ties. G The 75! percentile (Qs) is the value below which 75% of the data ties. These quartiles divide the data into € parts, making percentiles useful {Deseret at te stg eet aate eascee Set tg % Percentile Calculation -> (Find Qs (5 percent For a dataset with n values sorted in ascending order: i= 04) 1. Compute the rank using index i = 765%. be pete 2. gi is an integer, the p-" percentile is the vatue at ITS position i. 10 3. If (snot on integer, interpolate between the nearest Different methods of Quartile calculation a Box Plots for grouped data >} ‘method | formula | used By / Best for Visualization of continuous data Box plots are an excellent too! for visualizing and summarizing the ) distribution of continuous data, statistics | continoous Pox(nsl) | software | distributions, 00 (Numpy, R, | small datasets SPSS, ete) excel, Googe| loge datasets, | They provide a quick snopshat of the data’s spread, central tendency, and variability. heete | business epplications Mh. A box plot (also called a box-and- whisker plot) summarizes dato by fi showing its quartiles, spread, and potential outliers. dL I S The bok represents the as ay , interquartile range (CR) - oof the bx presen re mide the 25th percentile (ai) ond The ends of the box represent Theme i & het perce fond — A oO 75th percentile (Q3) ‘7th percentile (Q3) ~ Dots beyond the whiskers 7 Indicate outliers, which ore } 1 2A) datapoints that fll outside 7 ' the pected range. These might represent unusual or / Se ‘extreme Blood pressure values Tet indo ta (2 Sceeeskaegad | ineenrer ‘box to show the renge of deshed horizontal ne elves within 15 times the ‘The solid ing inside inside the bor, he 1QR from the quartiles. Values the borindicates the compare it with the median beyond the whiskers are median (50th percentile) Yor symmetry 1 ealined atta Sell, Large TOR = de vert n vale, y y The positon of the median within the box provides insights Into the syramet { and Skewness of the distribution) Types of Correlation © Correlation is a measure of how two variables change together. 1719/45 * A correlation coefficient quantifies this relationship's strength and ..--..-. at ip ee lation between variables. os ofc tf Sparen Rank Ordr carlton fi Kandalls tay, Pont sear at dfn yas of colton aft used in statics, Wen the coact Tople () Pear i biserial eae o Phi coefficient. Pearson's correlation measures the linear Spearman's correlation assesses the 3 Pearson's Correlation Coefficient Spearman Rank Order Correlation 3 ralatorahe between two continuous ‘monotonic relationship between two % it variables. It assumes that the data is variables based on rank orders rather normally distributed and that the than raw values. It is useful when data relationship is linear. is ordinal or when the relationship is not strictly linear. LK -|Yi-9) } Does not assume normality; less affected nee ee by outliers -9). - VE) = VEG -3 6d? where di eg, study hours - exam scores, p=1- the difference ¢ exercise hours - heart rates. n= 1) between ranks + * Students Group study | Peer Quality e hours (x) Measures ordinal association between two ranked variables. Counts concordant and discordant pairs. Preferred for emall datasets or with many tied ranks. a es es nv? - 1) smber of concordant pairs. umber of discordant pairs 2 SE CVD. Rank in math (2) | Rank in sci wo” 5(S?- 1) T 2 2 3 3 1 4 4 4 pair of observations (ij) is concordant if the ranks of both elements move in the same direction. iF x, > xj and y; > y Student P scores higher in both math if A and science than Student Q -> Concordant " iF x, ) (* Gcample: drawing two cords without replacement. Mutually exclusive events: cannot happen’) |) (( Exhaustive events: cover the entire sample at the same time, space, Example: drawing a red or black card Example: Rolling any number from I-6 on a die. from a single deck. Quick Recap of Probability-based problems ») A bag contains 3 red, 2 green, and | blue marble, P(Green) = 2/6 = 1/3. Product quality control uses probobility to predict defect rates across mixed batches, % The principle of counting: if event A has m outcomes and B has n outcomes: "Total outcomes = m x n. Example: 3 shirts x 2 pants = 6 outfits permutations (order matters) Venn diagram Representation ‘arrange a cet of items when the order Railerateta tak vecnscaiaaans Assoclative law: (EUFIUG = EU(FUG) | (EFIG = EFS) Distributive law: (EUF)G = EGUFG | EFUG =(EUG)(FUG) v * Pin) =o Sot ‘Axioms of Probability is the total umber of items ris ee rece Axiom 1: 0S PCE) $1 y the number of items to arrange ‘Axiom 3: for any sequence of mutually exclusive FF asthe ogee aS |s¢ Fesiisi vats pidud |Bece Eq 28 alei2 jz iltgas Seegs gees to 4 assis sa 2 Beecs Shee ep ES esp S/g heb palin Bees a: |Z. 285 Lapel Bad zys Eg Ese Wis gorges eS e¢ 2Es pileiis Br eey * The probably of event Ei a weahted average of the conionl pbablly of € given thal F has occured ‘and the conditional probability of € given that F has not occured, with each conditional probability being given ‘28 much weight as the event itis conditioned on has of occuring. %® Bayes’ theorem -> for two events A and B, where P(B) > 0: a) * rir Probctity | | Laatnoos re a _, at yee eR likeli is Pia) = ACE) cy fete = AIBA ee PCalA) = PRA (a r = er P(AnB) = P(BNA) -> because intersection is commutative. Baden (or wars |_| Posterior Probab stitute -> PAN B) = = PIBIA) P(A) obey) UB) The || PIALB) our ‘Substitute > P(ANB) = P(BIA)P(A) -> P(AIB) Beye) eet ot | | ar ys’ theorem isa fundamental formula hr probeity theory nd statics [Seat saat) | Ater that describes how to update our belief in'a hypothesis after observing new |s nomaliing feror. cl ‘evidence. It provides @ relationships between four key terms. =u % Law of total probability -> We gr calesifing the rcnl probity of on event (t's say event A) tht can happen ander several iret Set of muiualy exclusive ond exhaustive events By, By ~ 8, (ie, exactly one ofthese 8, must be true) POA) = SPLAIBP(B,) | —* Hi. enna en «partion ofthe Applying Bayes’ Theorem (Example) fea Example > Quality control in a factory A fact that produce widgets. Machine has two machines, Machine A and Machine B, A produces 60% of the widgets: machine B produces 40%. 2% of widgets from machine A are defective. 5% of widgets from machine B are defective. If a widget is selected at random and found to be defective, what is the probability it was produced by machine B? D = defective widget ~ B = widget produced by machi We wont % fhe PID)” PO) P(D) = POIANA) + P(D|B)P(B) = (0.02)(0.60) + (0.05)(0.40) = 0.012 + 0.020 = 0.032, cat ‘theorem: Pel "+ Random variables and probability lstibtion functions > 1 ‘Random Varies > 1 Arann tea ction til mapteanes ta! pads peat enema nee et yn Ih pra Stat er Stas Hal ape piel tas : Ean: The rec loath random raex —t sg the ve eis oe 1 Type of Rando Vries > I Cute Cuties 1 Tit spre coaiiavaew_| Toes van arange Es tbr soe art | Tinea ck — 1 Nbr of set sat ' apie Let be tember of den 3 can tnes 1 Pos ucones HAW. HVT THC, THAIN TIN TTT | lagi ee ; wits 3 t Ht TH, TH 2 t HHT, HTH, THH 1 : ATH TH 1 t * Expectation (Expected Value) > D) = P(O|B) P(B) . {(0.05)(0.40)) . 0.020 RO ade * ooae * Bez) ‘Probability density function (PDF) > Faction that describes the iketbod fa continuous random betaling ona prtelr value ‘arable the probability density function (9) satis: 1) fx) 20 forall 2) The total area under the cre is: 3) The probability that x sn an interval [a,b is Pa sx st)= fF fora continuous random variable x with PDF f(x), the COF, enoted F(), is defined as Fos) = POR 9) = [bat POF: instantaneous deny ata point (hight of he ere) Cor: Accumated probably op toa pint (ares under te cir) The expected value or mean of a random variable represents its long-run average outcome if the experiment is repeated many times. Serves as an indicator about the center of the. Expected value is def possible outcomes, weighted by their probabi ed as a weighted average of all ies For a discrete random variable x that takes x,,%2...%n with probabilities P,, P,... Py the expected value is: Each possible value x; is multiplied by how likely it is to occur (its probability A)| | Why weighted average? Suppose you buy = 3 pens at %10 each ~ 2 notebook at &50 each ‘Tn continuous _probal uous random variable) > ity (like time or weight), probabilities are defined sTHems ops eeaces Spending by a probability density function (PDF). The expected value becomes: %* Probability and repeated - Image a game where: = You win 1 with 80% chance - You win 10 with 20% chance over 100 games, - you win ®1 about 80 times - you win 210 about 20 times Elx] = 1(0.8) + 10(0.2) = 0.8 +2 = 2.80% feeceiedne tle us the average outcome we would expect if we repeated random process many times. Ina game where 99% chance = 20 and 1% chance = 21000, the mode is 20, but the expected value is 210, That you're evaluating whether to play. butions, the expected value helps you plan for not just the most common one. 3) The ‘law of large numbers’ states that as the number of trials of a random experiment increases, the average (m gets closer and closer to the expect Probability. distributions -> ‘A Probability density function (PDF) is function describing the probability density of a continuous variable Let x have POF: f(x) = {7% + OSxS4 ‘The probability that x les in on interval (,8) is find: i Ol Plasxsb)= {"Flddx 1. P(02 ~][2) Soil divin > ~)(3) sn detin > (9) Gon dttton > SE Meany: Raed voter of || us ounber of re eet | Us eu Nmber of als | Ueccase ne il with || es arll s, || fed ler! Gimeipes|| i he ft scene. Parmeter: P= Probaiy of Secs] | Parmeter: n= umber of rl, Parameter: Xe average number || Paramater: P= Probability of BPrebaiityo sucess of event pe intercal ae Po ifxet PME: rel 2-H pwr fPoeida(E) PUA asa PMEPOH) = kao _|| Me 2p ke. mail u z Example: tossing a coin until 0 otherwise || Example: Tossing a coin 5 times, | | Example: A server erashes 2 cn tg cae” | pees [fam eset? lie an aT X~ Bins,05) | pobabity 3 na day Ptx=3)=(3)(05)%05)" Raa se 1001254025 =0.3125, _} [P(e 0.80. a %* Types of continuous probability distributions —> 1) Uniform distribution -> \- Use ease: All outcomes in interval [ob] are equally likely 4 por: #0) = php. forasxsb (Examples Time of arrival between 10:00 and 11:00 is uniform -» X~ U(1O,11) and 10:45: P(0.25 « x < 1075) = 0/1 = 0.5 2) Normal (Gaussian) distribution > oe eiatdS fees exe eeoarnone ay mel arate 20) exon » Parameters: 4 (mean), a (variance | pF: |f60 = phew ay | |: Properties: — ~ Bell-sheped, symmetric obout Plu +0) » 68%, Plu # 20) ~ 95% In the normal distribution, 2 refers to a standardized value known as 0 2-score, The 2-score tells you how many standard deviations a value x is from the mean 41 of the distribution. Example: Suppose exam scores are normally distributed with mean 4.=70, standard deviation What is the 2-score for a student who scored 85? 2 = $5570 = 1.5 A score of 85 Is |.S standard deviations above the mean, 3) Exponential Distribution > S Use case: Time between successive random events. 4 Parameter: Exponential distribution models va the time between independent events occuring at a constant rate, Example > Tf calls come at a rate of 3 per minute, what is the probability that the next call comes within 1 minute? P(X < 1) = 1-8) = 1-9 = 1 - 0.0498 = 0.9502. Conclusion: 95.02% chance that the next call will come within 1 minute ‘if calls are arriving at an average of 3 per minute. Question: The 1Q scores of a large population fot @ normal distribution with a mean 0 and a standard devic 1, Calculate the z-score for an individual with an IQ of 130. Ans: z-score: z = (130-100) 2. Using the standard normal PDF, estimate the ee density at that 2-score. Ans: PDF of standard normal at 2=2: f(z) = Ae ~ 0.05399. 3, Interpret what the PDF value at that z-score represents. Ans: Interpretation: This value (0.054) is the height of the normal curve at 2=2. It indicates the density (not probability) ~ how tightly packed values are around 2=2. this not the probability of getting an IQ score exactly equal to 130? not a probability: In continuous distributions, the probability of any exact value (e.g., P(X=130)) is zero. The PDF valve is only used to compute probabilities over intervals, like P(128 < X < 132). A hypothesis is a formal statement or assumption about a population parameter (like mean or proportion) that we want fo test using sample data. Null and Alternate Hypothesis: In hypothesis testing, we always start with two competing claims. ~ Null hypothesis (Ho): This is the default assumption. It typically says *no effect’, “no difference’, or "no change’. = Bite naive hypothesis (H, or H,): This is what we suspect might be true instead. Example Scenar “=H: Hd = 20 hours (battery life is as claimed) - Hy: 1 < 20 hours (battery life is actually less) We collect data and use it to decide whether to reject the null hypothesis. ‘statistical significance When you compute a correlation coefficient (ike Pearsons r) between two variables (Gay, study time and exam score), you're mmr how shay ey move together. ple might hay P-Value Definition The probably of cbtcning« coreltion as extreme as the one observed ifthe true corelation in the population were actually zero (i,e no correlation), If there is no actual correlation in the population (ie 'p=0), then what is the probability that we would randomly get a sample correlation (r) as large (positive or negative) as the one we just observed? Interpreting P-Values: Low Bale fe, < 0.05) ~The sberved correlation i wily to have ocured by chance > static signa = High P-Vaue (0, > 0.05) > The sbered correlation could ently happen by chance > ot static sigan ‘Alpha (a): 1s the significance level -> the threshold we set before doing a test to decide how much ‘evidence we need to reject the null hypothesis. eck Ras Wo oaiece Ge pace oa ea eee eet oe Example: Fear 0.05, 1F th pa fr your bt 0.08 We 008 < 005 > Reject the ml fate, You accept the risk that there is a 3% chance you might be wrong. % Testing Correlation NNul hypothesis Ho: p= 0 (no correlation in the population). Alternative hypothesls Hj p#0. (some correlation exists). Example: Phone screen time and sleep duration. Let us soy, n=50, and r= -0.35. p= 0.02. Since p S ct, we [reject the null hypothesis" ‘There is a statistically significant negative correlation ~ as screen Time: increases, sleep duration tends to decrease, “Tnsights about P-Value® 1) The stronger the correlation and the larger the sample, the lower the p-volve tends to be. 2) Even a high r-value (e.9. 0.6) may not be staistcaly significant if the sample size is too small 3) The p-value doesn't tell you how big the correlation is ~ only whether it's likely to be real or dye to random variation. Calcd the (Compute the |__4| Oba sample correlation, r >| t-statistic The t-test > A statistical test used to compare a sample with a Known value or to compare two samples, especially when the sample size is small and the population standard deviation is unknown. ‘Two-sample t-test (Example): Compores the means of two Scenario: Does a new teaching method improve scores compared to a tradition ‘Group A (Traditional): mean X= 72, |” Group B (new method): mean X= 78, standard deviation \= 6, standard deviatc sample size n= 15. sample size na = 5. Tig fest oskg: "Ls this 6-poin difference res, or could it have happened jst by random chonce?=] * Confidence interval ‘A confidence interval (Cl) estimates a population parameter (\ike @ mean or a proportion) using a range of pleusible values bull fam sample data ‘What « 95% CI means: If you repeated the same sampling proces infinitely many ines and bul a I each time using the same method, about 95% of these intervals would contain the true parameter Tt des NOT mean ‘there's 0 95% probability the true mean isin this specif interval” The paramoer I ined; the interval i rand, Steps to compute confidence interval (Choose a confidence inten (OY, 95%, 99 (Compute a point eaimate (eg, ale mach, Hat 1 etn $00) ae, oe (SE) of Wal eaimale]— } In unicoum hen $= et a lwimntachath dont iin ea a ‘Assumptions (for mean CI) Dat ae from a rendom sample (or epresenttive, independent cbservation). Population {3 opproimately normal or sample size Is madertely large (ty CLT) ond data have no severe outers / aay shaw. "Standard error for proportions ‘A proportion i.a fraction of the sample with cera characterise (e120 out of 200 sludents passed. = (0.6)"SE measores the spreod of somple propertions around the Irve population proportion Far c= pase) ‘Shaler SE = Sonlepropaton «Sgr = Sante Boro ete Wale Gore vib),

You might also like