0% found this document useful (0 votes)
4 views30 pages

Stat - Notes

The document provides an introduction to statistics, covering key concepts such as data types, sampling methods, and the distinction between populations and samples. It explains the importance of statistical significance, probability, and the research process, along with descriptive and inferential statistics. Additionally, it discusses the normal probability distribution and the central limit theorem, emphasizing the significance of sample size in achieving normality in data distribution.

Uploaded by

emmy
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views30 pages

Stat - Notes

The document provides an introduction to statistics, covering key concepts such as data types, sampling methods, and the distinction between populations and samples. It explains the importance of statistical significance, probability, and the research process, along with descriptive and inferential statistics. Additionally, it discusses the normal probability distribution and the central limit theorem, emphasizing the significance of sample size in achieving normality in data distribution.

Uploaded by

emmy
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Week 1 + 2: Intro to statistics

Thursday, September 21, 2023 9:01 AM

What are the data like? How would you record the answers from the respondents.
- Data: are collection of observations, such as measurements, genders, or survey responses ( excel sheet) (what you would like to observe?) => we look for the
NUMBERS
+ Example: Name? Students' year? How much time you spend on studying? How many…? => The responses can answer for the question HOW MANY?, help
us to measure the quantity
- For quantitative research, you should ask CLOSE-ENDED questions: "Which one do you prefer?"/ Yes-No questions

Statistics: the science of planning studies and experiments, obtaining data, and organizing, summarizing, presenting, analyzing, and interpreting (make the
number sense) those data and then drawing conclusions based on them => to clarify WHAT DO THEY MEAN?

How can we interpret the number? What conclusion can we make based on the numbers/ results?
1. Study population + sample
- A population is the complete collection of all data that are being considered to make inferences about.

@KIEUANHNG
- Sample: is the subcolletion of members selected from a population, carries all the characteristics of the population.

L2 learners
Sample is the smaller size of population

Population of interest

Why poplation and sample?

PARAMETER vs. STATISTIC

Parameter is term for


characteristics of population

Statistic is term for


characteristics of sample

Parameter Statistic
A population parameter is a numerical A statistic is a numerical measurement
measurement describing some characteristic of a describing some characteristic of a Sample
Population
Example:
Population mean: 𝜇 (𝑚𝑖𝑙) Sample mean: 𝑋 (exboa)
Population standard deviation: 𝜎 (𝑠𝑖𝑔𝑚𝑎) Sample Standard deviation: s
Population variance: 𝜎 (𝑠𝑖𝑔𝑚𝑎 𝑠𝑞𝑢𝑎𝑟𝑒) Sample variance: 𝑠
Total size of population: N (big N) Total size of sample: n (small n)

RESEARCH PROCESS:

1. Prepare
2. Analyze
3. Conclude

----------------------------------------------------------------------------------------------------------------------------------------------
PROBABILITY
The relationship between probability and possibility?
Probability is science while possibility is fate

Classic samples of probability:


- A dice: 6 faces
A) P(1) = 1/6 = 16.6 %
b) P (1 or 5) = 1/3 + 1/3 = 0.6667 = 33.3%
@KIEUANHNG
C) P(1 and 5) = 0 because the dice only can show 1 face, cannot show 2 faces at the same time.

Definitions of Probability
P(A): probability of event A occuring
The most common approach: Relative frequency approximation of probability

P(A) =

STATISTICAL SIGNIFICANCE
A finding is very unlikely to occur by chance.

Statistical significance = if the likelihood of an event occurring by chance is 5% or less.


A significance level of 0.05 indicates a 5% of risk of concluding that the finding is untrue or due to chance.
=> The percentage for sth to happen for sure is 95%.
A finding is very unlikely to occur by chance.

Statistical significance = if the likelihood of an event occurring by chance is 5% or less.


A significance level of 0.05 indicates a 5% of risk of concluding that the finding is untrue or due to chance.
=> The percentage for sth to happen for sure is 95%.

If we expect sth to happen, this can happen or not, but there are some chances for it to happen.

VARIABLES
- Definition:
+ an image, perception or concept that is capable of measurement - hence capable of taking on different values.
=> 1 variable may have different values
+ A characteristic or attribute of an individual or an organization that can be measured or observed and that
varies among the people/ organization being studied.
+ Variables ~ constructs
(Khi vẽ biểu đồ, có bao nhiêu trục thì bấy nhiêu variables)
V1: Gender

Boy Girl

@KIEUANHNG
V2 Blue 1 2 => Boy, girl, pink, blue are values.
Color
Pink 3 4

Concept Variable
- Can not be measured - Can be measured
- Subjective meanings of mental images/ - Degrees of accuracy vary upon a
perceptions vary markedly among different measurement scale
people

- Ex: gender stereotypes; effective - Ex: gender and color

Satisfaction, preference, anxiety, attitudes, motivation,


translation
Week 3:
Thursday, October 5, 2023 8:19 AM

Quantitative data/ Numerical Data/ Continuous Qualitative data/ Atribute Data/ Categorical Data
Data
- Consists of numbers (counts, frequencied, - Consists of names or labels
measurements,…)

NOTE: Numbers such as IDs, codes, age groups or names are categorical data
(qualitative)

TYPES OF DATA: (exercise)


1. Proficiency level (low, mid, advance): => qualitative (don’t have specific numbers)
2. Word count in essays: => quantitative

@KIEUANHNG
3. Attitudes (e.g, strongly agree, agree, neutral,..) => qualitative

Quantitative DISCRETE vs. Quantitative Continuous Data:

DISCRETE CONTINUOUS
- Countable - Measureable
- Contain distinct or separate values - Contain any value within range
- Counts in whole numbers/ integers (so nguyen) - A scale can take any numeric value
- Cannot be broken into fraction/demical - Can be broken down into fraction/demical (so
thap phan)
- Even if being rounded up/down => still
continuous

LEVEL OF MEASUREMENTS:

Class level
@KIEUANHNG

(
-------------------------------------------------------------------------------------------------------------------------------------------------------

COLLECTING SAMPLE DATA (p.27)


Sampling methods:
1. Simple random sample: (ideal method)
2. Systematic Sample
3. Convenience sample
4. Stratified sample (divide into group => select)

@KIEUANHNG
5. Cluster Sample

ERRORS IN SAMPLING
- Outliers ~ extreme values: unsual observation of data that does not fit the rest of the data.
- Possible causes of outliers:
+ a defective counting device
+ a participant who did not understand the instructions min Max (number of participant)
+ a participant that does not belong to the study population of interest
+ data entry errors or typos
- Solutions:
+ Check outliers
+ Discard outliers with justification

- Sampling bias: is created when a sample is collected from a population and all members of the population are not equally
likely to be chosen as others.
- A sample should be representative of the population
=> Incorrect conclusions drawn abt the population that is being studied. The purpose of research is to find underlying truth
WEEK 4 + 5: DESCRIPTIVE and INFERENTIAL STATISTICS
Thursday, October 12, 2023 8:30 AM

Descriptive Inferential
- To organize and summarize the sample data - To make generalizations abt the unknown
population from sample
- Just focus on sample, cannot tell any thing much
about Popu - Can tell sth about pp from the sample

- HOW? Describe data visually (table, chart,…) and - HOW?: Use statistical tests (e.g. t-tests, chi-square
numerically by measures (chỉ số đại lượng): e.g tests) to estimate the PP parameter from sample
measures of data center, measures of data variation data, and hypothesis testing.

1. DESCRIPTIVE DATA

Definitions
@KIEUANHNG
- Def: to summarize or describe relevant characteristics of data:

- 3 Groups of characteristics:
1. Measures of center: mean, median, mode

The number or average of the numbers in the middle median


Terms Data set A: 1 2 3 4 5 5 Data set B: 1 2 3 4 5 5 6
3.5 4
The number that occurs most mode 5 5
An average mean 3.33 3.7

+ Being "resistant": A statistic is resistant if the presence of extreme values (outliers) does not cause it to change very much.
Q: Which measure of center (mean, median, mode) is resistant to outliers? Which is not? => median, mode (do not change); mean (not
resistant).

2. MEASURES OF VARIATION:
E.g:
Set A: -10, 0, 10, 20, 30 Set B: 8,9,10,11,12
Mean 10 10
Range = max - min 40 4

Derivation from mean (population)


For each data point: 𝑋 − 𝜇
∑( )
For all data:
∑( )
+ Population Variance: 𝜎 =
∑( )
+ Population Standard Deviation: =
𝑋 Mean (average)
Derivation from mean (Sample) 𝑋 Giá trị chạy từ số đầu tiên -> số
For each data point: 𝑋 − 𝑋 cuối
∑( )
For all data: = ∑ tổng
∑( ) s Sample standard variation
+ Sample variance: 𝑠 =
𝜇
∑( ) 𝜎
+ Sample standard Deviation: s =

Exercise:
Sample Standard Deviation
Set A
S= = 15.8

Set B ( ) ( )
S= = 1.58
DEVIRATION OF RELATIVE STANDINGS ( sgk/112)
- Z score:
Scenario: The IU to recruit new students. The VNU SAT and NHE (National Highschool Exam) are accepted.
VNU SAT NHE
Questions: 120 items 3 subjects: 10pts/each
Score range: 0 - 1200 Score range: 0-30
Student 1 with 800 points Student 2 with 24 points
=> SOLUTION: put them on the same scale. The pretest and posttest should have the same scale, should have equivalence, format.
=> convert them into Z SCORE
Apply z scores:
VNU SAT NHE
Ques: 120 items 3 subjects: 10pts/each
Score range: 0 - 1200 Score range: 0 -30
Mean: uSAT Mean: uNHE
SD: oSAT SD: oNHE
Student 1 with 800 (X1) Student 2 with 24 (X2)

VNU SAT: NHE:


Mean = 700 Mean = 22
SD= 2.5 SD= 1.5
=> Z1 = (800-700)/2.5 = 40 => z2 = (24-22)/2.5=13

@KIEUANHNG
Percentiles:
- A percentile indicates the position of a value in a distribution
- Percentiles range from 1 to 99.
Example: You score better than 90% of the people on the exam => you are at the 90th percentile for score. Your percentile rank is 90

𝒏𝒖𝒎𝒃𝒆𝒓 𝒐𝒇 𝒗𝒂𝒍𝒖𝒆𝒔 𝒍𝒆𝒔𝒔 𝒕𝒉𝒂𝒏 𝒙


Percentile of value x = . 𝟏𝟎𝟎
𝒕𝒐𝒕𝒂𝒍 𝒏𝒖𝒎𝒃𝒆𝒓 𝒐𝒇 𝒗𝒂𝒍𝒖𝒆𝒔

QUARTILES AND BOXPLOT


1. QUARTILES:
WEEK 6: NORMAL PROBABILITY DISTRIBUTION
Thursday, October 26, 2023 8:28 AM

Population Sample
Properties to describe parameter statistic
Mean 𝜇 (𝑚𝑖𝑙) 𝑋 (exboa)
Standard Deviation 𝜎 (𝑠𝑖𝑔𝑚𝑎) s
Number N n

HISTOGRAMS

- What does the horizontal scale represent? Classes/ categories (in the ex is how much time)
- What does the vertical scale represent? - frequencies

• Purpose: => phải viết chi tiết giống như ielts


Horizontal line represent what?
Vertical line represent what?
What info the histogram can show us?

sec @KIEUANHNG
- To visually display the shape of data distribution - not symmentric or symmentric
- To show the location of the center of the data such as the mean/mode - Around 160 sec - 170

- To show the SPREAD of the data - wide spread from the mean: tính range = max - min (check
by the dot diagram => nêú nó dàn trải thì wide spread, còn nếu nó cao và hẹp (chiều ngang) =>
narrow spread)
- To show the extreme values in data (outliers): No outliers: there is any value outside - do not
fit/ or not in the data
- To show changes in data distribution: No changes Narrow spread Wide spread
Example:

1. Shape of data distribution: not symmentric


2. The location of the center: (we can not know the mean => use mode to define the center): 3.2 -
3.4
3. Wide spread
4. No outliers
5. No changes

=> CHANGES: increase - decrease -


increase - decrease (many times go ups
and downs)
It could be 2 groups of students: students
in rural areas, and students living in the
city (who are offered better education)

=> DIVIDE INTO 2 GROUPS:


Mode 1: 3.8 => center group 1
Mode 2: 9.0 => center group 2

Group 2
Group 1

SHAPES OF HISTOGRAM

1. Normal distribution - symmentric shape


2. Skewness: positively skew/ negative skew (skew means lack of data) => nhìn chân
@KIEUANHNG
WEEK7: STANDARD NORMAL DISTRIBUTION
Thursday, November 2, 2023 8:22 AM

P.231/ A special type of normal distribution:


Three properties of the standard normal distribution:
+ Bell shaped + symmetric
+𝝁=𝟎
+𝝈=𝟏

CENTRAL LIMIT THEOREM ( CLT) - page 265


For all samples of the same size n with n > 30, the sampling distribution of 𝑋 (sample mean) can be
approximated by a normal distribution with mean and standard deviation 𝝈/√𝒏
If your sample size is greater than 30, you data distribution will be normal
If n < = 30, you have to check your data is normal or not => assess

TRUE/FALSE:

@KIEUANHNG
1. The sampling distribution of a sample means tends to be a normal distribution as the sample
sizes increases => T
2. There are four components of the CLT. => F
3. A sample size of 20 (n=20) is considered large enough for being normally distributed => F
4. As a sample size decreases, a test for normality should be conducted. => T

Distribution of sample means for an intendent random variables will get closer to a normal
distribution as the sample sizes get bigger.

ASSESSING NORMALITY (incase your sample size is less than 30 => not normal data)
Procedure for assessing if sample data are distributed normally
1. Histogram: bell-shaped
2. Outliers: + No outliers
+ Outliers are due to error or chance variation
3. Normal quantile plot:
+ The points lie close to a straight line
+ No systematic pattern that is not a straight line
WEEK 9: INFERENTIAL STATISTICS
Tuesday, December 5, 2023 3:56 PM
descriptive (cannot state strong inferential statistics (strong claim)
claim)
to organize and summarize (sample) to make generalizations about an
data unknown population from sample
data.
describe data visually (like chart, use statistical tests (e.g., t-tests, chi-
blocks,…) n numerically by measures square tests) to estimate the
(= đại lượng): population parameter (such as
e.g, measures of data center, mean, sd, proportion) from sample
measures of data variation, measure data, and hypothesis testing
of relative standing (z score,
percentile, quartile, boxplot)

I. Basic concepts of inferences from 2 samples


1. Independent:
Two samples are independent if the sample values from one population are not related to or somehow naturally paired
or matched with the sample values from the other population.
eg: whether e-learning effect or not → 2 groups:

Group 2: only f2f


@KIEUANHNG
Group 1: only online learning

⇒ 2 gr k trộn lẫn, ng nhóm nào ở yên không nhảy qua lại


• dùng để compare independently

2. Dependent (p429)
Two samples are dependent (matched pairs) if the sample values are somehow matched, where the matching is based on some
inherent relationship.
eg: a class: on-going, mid, final
ID mid final
1 7 8
2 6 5
7-8: scores of same student 1 ⇒ 7 and 8 are related data
⇒ sample are pair
⇒ chỉ có 1 group
⇒ thường là pre+post test

3. PRACTICE
Which type of samples are they?
- Writing scores of group A are compared with writing scores of group B
→ independent b/c student in group A is different from group B
- Pre-test writing scores of group A are compared with post-test writing scores of group A
→ dependent
- Relationship between parents’ education and the academic success of children
→ dependent b/c participants having relationship (parents-children)

Group A Group B
Pretest 1 2
Posttest 3 4
Delayed posttest 5 6

within gr: 1-3, 3-5, 1-5, 2-4, 4-6, 2-6 → dependent


between gr: 1-2, 3-4, 5-6 → independent
Pairwise: chỉ so sánh được 2 group

A B
Pretest 1 2
Posttest 3 4

II. Hypothesis
A hypothesis is a claim or statement about a property of a population.
= claim, prediction
→ dont have hypothesis for sample
prediction: right or wrong ⇒ phải test mới biết

hypothesis test (or test of signigiance)


is a procedure for testing a claim ab a property of a population

Type of hypothesis
2 tails: chỉ biết khác nhau thằng nào lớn hơn thì k biết
one tail: nhiều info hơn, biết nào lớn hơn cái nào
directional hypothesis = alternative hypothesis

Procedure for hypothesis testing (data analysis, proposal)

reject = not true


fail to reject = true
thường ngta muốn reject H0 tại H0 true thì còn kiếm chi nữa
H0: T ⇒ mu1 = mu2 ⇒ The [means] are equal/ there is no difference in [mean]
H0: F ⇒ mu1 ≠ mu2 ⇒ The [means] are not equal/ there is some difference in [mean]/ a difference exist
significance level alpha of 0.05 indicates a 5% risk of concluding that a difference exists; in fact, they are equal/ there is no actual
difference

Significance level a of 0.01 indicates a 1% risk of concluding that a difference exists; in fact, there is no actual difference.
WEEK 10:
Thursday, December 7, 2023 8:38 AM

PROCEDURE FOR HYPOTHESIS TESTING

To compare and contrast


2 groups
(Control gr vs experiment
gr in terms of 1 variable)/
2 samples
Q: What are the
differences between 2
sample/ 2 groups

= correlation,
connection
Q: What is the
relationship bwt 2
variables?/ Is there any

@KIEUANHNG relationship btw…?

Step 5: Calculate estimate statistics of popu parameter:


Step 06: Find the value of the test stat & p-value method
Statistics procedure = statistics test

After a run a statistics (such as t-test, pearson correlation, etc.), we need to report two most fundamental values:
1. Test statistics (362): is value in making decision (to reject - not true or fail to reject - true) abt the null hypothesis
=> make conclusion about whether it is equal to some claimed values or not
2. P-value (365)
- Decision criteria for the p-value method:
+ p-value ≤ significant level α, reject H0 ("If the P-value is low, the null must go")
+ p-value > α, fail to reject H0

@KIEUANHNG
(Verbs for statistic procedure in quantitative rs, hypothesis testing: "test", "examine")

Example:
α = 0.05
H0: µ = 50
HA: µ ≠ 50
Decide if you can reject or fail to reject H0:
A) p = 0.025 => reject H0
B) p = 0.3 => fail to reject
C) p = 0.000 => reject
D) p = 0.431 => fail to reject
E) p = 0.05 => reject

Final Conclusion for Hypothesis Test


E.g: H0, µ1 = µ2
P-value decision conclusion
P-value ≤ α Reject H0 There is sufficient evidence to support the rejection of the
claim that …. (H0)
P-value > α Fail to reject H0 There is not sufficient evidence to support the rejection of
the claim that …. (H0)

Find the value of the test stat, the critical values, and critical region

Step 07: Construct a confidence interval (for population)


A confidence interval estimates a population parameter contains the likely values of that parameter.
Reject H0 Fail to reject H0 Reject H0
- vô cùng + vô cùng
Lower limit Upper limit

@KIEUANHNG
REPORTING TEST STATISTICS & P-VALUE
(LL)

E.g. Independent sample t-test (α = 0.05)


(UL)

H0: µ1 = µ2 or µ1 - µ2 = 0
HA: µ1 ≠ µ2 or µ1 - µ2 ≠ 0

95% CI [1.5; 4.1]


- The CI of the difference 95% CI [1.5; 4.1] does not include the claimed value of a population difference (µ1 - µ2 =
0), reject H0
- There is sufficient evidence to support the rejection of the claim that the population mean scores are equal.
- There is sufficient evidence to support the alternative hypothesis that the population mean scores are different.

APA style report: The research can be 95% confident that the difference (hiệu số) between the population mean
scores ranges from 1.5 to 4.1.

----------------------------------------
CHI-SQUARE TEST FOR INDEPENDENCE
Cột dọc:

Class level: Nominal: cùng năm nhất có ng học ie có ng học ae, năm 2 năm 3 có thể học chung lớp, ko thể
so sánh group nào level cao hơn group nào
Type equation here.
@KIEUANHNG
@KIEUANHNG
WEEK 12:
Thursday, December 21, 2023 8:44 AM

PROPORTION

Example: At IU with n students


Number of boys: x1
Number of girls: x2
1. Propotion of boys in at IU = =𝑝
2. Proportion of girls in at IU = = 𝑞
=> q is complement of p: q = 1 - p

POPULATION PROPORTION (p) vs. SAMPLE PROPORTION (p


1.
1. The proportion of male students at all uni in HCMC => population
2. The proportion of male students at IU => sample
2.
1. The proportion of male students at Le Hong Phong High School => sample
2. The proportion of male students at all highschools in HCMC => population

A. Group 1 (IU) with 2600 students


B. Group 2 (LHP) with 2600 students
1. In group 01, 1950 males => proportion = 0.75
2. In group 02 (LHP) 1560males => proporion = 0.6

H0: p1 = p2 or p1-p2 = 0 or =1
HA: p1 ≠ p2 or p1 - p2 ≠ 0 or ≠ 1
H1: p1 > p2 or p1 - p2 > 0 or >1
H2: p1 < p2 or p1 - p2 < 0 or <1

The population proportion of males at universities is different from the population


proportion of males at high schools?

@KIEUANHNG
Requirements:
1. Sample proportion are from two simple random samples
2. The two samples are independent
3. For each of two samples, there are at least 5 data points (np>= 5, nq>=5)
Nếu chỉ đơn giản so sánh boys and girls trong 1 trươờng thì chỉ cần descriptive là được
rồi, khi nào cần so sánh giữa 2 trường => inferential

POOL SAMPLE PROPORTION/ OVERALL PROPORTION


CI for R
So sanh CI với 1

=> DECISION: fail to reject H0


=> THE CONCLUSION:
@KIEUANHNG
present
Thursday, January 4, 2024 8:46 AM

Group 1: test normality: normal distribution: requirements for each group

RQ:
How effective…
What is the effectiveness

Causal -> use t-test


Muốn so saánh đâầy đủ để khăẳng điịnh kết quả của study => 4 caái so sánh:
Pretest - pretest in group
Post-test - posttest
Pretest - posttest in experimental group
@KIEUANHNG
Pretest posttest in control group

Đừng lấy ordinal variables để tính mean, ordinal thì chi-square


Những số có số thập phân đâu có nghĩa => use continuous data not discrete data

Xeét female hay male thì chi-square.

You might also like