0% found this document useful (0 votes)
11 views10 pages

Understanding Statistics: Key Concepts Explained

The document provides a comprehensive overview of statistics, including definitions of key terms such as population, sample, parameter, and statistic, as well as types of data (quantitative and qualitative). It covers various statistical methods, including descriptive and inferential statistics, correlation, and hypothesis testing, along with their applications and limitations. Additionally, it discusses different types of variables, measurement scales, and correlation coefficients, emphasizing the importance of understanding relationships between variables in research.

Uploaded by

Asia Khan
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
11 views10 pages

Understanding Statistics: Key Concepts Explained

The document provides a comprehensive overview of statistics, including definitions of key terms such as population, sample, parameter, and statistic, as well as types of data (quantitative and qualitative). It covers various statistical methods, including descriptive and inferential statistics, correlation, and hypothesis testing, along with their applications and limitations. Additionally, it discusses different types of variables, measurement scales, and correlation coefficients, emphasizing the importance of understanding relationships between variables in research.

Uploaded by

Asia Khan
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Statistics - science of collecting, analyzing, interpreting, presenting and organzing data.

Meant
for quantitative data.
CAIO

Population - all set of individuals of interest on which information or data is being gathered on
Sample - selected individuals from the population which become representative of your
population. Should be according to aim of research/study. Should be representatice of
charactertistics/traits of your pop.

Parameter - the numerical value representing the population.

Statistic - the numerical value representing the sample.

Quantitative data - in forms of numbers


Qualitative data - in forms of interviews/words

Two categorisation of stats


Descriptive statistics - to organize, summarise and simplify the data e.g mean, percentile,
median, mode, frequency, polygon, histogram.
OSS

Inferential statistics - all statistical analysis techniques i.e t-test, anova. Used to study sample,
draw conclusion and generalise results to population.

Median = 50th percentile (n+½)


Mean
Standard dev

Percentile - relative position or ranking in specific group.

Sampling error - the difference between a sample statistic and population parameter due to
sampling variability. the more the difference between population parameter and sample statistic,
the more the sampling error. Happens when you deviate the sample, or a different population is
included other than the targeted one so sampling error increases. Sampling error is minimized
but removing completely is difficult.

Variable - a condition or characteristic that changes or varies from individual to individual and
within individual as well.

Constant - a condition or characteristic that does not vary or change.

Extraneous variable - any variable other than the study variable and has potential to confound
the result of the study
Confounding variable - an extraneous variable that disturbs the results of the research due to
not being constant, matched, or controlled.

Independent variable - the variable that is manipulated in an experiment to observe its effects
on dependent variable, predictor variable

Dependent Variable - the outcome variable that is measured in an experiment affected by the
IV. outcome/criterion variable

Construct - an abstract or hypothetical concept that is measured through indirect measrues


such as depression.

Conceptual definition - a basic general description of the construct.

Operational definition - a specific, measurable definition of a construct describing how the


construct will be measured.

Discrete variable - a variable that can only take specific separate values with no intermediate
values. It is not divisibile..

Continuous variable - a variable that can take any value in a given rage and is divisible for
example scores on a test.

scales/tools/assessments - instruments used to measure variables where scoring of the scale


generates a result. A scale can also include subscales.

Reverse scoring/lie detectors - are items which identify inconsistencies and deceptive
responses, whether participant has done the test attentively and not just gotten done with it
without reading the items and responding accordingly.
Scoring strongly agree as 1 instead of 5 for some items.

Correlation - statistical measure that describes the strength and direction of a relationship
between two variables.

Actual range - the range of observed values in a dataset, from the smallest to the largest
observed value. This should lie between potential range.
Potential range - the minimum and maximum score possible of a questionnaire itself. E.g,
potential range of test score is from 0 to 100 and actual range is 16-88.

Dichotomous variable - a variable with only two categories. Male female, yes no.
Discrete dichotomous - no numerics involved, two qualitative categories such as male female,
dead alive.
Continuous dichotomous - numerics involved in forming the categories e.g pass/fail.
Scales of measurements
Nominal - categories that describe only qualitative nature of the variables with no order or rank.
and without any difference in them e.g gender,.
Ordinal - difference in the categories is seen in the form of ranks. The distance between
categories is not known e.g race positions.
Interval - ordered set of categories with equal distance between them but no true zero, arbitrary
zero. E.g temperature.
Ratio - has a true zero point or an absolute zero e.g height weight.

Hypothesis - a tentative statement or prediction that can be tested.


Null - hypothesis of no difference or no relationship. E.g there is no rship bw study time and test
scores.
Alternative - suggests that there will be a relationship between the variables being studied.
Directional - specifies the direction of the effect for e.g increasing study time increases test
scores.
non directional - does not specify the direction but only that an effect exists, e.g study time
affects test scores.

N=sample size - the number of participants in a study

UL= upper limit - the maximum value for a range of data


LL=lower limit - the minimum value for a range of data

Skewness - measure of the asymmetry of distribution of values in a data set. (probaility of data
being more or less than the mean is higher. -1 to +1 is okay
Positively skewed distribution means right tail is longer, indicating more low values. Mode
median mean
Negatively skewed distributions means left take is longer indicating more high values. Mean,
median mode.
Symmetrical - mean medium mode values equal.

Kurtosis - kurtosis tells us about the peakedness of the data distribution. High kurtosis means
more data points are in the tails (extreme values), and low kurtosis means the data is more
evenly spread out. -2 to +2 is okay.

p=significant value - the probability that the observed results are due to chance. Smaller p value
i.e less than 0.05 indicates stronger evidence. 0.01 very high, 0.03 medium 0.05 significant.

Tests of normality - tests used to determine if a dataset follows a normal distribution.

Shapiro wilk test - helps check if a dataset follows normal distribution by comparing observed
data distribution to a normal distribution. If the p value is small, less than 0.05, then data
significantly deviates from normality.
Kolmogorov smirnov test - compares a dataset to a reference distribution.

We need non-significant results on both the tests as significant results will mean that the differences are
too much or data is not normally distributed, so more than 0.05.

Correlation
- Direction - positive, both increase together or negative, one increases other decreases
- Degree - strength of the correlation with values closer to -1 and +1 indiciating stronger
rship.
- Linear to non linear - linear corr follows straight line corr while non linear ocrr do not.

Effect size
Measure of the magnitude of rship or diff between variables, independent of sample size.
Cohens d common effect size measure.

Standard dev - tells how spread out the values are in a data set from the mean. If sd low, thern
data points close to mean, less variability. If sd high, data points further than mean, high
variability.

Mean - avergae of a set of numbers


Median - middle value in a data set when numbers arranged in an order.
Mode - most frequent value in a data set.

Replace missing values with mean value if only few values are missing

For outliers
Steam and leaf, histogram, normality plots.

Case processing summary = 100% = no missing data.


95% conference interval lower bound and upper bound - 95% of the time, the value of the target
variable falls between the lower and upper bound value. Sig should be 0.5 or lower, same signs
on upper and lower bound = results significant.

5% trimmed mean - 5% of higher and 5% of lower end scores are trimmed. The mean for the
data is calculated then after excluding these scores. Shows diff after eliminating extreme scores
by seeing diff between mean and trimmed mean. If not too much diff, then shows that outliers
on higher and lower ends are not affecting data much. If too much diff, then outliers are creating
issue in the data.
Listwise when want to exclude entire cases, row that has missing values, participants whole
data excluded.
Pairwise - when you only want to exclude the missing values in each pair of variables,
correlation regression analysis.

Corr interpr

0.1 - very weak

0.3 - weak

0.5 - moderate

0.7 - strong

1 very strong
Correlation

A statistical measure that describes the strength and direction of a relationship between two
variables.

Characteristics of relationship

Direction - positive and negative - both together, or opposite


Degree - 0 = no relationship to 1 = perfect positive relationship
Form - Linear = straight line plotted on graph. or non linear = may involve curves or more
complex patterns.

Correlation coefficient has t lie bw -1 and +1.


+1 - two variables perfectly positively correlated, as one increases so does the other by a
proportionate amount

-1 = two v perfectly negatively correlated, as one increases, the other decreases by a


proportionate amount.

0 = no relationship at all, if one variable changes the other stays the same.

If more than +1 or -1, then error in data

Significance value P- value = less than or equal to 0.05 indicates statistically significant
relationship.
the probability that the observed results are due to chance
The lower, the more highly statistically significant

Uses of correlation
Prediction - helps predict value of one variable based on value of another valuable.
Validity - assesses validity of measures by examining relationships between difference
constructs, making sure that a test measures what it is intended to.
Reliability - measures the consistency of results across different instances, helps establish
reliability of assessment or measurement tool.
Theory verifications - helps test theoretical relationships, providing evidence for or against
hypotheses in research studies.

Limitations of correlation
Correlation does not give every insight of the data
Correlation between patients bp and medication used
Time spent on ecommerce website vs money spent by customer.

Correlation does not indicate the direction of causality for example.

Correlation does not mean causation e.g ice cream sales and drowning incidents go up in
summer. doesnt mean that ice cream causes drowning. rathey they are related to warmer
weather

Context matters: The cause-effect relationship can change depending on the situation.

2 problems
The third variable/tetrium quid - causality bw two variables can not be assumed because there
might be other measured or unmeasured variables affecting the results.

Direction of causality - correlation coefficient does not say which variable causes the other to
change, doesnt indicate in which direction causality operates. E.g theres a correlation between
amount of tv watched and number of hours of sleep. Does more tv cause less sleep or does
less sleep cause more tv?

Direction of causality and third variable

Corr is not = to causation, corr doesnt state the direction of causality, context matters, terrarium
quid

Bivariate correlation - measuring correlation between two variables. E.g height anf weight.

- Pearson product moment correlation (parametric test) - when both variables are
measured on interval scale or are continuous. Measures the degree and direction of
linear relationship bw two variables. Rship Age and income in sample of adults, if corr
high, as age increases, income increases (if positive, neg)
Both variables interval scale or continuous. Measure degree and direction of linear rship
bw 2 vrbs. Age and income

- Spearmans rho/Spearmans correlation-rs (non parametric test) - used when the


variables are ordinal. Useful to minimize effects of extreme scores. Rank students by
how fast they run and then how well they do on a test. Spearmans rho will tell if those
who run faster also perform better on the test.

- Kendalls tau (non-parametric) - used when small data set with large number of tied
ranks. If you rank all scores and many scores have the same rank then Kendall tau
used. For example,
Spearmans statistic more popular than kendalls tau but kt statistic is a better estimate of the
correlation in population.
Bivariate corr - pearson product moment corr, spearhmans rho/corr, kendalls tau.

height and weight pearsons

Spearmans - students grade level e.g 9,10,11 and social media usage in categories low
medium high

Kendalls tau - education level and income bracket.

Point biserial and biserial correlation

Gender and test scores


Marital status and income levels

Point biserial correlation rpb - quantifies relationship bw continuous variable and discrete
dichotomy, no continuum in the two categories such dead or alive, you are either one or the
other. Height and gender.
Discrete dichotomy refers to a situation where a variable can take on only two distinct and
mutually exclusive values or categories. Example relationship between gender and height, point
biserial corr will tell if on average male or female differ in height.

Point biserial - continuous and discrete dichotomy, gender and height


Biserail corr - continuous and continuous dichotomty, job satisfaction and monthly salary.

Height and clothing size


Age (continuous) and fitness level (varying levels), poor fair goof excellent

Biserial correlation, rb - quantifies relationship between a continuous variable and a variable that
is continuous dichotomy, there is a continuum underlying the two categories e.g pass or fail.
Range of values between the two variables, introversion/extroversion. Job satisfaction ( and
monthly salary (continuius)

Partial and semi partial correlations. Watch video on corr

Partial corr - correlation between two variables in which the effects of other variables are held
constant. The effect of third variable is controlled on both of the variables. E.g corr between
study hours and exam scores while controlling the effects of sleep hours..

Control for one variable is known as first order partial correlation. For two variables = second
order partial correlation, third order partial correlation.
E.g

Semi-partial correlation
Correlation between two variables where the effects of a third variable are controlled. The effect
of the third variable is controlled on only one of the variables in the correlation

E.g in partial correlation, the calculation takes into account not only the effect of revision (third
variable) on exam performance but also on anxiety, on both variables.

Whereas in semi partial correlation would control for only the effect of revision on exam
performance and effect of revision on exam anxiety is ignored, only one variable.

Partial corr useful for when looking at the unique relationship between two variables when other
variables are ruled out.

Semi partial correlation is useful when trying to explain the variance in one particular variable
(an outcome) from a set of predictor variables.
Test scores based on study hours

Effect size
Correlation coefficient are effect sizes. So no ther calculations necessary.
.1 - small effect
.3 - medium effect
.5 large effect

Calculation of bivariate correlation

Analyse --------- correlate --------- bivariate


Choose which correlation statistics. Default is pearsons product moment correlation but can
also choose spearmans corr and kendalls corr

Comparing dependent correlation coefficients


The split file command is used to compute the corr coefficient between exam anxiety and exam
performance in men and women.

Reporting correlation video

A Pearson correlation was conducted to assess the relationship between hours of study and
exam scores. The results revealed a moderate, positive correlation between hours of study and
exam scores, r(98) = .45, p = .002. This indicates that as the number of hours spent studying
increases, exam scores tend to improve. The correlation is statistically significant (p < .05),
suggesting that this relationship is unlikely to have occurred by chance. With an effect size of r
= .45, the relationship between hours of study and exam scores is moderate in strength.

You might also like