0% found this document useful (0 votes)
8 views9 pages

Central Tendency and Variance Explained

Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
8 views9 pages

Central Tendency and Variance Explained

Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Central Tendency and Data Distribution

 There are three main measures of central tendency:

Mode

median

mean

 Mode

The mode is the most commonly occurring value in a distribution.

Consider this dataset showing the retirement age of 11 people, in whole years: 54, 54, 54, 55, 56, 57,
57, 58, 58, 60, 60

This table shows a simple frequency distribution of the retirement age data.

The most commonly occurring value is 54, therefore the mode of this distribution is 54 years.

 Median

The median is the middle value in distribution when the values are arranged in ascending or
descending order.

Looking at the retirement age distribution (which has 11 observations), the median is the middle value, which
is 57 years:

54, 54, 54, 55, 56, 57, 57, 58, 58, 60, 60
When the distribution has an even number of observations, the median value is the mean of the two middle
values. In the following distribution, the two middle values are 56 and 57, therefore the median equals 56.5
years:
52, 54, 54, 54, 55, 56, 57, 57, 58, 58, 60, 60

 Mean

The mean is the sum of the value of each observation in a dataset divided by the number of
observations.
Looking at the retirement age distribution again:

54, 54, 54, 55, 56, 57, 57, 58, 58, 60, 60

The mean is calculated by adding together all the values (54+54+54+55+56+57+57+58+58+60+60 = 623)
and dividing by the number of observations (11) which equals 56.6 years.

 Finding the Variance In Your Sample


The variance is a figure that represents how far the data in your sample is clustered around the
mean.
Step 1 : Subtract the mean from each of your numbers in your sample. This will give you a figure
of how much each data point differs from the mean.

For example, in our sample of test scores (10, 8, 10, 8, 8, and 4) the mean or mathematical
average was 8.
10 - 8 = 2; 8 - 8 = 0, 10 - 8 = 2, 8 - 8 = 0, 8 - 8 = 0, and 4 - 8 = -4.

Step 2 : Square all of the numbers from each of the subtractions you just did.

Step 3 : Add the squared numbers together. This figure is called the sum of squares.

Step 4 : Divide the sum of squares by (n-1). Remember, n is how many numbers are in your sample. Doing this
step will provide the variance. The reason to use n-1 is to have sample variance and population variance
unbiased.

 Calculating the Standard Deviation


Step 1 : Find you variance figure
Step 2 : Find the square root of you variance

 Coefficient of Variation
the ratio of the standard deviation to the mean

general steps to find the coefficient of variation are as follows:

 Step 1: Check for the sample set.


 Step 2: Calculate standard deviation and mean.
 Step 3: Put the values in the coefficient of variation formula, CV =σ/μ × 100, μ≠0,
Answer: Hence plant D has greater variability in individual wages.

 Frequency distribution of ungrouped data:

Given below are marks obtained by 20 students in Math out of 25.

21, 23, 19, 17, 12, 15, 15, 17, 17, 19, 23, 23, 21, 23, 25, 25, 21, 19, 19, 19
 Probability
ratio of the number of favorable outcomes to the total number of outcomes of an event.

For an experiment having 'n' number of outcomes, the number of favorable outcomes can be denoted
by x. The formula to calculate the probability of an event is as follows.

Probability(Event) = Favorable Outcomes/Total Outcomes = x/n


Z Test
 The z-test can be performed on one sample, two samples, or on proportions for hypothesis
testing.
 It checks if the means of two large samples are different or not when the population variance is
known.
 A z-test can further be classified into left-tailed, right-tailed, and two-tailed hypothesis tests
depending on the parameters of the data.
 For this purpose, the null hypothesis and the alternative hypothesis must be set up and the value
of the z test statistic must be calculated. The decision criterion is based on the z critical value.

 One-Sample Z Test

The formula for the z test statistic is given as follows:

The algorithm to set a one sample z test based on the z test statistic is given as follows:

Left Tailed Test:

Null Hypothesis: H0𝐻0 : μ=μ0𝜇=𝜇0


Alternate Hypothesis: H1𝐻1 : μ<μ0𝜇<𝜇0
Decision Criteria: If the z statistic < z critical value then reject the null hypothesis.

Right Tailed Test:

Null Hypothesis: H0𝐻0 : μ=μ0𝜇=𝜇0


Alternate Hypothesis: H1𝐻1 : μ>μ0𝜇>𝜇0

Decision Criteria: If the z statistic > z critical value then reject the null hypothesis.

Two Tailed Test:

Null Hypothesis: H0𝐻0 : μ=μ0𝜇=𝜇0


Alternate Hypothesis: H1𝐻1 : μ≠μ0𝜇≠𝜇0

Decision Criteria: If the z statistic > z critical value then reject the null hypothesis.

 Two Sample Z Test

The two-sample z test can be set up in the same way as the one-sample test. However, this test will be used to
compare the means of the two samples. For example, the null hypothesis is given as H0: μ1=μ2
 Correlation
correlation tells us how related two variables are
Positive correlations – occur when both variables move in the same direction (e.g., as SAT scores
increase, so to do GPAs).
Negative Correlations – occur when one variable increases, the other decreases (e.g., as age
increases, the number of speeding tickets decrease).

Common questions

Powered by AI

Null and alternative hypotheses provide a framework for testing assumptions about population parameters using sample data. The null hypothesis typically states no effect or difference, serving as a baseline. The alternative hypothesis posits the presumed effect or difference. In a Z-test, evaluating the test statistic against the critical value helps determine whether to reject the null hypothesis, guiding conclusions on statistical significance and validity of effects claimed .

To calculate sample variance, first subtract the mean from each data point, then square these differences. Next, sum the squared differences and divide by (n-1), where n is the number of observations. This adjusts for bias in estimating population variance from the sample. Variance quantifies the extent of data spread around the mean, providing insight into data variability; higher variance indicates greater dispersion .

The Z-test evaluates whether the difference between sample means is statistically significant by comparing the z statistic to a critical z value. In a one-sample test, the null hypothesis posits that the sample mean equals a known population mean. Decision criteria involve rejecting the null hypothesis if the z statistic exceeds the critical z value, depending on whether it's a one-tailed or two-tailed test. Two-sample tests compare means of two samples, utilizing similar rejection criteria .

The mean is the arithmetic average of a dataset and is sensitive to extreme values, making it less reliable in skewed distributions. The median, being the middle value, is resistant to outliers and thus a more robust measure in the presence of skewed data. The mode, the most frequently occurring value, provides insight into the distribution's peak. In a skewed dataset, the mean moves towards the tail, the median lies between the mode and mean, and the mode remains at the peak of the distribution .

To calculate the probability of an event, determine the number of favorable outcomes and the total number of possible outcomes, then divide the former by the latter. This ratio quantifies likelihood, guiding decisions in uncertain conditions. Probability theory underpins risk assessment and decision-making by providing frameworks for predicting event likelihoods, aiding in strategy development and anticipatory planning .

Standard deviation quantifies the average deviation of each data point from the mean, offering insights into data spread. A small standard deviation suggests tightly clustered values near the mean, indicating consistency or predictability, whereas a large standard deviation implies more spread and variability. This measure is integral for understanding normal distribution properties and likelihoods of values falling within certain ranges .

The coefficient of variation (CV), calculated as (standard deviation / mean) × 100, expresses variability as a percentage of the mean, making it unitless. This allows for direct comparison of variability across datasets of different scales or units. A higher CV indicates greater relative variability or risk in relation to the mean, useful in contexts like finance where varying investment scales exist .

The median is preferable when dealing with skewed distributions or datasets with outliers, as it is less affected by extreme values than the mean. In situations with non-normal distributions, such as income or real estate prices, where there's significant right-skew, the median provides a better central value representation, avoiding distortion from atypical data points .

Correlation quantifies the direction and strength of a linear relationship between two variables, influencing how hypotheses are formulated in studies. A positive correlation indicates variables move in tandem, while a negative correlation implies inverse movement. Strong correlations suggest potential causal relationships or areas warranting deeper investigation, impacting hypothesis structures and leading to targeted testing or observational study designs .

In multimodal distributions, multiple modes represent the dataset's local maxima, each indicating a peak value. Using mode as a central tendency measure in such datasets may suggest clusters within the data, complicating interpretations and potentially indicating underlying subgroupings or processes. Identifying all modes can inform about distinct population features or heterogeneity, guiding more nuanced analyses .

You might also like