0% found this document useful (0 votes)
7 views7 pages

Understanding Normal Distribution Basics

USMLE

Uploaded by

singhaamrita688
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views7 pages

Understanding Normal Distribution Basics

USMLE

Uploaded by

singhaamrita688
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

2 5312451306189753053

Transcribed by [Link]. Go Unlimited to remove this message.

Hi everyone and welcome to our module on basic statistics. Let's suppose we randomly
selected some healthy subjects and measured their blood glucose levels. We would get a
collection of readings kind of like what I've shown on the screen here.

There'd be a few people with low numbers like 85, and there'd be a couple of people with
higher numbers like there are some 112s in here, and then there'd be a lot of people in the
middle. And if we plotted the blood glucose level on the x-axis and the number of subjects that
had each level of blood glucose on the y-axis, we would get something that looks like this.
There'd be some number in the middle here that was very common.

A lot of people had this number in the middle, and as you went above it this way or below it
that way, you would find less and less subjects had each measurement until you got down to
some number that nobody had on either side. And it turns out a lot of things in nature have a
distribution like this. We call this a normal or a Gaussian distribution, and it's clustered around
some central number, and then it tapers out evenly as you go away from that central number,
either higher or lower.

And many things are like this. For example, if we looked at height of people, we could find a
distribution like this. If we looked at weight, if we looked at the number of days people stay in
the hospital for treatment for pneumonia, we might find a cluster around some central number
of days, and then tapering out as you go above or below that.

So many things in nature have this distribution, and therefore we spend a lot of time talking
about how to characterise it in statistics. So the first characterisation of that normal distribution
I want to talk about are measures of central tendency. So you notice that the curve looked
something like this, and it tended to collect around a central number, and we call that a central
tendency.

This is the centre of a normal distribution, and there are three ways to characterise that central
tendency of the distribution. The first is the mean, which is the average of all the numbers. The
second is the median, which is the middle number of the dataset when they're all lined up in
order, and the third is the mode.

It's the most commonly found number, and we'll talk about all three of these now. So just to
keep this simple, instead of all those blood glucose levels, let's imagine that we had six
recordings of systolic blood pressure, and they were 90, 80, 80, 100, 110, and 120. Well, the
mean would be the sum of all six of those numbers divided by six, so that's very easy to get.

You've probably done this many times when you average things out, and that would be 96.7 for
this dataset. The mode is the number that comes up most frequently, and in this dataset there
are two examples of the number 80, so the mode would be 80. The median is the middle of the
dataset, and the way we determine the medium is a little bit different depending on whether
we have an odd number or even number of data elements.

If we have an odd number of data elements, like suppose we just had three, 80, 90, and 110,
then the middle number is the median, and that would be 90. If we have an even number of
data elements, let's suppose we had 80, 90, 110, and 120, then the median is going to be
halfway between the middle two numbers, so it's going to be 100. So just remember that the
median is a little bit tricky to measure.

It's different from the other two, and it depends on whether there's an odd or even number of
data elements. And very importantly, you have to put all the data in order in order to find the
median. First thing you have to do is start with the lowest number and work your way up to
highest number.

Once you have them laid out, you can figure out the median value. So let's get back to this chart
of blood glucose levels on the x-axis and the number of subjects with each recording on the y-
axis, and let's figure out where the mean, median, and mode are for this distribution. Now I
haven't given you any numbers down here on the x-axis or on the y, but we can still figure out
where the mean, median, and mode are just based on the shape.

So the mode is very simple. It's always the highest point. Remember, the mode is the number
that occurs most commonly, and we have number of subjects over here on the y-axis, so that's
the number of times that recording occurs.

So whatever this blood glucose level is is the most common, and that's why it's the highest
point on the curve. If the data are evenly distributed, like in this example here where there's an
even number of measurements on the left side and the right side of the mode, then the mean
and median are equal to the mode. So anytime you see a nice, normal, even distribution like
this, the mean, median, and mode are always at the top right in the middle.

Now data sets aren't always evenly distributed. Sometimes they are what we call skewed, and
data sets can have either a positive or negative skew, and the way to keep straight what's a
negative skew and what's a positive is to look for the tail. So the tail of this data set is pointing
in this direction, and that's towards the negative numbers over here.

So when the tail points that way, it's called negatively skewed, and the way you figure out the
mode, median, and mean in a negatively skewed data set is as follows. First of all, remember
that the mode is always the highest point, so that's the easiest thing to find. Next, remember
that the mean is always furthest away from the mode towards the tail, and then the median
occurs in the middle, and these rules will still hold up when the data are positively skewed.

So here's an example of a positively skewed data set, and you can see that the tail is pointing in
this direction. It's pointing away from the really highly positive numbers over here. So once
again, the mode is the highest point.

That's always easy to find. The mean is always the furthest away from the mode, and then the
median is in the middle right here. So a couple of key points to remember on these measures
of central tendency.

If the distribution is equal, if it's normally distributed, then the mean, mode, and median are all
equal to one another. Remember that the mode is always at the peak, and then in skewed data
sets, remember that the mean is furthest away from the mode toward the tail, and the median
is in the middle. And then a final point, and something that's very high yield to know for your
exams, is that the mode is the least likely to be affected by outliers.

When you add an outlier, it means you add one measurement that's very far away from all the
others. This will have a slight effect on the mean and the median, but it's only going to affect
the mode if it changes the most common number, and it's very unlikely that an outlier is going
to change the most common number. That's why it's called an outlier, because it lies out and
away from all the other numbers.

So it's very unlikely that it's going to affect the mode. So the mode tends to stay in place when
you add some unusual measurements to a data set. The mean and median are slightly affected
by adding one outlier.

So now I'm going to switch gears from talking about measures of central tendency and talk
about measures of dispersion. Let's look at these two data sets charted on this slide. They both
have a mean, median, and mode equal to 10.

Their measures of central tendency are identical, but you can tell from looking at them that
they're very different data sets, and what's different about these two data sets are the
measures of dispersion that characterise them. So there are many different measures of
dispersion that people use, but the one that you're most commonly going to see is the standard
deviation. You may also see variance or standard error of the mean.

Sometimes you will see z-scores, and sometimes you'll see confidence intervals, and we'll go
over all five of these measures in the next few slides. So let's talk about standard deviation first.
So standard deviation is represented by this symbol sigma right here, and this is the equation
for standard deviation.

Now you're very unlikely to be asked to calculate a standard deviation on step one of the
boards, but you do need to know how to interpret standard deviations. So let me show you how
you get the number, and then we'll talk about how you use it. So first let's talk about the
elements of this equation right here for standard deviation.

x minus x mean is the difference between each data point and the mean, and the sum of all the
x minus x means is the sum of all the differences. So you're adding up how far away each data
point is from the mean. Then you're going to square those differences, and then you're going
to divide by the number of samples minus one, and then you're going to take that entire thing
and take the square root of it right here.

So let me show you how you would do this for two different data sets. So we've got group one
over here represented by these numbers right here, and the mean of all those numbers is 10,
and then we've got group two over here represented by these numbers. So the first thing that
we're going to do is determine the difference from the mean of 10 for each number.

So 9 for example is negative one from that mean, and 12 for example is plus two away from
that mean. And we're going to do the same thing over here for group two, and so we get a
collection of numbers that represent the difference of each data point from the mean. Next
thing we do is square all those numbers, so this is going to eliminate all the negative numbers
and give us only positive numbers, and then we add them up.

So for group one they add up to 11 because most of the data points in group one are close to
the mean. In group two they add up to 96 because the data points in group two are further
away from the mean. And now we're going to plug it into this equation up here in the top right,
and we're going to take that summed value, divide by the number of samples minus one, which
is seven in both groups, and then take the square root of it.

And what we see is that the standard deviation for group one is 1.24, and the standard
deviation for group two is 3.7. This means that group one is much more tightly orientated
around that mean compared to group two, which is much more spread out. Even though they
have the same mean, the data are more dispersed in group two, and that's why you have a
larger standard deviation in group two. So one of the things you definitely need to know for
step one is how much of the data set you capture if you go one or more standard deviations
away from the mean.

So let's look at this data set right here, which has a mean right at this point I'm drawing here. If
we go plus one standard deviation in this direction and minus one standard deviation in that
direction, we will capture 68 percent of the data points within that range. If we go plus two
standard deviations and minus two standard deviations, we will capture 95 percent of the data
points.

And finally, if we go plus three and minus three, we will capture 99.7, virtually all of the data
points within that range. So let's go back to those two curves I showed you before. They both
have the same mean, median, and mode, but this one on the right is much more dispersed
than the one on the left.

Well, the rules still hold that if you go plus or minus one standard deviation on the left, you
capture 68 percent of the data set. Same thing on the right. It's just that plus or minus one
standard deviation represents more distance on the x-axis over here on the right than it does
on the left.
And the same is true for the 95 percent intervals and the 99.7 percent intervals. So let me give
you an example of how this kind of information might show up on an exam. Imagine that a test
is administered to 200 medical students.

The mean score is 80 with a standard deviation of five. The test scores are normally distributed.
How many students scored greater than 90 on the test? Well, the first thing we need to figure
out is how far away from the mean 90 is, and 90 is two standard deviations away from the
mean.

The standard deviation is five. 90 is 10 points away from that, so that's two standard deviations.
So we know that plus or minus two standard deviations is 95 percent of the data.

I've shown it down here on a picture. What you need to realise, though, is that means five
percent of the data are outside this range, and that five percent is split up so that 2.5 percent is
higher than this range and 2.5 percent is below this range. So this means that 2.5 percent of
students scored above 90.

That's in this top portion here right there, and 2.5 percent of 200 is five students. And this is a
type of question that often comes up on statistic exams or even in step one. They will give you a
data set and ask you to identify how many patients are in some range based on the number of
standard deviations away from the mean.

Most scientific literature you read will report a standard deviation, but once in a while you'll see
a study that reports a variance. So remember that the standard deviation sigma is equal to the
square root of all this stuff in here. Well, the variance is sigma squared, so you take away the
square root.

It's just the stuff under the square root sign. So sometimes this is reported instead, and it's just
a different measure from standard deviation, but represents similar things. Those rules about
68 percent and 95 percent and all that don't apply to the variance.

Those are only for the standard deviation. Next dispersion measure I'm going to talk about is
called the standard error of the mean, and this represents how precisely we know the true
population mean. In statistics, we often take a small sample and use its characteristics to
estimate the characteristics of a larger population.

For example, we might measure 10 systolic blood pressure readings and take the mean value
and hope that that represents the mean value of a larger population of 10,000 people. The
standard error of the mean gives a representation of how close we are to that true population
mean. The standard error of the mean is equal to the standard deviation of our sample set
divided by the square root of the number of samples.

What this means is that the more samples we take, the less standard error of the mean is, the
closer we are to the true mean. And this should make sense to you. If I take 10 blood pressure
readings and use those to estimate the characteristics of 10,000 people, I could be way off.
On the other hand, if I take 1,000 blood pressure readings and use those to estimate the
characteristics of 10,000, I am probably closer to the true mean. And that's why the standard
error of the mean gets lower the more samples you have. If you have a big standard deviation
in your sample data set, this means you're going to have a big standard error of the mean.

You're going to need a lot of samples in that case to make the standard error of the mean
small. That's because the standard deviation is on top here. So if that's a big number, you have
a lot of error.

And the only way to cancel that out is to have lots of samples down in the denominator. A small
standard deviation in your sample data set means a small standard error of the mean. And you
don't need quite as many samples to still get a result with a low standard error of the mean.

The next dispersion measure is the z-score. And the z-score is very simple. It's equal to the
number of standard deviations you are away from the mean.

So if a sample has a z-score of 0, that sample is equal to the mean. If the sample has a z-score
of plus 1, it's one standard deviation above the mean. And if it has a z-score of minus 1, it's one
standard deviation below the mean.

So all these dotted lines here on this chart down at the bottom, which represent standard
deviations away from the mean, can also be changed into z-scores by just writing minus 1 or
positive 1. So just to make this very clear, in case you're ever asked to calculate it, suppose a
group of students had a test grade average or mean of 79 and the standard deviation was 5. If
your score on the test was 89, then your z-score would be 89 minus 79, which is 10, divided by
5, or plus 2. You would be two standard deviations away from the mean. And your z-score is
plus 2. The final measure of dispersion I want to discuss are the confidence intervals. So when
you read the scientific literature, you will see most mean values reported with something called
95% confidence intervals.

For example, a study of diabetic patients might report that the mean blood glucose level was
120 plus or minus 5. That plus or minus 5 is the 95% confidence intervals. And the authors are
saying that this is the range in which 95% of repeated measurements would be expected to fall.
The authors are saying that if you took our study of diabetic patients and repeated it, we
believe that you would get a mean value of blood glucose that is similar to ours within plus or
minus 5 of the mean that we got.

So these confidence intervals are really for estimating population means from a sample data
set. When the authors report their findings on a group of, say, 10 or 100 diabetics, they're
trying to give you an estimate of what the mean would be like in the entire population of similar
diabetics. So just to make this clear, suppose we take 10 samples of a population in the world
that's represented by 1 million people.

The mean of our 10 samples is x. How sure are we that the mean of the 1 million people is also
x? The confidence intervals are what help us answer this question. So this is the equation for
the 95% confidence interval down here in the bottom of the screen. It's equal to the mean value
plus or minus 1.96 times the standard error of the mean.

And let me show you how you would calculate that for a sample data set. So suppose we have a
group of measurements and the mean value is 10, and the standard deviation is 4, and the
number of measurements is 16. Well, recall from a few slides ago that the standard error of the
mean is equal to the standard deviation 4 divided by the square root of the number of samples.

So that would be the square root of 16. Well, the square root of 16 is 4. So this would be 4
divided by 4. And the standard error of the mean, therefore, is 1. This means our confidence
interval is 10 plus or minus 1.96 times 1. And 1.96 times 1 is basically 2. So our confidence
intervals are 10 plus or minus 2. So what we are saying is that we believe, based on this data set
here, that 95% of repeated means would fall between 8 or 12. That's 10 plus or minus 2. And
sometimes this is stated as the upper confidence limit being 12 and the lower confidence limit
being 8. Now, it's very important that you don't confuse the 95% confidence interval with the
standard deviation.

Remember that I said that plus or minus 2 standard deviations includes 95% of a data set.
That's completely different from the confidence intervals. The standard deviation is for a given
group of data points.

So suppose we have 10 samples. These 10 samples have a mean and a standard deviation. And
95% of these samples are going to fall between the mean plus or minus two standard
deviations.

This is purely a descriptive characteristic of the data set and all the samples. The confidence
intervals, on the other hand, do not describe the sample set at all. These are an inferred value
of where we think the true mean lies for a population.

The confidence intervals help us tell readers of our clinical data how close we think our data set
represents the population outside of our experiment. So when you see the number 95% in a
question stem, don't get confused. Read the question very carefully.

What are they asking for? If they say what is the range in which 95% of the measurements in a
data set fall, then they're looking for the mean plus or minus two standard deviations. On the
other hand, if they say what is the 95% confidence interval of the mean, then they are looking
for the mean plus or minus 1.96 times the standard error of the mean. These two things are
different.

This is often confusing for students, and make sure you have them straight in your mind. And
that concludes our module on basic statistics.

Transcribed by [Link]. Go Unlimited to remove this message.

You might also like