Module 5
Lesson 5
Welcome,
In this lesson I will introduce a general and important issue with disproportionality analysis and one
way to resolve that issue.
To get us started, consider these two different situations:
In the first one we have a drug event combination with two reports observed and 0.01 expected, and
in the other one we have a drug event combination with 20 reports and 0.1 expected. As you see,
the observed to expected ratio is the same for both of these drug event combinations, it is 200. So
the reporting of both of these combinations appears to be highly disproportional. However, in
practice these situations are quite different. First of all, it's generally much less likely that a
combination with two reports would end up as a signal compared to one with twenty reports. And
secondly, the weight of the evidence from just two reports is in general not particularly high even
though it can yield very high disproportionality, as in this example. This problem is due to something
we might call random variability.
If the expected number of reports is really low, like 0.01 in this example, it's sufficient with just a few
reports, like two here, to obtain a really high value of disproportionality. But in reality we could
obtain a few reports on any drug-event combination just by chance even if there is no causal link
between the drug and the event.
So, if we take observed-to- expected ratios at the crude face value, we run a clear risk of detecting a
lot of false associations just by chance.
There are different ways to handle this issue and the one that we're going to present here is the one
that we recommend and the one that we've also been using in our own disproportionality analysis
on VigiBase for a long time.
The underlying theory is mathematically quite complex but actually the end result is rather
straightforward. One part of the solution is something we call shrinkage, which we can think of as
adding half an imaginary report both to the observed and to the expected numbers of reports. And
as you see, with this little trick we pull down the observed-to-expected ratios for both of these
combinations, and the pull is much stronger for the combination with just two reports, which is the
one that has less data. And these new numbers, 4.9 and 34 respectively, much better reflect our
intuition that probably the strength of the evidence is stronger for the combination with 20 reports.
And now we´ll use another little trick and move these values to a log scale, and the main purpose of
doing that is that it becomes easier to compare high and low values of disproportionality. If we use a
log2 scale like I've done here, we get the values 2.3 and 5.1 respectively. With a log2 scale and with
an expected number of reports, that is computed as we do for the “relative reporting ratio”. The
result we get is something we call an IC value, where IC is for “information component”. This is one
of the common measures of disproportionality and the one that we use here at UMC. This scale may
seem confusing at first but it's actually not that difficult to interpret. If the IC is equal to 0, that
means the observed is equal to the expected and if the IC is 1, that means the observed is twice the
expected. And then for each step that the IC increases the observed-to-expected ratio is doubled.
So, if the IC is 2 that means the observed is 4 times the expected and if it is 3 that means the
observed is 8 times the expected and so on.
Now, let's get back to our example. You may recall that we have two drug-event combinations. One
that has two reports and one that has 20 reports, and that the observed-to-expected ratio is the
same for both of them. It's 200, which is just below 8 on a log2 scale. If we then apply the shrinkage,
Module 5
Lesson 5
which means we compute IC values, both combinations are pulled down as we noted before. And
again, the less data there is the stronger the pull, which is very obvious in this example.
However, the shrinkage alone is not enough. Even if we don't quite trust the evidence for the
combination, which is two reports, it still looks disproportional. As you see, the IC is still above zero
so the observed is still greater than the expected.
And this is where the second part of our solution comes in, the uncertainty intervals. For normal use
we recommend to compute 95% uncertainty intervals around the IC, which is what I've done here as
well. The most important thing to know about these intervals is that they provide additional
protection against detecting false associations. The more data there is, the more we trust the
disproportionality analysis calculation, which will result in shorter intervals. And in this example, you
see that the combination with 20 reports has more data and therefore a shorter interval. The other
combination, the one with two reports, you see now has a lower endpoint of its uncertainty interval
that is below zero. This means that after we have handled the issue of random variability we are no
longer confident that the data is sufficient to say that it is disproportionately reported.
So in short, the IC or the “information component” has two main features: shrinkage and uncertainty
intervals, that together provide good protection against detecting false associations when
disproportionality looks high but is based on just a few reports.