GRADE Guidelines on Network Meta-Analysis Imprecision
GRADE Guidelines on Network Meta-Analysis Imprecision
ORIGINAL ARTICLE
GRADE guidelines 33: Addressing imprecision in a network
meta-analysis
Romina Brignardello-Petersen a,∗, Gordon H. Guyatt a, Reem A. Mustafa a,b, Derek K. Chu a,
Monica Hultcrantz c, Holger J. Schünemann a,d, George Tomlinson e
a Department of Health Research Methods, Evidence and Impact, McMaster University, Hamilton, Ontario, L8S 4L8, Canada
b Department of Internal Medicine, Division of Nephrology and Hypertension, University of Kansas Medical Center, Kansas City, KS 66160, United
States
c Swedish Agency on Health Technology Assessment and Assessment of Social Services (SBU), Stockholm, Sweden
d Department of Medicine & Institut für Evidence in Medicine, Medical Center & Faculty of Medicine, University of Freiburg, Freiburg, 79110,
Germany
e Department of Medicine, University Health Network, Toronto, Ontario, M5G 2C4, Canada
Received 11 December 2020; Received in revised form 11 June 2021; Accepted 15 July 2021; Available online 19 July 2021
Abstract
Objective: This article describes GRADE guidance for assessing imprecision when rating the certainty of the evidence from network
meta-analysis.
Study Design and Setting: A project group within the GRADE working group conducted iterative discussions, computer simulations,
and presentations at GRADE working group meetings to produce and obtain approval for this guidance.
Results: When addressing imprecision of a network estimate, reviewers should consider the 95% confidence or credible interval,
and the optimal information size. If the 95% confidence or credible interval crosses a pre-specified threshold, reviewers should rate
down the certainty of the evidence. If the 95% confidence interval does not cross any pre-specfied threshold, reviewers should consider
the optimal information size. Because addressing the optimal information size may be challenging, reviewers can use the effect size to
decide if any calculations are necessary. When the size of the effect is modest or the optimal information size is met, reviewers should
not rate down for imprecision.
Conclusion: Reviewers should use this guidance when to appropriately address imprecision in the context of the assessment of
certainty of evidence from network meta-analysis. © 2021 Elsevier Inc. All rights reserved.
Keywords: GRADE guidance; Certainty of evidence; Network meta-analysis; Imprecision; Optimal information size
[Link]
0895-4356/© 2021 Elsevier Inc. All rights reserved.
50 R. Brignardello-Petersen et al. / Journal of Clinical Epidemiology 139 (2021) 49–56
What is new?
Key findings
• When addressing imprecision of network estimates,
reviewers should consider the 95% confidence in-
terval and its relation with thresholds of interest and
the optimal information size.
• The imprecision of network estimates can be seri-
ous enough to rate down the certainty of the evi-
dence 3 levels. Fig. 1. The relationship between the confidence interval (CI) of an
What this adds to what was known? effect and the thresholds determine if reviewers should rate down for
• We describe a sequential process for reviewers to imprecision. Vertical lines represent thresholds, which delimit ranges
follow when making this assessment, to maximize of magnitudes of effect labeled as per in the text between the square
brackets. Using a minimally contextualized approach, reviewers would
efficiency and ensure that no necessary step is ig- not rate down due to imprecision in A, B, or C if they are rating the
nored. certainty that there is any benefit (the CI does not cross the threshold
• We provide a previously unavailable approach to of null effect). They would, however, rate down in A but not in B or
address optimal information size in the context of C if they are rating the certainty that there is at least a small benefit
network meta-analysis. (the CI in A crosses the small effect threshold, but the CI in B or C
does not). Using a partially contextualized approach, reviewers would
What is the implication and what should change rate down the certainty of the evidence due to imprecision in A or C
now? but not in B if they are rating their certainty that the effect is small
• When assessing the certainty of network estimates, (the CI of A crosses the small effect threshold and the CI of C crosses
systematic reviewers and guideline developers ad- the moderate effect threshold, but the CI of B does not cross either
dressing imprecision should use this approach and the small effect or the moderate effect threshold).
be more transparent in the justification of their
judgments.
2. Addressing imprecision in traditional pairwise
meta-analysis
When addressing imprecision in traditional pairwise
meta-analysis, systematic reviewers and guideline devel-
opers consider the bounds of the confidence interval (CI)
1. Introduction of the absolute estimate of effect, and the optimal infor-
mation size (OIS) [7]. The certainty of the evidence is
The Grading of Recommendations Assessments, Devel- decreased when the CI crosses specific thresholds of inter-
opment, and Evaluation (GRADE) working group has pre- est or when the OIS is not met. We explain both concepts
viously presented guidance for evaluating the certainty of below and for simplicity, we refer to thresholds of interest
the evidence (confidence in estimates of effect, quality of as thresholds.
evidence) in network meta-analysis (NMA) [1-4] and how If the CI lies completely on one side of a threshold
to interpret its results and draw conclusions [5,6]. This arti- or between two thresholds, we can make more confident
cle provides guidance for addressing imprecision of NMA statements regarding the patient-importance of the effect
estimates. The discussion assumes familiarity with the ba- [8]. The thresholds depend on the purpose of the system-
sic concepts of NMA and of addressing imprecision in the atic review, and thus on the degree of contextualization [9].
context of traditional pairwise meta-analysis, and consti- Fig. 1 illustrates the fundamental logic of these decisions.
tutes official guidance from the GRADE working group. The CIs estimated based on small numbers of trial par-
Several factors play a key role in the assessment of im- ticipants or few events can be fragile, in the sense that
precision both in traditional pairwise meta-analysis [7] and small amounts of additional data can substantially change
NMA. These include the degree of contextualization of the the location of the interval. The OIS helps address this con-
rating of certainty of the evidence and the closely related cern. The OIS is the number of participants obtained with
issue of the target of the certainty rating (i.e., what it is a conventional sample size calculation for a trial address-
in which are we rating our certainty); decisions regarding ing the question of interest. When the number of patients
the thresholds of interest necessary for this assessment; included in a meta-analysis for an outcome is larger than
and value judgments regarding the magnitude of effect es- the OIS, important changes in the results in the face of
timates (e.g., what are a small but important effect, a mod- considerable amounts of new evidence are unlikely. Initial
est effect, and a large effect). The scope of this article is GRADE guidance advises systematic reviewers to always
to illustrate how to assess imprecision after making these consider the OIS and for guideline developers to consider
decisions. OIS when the CI does not cross a threshold[7]. Further
R. Brignardello-Petersen et al. / Journal of Clinical Epidemiology 139 (2021) 49–56 51
developments of this guidance [9] provide the conceptual effect to rate their certainty that there is a benefit or a harm.
underpinning for this paper. Alternatively, they may choose a small but important ef-
fect as their threshold. Using a partially contextualized ap-
3. The guidance for pairwise meta-analysis may not proach, reviewers use thresholds for small, moderate, and
large effects and rate their certainty that the true effect has
always be helpful in the context of NMA
a particular magnitude (e.g., trivial or none, small, mod-
When considering the OIS in the context of pair- erate, or large) [10]. Guideline developers using a fully
wise meta-analysis, systematic reviewers calculate the total contextualized approach use a threshold that represents an
number of participants included across studies. In the con- effect large enough to mandate a decision between alterna-
text of NMA, because network estimates are calculated tives. Whatever the degree of contextualization, when the
using direct and indirect evidence, this approach for de- CI crosses one or more of the thresholds, reviewers should
termining whether the total sample size meets the OIS of- rate down the certainty of the evidence (i.e., whenever a CI
ten does not work. For a specific comparison, if the CI crosses a threshold, we are less certain that the true effect
of the network estimate is similar to the CI of the direct lies above or below the threshold, or within two thresholds)
estimate, then there is little information gained from the [9]. This principle applies to both pairwise meta-analysis
indirect estimate, and reviewers could add up the number and NMA estimates (Fig. 1).
of participants in the trials that compare the two interven- Ideally, reviewers would decide a priori where to set
tions directly to determine if the sample size is sufficient to the thresholds based on absolute effects. In addition, these
meet the OIS. When, however, the indirect evidence makes thresholds should be consistent with the approach that re-
a non-negligible contribution to the network estimate, and viewers will use when interpreting the results of the NMA
the direct and indirect evidence yield similar results (i.e., [5,6]. Because the choice of thresholds must be done us-
are coherent), considering only the direct evidence would ing absolute estimates of effect, assessing whether the CI
underestimate the number of participants contributing to crosses a threshold requires using the CI of the absolute
the network estimate. In the presence of incoherence, the estimate of effect.
network estimate will have a wider CI than the direct esti- For example, in an NMA of methotrexate monotherapy
mate, and counting the number of people contributing the vs. methotrexate combinations for patients with rheuma-
direct evidence will overestimate the effective sample size toid arthritis, the absolute difference between the num-
of the network estimate. ber of patients who achieved a response when comparing
Fortunately, in many situations considering the OIS may methotrexate plus ciclosporin vs. methotrexate alone was
not be necessary. First, if the CI of the network esti- 134 more per 1,000 patients, with a 95% CI from 35 fewer
mate crosses a threshold, reviewers or guideline developers to 290 more [11]. When rating the certainty that there is a
should rate down the certainty of the evidence for impre- benefit (i.e., higher proportion of response with the com-
cision; considering the OIS will not change this decision. bination therapy), because the CI crosses the null effect,
Second, if the point estimate of the relative effect indicates the authors rated down the certainty of evidence due to
a modest effect and the CI does not cross a threshold, one imprecision (Fig. 2, scenario 1a).
can infer that the sample size is large enough that new It is important to highlight that whether a threshold is
evidence is unlikely to result in important changes in in- crossed and the certainty of the evidence should be rated
ferences. In the subsequent discussion, we elaborate on down for imprecision depends on the degree of contextu-
both these situations. alization, the target of the certainty rating, and thresholds
Considering the above, specific guidance for addressing chosen [10]. For example, crossing the null effect thresh-
imprecision in the context of NMA is needed. This guid- old is not a reason to rate down the certainty of evidence
ance will facilitate assessments, help authors of NMAs to when the target of the certainty rating is to learn if there
achieve explicitness regarding their methods, and users to is trivial to no effect, and the 95% CI is entirely within
understand ratings of certainty of evidence. the thresholds of small benefit and small harm. Similarly,
when rating the certainty of the same estimate, systematic
4. Addressing imprecision in NMA review authors who choose to not make value judgments,
those who choose a minimally important difference, and
Fig. 2 summarizes GRADE guidance on how to address guideline developers who choose a decision threshold may
imprecision for each network estimate. end up with different judgments with regards to impreci-
sion.
4.1. If the CI crosses a threshold, we should rate down
for imprecision 4.2. When the CI crosses a threshold, we can rate down
the certainty of the evidence one, two, or three levels
Systematic reviewers should choose thresholds based on
the degree of contextualization of their review [9]. Using a Whether reviewers’ rate down the certainty of the ev-
minimally contextualized approach, reviewers use the null idence one or two levels due to imprecision remains a
52 R. Brignardello-Petersen et al. / Journal of Clinical Epidemiology 139 (2021) 49–56
Fig. 2. Process for addressing imprecision for each network estimate. Although the effect size (2) is another way of addressing whether the OIS
is met, (3) we illustrate it as a different step because it may help reviewers avoiding to go through the calculations required when addressing
2b. What is a “modest” effect (2) requires value judgment specific to each scenario (an example is 30% relative risk reduction). Step 3 can be
addressed by calculating the sample size underlying a network estimate and contrasting this with the optimal information size, or, for OR, and RR
effect measures by looking at the ratio of the upper to the lower extremes of the confidence interval.
matter of judgment. It depends on where the extremes of Even assuming no other serious concerns in this body of
the CI are and how many thresholds it crosses, as well as evidence, concluding that there may be a benefit would
the interpretation of these extremes in absolute terms. For be misleading. Thus, reviewers could rate down the cer-
continuous outcomes, it also depends on the outcome, and tainty of the evidence 3 levels resulting in a rating of very
the scale of measurement. low certainty [15]. Such wide confidence intervals are not
It may be helpful, however, to consider the conse- uncommon in sparse networks [3].
quences of rating down one vs. two levels. For exam-
ple, in an NMA of pharmacological interventions for acute
diarrhea and gastroenteritis in children, when comparing
Kaolin Pectin vs. standard care, the reduction in the num- 4.3. If the CI does not cross a threshold, we need to look
ber of hours of diarrhea duration was 5 (95% CI, reduction at the effect size
of 34 hours to increase of 23 hours) [12]. Assuming that Evidence from meta-epidemiologic studies [16-19], as
imprecision was the only serious concern, reviewers fol- well as numerous dramatic individual examples [20-25],
lowing GRADE guidance on how to communicate results have documented that large relative estimates of effects
[13] would, if they rated down one level, have concluded from relatively small randomized trials with few events
that there is “probably” or “likely” a reduction in the du- are often refuted by subsequent trials. A systematic review
ration of diarrhea when using Kaolin Pectin. If they rated and NMA of such early trials may result in a large pooled
down two levels, however, they would have concluded that treatment effect and a CI that does not cross the thresholds,
there “may” be a reduction in this outcome. The latter but which has wide CIs. The result will be dissemination
seems more appropriate, thus illustrating how bearing in of large overestimates of effect, and patients being exposed
mind the consequences of rating down one vs. two levels to unnecessary harms on the basis of overly optimistic es-
can facilitate the judgment. timates of treatment benefit.
In addition, there are situations in which rating down When the relative effect size is modest (e.g,. less than
three levels due to imprecision may be appropriate. In one 30% relative risk reduction or increase), a CI calculated
of the first iterations of a living NMA of pharmacologi- based on a small number of patients is likely to be wide
cal treatments for Coronavirus Disease 2019 [14], several and cross a threshold, resulting in reviewers rating down
network estimates had extremely imprecise CIs. For ex- for imprecision (Fig. 2, scenario 1a). When, however, the
ample, the risk difference in mortality per 1,000 patients CI does not cross any threshold, reviewers should look
was 329 fewer (95% 330 fewer to 670 more) when com- at the size of the effect represented by the point estimate
paring ribavirin to standard care. This CI provides very (Fig. 2, scenario 1b). We describe below two general cases
little information about the effect of ribavirin on mortal- to illustrate this concept. Because these are considerations
ity because it crosses multiple thresholds and its bounds relevant to the number of patients contributing information
reflect dramatically different effects (the lower bound of to a network estimate, typically calculated using relative
330 fewer deaths represents a large benefit, and the up- estimates of effect, all assessments, and steps below fo-
per bound of 670 more deaths represents a large harm). cus on the relative estimate of effect (as opposed to the
R. Brignardello-Petersen et al. / Journal of Clinical Epidemiology 139 (2021) 49–56 53
assessments related to the thresholds, which are based on the estimate and its CI. With a small number of assump-
absolute estimates of effect). tions, it is possible to work backwards from the estimate of
the treatment effect and its CI to calculate the equivalent
4.3.1. When the effect size is modest and plausible, we sample size for the single study that would have generated
should not rate down for imprecision those values (which we will call the “the effective sample
When the effect size is modest and plausible, a CI that size”). The details of the calculations depend on the type
is narrow and does not cross a relevant threshold, indicates of outcome and effect measure. Boxes one, two, and three
a sufficient sample size underlying the network estimate show the details for the RR, OR, and mean difference as
(we will refer to this henceforth as the “effective sample treatment effect measures.
size”). Therefore, when the CI does not cross a threshold BOX 1.
and the effect size is modest, reviewers should not rate The standard error of the log(RR) in a two-arm trial
down due to imprecision (Fig. 1, scenario 2a). with equal arm sizes can be estimated as
For example, a systematic review comparing the effects 1 1 1
of uterotonic agents for preventing postpartum hemorrhage SEtrial = SE(log [RR]) = + −2
n pc RR × pc
[26] reported a risk ratio of 0.72 (95% CI, 0.56–0.97) for
the comparison of carbetocin with oxytocin for the preven- Where n is the per-arm sample size, pc is the observed
tion of postpartum hemorrhage equal or larger than 500 proportion with a response in the reference group and RR
milliliters. Because the CI does not cross the null effect is the observed relative risk with treatment.
and the size of the effect is modest, reviewers taking a The SE for the NMA estimate of the log(RR) can be
minimally contextualized approach, and rating the certainty calculated from the upper and lower limits of the reported
that there is a benefit of carbetocin over oxytocin would CI
not rate down the certainty of this evidence due to impre- log(CIupper ) − log(CIlower )
cision. Indeed, 30,633 participants contributed to the direct SEN M A =
3.92
estimate informing that comparison. Setting these two estimates of the SE to be equal, we
can solve for n as
4.3.2. When the effect size is large, we should assess if 1
pc + 1
RR×pc − 2
the optimal information size is met n= 2
A CI may be wide and not cross a threshold only be- (SEN M A )
cause the point estimate is large and potentially implau- The value of pc should be estimated from the relevant
sible. In this case, it is important to assess the effective arms in the network, ideally with a meta-analysis of pro-
sample size of the network estimate and determine if the portions. But as this is an approximate approach, it may
OIS is met (Fig. 1, scenario 2b). be acceptable to simply count of the number of responses
For example, in an NMA of methotrexate monotherapy and number of patients across these arms to estimate pc
vs. methotrexate combinations for patients with rheumatoid BOX 2
arthritis, the odds ratio associated with the comparison be- The standard error of the log(OR) in a two-arm trial
tween methotrexate plus tofacinib vs. methotrexate alone with equal arm sizes can be estimated as
was 3.0 with a 95% CI from 1.1 to 9.4. Even though the
CI does not cross the threshold (i.e., null effect), because SEtrial = SE(log [OR])
the effect size is large, reviewers should determine if the
1 1 1 1 1
OIS is met before deciding whether to rate down for im- = + + +
n pc 1 − pc pt 1 − pt
precision.
Where n is the per-arm sample size, pc is the observed
proportion with a response in the reference group and pt
4.4. The confidence interval of a network estimate can is the observed proportion with a response in the reference
help determine if the optimal information size is met and group. Pt can be replaced by pt=pc∗ OR/(1-pc +pc∗ OR)
whether to rate down for imprecision The SE for the NMA estimate of the log(OR) can be
4.4.1. We can use the confidence interval of a network calculated from the upper and lower limits of the reported
estimate to calculate the number of patients underlying CI
that estimate log(CIupper ) − log(CIlower )
SEN M A =
When presented with an estimate and CI for a treatment 3.92
comparison from an NMA, we can inquire regarding the Setting these two estimates of the SE to be equal, we
size of a single study that would give an estimate with that can solve for n as
CI. This question puts aside the structure of the network
1 1 1 1
and the combinations of direct and indirect evidence that pc + 1−pc + pt + 1−pt
n= 2
are used to generate the network estimate and focuses on (SEN M A )
54 R. Brignardello-Petersen et al. / Journal of Clinical Epidemiology 139 (2021) 49–56
The value of pc should be estimated from the relevant to detect, with 80% power at alpha of 0.05, and a control
arms in the network, ideally with a meta-analysis of pro- group risk of 77% results in 92 patients per group or 184
portions. But as this is an approximate approach, it may in total. Therefore, the network estimate, with an effective
be acceptable to simply count of the number of responses sample size of 62, does not satisfy the OIS criterion. In a
and number of patients across these arms to estimate pc case like this, reviewers should rate down for imprecision
BOX 3 (Fig. 2, scenario 3a)
The standard error of the mean difference in a two-arm The same process applies to the second example (5%
trial with equal arm sizes and equal standard deviations in NaF varnish vs. usual care). The overall risk of caries in
each group can be estimated as the reference groups for this comparison was 37%. The
sample size for a single trial with 80% power at alpha of
2
SEtrial = SD × 0.05 in which researchers want to detect a relative risk in-
n crease of 1 of 0.75 (so that the RR is 1.33) in the arrest of
Where n is the per-arm sample size, and SD is the caries with a control group risk of 37% is 255 per group
pooled within-group standard deviation or 510 in total, so the network estimate with an effective
The SE for the NMA estimate of the mean difference sample size of 426, does not satisfy the OIS criterion and
can be calculated from the upper and lower limits of the reviewers should rate down for imprecision (Fig. 2, sce-
reported CI nario 3a). However, if reviewers determine that a relative
CIupper − CIlower increase of 40% is the minimum important effect, then the
SEN M A = required single trial sample size would be 176 per group,
3.92
Setting these two estimates of the SE to be equal, we the OIS criterion would be met, and reviewers should not
can solve for n as rate down for imprecision (Fig. 1, scenario 3b).
Because the optimal information size depends on the
2 × SD2
n= 2
minimal effect to detect used in the calculations, the same
(SEN M A ) estimate may or may not meet the optimal information
The value of SD can be estimated as the pooled SD size based on the degree of contextualization and what this
across all k arms relevant to the comparison minimal effect is [28] Therefore, systematic reviewers who
For example, an NMA of treatments to arrest dental choose a biological plausible effect or a minimal important
caries [27] reported a network estimate of 0.67 for the RR difference, and guideline developers who use a specific ef-
comparing 1.23% APF gel to 5% NaF Varnish, with a 95% fect size based on full contextualization may come up with
CI for the RR of 0.45–0.99. The overall risk of caries in different judgments with regards to imprecision associated
the control groups (i.e., baseline risk) for this comparison with optimal information size.
was 77%. Using the formulas in Box one, we find that a
single study with that observed RR, baseline risk and CI
would have had 31 in each of the two arms, so the effective 4.4.3. There may be scenarios in which calculating the
sample size of the network estimate is 62 patients. That effective sample size is not necessary
same NMA reported a network estimate of 1.97 for the RR There may be scenarios in which the width of the CI
comparing 5% NaF varnish to usual care, with a 95%CI is such that it is not necessary to calculate the effective
for the RR of 1.63–2.40. Using the formulas in Box one, sample size because the OIS will not be met. For relative
we find that a single study with that observed RR, control risks and odds ratios, the CI width is represented by the
group risk (37%) and CI would have 213 patients in each ratio of the upper limit of the CI to the lower limit. When
of the two arms, so the effective sample size of the network the ratio of the upper to the lower limits of the CI of the
estimate is 426 patients. relative risk estimate is higher than three, the OIS of a
RR estimate will not be met regardless of the effect size,
4.4.2. If the effective sample size does not meet the any reasonable minimally important difference to detect,
optimal information size, we should rate down for and baseline risk (appendix). In these cases, reviewers can
imprecision confidently rate down due to imprecision without making
Focusing on the first example of the NMA of caries any calculations.
arrest (1.23% APF gel vs. 5% NaF Varnish), if reviewers In the two examples from the NMA of caries, these ra-
judge the relative risk reduction of 33% to be large, they tios are 2.2 (the result of dividing the upper limit of the
should assess whether the effective sample size of the data CI, 0.99, by its lower limit, 0.45) and 1.5 (2.40 divided by
contributing to network estimate (62 patients) is at least as 1.63), so the OIS calculation would be necessary. But in
large as what we would require in a single trial: in other the example of methotrexate monotherapy vs. methotrex-
words, whether the OIS is met (Fig. 2, step 3). ate combinations for patients with rheumatoid arthritis, the
A sample size calculation for a single trial, assuming relative effect estimate was an odds ratio of 3.4, the upper
that the reviewers consider a relative reduction of 25% limit of the CI was 9.4 and the lower limit was 1.1. The
(relative risk 0.75) as the minimal effect in arrest of caries ratio of the upper to the lower limit of the CI is 9.4 of
R. Brignardello-Petersen et al. / Journal of Clinical Epidemiology 139 (2021) 49–56 55
The width of the CI can guide the decision whether to 8. Author contributions
rate down one or two levels due to imprecision. If review-
ers had to calculate the effective sample size and deter- RBP, GG, and GT conceptualized this work and col-
mine if the OIS was met, this means that the CI is narrow lected examples. GT developed the statistical concepts re-
enough to rate down one level only. In both examples of lated to the calculation of optimal information size in net-
the treatments to arrest caries, if rating down, reviewers work estimates. RBP, GG, and GT drafted the manuscript.
should rate down only one level. However, if the CI is RM, DC, MH, and HS provided input that resulted in
wide enough to not have to bother calculating the effec- important modifications. All authors approved the final
tive sample size, reviewers may rate down two levels, as manuscript.
illustrated in the example of methotrexate monotherapy vs.
methotrexate combination. Acknowledgments
If the CI is extremely wide, reviewers should consider
rating down 3 levels [15]. For example, in the first itera- We would like to thank the authors of some of the
tion of the living NMA of pharmacological treatments for systematic reviews we used as examples throughout this
Coronavirus Disease 2019 [14] for the comparison between paper: Olivia Urquhart, and Glen Hazlewood. We would
Baloxavir Marboxil and standard care, the CI around the also like to thank all GRADE working group members for
relative estimate of effect on mechanical ventilation ranged their input and feedback, in particular M Hassan Murad.
from 1.71 to 1.79 × 1026 . Such a CI does not require deter-
mining if the OIS is met: it is so wide that the uncertainty Supplementary materials
on the true estimate due to imprecision warrants rating
down three levels. Supplementary material associated with this article can
be found, in the online version, at doi:10.1016/[Link].
2021.07.011.
5. Other considerations
Finally, previous GRADE guidance suggests that re- References
viewers should avoid rating down the certainty of a net- [1] Puhan MA, Schunemann HJ, Murad MH. A GRADE Working
work estimate when imprecision is associated with another Group approach for rating the quality of treatment effect estimates
limitation for which they already rated down: i.e., they from network meta-analysis. BMJ 2014;349:g5630. doi:10.1136/
bmj.g5630.
should not double count. For instance, when imprecision
[2] Brignardello-Petersen R, Bonner A, Alexander PE. Advances in the
is the result of serious incoherence between direct and in- GRADE approach to rate the certainty in estimates from a net-
direct evidence, reviewers should only rate down for one work meta-analysis. J Clin Epidemiol 2018;93:36–44. doi:10.1016/
of those domains [4]. [Link].2017.10.005.
56 R. Brignardello-Petersen et al. / Journal of Clinical Epidemiology 139 (2021) 49–56
[3] Brignardello-Petersen R, Murad MH, Walter SD. GRADE approach [16] Pereira TV, Horwitz RI, Ioannidis JP. Empirical evaluation of
to rate the certainty from a network meta-analysis: avoiding spuri- very large treatment effects of medical interventions. JAMA
ous judgments of imprecision in sparse networks. J Clin Epidemiol 2012;308(16):1676–84. doi:10.1001/jama.2012.13444.
2019;105:60–7. doi:10.1016/[Link].2018.08.022. [17] Rerkasem K, Rothwell PM. Meta-analysis of small random-
[4] Brignardello-Petersen R, Mustafa RA, Siemieniuk RAC. GRADE ized controlled trials in surgery may be unreliable. Br J Surg
approach to rate the certainty from a network meta-analysis: address- 2010;97(4):466–9. doi:10.1002/bjs.6988.
ing incoherence. J Clin Epidemiol 2019;105:77–85. doi:10.1016/j. [18] Trikalinos TA, Churchill R, Ferri M. Effect sizes in cumulative meta-
jclinepi.2018.11.025. analyses of mental health randomized trials evolved over time. J Clin
[5] Brignardello-Petersen R., Florez I., Izcovich A., GRADE ap- Epidemiol 2004;57(11):1124–30. doi:10.1016/[Link].2004.02.018.
proach to drawing conclusions from a network meta-analysis using [19] Nagendran M, Pereira TV, Kiew G. Very large treatment effects in
a minimally contextualized framework. MBJ 2020;11(371):m3900 randomised trials as an empirical marker to indicate whether sub-
doi:10.1136/bmj.m3900 sequent trials are necessary: meta-epidemiological assessment. BMJ
[6] Brignardello-Petersen R., Izcovich A., Rochwerg B., GRADE ap- 2016;355:i5432.
proach to making conclusions from a network meta-analysis us- [20] Desai K, Carroll I, Asch S. Extremely large outlier treatment effects
ing a partially contextualized framework. BMJ. 2020;10(371):m3907 may be a footprint of bias in trials from less developed countries:
doi:10.1136/bmj.m3907 randomized trials of gabapentinoids. J Clin Epidemiol 2019;106:80–
[7] Guyatt GH, Oxman AD, Kunz R. GRADE guidelines 6. Rat- 7. doi:10.1016/[Link].2018.10.012.
ing the quality of evidence–imprecision. J Clin Epidemiol [21] Bosch J, Yusuf S, Gerstein HC. Effect of ramipril on the inci-
2011;64(12):1283–93. doi:10.1016/[Link].2011.01.012. dence of diabetes. N Engl J Med 2006;355(15):1551–62. doi:10.
[8] Guyatt G, Montori V, Devereaux PJ. Patients at the center: 1056/NEJMoa065061.
in our practice, and in our use of language. ACP J Club [22] Devereaux PJ, Beattie WS, Choi PT. How strong is the evidence
2004;140(1):A11–12. for the use of perioperative beta blockers in non-cardiac surgery?
[9] Hultcrantz M, Rind D, Akl EA. The GRADE Working Group Systematic review and meta-analysis of randomised controlled trials.
clarifies the construct of certainty of evidence. J Clin Epidemiol BMJ 2005;331(7512):313–21.
2017;87:4–13. doi:10.1016/[Link].2017.05.006. [23] Bangalore S, Wetterslev J, Pranesh S. Perioperative beta block-
[10] Zeng L, Brignardello-Petersen R, Hultcrantz M. GRADE guide- ers in patients having non-cardiac surgery: a meta-analysis. Lancet
lines 32: GRADE offers guidance on choosing targets of GRADE 2008;372(9654):1962–76. doi:10.1016/s0140-6736(08)61560-3.
certainty of evidence ratings. J Clin Epidemiol 2021;137:163–75. [24] Yusuf S, Collins R, MacMahon S. Effect of intravenous ni-
doi:10.1016/[Link].2021.03.026. trates on mortality in acute myocardial infarction: an overview of
[11] Hazlewood GS, Barnabe C, Tomlinson G. Methotrexate monother- the randomised trials. Lancet 1988;1(8594):1088–92. doi:10.1016/
apy and methotrexate combination therapy with traditional and bio- s0140-6736(88)91906-x.
logic disease modifying antirheumatic drugs for rheumatoid arthri- [25] Finfer S, Bellomo R, Boyce N. A comparison of albumin and saline
tis: abridged Cochrane systematic review and network meta-analysis. for fluid resuscitation in the intensive care unit. N Engl J Med
BMJ 2016;353:i1777. 2004;350(22):2247–56. doi:10.1056/NEJMoa040232.
[12] Florez ID, Veroniki AA, Al Khalifah R. Comparative effectiveness [26] Gallos ID, Papadopoulou A, Man R. Uterotonic agents for prevent-
and safety of interventions for acute diarrhea and gastroenteritis in ing postpartum haemorrhage: a network meta-analysis. Cochrane
children: A systematic review and network meta-analysis. PLoS One Database Syst Rev 2018;12(12):CD011689.
2018;13(12):e0207701. doi:10.1371/[Link].0207701. [27] Urquhart OA-O, Tampi MP, Pilcher L. Nonrestorative treatments
[13] Santesso N, Glenton C, Dahm P. GRADE guidelines 26: Informa- for caries: systematic review and network meta-analysis. J Dent Res
tive statements to communicate the findings of systematic reviews 2019;98(1):14–26.
of interventions. J Clin Epidemiol 2020;119:126–35. doi:10.1016/j. [28] Schunemann HJ. Interpreting GRADE’s levels of certainty or qual-
jclinepi.2019.10.014. ity of the evidence: GRADE for statisticians, considering review
[14] Siemieniuk R, Bartoszko JJ, Ge L. Drug treatments for information size or less emphasis on imprecision? J Clin Epidemiol
covid-19: living systematic review and network meta-analysis. Bmj 2016;75:6–15. doi:10.1016/[Link].2016.03.018.
2020;370:m2980.
[15] Piggott T, Morgan RL, Cuello-Garcia CA. Grading of recommenda-
tions assessment, development, and evaluations (GRADE) notes: ex-
tremely serious, GRADE’s terminology for rating down by three lev-
els. J Clin Epidemiol 2020;120:116–20. doi:10.1016/[Link].2019.
11.019.