Beyond Effect Size
Abstract
It is our contention that the authors of many clinical trials in anesthesia are not fully considering the implications of the minimum effect size of interesta entered in their power calculations, and are therefore making conclusions that are not supported by their findings. In this paper we use hypothetical examples to explain how the choice of minimum effect size of interest in the design of clinical trials sets conditions on their interpretation, including the threshold for clinical relevance, the magnitude of effect sizes that can be confidently excluded, and the discrimination between primary and secondary outcomes. We then use examples from a sample of recent highly cited anesthesia trials to show that this is a problem, which as a specialty we should be addressing. THE MINIMUM EFFECT SIZE OF INTEREST IN THE DESIGN OF CLINICAL TRIALS Let us suppose that a group of researchers wishes to investigate whether a new short-acting antihypertensive drug has a clinically worthwhile treatment effect for the attenuation of the arterial blood pressure response to tracheal intubation. The first step in planning their study is to estimate the sample size required to detect a clinically worthwhile difference for their primary outcome (in this case the blood pressure response), with adequate power.1–4 The alternative, using a confidence interval (CI) approach, is to predefine an acceptable CI width for the primary outcome.b5,6 After careful consideration of all factors, including drug costs, they decide that the minimum clinically worthwhile difference or “threshold for clinical relevance” is 10 mm Hg. This is their minimum effect size of interest. They perform a t test with a null hypothesis that there is no difference between the new drug versus control. They accept a type I error rate of 5% (α = 0.05; i.e., P < 0.05 will be considered significant), and a type II error rate of 20% (β = 0.2, i.e., power = 80% [power = the probability of getting a statistically significant result if there is a true difference ≥ the specified minimum effect size of interest]).1–4 They anticipate that the SD in both groups will be about 10 mm Hg. With these variables they require n = 16 in each group.7 (If they had chosen a minimum effect size of interest = 5 mm Hg, they would have required n = 63 in each group; alternatively, with n = 16 in each group, their power would be reduced to around 29%.)7 THE ROLE OF THE MINIMUM EFFECT SIZE OF INTEREST IN THE INTERPRETATION OF CLINICAL TRIALS Scenario 1: No Significant Difference In a hypothetical scenario, the authors find that the mean difference is 6 mm Hg (95% CI −0.4 to 12.4 mm Hg), P = 0.06. They accept (or more correctly, fail to reject) their null hypothesis and conclude that there is “no difference” between the groups. This conclusion is not correct. A nonsignificant finding does not support a conclusion of no difference without qualification.2–4,6 Nonsignificant findings are qualified by the minimum effect size of interest entered in the power calculation, and the power. This is because the minimum effect size of interest entered in the power calculation is also the “minimum detectable difference” of the trial.1–4 The trial does not exclude or confirm a difference up to this value (in this case 10 mm Hg). Moreover, the power (1 – β, in this case 80%) defines the likelihood of a type II error (β, in this case 20%). In other words, in this scenario there is still a 20% chance that there is a true difference ≥10 mm Hg, even though the investigators failed to find a statistically significant difference. This is very different from an unqualified conclusion of “no difference”! Moreover, while the null hypothesis may not have been rejected, the actual P value continues to provide important information about the probability of observing the finding (or one more extreme) given that the null hypothesis is true. For example, in this case the probability is 6%; in contrast, if the P value were 0.6 rather than 0.06, it would be 60%. Furthermore, the observed point estimate and its 95% CI remain the best estimate of the difference between groups, whether or not statistical significance is reached. For example, an observed 95% CI of −5.4 to 7.4 mm Hg would likewise have been “not statistically significant,” but would have suggested that the true point estimate is closer to 1 mm Hg, which would be very different than the actual example, where the most likely point estimate is about 6 mm Hg. The importance of the minimum effect size of interest and power when interpreting nonsignificant findings becomes clearer when we consider the possibility of smaller true effect sizes. For example, let us say that the drug cost turns out to be lower than expected and that most clinicians would accept a 5 mm Hg difference as clinically worthwhile, rather than the 10 mm Hg used by the authors. How then does the information from this trial help them in their decision whether to use the drug? The answer is very little. The trial was designed to have sufficient probability of detecting a difference, given that the true difference is ≥10 mm Hg. It provides little information on the probability of detecting a difference, given a true difference <10 mm Hg. Scenario 2: Significant Difference and Observed Effect ≥ Minimum Effect Size of Interest In another hypothetical scenario, the authors find a mean difference of 12 mm Hg (95% CI 0.5 to 23.5 mm Hg), P = 0.04. They reject their null hypothesis (acknowledging a 5% chance that they are making a type I error) and correctly conclude that there is a statistically significant treatment effect. They also correctly conclude that this treatment effect is likely to be clinically worthwhile, because the point estimate is at least as large as their predefined minimum clinically worthwhile difference (= minimum effect size of interest stipulated in their power calculation). Scenario 3: Significant Difference and Observed Effect < Minimum Effect Size of Interest In yet another hypothetical scenario, the authors find a difference of 6 mm Hg (95% CI 0.3 to 11.7 mm Hg), P = 0.04. They correctly conclude that there is a statistically significant treatment effect. But what should they conclude about whether the effect is clinically worthwhile? The hypothesis they tested was that there was no difference between the groups (null). The P value <0.05 supports rejection of this hypothesis. However, it does not provide information on the likely magnitude of the effect or whether it is clinically worthwhile. To assess whether the observed effect size is worthwhile, it is necessary to refer to the minimum clinically worthwhile difference decided before the trial began, and which was entered as the minimum effect size of interest in the power calculation (in this case 10 mm Hg). Therefore, the correct conclusion is that while there is a statistically significant effect in this particular trial (i.e., P < 0.05), the observed effect (6 mm Hg) is too small to be clinically worthwhile (i.e., <10 mm Hg). The authors may not accept this conclusion. They may argue that the 95% CI around their point estimate of 6 mm Hg includes 10 mm Hg, so their finding is still compatible with a clinically worthwhile difference. However, they would have to concede that the same 95% CI would be compatible with a range of effect sizes as small as 0.3 mm Hg, and that the true effect size is most likely closer to the point estimate of 6 mm Hg (i.e., <10 mm Hg). The authors may also argue that their original 10 mm Hg was not a true minimum clinically worthwhile difference, and was chosen only for pragmatic reasons to limit the required sample size. They may argue that 5 mm Hg is a more realistic value in any case. However, their trial did not have adequate power (i.e., only around 29%) to detect a ≥5 mm Hg difference.7 The authors may argue that the power is now irrelevant, because they have observed a statistically significant difference. This is a misconception, because there is no guarantee that a repeat trial with the same sample sizes would provide another significant result.8,9 The likelihood of reproducing a significant result (assuming that a true difference ≥ the stipulated minimum effect size of interest exists) is equal to the power of the trial.8 For example, with 80% power, if the trial were repeated using different samples of the same size, there would be an 80% chance of again observing a significant difference.8 In contrast, with 29% power, the likelihood would be only around 29%.8 It would be difficult to be conclusive about any finding with this low level of replicability. For this reason, the power of a study is important for both negative and positive findings. Reducing the minimum effect size of interest after the fact reduces the power, thereby reducing the likelihood of replicating a finding. In effect, the authors are faced with a dilemma. Either they accept that the observed effect size is too small to be clinically worthwhile, or they accept that they cannot be confident that the finding has >80% chance of being repeated. Neither of these supports a conclusion of a reproducible clinically worthwhile effect. THE MINIMUM EFFECT SIZE OF INTEREST AND PRIMARY VERSUS SECONDARY OUTCOMES Let us say the authors also assessed the heart rate (HR) response, but because this was a secondary outcome, they did not perform a power calculation. They found that the mean difference in HR was 6 beats per minute (bpm) (95% CI −1 to 13 bpm), P = 0.10. How should they interpret this finding? To interpret it correctly they need to refer to the minimum effect size of interest (= minimum detectable difference) in the power calculation and the power. Clearly, if these values are not presented, it is not possible to make a meaningful interpretation. Similarly, it would be difficult to interpret a P value <0.05, because without knowing the power and the minimum effect size of interest, it would not be clear how likely the result could be replicated.8,9 Unfortunately, it is not possible to extrapolate power from other outcomes.8,9 Moreover, it is not appropriate to perform a power calculation once the results are already known.6 The authors may argue that a power calculation is not necessary, because the CI alone provides sufficient information. This is another misconception. The CI does not provide sufficient information on the adequacy of the sample size, a major determinant of the CI width.5,6 Perhaps a larger sample size for the HR outcome (with the same sample variability) would have reduced the CI width around the same point estimate sufficiently for the lower limit to be above zero? In fact, ensuring an adequate sample size is equally important using CI as it is using inferential tests, as is predefining a clinically worthwhile difference.5,6 For these reasons, it is not possible to make conclusions about secondary outcomes, unless they are accompanied by this information.9,10 They might still be important (depending on the point estimate and the actual P value or the 95% CI), but remain as “observations” until confirmed or excluded in future studies.9,10 Nevertheless, let us say that in this particular trial the authors made conclusions about both blood pressure and HR responses without specifying which was the primary outcome. A quick check of the minimum effect size of interest would identify the blood pressure response as the primary outcome (i.e., the outcome for which the power had been calculated, the sample size estimated, and the threshold for a clinically worthwhile difference set).9 In this way the minimum effect size of interest discriminates between primary and secondary outcomes. A SAMPLE OF HIGHLY CITED ANESTHESIA TRIALS We identified 20 highly cited prospective anesthesia trials by interrogating the ISI Web of knowledge (http://apps.isiknowledge.com/, accessed December 2010) using the following search strategy: topic = anesthesia or anaesthesia; journal = Lancet, New England Journal of Medicine, Anesthesiology, Anesthesia and Analgesia, or British Journal of Anaesthesia; year of publication = 2001 to 2010, with ranking of trials by number of citations.11–30 Publications other than prospective clinical trials were excluded. We scrutinized the top 20 most cited trials for conclusions that were not supported by the minimum effect size of interest stipulated in their power calculations. We did not recheck any statistical analysis or assess any other aspect of the trials. Conclusions Based on Secondary Outcomes There were 10 trials that based 1 or more conclusions on secondary outcomes (for which no minimum effect size of interest or power was provided) (Table 1).15,17,18,21,22,24,27–30 Three trials even included the findings of a secondary outcome in their title (Table 1).15,17,27 In many cases, it was not possible to differentiate between primary and secondary outcomes without reference to the minimum effect size of interest in the power calculation.Table 1: Studies that Include Secondary Outcomes in Their ConclusionsConclusions Based on Statistically Significant Findings too Small to Be Clinically Worthwhile There were 5 trials with statistically significant findings for their primary outcome, but with an observed effect size less than their minimum effect size of interest (Table 2).13,17–19,30 For example, Myles et al. chose a minimum effect size of interest of 0.9% “because uptake into routine practice would require convincing proof of benefit.“13 Yet they observed a mean effect size of only 0.74%.13 Similarly, Carli et al. specifically chose a “minimum effect size of interest” of 36 m walked in 6 minutes, because this difference produced “a meaningful impact” on long-term exercise capacity.17 Yet they observed a mean difference of only 33.6 m at 3 weeks, and 18 m at 6 weeks.17 Neither Myles et al. nor Carli et al. concluded that their observed effect was too small to be clinically worthwhile. Similar considerations apply to the other 3 trials in this category (Table 2).18,19,30 Only Myles et al. presented the 95% CI for their observed effect size. The remainder either presented no CI for the observed effect size, or presented CI in a different metric to the minimum effect size of interest (Table 2).Table 2: Studies with a Statistically Significant Primary Outcome but an Observed Effect Size Less than the Minimum Effect Size of InterestConclusions Based on Nonsignificant Findings—Unable to Exclude All Clinically Worthwhile Effect Sizes Four trials had nonsignificant findings for their primary outcomes (Table 3).11,22,26,27 Scrutiny of their minimum effect size of interest indicated that none could confidently exclude all clinically worthwhile effect sizes. For example, they were powered to detect differences in the incidence of morbidity and mortality ≥10%, length of stay ≥2.5 days, awareness incidence ≥0.9%, and block success rate ≥23%, respectively. Yet effect sizes below these ranges might still be considered clinically worthwhile. (e.g., mortality and morbidity reduction of 9%, length of stay reduction of 2 days, incidence of awareness reduction of 0.8%, block success rate improvement of 22%). Only Rigg et al. explained that they could not confidently exclude the possibility of a worthwhile true effect size less than the minimum effect size of interest stipulated in their power calculation.11 Only Avidan et al. presented the 95% CI for their observed effect size (which was instead of a P value from an inferential test).26Table 3: Studies with Nonsignificant Findings for the Primary OutcomeConclusions Based on Findings with no Power Analysis or Stipulation of Minimum Effect Size of Interest There were 4 trials with no power calculation or minimum effect size of interest for any outcomes.12,16,23,25 None of these explained that their findings could not be fully interpreted without this information. CONCLUSION To fully interpret a clinical trial in which inferential statistics are used, it is necessary to go beyond effect size, and consider also the minimum effect size of interest stipulated in the power calculation. This is an important value, which not only has a major influence on the required sample size, but also defines the threshold for clinical relevance for positive findings, and the minimum detectable difference for negative findings (Fig. 1). Readers should also scrutinize the value chosen by the authors, to determine if it is appropriate. Failure to consider the minimum effect size of interest may result in erroneous conclusions, such as conclusions based on secondary outcomes, on outcomes that are statistically significant but not clinically worthwhile, or on nonsignificant findings that do not exclude the possibility of a smaller, but nevertheless true clinically worthwhile treatment effect. We have provided examples of such conclusions in a sample of highly cited anesthesia trials in a selection of high-impact-factor journals. Given their criteria for selection, it is unlikely that these trials represent a negatively biased sample in terms of quality of statistical reporting. We suspect that similar findings would be found in any sample of anesthesia trials. To address this situation, we recommend greater rigor in the design and interpretation of clinical trials, with closer scrutiny of the minimum effect size of interest (by both authors and readers), adequate power, and a focus on primary rather than secondary outcomes. For key secondary outcomes, we recommend that additional a priori power calculations be provided, along with their minimum effect sizes of interest. The use of CI for the observed effect size has advantages, because CI provide information on the most likely true effect size and the range of likely true effect sizes for both primary and secondary outcomes. Nevertheless, the same principle of defining the minimum clinically worthwhile effect size before the trial commences applies, as well as holding to this value when interpreting outcomes, and ensuring that an adequate sample size was used.Figure 1: The central role of the minimum effect size of interest in the design and interpretation of clinical trials. Once chosen, the minimum effect size of interest determines the sample size required (for any given level of power, α, and sd of the samples). The threshold for clinical relevance for the observed effect size and the minimum effect size detectable (given the power) are mathematically equal to this value. As sample size cannot be changed once the trial is completed, none of the values can be altered post hoc without affecting the trial's power.DISCLOSURES Name: Neville M. Gibbs, MD, FANZCA. Contribution: Study design, conduct of study, data analysis, and manuscript preparation. Name: William M. Weightman, MB, FANZCA. Contribution: Study design, conduct of study, data analysis, and manuscript preparation. This manuscript was handled by: Franklin Dexter, MD, PhD.
Community
0 commentsNo discussion yet
Be the first to share a question or observation.