Apr 20, 2019·MaxEnt 2019 - Proceedings of the 39th International Workshop on Bayesian Inference and Maximum Entropy Methods in Science and Engineering, Garching, Germany, 30 June - 5 July 2019
Randomization is an integral part of well-designed statistical trials, and is also a required procedure in legal systems. Implementation of honest, unbiased, understandable, secure, traceable, auditable and collusion resistant randomization procedures is a mater of great legal, social and political importance. Given the juridical and social importance of randomization, it is important to develop procedures in full compliance with the following desiderata: (a) Statistical soundness and computational efficiency; (b) Procedural, cryptographical and computational security; (c) Complete auditability and traceability; (d) Any attempt by participating parties or coalitions to spuriously influence the procedure should be either unsuccessful or be detected; (e) Open-source programming; (f) Multiple hardware platform and operating system implementation; (g) User friendliness and transparency; (h) Flexibility and adaptability for the needs and requirements of multiple application areas (like, for example, clinical trials, selection of jury or judges in legal proceedings, and draft lotteries). This paper presents a simple and easy to implement randomization protocol that assures, in a formal mathematical setting, full compliance to the aforementioned desiderata for randomization procedures.
Although the problems identified in the statement have been known for several decades, previous expressions of concern and calls for action have not fostered broad improvements in practice.2 A P value of 0.05 carries a 5% risk of a false positive result (i.e. there is no true difference between treatments). If a trial is meant to provide proof of a genuine treatment difference beyond reasonable doubt, a much smaller P value – say p < 0. 001 – is required.5 We disagree ….that our statement… is erroneous. According to the null hypothesis, P < 0.05 will occur 5% of the time.6 No editorial corrigendum has appeared. A P-value is the area under the curve of a probability distribution defined by a mathematical model. The model, usually presented graphically, describes the expected distribution of a sample statistic around a central measure, the parameter or theoretical ‘true’ value, for example the population mean, μ. Under the central limit theorem, this would be the standard normal distribution of sample means generated by repeat sampling of a population variable of interest. The mean of the sample means would equal the ‘true’ population mean, μ. In medicine, it is rare for us ever to know the true value of the variable of interest. However, we can usefully assign a value in the special case of a difference statistic, for example the difference in mean outcome variables in a placebo-controlled drug trial. In this case, the sampling distribution would represent that of the difference statistic. In this case, if the value we assign μ is zero then the mathematical model becomes the null hypothesis used in NHST. By way of contrast, non-inferiority drug trials require a non-zero value to be assigned. The cumulative AUC of the sampling distribution of a continuous variable is represented by a mathematical function called the cumulative density function. In medical science, most study variables are continuous or, if categorical, are transformed using the logit model. As the P-value is a mathematical integral, that is the cumulative AUC, it cannot take on a precise value as there is no AUC defined by a single point on the curve, for example the P-value ≤ 0.05, but not P = 0.05. While this may seem pedantic, the semantics of statistical inference are influential in thinking and decision-making yet misinterpretation and misuse of terminology are commonplace. Under the null hypothesis, one sample mean that happens to fall within an extreme region of the standard normal distribution may be expected to occur with a low frequency, say P ≤ 0.05 meaning such a sample mean or one more extreme would be expected to occur with a frequency of 5% or less. To be valid, the assumptions of independence and random selection of each sample mean selected from the normal distribution of sample means must be assumed. Another way of stating this is as a conditional probability: . Note: | means ‘given’. It is important to understand that the P-value is a measure conditional on the assumption that the mathematical model describes the distribution of sample means and is not a measure of the probability of the ‘truth’ of the mathematical model. To make this claim would invert the conditional probability statement and commit an error of reasoning called transposing the conditional7 aka the prosecutor's fallacy: . In reasoning from NHST, the commonly used definition of the P-value as ‘a measure of evidence against the null hypothesis’ is potentially misleading in that it seems to legitimise transposing the conditional as if it were a mathematically valid function rather than a matter of intuition. It was the intuitive interpretation that Fisher used in his a posteriori model of NHST.8, 9 His aim was to use the P-value as an aid in deciding which experiments to repeat. If on several repetitions, a consistent extreme P-value for the sample statistic was obtained then that would accumulate evidence for a true experimental effect. If no such effect was present, regression to the mean parameter (μ) would be expected (P ≥ 0.05). In real-life scenarios, many factors inhibit repetition and replication of experiments; however, modelling can give us insight into the precision and reproducibility of extreme P-values10, 11 and hence the intuitive weight we place on the P-value ‘as a measure of evidence against the null hypothesis’. Table 2 is a reproduction.10 It describes the results of simulating repeat experimentation and the probability of producing a P-value ≤ 0.05 under the prescribed conditions of the simulated experiment. It may be surprising to many how poorly reproducible the P-value is as a bright line test (a bright line test is a clearly defined rule or standard, the purpose of which is to produce consistent and predictable results). For example, if in the first experiment P ≤ 0.05 was produced there would be a 50% probability of reproducing P ≤ 0.05 in a repeat experiment; if P ≤ 0.01was produced in the first experiment the probability of producing P ≤ 0.05 in a repeat experiment, would be 73%; and if P ≤ 0.001 was produced in the first experiment the probability of P ≤ 0.05 in a repeat experiment would be 91%. The magnitudes of a number of these first experiment P-values are those commonly used in pharmaceutical trials and other medical analyses. The P-value is also sensitive to sample size. Irrespective of the effect size, with increasing sample size (n) the P-value can be made as small as you wish12 because the standard error is proportional to the inverse of n. If statistical significance is substituted for ‘clinical significance’ even small irrelevant differences may be regarded as worthy of investment. Large sample sizes are often a feature of pharmaceutical trials of secondary and primary prevention interventions such as preventive therapies in atherosclerotic diseases and osteoporosis. The quoted extract from the article on clinical trials mistakenly promotes the P-value as a measure of error and further states that the error rate can legitimately be adjusted depending on the magnitude of the P-value thus providing ‘proof of a genuine treatment difference beyond reasonable doubt’. This erroneous interpretation has arisen from the illusion of coherence resulting from the conflation of the dominant models of hypothesis testing.8, 9 The setting of theoretical type 1 (α) and type 2 (β) error rates in the Neyman and Pearson model envisions the frequency of error ‘in the long run of experience’ (experimental repetition) given randomness and independence of sample means from two juxtaposed probability distributions. A priori two identical populations are imagined except that they differ in mean parameters, null μ0 and alternative μA. This model is valuable in providing a rationality to sample size selection. However, the conflation has resulted in confusion between Fisher's P-value and Neyman's α giving the P-value an apparent legitimacy as an a posteriori ‘sliding’ type 1 error rate. Even if this were logical, decreasing α would increase β, resulting in a decrease in power (1-β). Also the dichotomous approach of pitting null hypothesis against alternative hypothesis carries the risk of blinding the researcher or the consumer to other explanatory hypotheses. For those who think the use of confidence intervals (CI) overcomes the problems described, think again. Although it has greater intuitive value especially with respect to estimating effect size, the CI relies on the same premises as the P-value. For example the CI of juxtaposed probability distributions can be made as large or as small as can be paid for by increasing the sample size such that for any small difference the CI can be made not to overlap. Statistical analyses are very valuable tools for extracting information from data. However, the reliability of the knowledge generated is dependent on many more important factors inter alia, evidential justification of the experimental hypothesis, study design, study conduct and data collection and cleansing, competence in choice of statistical model, valid reasoning, reviewer bias, publication bias and replication. Much of the criticism of medical science centres on its overemphasis on the importance of the P-value, NHST and statistically defined effect sizes. A better understanding of how sound statistical inferences are made and how they influence decision making will be key elements to improving all aspects of healthcare. This is critically important in acknowledgement of individuals as complex adaptive systems with characteristics of emergence, adaptability, non-linearity and unpredictability13 rather than as static population averages. Surveys suggest statistical literacy amongst doctors is low.14, 15 Teaching and assessing knowledge and application of statistical inference, critical appraisal and decision-making skills should be a primary focus of medical schools and specialist colleges. Difficult concepts underpinning statistical inference may be more effectively and efficiently taught using computer simulation whereby the learner can manipulate effect sizes, sample sizes and other statistics in order to see how parameter estimates, P-values and CI change with reproduction and replication.16 This will foster a more in-depth understanding of the limits of statistical inference, making clinicians better able to choose wisely amongst the myriad of investigations and treatment options on offer. Subsequent to article submission and review the author attended the referenced ASA conference.2 A special issue of the ASA journal reporting the conference proceedings is planned for 2018. In the opening addresses, the 400 participants were encouraged to devote their energies to developing proposals and goals to address the long standing yet stubbornly persistent errors in statistical inference described in this article. While concrete proposals are yet to be endorsed by the ASA, many speakers emphasised the need to place greater emphasis on teaching the conceptual framework of the different philosophical approaches to science (mastering the concepts as a priority rather than the mechanics of statistical inference). The need for better understanding of statistical semantics on the part of non-statistician scientists was also highlighted. Further that the best way to achieve understanding would be to develop context-specific learning modules. An aspect of the conference that resonated with the author with respect to prediction in medical science was the idea that science defines degrees of uncertainty (not certainty) apropos caution must be applied to the use of prediction models in medical practice lest they be over-extended.
The scientific credibility of findings from clinical trials can be undermined by a range of problems including missing data, endpoint switching, data dredging, and selective publication. Together, these issues have contributed to systematically distorted perceptions regarding the benefits and risks of treatments. While these issues have been well documented and widely discussed within the profession, legislative intervention has seen limited success. Recently, a method was described for using a blockchain to prove the existence of documents describing pre-specified endpoints in clinical trials. Here, we extend the idea by using smart contracts - code, and data, that resides at a specific address in a blockchain, and whose execution is cryptographically validated by the network - to demonstrate how trust in clinical trials can be enforced and data manipulation eliminated. We show that blockchain smart contracts provide a novel technological solution to the data manipulation problem, by acting as trusted administrators and providing an immutable record of trial history.
<ns4:p>Trust in scientific research is diminished by evidence that data are being manipulated. Outcome switching, data dredging and selective publication are some of the problems that undermine the integrity of published research. Methods for using blockchain to provide proof of pre-specified endpoints in clinical trial protocols were first reported by Carlisle. We wished to empirically test such an approach using a clinical trial protocol where outcome switching has previously been reported. Here we confirm the use of blockchain as a low cost, independently verifiable method to audit and confirm the reliability of scientific studies.</ns4:p>
Committee for Proprietary Medicinal Products (CPMP)
A number of recent applications have led to CPMP discussions concerning the interpretation of superiority, noninferiority and equivalence trials. These issues are covered in ICH E9 (Statistical Principles for Clinical Trials). There is further relevant material in the Step 2 draft of ICH E10 (Choice of Control Group) and in the CPMP Note for Guidance on the Investigation of Bioavailability and Bioequivalence. However, the guidelines do not address some specific difficulties that have arisen in practice. In broad terms, these difficulties relate to switching from one design objective to another at the time of analysis. The types of trials in question are those designed to compare a new product with an active comparator. The objective may be to demonstrate: the superiority of the new product the noninferiority of the new product or the equivalence of the two products. When the results of the trial become available, they may suggest an alternative interpretation. Thus the results of a superiority trial may only appear to be sufficient to support noninferiority, while the results of a noninferiority trial may appear to support superiority. Alternatively, the results of an equivalence trial may appear to support a tighter range of equivalence. A satisfactory approach to this subject requires an understanding of confidence intervals and the manner in which they capture the results of the trial and indicate the conclusions that can be drawn from them. Such an understanding also leads to an appreciation of why power calculations are of relatively little interest when a trial is complete. For simplicity, this paper addresses the issues of superiority, noninferiority and equivalence from the perspective of an efficacy trial with a single primary variable. Some comments on other situations are made in Section VI. It is assumed throughout this document that switching the objective of a trial does not lead to any change in the selection or definition of the primary variable. A superiority trial is designed to detect a difference between treatments. The first step of the analysis is usually a test of statistical significance to evaluate whether the results of the trial are consistent with the assumption of there being no difference in the clinical effect of the two treatments. In a trial of good quality, the degree of statistical significance (P value) indicates the probability that the observed difference, or a larger one, could have arisen by chance assuming that no difference really existed. The smaller this probability is, the more implausible is the assumption that there really is no difference between the treatments. Once it is accepted that the assumption of ‘no difference’ is untenable, it then becomes important to estimate the size of the difference in order to assess whether the effect is clinically relevant. This has two aspects. First there is the best estimate of the size of the difference between treatments (point estimate). For normally distributed data this is usually taken as the observed difference between the mean values on each. Next, there is the range of values of the true difference that are plausible in the light of the results of the trial (confidence interval). It is clear that this range should not include zero since the possibility of a zero difference has already been rejected as unreasonable. The method of constructing confidence intervals generally ensures that this is so, provided it corresponds to the choice of significance test. Thus the following two statements are usually equivalent: The two-sided 95% confidence interval for the difference between the means excludes zero. The two means are statistically significantly different at the 5% level (P < 0.05) two-sided. The above text addresses the situation where the difference between two mean values is the statistic of interest and a zero difference represents no effect. In practice a number of other summary statistics are used for the evaluation of differences between treatments, for example the odds ratio for proportions or the ratio of geometric means in bio-equivalence studies. (The latter arises from the logarithmic transformation used for bioavailability data.) In such cases the same principles apply but ‘no difference’ may be represented by a value other than zero – a value of 1 in both the examples quoted here. In these cases it is the position of the confidence interval for the test statistic relative to this ‘no difference’ value that is of interest. When significance tests are carried out in practice, precise numerical values of probabilities are usually quoted, for example P = 0.032, because this is more informative than P < 0.05. This allows judgement to be based more precisely on the extent of the disagreement between the null hypothesis and the observed data rather than on the approximations implied by using cut-off points of 0.05, 0.01 and 0.001. However, confidence intervals have to be associated with a specific probability value (coverage probability) and this is nearly always taken as 95% (0.95). When a difference is statistically significant at a more extreme level, e.g. P = 0.002, the two-sided 95% confidence interval will exclude zero by a wider margin. Figure 1 illustrates these points. Relationship between significance tests and confidence intervals. Whether the observed difference is indeed clinically relevant is a matter of judgement. In contrast to an equivalence or noninferiority trial where clinical relevance is addressed through the prestudy choice of Δ (see II.2 and II.3), in a superiority trial clinical relevance requires separate consideration: a statistically significant difference may not be clinically relevant. The difference taken as the basis of the power calculation in a superiority trial cannot be assumed to provide a suitable value. Note that in Figure 1, and throughout the rest of the document, it is assumed that values to the right of zero correspond to a better response on the new treatment so that values to the left are worse, i.e. better on the control treatment. An equivalence trial is designed to confirm the absence of a meaningful difference between treatments. In this case it is more informative to conduct the analysis by means of the calculation and examination of the confidence interval although there are closely related methods using significance test procedures. (See also II.3.) A margin of clinical equivalence (Δ) is chosen by defining the largest difference that is clinically acceptable, so that a difference bigger than this would matter in practice. There are well-recognized difficulties associated with this task which will not be discussed in any detail here. If the two treatments are to be declared equivalent, then the two-sided 95% confidence interval – which defines the range of plausible differences between the two treatments – should lie entirely within the interval −Δ to + Δ, see Figure 2. There are situations in which the equivalence margins may be chosen asymmetrically with respect to zero. Confidence interval approach to analysis of equivalence trial. In the case of bioequivalence studies a coverage probability of 90% for the confidence interval has become the accepted standard when evaluating whether the average values of the pharmacokinetic parameters of two formulations are sufficiently close. Clinical equivalence trials, with two-sided 95% confidence intervals, may be carried out when conventional bio-equivalence trials are impossible, for example in the case of a generic inhaled or topically applied product. In Phase III drug development, noninferiority trials are more common than equivalence trials. In these we wish to show that a new treatment is no less effective than an existing treatment – it may be more effective or it may have a similar effect. Again a confidence interval approach is the most straightforward way of performing the analysis but now we are only interested in a possible difference in one direction. Hence the two-sided 95% confidence interval should lie entirely to the right of the value −Δ, see Figure 3. Non-inferiority trials are sometimes mistakenly referred to, and designed as, equivalence trials. This distinction is important and can be a source of confusion. Confidence interval approach to analysis of non-inferiority trial. Note also that by using the closely related significance testing procedures referred to in II.2, it is possible to calculate a P value associated with the null hypothesis of inferiority. This is a valuable further aid to assessing the strength of the evidence in favour of noninferiority. It will be assumed throughout this document that two-sided 95% confidence intervals are to be used for all clinical trials whatever their objective. Among other benefits, this preserves consistency between significance testing and subsequent estimation. It is also consistent with the guidance provided in the ICH E9 Note for Guidance. If one-sided intervals are used, then they should be used with a coverage probability of 97.5%. In the special case of bioequivalence studies, two-sided 90% confidence intervals have been established as the norm as recommended, for example, in the CPMP Note for Guidance on the Investigation of Bioavailability and Bioequivalence. A conclusion of equivalence or noninferiority clearly depends upon the value of Δ chosen as the maximum acceptable difference. It is always possible to choose a value of Δ which leads to a conclusion of equivalence or noninferiority if it is chosen after the data have been inspected. Since the choice of Δ is generally a difficult one, there is ample room for bias here, however, well intentioned the researcher may be. Plausible arguments may often be advanced for a retrospective choice. In the design of equivalence and noninferiority trials, this reason (amongst others) makes it necessary for the choice of Δ, and the reasoning behind the choice, to be set down in advance by the researcher in the study protocol. The corresponding coverage probability for the confidence interval (usually 95%) should also be chosen at this time. (See Section IV.2 for how these requirements apply when objectives are changed.) The question of how to choose an appropriate Δ will be addressed in a subsequent CPMP Points to Consider. Pre-definition of a trial as a superiority trial, an equivalence trial or a noninferiority trial is necessary for numerous reasons including the following: to ensure that comparator treatments, doses, patient populations and endpoints are appropriate (see ICH E10) to allow sample size estimates to be based on the correct power calculations to ensure that equivalence and noninferiority criteria are predefined to permit appropriate analysis plans to be described in the protocol to ensure that the trial has sufficient sensitivity to achieve its objectives (see ICH E10) If the objective of a trial is switched from superiority to noninferiority, or vice versa, these aspects may lead to greater difficulty than the interpretation of significance tests and confidence intervals. The only switching which is likely to have any practical relevance is switching between superiority and noninferiority. The place of equivalence trials is so specific that they stand alone. If the 95% confidence interval for the treatment effect not only lies entirely above −Δ but also above zero then there is evidence of superiority in terms of statistical significance at the 5% level (P < 0.05). See Figure 4. In this case it is acceptable to calculate the P value associated with a test of superiority and to evaluate whether this is sufficiently small to reject convincingly the hypothesis of no difference. There is no multiplicity argument that affects this interpretation because, in statistical terms, it corresponds to a simple closed test procedure. Usually this demonstration of a benefit is sufficient on its own, provided the safety profiles of the new agent and the comparator are similar. When there is an increase in adverse events, however, it is important to estimate the size of the effect to evaluate whether it is sufficient in clinical terms to outweigh the adverse effects. Non-inferiority to superiority. There are a number of other factors that might be affected by this changed objective. If the comparator was suitable for a demonstration of noninferiority, then there should be well-controlled data to show that it is an effective treatment. Hence, for proof of efficacy, a clear demonstration of superiority to the comparator in terms of statistical significance should be acceptable. Non-inferiority trials are generally large because of their need to exclude the possibility of a small degree of inferiority of a new agent relative to an active control. However if the new agent is actually superior to control by a small amount, then the power to show its noninferiority is increased. Demonstrating the small amount of superiority to control might in principle require the planning of an even larger trial. When the trial is completed, however, the results provided by the confidence interval supply a concrete assessment of the precision actually achieved, superseding any calculations of power carried out before the trial was undertaken. Since the comparator in a noninferiority trial must be an effective agent, any superiority to that agent should carry the implication of acceptable superiority to no treatment (placebo). For this reason the size of the additional clinical benefit demonstrated is not likely to be relevant to a claim of efficacy except in relation to any increase in adverse effects and hence relative risk/benefit. However, when the proposed licence includes a claim of superiority to the comparator, the size of the additional benefit should be discussed in clinical terms. In a superiority trial the full analysis set, based on the ITT (intention-to-treat) principle, is the analysis set of choice, with appropriate support provided by the PP (per protocol) analysis set. In a noninferiority trial, the full analysis set and the PP analysis set have equal importance and their use should lead to similar conclusions for a robust interpretation. A switch of objective would require this difference of emphasis to be recognized. More details of the relative importance of these two analysis sets in superiority and noninferiority trials can be found in the ICH E9 Note for guidance. A trial to show equivalence or noninferiority must show a high degree of consistency with protocolled plans if it is to be reliable. Deviations from the inclusion criteria, from the intended treatment regimen, from the schedule, manner and precision of taking measurements, and so on, all tend to reduce the sensitivity of a trial and to make a conclusion of ‘no difference’ more likely, even when the deviations are of an unsystematic or random nature. The size of the bias associated with these and other departures from the protocol is generally unknown and may render such a trial uninterpretable. Failure to show a difference between two treatments can also arise when both treatments are inefficacious, perhaps as a result of being inappropriately administered. This problem does not affect superiority trials to the same extent because the demonstration of a difference is itself validation of the sensitivity of the trial. The estimate of the size of the effect may however, be similarly affected. For these reasons, switching from noninferiority to superiority is likely to carry with it a greater degree of confidence in the conclusion. Switching the objective of a trial from noninferiority to superiority is feasible provided: The trial has been properly designed and carried out in accordance with the strict requirements of a noninferiority trial. Actual P values for superiority are presented to allow independent assessment of the strength of the evidence. Analysis according to the intention-to-treat principle is given greatest emphasis. If a superiority trial fails to detect a significant difference between treatments, there may be interest in the lesser objective of establishing noninferiority. If the results of the superiority trial are summarized by means of a 95% confidence interval for the treatment difference, the lower end of that confidence interval provides a quantitative estimate of the minimum estimated effect of the new treatment relative to the comparator. When the study protocol an acceptable, margin −Δ for noninferiority, the objective less a noninferiority margin would appear only to make in trials with noninferiority as an However, in any superiority trial where noninferiority may be an acceptable for it is to a noninferiority margin in the protocol in order to the difficulties that can arise from such it is also to design to the possible need to that the study sufficient sensitivity to detect the drug effects of interest (see It is important to that there are of where noninferiority to an active control is to be acceptable as the or evidence of efficacy, and trials are to In trials where there is no noninferiority such a has to be after the and in situations this will not be It is likely that the will have to be after the results have been and there may be little basis for an objective choice of margin. there does not appear to be a statistical multiplicity related to this switch of that does not the difficulties associated with the definition of A number of other issues require A comparator chosen for a demonstration of superiority may not be acceptable for a conclusion of noninferiority. In order for it to be acceptable, it will be necessary to that there are data from good superiority trials consistent evidence that the comparator is an effective treatment with and establishing the size of its effect relative to no treatment. There should also be a basis for that the same degree of efficacy would be in the trial. For example, the patient and the endpoints should be similar. These issues are covered in ICH in the results provided by the confidence interval supply a concrete assessment of the precision actually by a clinical trial, superseding any calculations of power carried out before the trial was undertaken. The position of the lower end of the confidence interval relative to the of noninferiority provides the for noninferiority. In a superiority trial the full analysis set, based on the ITT (intention-to-treat) principle, is the analysis set of choice, with appropriate support provided by the PP (per protocol) analysis set. In a noninferiority trial the full analysis set and the PP analysis set have equal importance and their use should lead to similar conclusions for a robust interpretation. A switch of objective would require this difference of emphasis to be recognized. More details of the relative importance of these two analysis sets in superiority and noninferiority trials can be found in the ICH E9 Note for Guidance. A trial to show equivalence or noninferiority must show a high degree of consistency with protocolled plans if it is to be reliable. Deviations from the inclusion criteria, from the intended treatment regimen, from the schedule, manner and precision of taking measurements, and so on, all tend to reduce the sensitivity of a trial and to make a conclusion of ‘no difference’ more likely, even when the deviations are of an unsystematic or random nature. The size of the bias associated with these and other departures from the protocol is generally unknown and may render such a trial uninterpretable. Failure to show a difference between two treatments can also arise when both treatments are inefficacious, perhaps as a result of being inappropriately administered. This problem does not affect superiority trials to the same extent because the demonstration of a difference is itself validation of the sensitivity of the trial. For these reasons, switching from superiority to noninferiority is likely to carry with it a lesser degree of confidence in the It will be necessary to to the sensitivity of the trial by or that the control treatment is its efficacy the trial with trials which demonstrated the efficacy of the control agent in of and of of and data that are at to those in the trials similar results from the full analysis set and PP analysis set. Switching the objective of a trial from superiority to noninferiority may be feasible provided: The noninferiority margin with respect to the control treatment was predefined or can be (The latter is likely to difficult and to be to cases where there is a accepted value for Analysis according to the intention-to-treat principle and PP confidence intervals and P values for the null hypothesis of similar The trial was properly designed and carried out in accordance with the strict requirements of a noninferiority trial (see ICH E9 and The sensitivity of the trial is high to ensure that it is of relevant differences if they There is or evidence that the control treatment is its level of A further related that has arisen in with equivalence and noninferiority trials to the equivalence margins when the trial is complete. that a bioequivalence trial a 90% confidence interval for the relative bioavailability of a new that from to we only that the relative bioavailability lies between the conventional of and because these the predefined equivalence can we that it lies between and The interval based on the data is the appropriate one to Hence, if the changed to this study would have satisfactory There is no question of a selection However, if the trial in a confidence interval from to then a change of equivalence margins to would not be acceptable because of the conclusion that the equivalence margin was chosen to the These apply to the 95% confidence intervals used for clinical equivalence and for noninferiority. The confidence interval based on the results of the trial is always the best summary of the It is the choice of equivalence margin that is subject to This should be chosen on the basis of and not chosen to the This Points to has been from the perspective of an efficacy trial active with a single primary variable. In practice some studies have more than one primary and most studies have respect to switching of these requires separate in the of the specific drug development, separate conclusions superiority or noninferiority for in judgement whether the trial as a has established the superiority or noninferiority of the new treatment will upon the requirements for that clinical and the of results all relevant The covered in these Points to can also be applied to specific safety when these have been as endpoints of a trial to compare active In practice the of switching objectives is not relevant to trials, even where noninferiority to is a valuable i.e. for safety The problem of switching objectives can be by a trial in the that both noninferiority and superiority are of value. In this case all the issues in this document should be addressed In the statistical analysis should be using an appropriate from noninferiority to superiority. The interpretation of superiority trials as noninferiority trials and vice is best by the results as a confidence interval for the difference between the test treatment and control. There is no problem associated with the use of this confidence interval as a basis for of interpretation. For a and trial, there are difficulties with the change from noninferiority to superiority that cannot be addressed by appropriate analysis. However, there are more difficulties associated with the switch from superiority to noninferiority because of the possible need to a basis and on, a margin of equivalence after the and because of the difficulties of noninferiority trials. There are for the design of a superiority trial in which noninferiority might be an acceptable When the results with respect to alternative of the equivalence margins the problem from to switch to wider acceptable that equivalence margins may be in this
Open access
Statistical Methods in Clinical Trials
Health Systems, Economic Evaluations, Quality of Life
. We give several efficient transformations for manipulating the statistical difference (variation distance) between a pair of probability distributions. The effects achieved include increasing the statistical difference, decreasing the statistical difference, &quot;polarizing&quot; the statistical relationship, and &quot;reversing&quot; the statistical relationship. We also show that a boolean formula whose atoms are statements about statistical difference can be transformed into a single statement about statistical difference. All of these transformations can be performed in polynomial time, in the sense that, given circuits which sample from the input distributions, it only takes polynomial time to compute circuits which sample from the output distributions. By our prior work (see FOCS 97), such transformations for manipulating statistical difference are closely connected to results about SZK, the class of languages possessing statistical zero-knowledge proofs. In particular, some of the transformation...