All that glisters… How to assess the ‘value’ of a scientific paper
Abstract
Doctors are expected to keep up to date with the literature, and to change their practice on the basis of what they have read. We believe that the skills needed for evaluating and interpreting scientific papers are often lacking. Even when they are taught, the emphasis is often focused on technical aspects of a paper without consideration of its value, a term we use here to encompass not only the quality of a paper's methods, but also its context and importance. We argue that critical appraisal, the term used to describe the assessment of a paper's methodological niceties [1], should never take place in isolation but must always occur in parallel with assessment of its value, since a paper may be methodologically sound but contribute little to a better understanding of the subject. Here, we discuss various aspects of scientific papers that contribute to or detract from their value, and suggest a stepwise approach for assessing them. The purpose of a scientific publication, of which there are numerous types, is to communicate information. Editorials are summary or personal views, perhaps commenting on a specific paper. Reviews are longer, in-depth analyses of the literature; they are most persuasive when their analysis is objective but their conclusions (as with editorials) are often subjective and personal. Specific forms of supposedly objective reviews are quantitative or systematic reviews, or meta-analyses[2]. Case reports describe one (or more) specific clinical cases, either so unusual as to be of interest, or from which it is hoped some particular lesson can be learned. A letter (correspondence) is a comment on, or criticism of, another's published paper, or seeks answers to a question. A letter may also convey information or data that do not constitute a full study but nonetheless are thought worthy of dissemination. Finally, there are experimental investigations. Broadly, these attempt to answer a question by the use of the ‘scientific method’, which involves the following process [3, 4]. First, the researcher formulates a hypothesis. This hypothesis may represent current ideology (the current theory, or paradigm), or it may be an idea developed de novo (e.g. suggested by preliminary observations). The hypothesis leads to an experimental prediction: if a certain experiment is conducted, then this predicted result should obtain if the hypothesis is correct. A result consistent with the prediction supports the hypothesis, but does not ‘prove’ it (thus scientific proof is very different from mathematical proof). If, on the other hand, the experimental result is not as was predicted, then either the hypothesis is incorrect (so disproved) or the conduct of the experiment was flawed. There are many types of experimental investigations: clinical studies or trials (the terms are often used interchangeably, although the latter is sometimes restricted to investigations of a treatment's efficacy); laboratory investigations; mathematical modelling of data; and some observational studies and audits. We focus here on experimental investigations since these are the papers that advance the knowledge base the most. There are generally two aspects to excellence in an experimental study: first, the study's conduct (of which handling of the data is an inherent part), and second, its presentation. Important aspects of presentation are covered by journals' instructions to authors and by specific publications [5, 6], and we will not discuss these in detail here. First, the study must consider a clear hypothesis that is stated unambiguously, ideally, illustrated by the results that would be expected if the hypothesis were true. The final results and conclusion of the study should relate to this prediction and must either refute or be consistent with the hypothesis. Second, the study must have appropriate ethical approval (or conform to animal research guidelines). We will not consider this issue further, since it has been addressed recently [7]. Third, the technical conduct of the experiment must be sound. Particular attention should be given to the avoidance of surrogate measures, appropriate measurement tools, proper randomisation and blinding, appropriate use of control groups, and appropriate application and interpretation of statistics. We consider each of these below. Surrogate measures or end-points have serious limitations. For example, a study of the control of minute ventilation may not actually involve measurement of this outcome at all, but instead may try to derive conclusions based on the measurement of another, related variable (say, arterial Pco2). The problem is that other factors may influence the related variable (e.g. ventilation is not the only factor influencing Pco2). Furthermore, measures that seem superficially related may not be; for example, although flecainide reduces cardiac arrhythmias, it increases mortality – a much more relevant end-point [8]. Surrogate measures are sometimes used in clinical studies because of the difficulty in measuring the desirable end-point, such as long-term survival. They might also give crude estimates of trends over time for certain variables [9], but they have very little (if any) place in studies that seek to question or overturn fundamental hypotheses in the underlying science. Measuring devices and assessment tools must be valid (measure what they are supposed to measure), accurate (measure the true quantity), and reliable (different users obtaining the same results) – aspects that often escape attention in manuscripts. This is not restricted to technical measurements; for example, assessment of ‘maternal satisfaction’ with the use of a simple visual analogue scale continues widely despite little evidence to support it [10]. When technology is used, coefficients of variation should be given, but rarely are. Further, there ought to be some confidence as to how the technology works. For example, much of the growing literature on the bispectral index (BIS) as a monitor of ‘awareness’ raises concerns. Since it is not known precisely what is being measured by the BIS [11], it can never be known whether an unexpected result has arisen because the BIS is invalid or inaccurate, or because the hypothesis being tested is incorrect [12]. Randomisation and blinding are intended to minimise the influence of bias. The hope with randomisation is that all ‘confounding factors’ (both known and unknown) that might influence the outcome will be distributed equally amongst the groups. If any differences are found they can therefore be attributed to the sole factor – the treatment under study – that has not been ‘shared out’ in this way. However, even with proper randomisation, groups may be unequal: chance alone might result in one group's subjects being older, heavier, younger or just luckier than those in the other group. Even when groups appear equal, small inequalities might combine to influence the results. For example, in a study by Greif et al. [13] into the possible anti-infective effect of peri-operative oxygen therapy, patients randomly allocated to receive extra oxygen were by chance more likely to be fitter and less likely to be smokers, to have inflammatory bowel disease, and to undergo rectal surgery than those in the ‘no oxygen’ group. Could these factors have combined to contribute to the dramatic reduction in infection seen in the ‘oxygen’ group, such that a subsequent study obtained completely the opposite results [14]? In addition, the human urge to guess or manipulate treatment allocations, or otherwise interfere with proper randomisation in studies, is well reported [15]. Blinding is present when the person treating the patient, or making the assessments, does not know which patients receive which treatment. Some studies fail to take even the simplest steps to ensure blinding, while others go to extraordinary lengths (for an example of the latter, see Smith and Thwaites's [16] commendable study comparing intravenous with inhalational anaesthesia). Occasionally, blinding is impossible (e.g. when comparing two different laryngoscope blades or bougies [17]), but the results of a blinded study are always more persuasive. Indeed, studies with insufficient blinding tend to report greater treatment effects than those with proper blinding procedures [15]. Control groups may be inappropriate because of poor randomisation or blinding. Even if these are sound, though, there may be other reasons why treatment effects may be masked or exaggerated by problems with control groups: lack of consideration of other possible ‘confounding factors’; use of historical, rather than contemporaneous, controls; comparison of a treatment against placebo instead of standard practice, or against an inappropriate treatment; or (at worst) lack of a control group at all. Much of the above might give the impression that the testing of hypotheses by close attention to the study's conduct will always give a clear-cut result. Unfortunately, this is not the case, because we can never achieve certainty; the best we can do is use statistical analysis to indicate the degree of uncertainty [18]. A full discussion of statistical methods is dealt with elsewhere [19], and here we consider only two related areas that commonly cause difficulties: significance and power. Significance. If, in a study, one drug appears more effective than another, the traditional approach is to ask the question: ‘What is the likelihood (or probability) that this result is a chance finding, and that these two drugs are in fact equivalent?’ This approach (testing the equivalence, as opposed to testing the difference) is known as testing the null hypothesis. It is important to stress that this null hypothesis assessed by the statistical test may not always be exactly the same as the underlying scientific hypothesis being examined by the study as a whole: the result of the former will help interpret the latter. Many statistical tests ultimately generate a ‘p-value’: the lower the p-value, the less likely it is that the null hypothesis is correct. A p-value of 0.03 indicates that if the two drugs are indeed equivalent, chance alone would be expected to yield the observed results three out of every 100 times one conducted the study. Conventionally, a p-value of < 0.05 is taken to represent ‘statistical significance’, although this value can and should be adjusted in certain circumstances, for example if multiple comparisons are made [20, 21]. Confidence intervals can be used as an alternative to testing the null hypothesis [22, 23]; nonetheless, the conclusions reached by using confidence intervals are invariably the same. The real problem lies in how p-values are interpreted rather than how they are calculated [24]. An entirely different approach is to interpret a study's results mathematically in the context of prior knowledge (Bayesian statistics) [25, 26]. Regardless of the method of calculating or presenting statistical significance, the smaller the p-value, or the further away from zero the difference in confidence intervals, the less likely the result is to be a ‘chance’ finding and therefore the more likely it is that the difference between the two groups is indeed ‘genuine’. But such a chance finding is still possible, albeit unlikely; as Counsell et al. [27] point out, chance ‘…doesn’t get the credit it deserves'. Furthermore, a low p-value does not exclude poor methodology in the conduct of the study. Power. If the p-value in a drug study is, say, 0.07, does this mean there is genuinely no difference between the two drugs? Or does it mean that there might be a true difference, but that the study has simply failed to show it? It is specifically to help answer such questions that a power analysis is useful. The power analysis estimates how likely it is that a negative result can be ‘believed’. One emerging problem is that some researchers (or their critics) are placing far too much emphasis upon power analysis [28]. One example demonstrates the type of misplaced faith in power analysis: ‘…at least 400 patients would be required to prove there is no statistically significant difference between the groups…’[29] (our emphasis). Such statements reveal a poor understanding of the scientific method and of the concept of scientific proof. One reason for our concern is that the concept of power analysis itself has very serious limitations. Power analyses are only crude estimates of a sample size (indeed, the word ‘crude’ is emphasised by statisticians [30]). For example, two main elements that contribute to power for normally distributed continuous data are the difference between the means that is deemed important and the expected standard deviation (SD) of the measure of interest. Both of these are subject to serious shortcomings. The choice of what constitutes an ‘important difference’ is almost entirely subjective, and small but arbitrary adjustments to its value can have a great impact upon a study's calculated power. Where no previous data exist, the expected SD is usually taken from a pilot study, often without a control group, and by definition always smaller and less robust than the planned substantive study. In reality, the power analysis itself is probably best expressed in terms of a confidence interval: for example, ‘Power analysis indicated that we would require 20–60 subjects to be 70–90% confident of detecting a difference between the means of 10–50 s’– although this is rarely done. The crudeness of the power analysis as a tool is reflected in the different sample sizes yielded by different methods. If we assume that for a hypothetical study, the important difference is 1.0 arbitrary units, with a standard deviation of 0.8 arbitrary units, then various calculations give a sample size per group (with 80% power at p < 0.05) of 10 [30], 11 [31], 13 [32] and 18 [33]. So at best, power analysis only gives an approximate estimate of sample size. Indeed, Bacchetti [34] has suggested that a study's power should only be criticised if the study has no other shortcomings; in other words, all other features of a study (especially relating to its conduct) are far more important than its power analysis. In practice, the actual sample sizes of studies are related to the type of outcome, the variability of the result and the statistical test used. We observe that in published studies, sample sizes tend to fall into three groups, though with considerable overlap in which the outcome is and or there is little variability with factors (e.g. laboratory or studies in which experimental can be with use smaller sample usually increases patients are since are to studies comparing have sample in which the outcome is variability is great factors are much sample sizes It is how often studies and their sample sizes fall into these groups. It is to whether this is a of proper power analysis the of each study, whether tend to use sample sizes because of the in their particular of or whether they even the power calculations to yield sample sizes that the study to be a There are ethical reasons for power calculations as well as scientific and we do not suggest that should but given the crudeness of these calculations we help but whether the crudeness of is any It can be seen from the above that studies, despite their not to subjects in advance – as they do on data – are than have to on surrogate measures, using tools that be with little of blinding and of the groups and their The one of studies is the with which very can be So much for a paper's We to a different question: what one paper more than another, proper and attention to their some the answer to this question will always be subjective and upon the interest, and However, we suggest that there are some elements that contribute to a paper's We these as being related to the the type of question addressed and the answer the of the evidence certain aspects of a paper. The of the type of question to its This will on the (e.g. of the person assessing the paper, the significance of the problem being and whether the and appropriate question is being For clinical studies, the of a paper can on the of the change in practice likely to from the the with which that change might and the of such a For example, the that the of and their more and than alternative drugs The problem it addressed of is relevant to a of and is and serious to it a very significant Furthermore, the question was more effective than was stated and appropriate for current knowledge at the The change in clinical practice by the study was and to was predicted to result in a reduction in and mortality Indeed, practice and as a result A particular of the is that it a clear and is more effective than its and we should use it to of The of for the of still in an of great clinical and significance, as many questions as it reduces the of but whether we should give it to all is less certain This example that sometimes the of a paper's conclusion can its as well as the of the question in the in which the answer may a paper's value to the of information For example, a study of the effects of on in the and of the reduction in and other but not the or In whether the reduction in might be or for a particular patient, knowledge of the likely of the best and might be more relevant than how the of patients might be the value of the study is by this the value of the published paper is as a result. For studies in it is perhaps the effect on which is likely to be rather than any For example, in the the of no but certain better than more traditional An example from the of papers that we is the that is a is the finding that is fundamental to this to in a all in which the of the is The in these but there are many less well known For the of a that specific and that is to by the to consider the that this is a if not by which may be at the It is possible indeed to clinical and scientific of The of in are to from to a clinical problem to the a approach to its results would be to consider whether there is a fundamental reason why better than other and that this reason will help more as a However, we must that in related finding the between and clinical is not The of to is not always to and can often be by the to control all the variables in clinical of the type of question which can equally to clinical and is its and – if the study's results can be into an are to the more than and this can sometimes to the of proper scientific by a A example is the of study of the between and in the impact factor – see and then the subsequent of this paper In the scientific a very of the of studies appears to have in upon the type of question This to have been by the with for and we discuss these further below. The of the evidence in a paper first, on those methodological aspects second, on the greater to certain types of study over a by and When assessing the of the would do well to ask (as indeed the should have there possible for these results other than the conclusions methodology and attention to the above aspects may or many of these alternative but the more that and the less the that should be to a of results. This approach may seem an negative but we to of it as a of and for a scale of the of evidence in studies is in The emphasis on as a tool has been much criticised Some have that is, in a of rather than that has no place in scientific studies to test hypotheses The to clinical and not to scientific argue that fundamental hypotheses in can only be addressed with the scientific as by a conducted experiment to test the prediction from the hypothesis, and not by simply the results of different of It is possible that is more for those studies of a drug or or for rather than problem with is that it reports and to the of the of evidence In many areas of in the more such as reports can have a very persuasive effect on an clinical practice, since they are based on clinical and not on trials in which may be masked by into and summary statistics. If a a drug and for example, then that drug may cause the results of a clinical Case reports are for very for very or when they describe unexpected clinical (or We the above term to those aspects of a paper not relating to its scientific or but which nonetheless seem to have an and influence on the in which papers are by the scientific or clinical The the was out is one of these For example, from the of or is likely to have greater impact than of say, a The can be criticised for this the obtain better and can more which in is as more since it from the and so on A related issue is the for authors to their to the quality and they to its importance. A quality is by a the impact factor index of how often papers published are by other the impact factors have some a paper published say, a will have very of its scientific quality and so is to be widely though, impact factors are a poor measure of a and are to by However, some authors and perhaps also the to a paper's by the in which it is The of the authors (or at least one of may also have an A study a current fundamental hypothesis is more likely to be and if a is an It is perhaps that an is to time on Finally, the of a study can influence its importance. A study's support from a (e.g. or gives the impression that the or at least its preliminary has some and that in a it has been discuss this further The of areas for research by the and superficially but may the influence of certain since it is that studies the current areas are more However, aspects can of that the impact of a paper (e.g. when has been obtained from to support a paper relating to research However, such may be and we have how an might critical a have developed by which they or a process of to their which research to the the study is conducted or the that the published from the study will have a value to this the published by which it research It is clear that particular emphasis is given to that have a impact on practice or scientific and to those that are important in terms of or knowledge of The for of each to support research and research and the is the method used to these to the the most The used by although not made seem to place emphasis on the of to by of as those to the in and to papers in impact factor In the they receive to using our impression of the types of paper which seem to more in this are to much criticism and whether they critical is to However, they de the means by which research is If a as on research which does not a value in these then these will ensure (as they are to that the as so will research and research and then it may impossible for that to conduct any research at all. Many authors have how is in such a So while critical of papers by an (or by a might to one by may to a different The for a in this is to consider the to which it to (or of from We have various aspects of a paper that can contribute to its We here by an approach to a paper, based on the a hypothetical of with the the most papers and the the The is to go each paper and at the of the be to place it in the appropriate of the above factors – aspects of the study's conduct of other aspects of its methodology not covered the of the question and the answers the of the evidence and any persuasive factors – should be at intervals when a paper, and each has the to the paper up or to the impression The steps are in this process the paper can be or many times up in its final in two First, every paper will have some value, even if it up in the any a paper should be for what is all a considerable in the current Second, different may place the same paper into different final – but this is a and appropriate of value in any of the paper with or the should be reasons for the paper to a particular and for being in various (or by all aspects of the paper. This it to to others how the process of and is a of discussion and which is a process in We also by our to some One of our is to the value of all any given paper, the should can we further on the approach used to answer the question The specific answer to this will in on the of the paper we have indicated above some of these might However, in to these by a better study or the researcher will the serious problems the as a those related to discussion above two important questions to First, if is so critical to studies to be conducted, then we should not to what we alone value, but also ‘What do the Second, any given which as a there are always a of possible questions we might It is here to specific questions or hypotheses this will the value of our It is possible that for the of to our others place value on hypotheses different from those the has We hope that our discussion above might generate a which answer these two is an of this in in all aspects of the conduct and of is the of the of with the report on the for and Both have of and clinical and have papers that they have been more The above are the and do not of the of the of of and or the of
Community
0 commentsNo discussion yet
Be the first to share a question or observation.