A Consideration of the Measurement and Reporting of Interrater Reliability
Abstract
Discussion Points1Cruz et al1Cruz C.O. Meshberg E.G. Shofer F.S. et al.Interrater reliability and accuracy of clinicians and trained research assistants performing prospective data collection in emergency department patients with potential acute coronary syndrome.Ann Emerg Med. 2009; 54: 1-7Abstract Full Text Full Text PDF PubMed Scopus (13) Google Scholar contains 2 parts, a comparison of the values gathered by trained research assistants and physicians about historical information in chest pain patients and the comparison of these participants' recordings with a “correct” value for each item.A. For each part, indicate whether the authors are studying reliability or validity and explain the difference between these concepts.B. What did the authors use as their criterion standard for the validity analysis?C. What are potential problems with their method of defining the criterion (gold) standard? Can you think of alternative approaches?D. The authors report crude agreement and interquartile range for their validity analysis. What part of a distribution is described by the interquartile range? List other statistics used to describe the validity of a measure and why they might be preferable to reporting crude agreement.2Tabled 1MD Recorded “Yes”MD Recorded “No”TotalRA recorded yes1176123RA recorded no18220Total1358143MD, Medical doctor; RA, research assistant. Open table in a new tab A. Calculate the crude percentage agreement for this table. What is the range of possible values for percentage agreement?B. Calculate Cohen's κ for this table. What is the formula for κ for raters making a binary assessment (eg, yes/no or true/false)? Discuss the purpose of Cohen's κ, its range, and the interpretations of key values such as –1, 0, and 1.C. What other measures can be used to measure reliability for binary, categorical, and continuous data? 3Cruz et al quote the oft-cited Landis and Koch2Landis J.R. Koch G.C. The measurement of observer agreement for categorical data.Biometrics. 1977; 33: 159-174Crossref PubMed Scopus (49675) Google Scholar article stating that a κ of “less than 0.2 represents poor agreement; 0.21 to 0.40, fair agreement; 0.41 to 0.60, moderate agreement; 0.61 to 0.80, good agreement; and 0.81 to 1.00, excellent agreement.” Consider studies of the agreement of airline pilots deciding whether it is safe to land and psychologists deciding whether interviewees have type A or type B personalities. the studies the κ the by Landis and Koch be 2 are in of a and to such as is a a or by a in the for and in the for are and the are to a for each that they percentage agreement is and κ is of are and are that the is the for a the the this the percentage agreement and κ for the the of the are and of the are by the of the are and of the are by the the are and of the are by the and of the are and of the are by the Discuss the of percentage agreement and κ in these Consider the 2 and percentage agreement and κ for is κ the What this that the table the described and that that to 2 in the raters are that are and the raters are that be with with or in with each of κ the in these 2 the of κ, that such that and or are for the and percentage agreement and κ for these is the measure for Consider the of the raters in the in this be reliability is might this be the percentage agreement κ for the in of et The are to indicate the in the and 2 Open table in a new tab A. in the table are with the the pain it to the it to the it to the for these Can you explain why these have percentage agreement you is the the Can you the between the of the in the table and the to κ percentage the problems with percentage agreement and κ in these you think it be the in the of each table of reporting the percentage agreement or et al contains 2 parts, a comparison of the values gathered by trained research assistants and physicians historical information in chest pain and the comparison of these participants' recordings with a “correct” value for each For each part, indicate whether the authors are studying reliability or validity and explain the difference between these part is assessment of and the is assessment of The between reliability and validity is the that the in a that a a The reliability of a to the agreement the the or assessment of validity a observer a or the criterion standard is to be validity studies report the of the observer statistics such as and or reliability such as percentage agreement or What did the authors use as their criterion standard for the validity the and the research it is that their is they a research the of the 2 is What are potential problems with their method of defining the standard? Can you think of alternative a standard for this is For can be 2 the and is For a you have pain in the might that is a for in the is a might that is its the with is the criterion standard for this the the have or the the information The of is that have to the emergency have the of reporting part of a to their and the the a be or the other of the physicians the in a that to the or or in the a the the or are or whether they are to the and the patients be in or to the authors have to the in the research and each and the of to a in accuracy with The authors report crude agreement and interquartile range for their validity analysis. What part of a distribution is described by the interquartile range? List other statistics used to describe the validity of a measure and why they might be preferable to reporting crude interquartile range to the of a of is a that represents the the to the this is the and the the can be by the to the that a distribution the is used to these The is the the the and the the The is the difference between the and is a by than the range of a and it is data are in the of a the and are to and and the the or you the to the and you the or a to in the research and did the authors report the percentage agreement with the “correct” by the criterion agreement is a for a reliability is the to describe this validity assessment of a observer with a criterion that are to a validity report statistics such as and or reliability such as percentage agreement or percentage agreement is a to report Consider the table for the the of the chest pain or 1MD Recorded “Yes”MD Recorded “No”TotalRA recorded recorded Open table in a new tab Calculate the crude percentage agreement for this table. What is the range of possible values for percentage 2 and 2 crude percentage agreement for this table is agreement can range between and Calculate Cohen's κ for this table. What is the formula for κ for raters making a binary assessment (eg, yes/no or true/false)? Discuss the purpose of Cohen's κ, its range, and the interpretations of key values such as –1, 0, and κ by A of agreement for Scopus Google Scholar in is to to is with the method as For it is to to the by with and the and and The and of the a to these the of the they are to as that the values of a to can The is in the the The agreement for this table the 2 raters recorded the or are and κ the to the percentage of agreement to for each agreement and these are to the the values κ is as The value of to The value of to The percentage agreement to to the of this that agreement or have that the agreement to is these 2 the κ formula as a of agreement for A of agreement for Scopus Google Scholar to measure agreement κ can range that agreement than by to agreement percentage of percentage agreement to A κ of that agreement is that by percentage of the κ is that the of the agreement table of the in that are and the these and agreement be a of the explain these in in What other measures can be used to measure reliability for binary, categorical, and continuous can be with a of excellent for Scholar that is about is and that method is for is to of data are a of such as or and in a such as to A binary is a categorical with 2 or or as can of values be and measurement accuracy continuous is to the reliability are for use with continuous and whether a is to their measure data with a a of indicate that the 2 are with each other are as in the of the or be a in the with is that 2 be as in the Open table in a new tab The A new measure of Scholar and The and measurement of between Google Scholar measure the of between 2 and is a to measure the of a in a of agreement of A of agreement in a value of a between 2 a of in a value of A of is that its range the a as this and between and a Scholar and is to in of to measure can be used to measure in PubMed Scopus Google Scholar The the raters a to the and that physicians use a to the of acute coronary in each of patients The 2 for each for each the the raters for each The in for is with the of the is in the patients than is in the A that the raters have good a the for each to be the each are the the be as the raters are A of have the of the statistics can to the that the the about agreement and are the by the is a of agreement data that the difference in for each their for agreement between of PubMed Scopus Google Scholar Consider a that measures 2 in et al quote the oft-cited Landis and Koch that a κ of “less than 0.2 represents poor agreement; 0.21 to 0.40, fair agreement; 0.41 to 0.60, moderate agreement; 0.61 to 0.80, good agreement; and 0.81 to 1.00, excellent agreement.” Consider studies of the agreement of airline pilots deciding whether it is safe to land and psychologists deciding whether interviewees have type A or type B personalities. the studies the κ the by Landis and Koch be their κ values to by Landis and Koch2Landis J.R. Koch G.C. The measurement of observer agreement for categorical data.Biometrics. 1977; 33: 159-174Crossref PubMed Scopus (49675) Google Scholar and by for and Scholar the of values of κ to the and excellent is with A κ of might be good the of is as than agreement is the be a κ of it safe to (eg, a of historical that are used to patients for might be their are (eg, a of and data that are used to patients with pain can be they are is be used by of agreement as or in between of the a 2 are in of a and to such as is a a or by a in the for and in the for are and the are to a for each that they percentage agreement is and κ is of of are and are that the is the for a the the this the percentage agreement and κ for the the of the are and of the are by the of the are and of the are by the the are and of the are by the of the are and of the are by the Discuss the of percentage agreement and κ in these is to that in κ can are and percentage agreement is that 2 raters of these are between and or the and 2 2 of the possible that whether the or or each that of the with possible percentage agreement and and 2 possible The of the each the κ a than the are and Open table in a new tab that in these 2 of the raters are to and The agreement be the in the value of percentage agreement to as for κ, κ is in than in raters are performing in percentage agreement the of are and are can be by the and and The in each about are the for the is that in and in and Open table in a new tab by the might in a or or agreement a or these to and the percentage agreement and κ range between and to and and Open table in a new tab a possible of are and are by the and The in raters that of the are might in κ is for this table the agreement is than that to and in and Open table in a new tab The might in a table. that that of the and of the as have that this raters the as the percentage agreement and κ range between and and and to and and and and Open table in a new tab of Open table in a new tab that the of the raters is the each raters can the they the percentage of can in a range of κ, agreement be the range of κ as the of The of the κ is that agreement to as the is of that that is and can to κ values that agreement is agreement the problems of Full Text PDF PubMed Scopus Google agreement the Full Text PDF PubMed Scopus Google Consider the 2 and percentage agreement and κ for is κ the What this the in The percentage agreement is the in of these κ is in the table its are κ the and this to the and to agreement by agreement is are with the are is the are of κ κ that are is to be this 2 raters or each of a fair with of and is they have of and and that their agreement to be have to than for to about or raters are that the be to raters to have of and and to by of the κ that indicate agreement to are in the in in the raters making it κ is used to measure the reliability of 2 that a or of The the that of the be to agreement in the the they be this be For that the percentage agreement is a and as reporting the table is the method of reliability the of κ, the that such that and or are for the and percentage agreement and κ for these is the measure for Consider the of the raters in the in this be reliability is might this be to between 2 the raters with a percentage whether the values of they are are or in is and the other are the be or that in this the of the raters the values of the and the the values of these the raters have they to of to to or they the of the they their of the to their a it is the value of the that that value of the and agreement is by is that is to table and have to it and The table and the are agreement and κ and and the of the the problems in raters to the and their of the their with the agreement is to the and κ is for this the percentage agreement κ for the et The are to indicate the in the in the table are with the the pain it to the it to the it to the for these Can you explain why these have percentage agreement you is the the percentage table κ its are than of the table. the of the and with the of the and for each and and and and each these 2 κ these values to percentage agreement to of reliability is that the reliability data in the agreement is to it is is in that it for agreement that by think of and κ a and are and the to use the values to that each and research these of the and they information about the or of chest pain that might their the of κ are and that the it to the than pain the agreement A B Open table in a new tab The in these to a the table agreement for these the data a of the agreement with for or raters to a that by to that agreement of agreement for the to with are with part of that be to assessment of and raters their in and to their in a is criterion standard to the validity of be that of the in this did in their of of these the of be a to defining and for the of agreement than Can you the between the of the in the table and the to κ percentage that κ is to a to percentage agreement percentage agreement is and is with a The of a with data that the are as in the to as agreement and κ a percentage that the in κ the in of the et al article have to with the of the than of the of in the of the are to have of the reliability of the the problems with percentage agreement and κ in these you think it be the in the of each of reporting the percentage agreement or that this of the and that can a table data is to a reliability such as of reporting the reliability data than percentage agreement or information in it is to in the Discussion Points1Cruz et al1Cruz C.O. Meshberg E.G. Shofer F.S. et al.Interrater reliability and accuracy of clinicians and trained research assistants performing prospective data collection in emergency department patients with potential acute coronary syndrome.Ann Emerg Med. 2009; 54: 1-7Abstract Full Text Full Text PDF PubMed Scopus (13) Google Scholar contains 2 parts, a comparison of the values gathered by trained research assistants and physicians about historical information in chest pain patients and the comparison of these participants' recordings with a “correct” value for each item.A. For each part, indicate whether the authors are studying reliability or validity and explain the difference between these concepts.B. What did the authors use as their criterion standard for the validity analysis?C. What are potential problems with their method of defining the criterion (gold) standard? Can you think of alternative approaches?D. The authors report crude agreement and interquartile range for their validity analysis. What part of a distribution is described by the interquartile range? List other statistics used to describe the validity of a measure and why they might be preferable to reporting crude agreement.2Tabled 1MD Recorded “Yes”MD Recorded “No”TotalRA recorded yes1176123RA recorded no18220Total1358143MD, Medical doctor; RA, research assistant. Open table in a new tab A. Calculate the crude percentage agreement for this table. What is the range of possible values for percentage agreement?B. Calculate Cohen's κ for this table. What is the formula for κ for raters making a binary assessment (eg, yes/no or true/false)? Discuss the purpose of Cohen's κ, its range, and the interpretations of key values such as –1, 0, and 1.C. What other measures can be used to measure reliability for binary, categorical, and continuous data? 3Cruz et al quote the oft-cited Landis and Koch2Landis J.R. Koch G.C. The measurement of observer agreement for categorical data.Biometrics. 1977; 33: 159-174Crossref PubMed Scopus (49675) Google Scholar article stating that a κ of “less than 0.2 represents poor agreement; 0.21 to 0.40, fair agreement; 0.41 to 0.60, moderate agreement; 0.61 to 0.80, good agreement; and 0.81 to 1.00, excellent agreement.” Consider studies of the agreement of airline pilots deciding whether it is safe to land and psychologists deciding whether interviewees have type A or type B personalities. the studies the κ the by Landis and Koch be 2 are in of a and to such as is a a or by a in the for and in the for are and the are to a for each that they percentage agreement is and κ is of are and are that the is the for a the the this the percentage agreement and κ for the the of the are and of the are by the of the are and of the are by the the are and of the are by the and of the are and of the are by the Discuss the of percentage agreement and κ in these Consider the 2 and percentage agreement and κ for is κ the What this that the table the described and that that to 2 in the raters are that are and the raters are that be with with or in with each of κ the in these 2 the of κ, that such that and or are for the and percentage agreement and κ for these is the measure for Consider the of the raters in the in this be reliability is might this be the percentage agreement κ for the in of et The are to indicate the in the and 2 Open table in a new tab A. in the table are with the the pain it to the it to the it to the for these Can you explain why these have percentage agreement you is the the Can you the between the of the in the table and the to κ percentage the problems with percentage agreement and κ in these you think it be the in the of each table of reporting the percentage agreement or et al1Cruz C.O. Meshberg E.G. Shofer F.S. et al.Interrater reliability and accuracy of clinicians and trained research assistants performing prospective data collection in emergency department patients with potential acute coronary syndrome.Ann Emerg Med. 2009; 54: 1-7Abstract Full Text Full Text PDF PubMed Scopus (13) Google Scholar contains 2 parts, a comparison of the values gathered by trained research assistants and physicians about historical information in chest pain patients and the comparison of these participants' recordings with a “correct” value for each item.A. For each part, indicate whether the authors are studying reliability or validity and explain the difference between these concepts.B. What did the authors use as their criterion standard for the validity analysis?C. What are potential problems with their method of defining the criterion (gold) standard? Can you think of alternative approaches?D. The authors report crude agreement and interquartile range for their validity analysis. What part of a distribution is described by the interquartile range? List other statistics used to describe the validity of a measure and why they might be preferable to reporting crude agreement.2Tabled 1MD Recorded “Yes”MD Recorded “No”TotalRA recorded yes1176123RA recorded no18220Total1358143MD, Medical doctor; RA, research assistant. Open table in a new tab A. Calculate the crude percentage agreement for this table. What is the range of possible values for percentage agreement?B. Calculate Cohen's κ for this table. What is the formula for κ for raters making a binary assessment (eg, yes/no or true/false)? Discuss the purpose of Cohen's κ, its range, and the interpretations of key values such as –1, 0, and 1.C. What other measures can be used to measure reliability for binary, categorical, and continuous data? 3Cruz et al quote the oft-cited Landis and Koch2Landis J.R. Koch G.C. The measurement of observer agreement for categorical data.Biometrics. 1977; 33: 159-174Crossref PubMed Scopus (49675) Google Scholar article stating that a κ of “less than 0.2 represents poor agreement; 0.21 to 0.40, fair agreement; 0.41 to 0.60, moderate agreement; 0.61 to 0.80, good agreement; and 0.81 to 1.00, excellent agreement.” Consider studies of the agreement of airline pilots deciding whether it is safe to land and psychologists deciding whether interviewees have type A or type B personalities. the studies the κ the by Landis and Koch be 2 are in of a and to such as is a a or by a in the for and in the for are and the are to a for each that they percentage agreement is and κ is of are and are that the is the for a the the this the percentage agreement and κ for the the of the are and of the are by the of the are and of the are by the the are and of the are by the and of the are and of the are by the Discuss the of percentage agreement and κ in these Consider the 2 and percentage agreement and κ for is κ the What this that the table the described and that that to 2 in the raters are that are and the raters are that be with with or in with each of κ the in these 2 the of κ, that such that and or are for the and Calculate percentage agreement and κ for these is the measure for Consider the of the raters in the in this be reliability is might this be the percentage agreement κ for the in of et The are to indicate the in the and 2 Open table in a new tab A. in the table are with the the pain it to the it to the it to the for these Can you explain why these have percentage agreement you is the the Can you the between the of the in the table and the to κ percentage the problems with percentage agreement and κ in these you think it be the in the of each table of reporting the percentage agreement or et al contains 2 parts, a comparison of the values gathered by trained research assistants and physicians historical information in chest pain and the comparison of these participants' recordings with a “correct” value for each For each part, indicate whether the authors are studying reliability or validity and explain the difference between these part is assessment of and the is assessment of The between reliability and validity is the that the in a that a a The reliability of a to the agreement the the or assessment of validity a observer a or the criterion standard is to be validity studies report the of the observer statistics such as and or reliability such as percentage agreement or What did the authors use as their criterion standard for the validity the and the research it is that their is they a research the of the 2 is What are potential problems with their method of defining the standard? Can you think of alternative a standard for this is For can be 2 the and is For a you have pain in the might that is a for in the is a might that is its the with is the criterion standard for this the the have or the the information The of is that have to the emergency have the of reporting part of a to their and the the a be or the other of the physicians the in a that to the or or in the a the the or are or whether they are to the and the patients be in or to the authors have to the in the research and each and the of to a in accuracy with The authors report crude agreement and interquartile range for their validity analysis. What part of a distribution is described by the interquartile range? List other statistics used to describe the validity of a measure and why they might be preferable to reporting crude interquartile range to the of a of is a that represents the the to the this is the and the the can be by the to the that a distribution the is used to these The is the the the and the the The is the difference between the and is a by than the range of a and it is data are in the of a the and are to and and the the or you the to the and you the or a to in the research and did the authors report the percentage agreement with the “correct” by the criterion agreement is a for a reliability is the to describe this validity assessment of a observer with a criterion that are to a validity report statistics such as and or reliability such as percentage agreement or et al contains 2 parts, a comparison of the values gathered by trained research assistants and physicians historical information in chest pain and the comparison of these participants' recordings with a “correct” value for each For each part, indicate whether the authors are studying reliability or validity and explain the difference between these The part is assessment of and the is assessment of The between reliability and validity is the that the in a that a a The reliability of a to the agreement the the or assessment of validity a observer a or the criterion standard is to be validity studies report the of the observer statistics such as and or reliability such as percentage agreement or What did the authors use as their criterion standard for the validity the and the research it is that their is they a research the of the 2 is What are potential problems with their method of defining the standard? Can you think of alternative a standard for this is For can be 2 the and is For a you have pain in the might that is a for in the is a might that is its the with What is the criterion standard for this the the have or the the information The of is that have to the emergency have the of reporting part of a to their and the the a be or the other of the physicians the in a that to the or or in the a the the or are or whether they are to the and the patients be in or to A the authors have to the in the research and each and the of to a
Community
0 commentsNo discussion yet
Be the first to share a question or observation.